EDBT 2026 Demo / reviewers in the wild / expert
Sayan Mukherjee 0001
dblp:52/5375-1 · also Shayan Mukherjee
· DBLP profile ↗
44ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-6715-3920ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Higher Order Bipartiteness vs Bi-Partitioning in Simplicial ComplexesabstractBipartite graphs are a fundamental concept in graph theory with diverse applications. A graph is bipartite iff it contains no odd cycles, a characteristic that has many implications in diverse fields ranging from matching problems to the construction of complex networks. Another key identifying feature is their Laplacian spectrum as bipartite graphs achieve the maximum possible eigenvalue of graph Laplacian. However, for modeling higher-order connections in complex systems, hypergraphs and simplicial complexes are required due to the limitations of graphs in representing pairwise interactions. In this article, using simple tools from graph theory, we extend the cycle-based characterization from bipartite graphs to those simplicial complexes that achieve the maximum Hodge Laplacian eigenvalue, known as disorientable simplicial complexes. We show that a $N$-dimensional simplicial complex is disorientable if its down dual graph contains no simple odd cycle of distinct edges and no twisted even cycle of distinct edges. Furthermore, we see that in a $N$-simplicial complex without twisting cycles, the fewer the number of (non-branching) simple odd cycles in its down dual graph, the closer is its maximum eigenvalue to the possible maximum eigenvalue of Hodge Laplacian. Similar to the graph case, the absence of odd cycles plays a crucial role in solving the bi-partitioning problem of simplexes in higher dimensions. Marzieh Eidi, Sayan Mukherjee 0001 |
SoCG | 2 |
| 2025 | Developing PeACE: A Pedagogical Agent for Children's Emotions
Tyler Colasante, Eric Roldán Roa, Doris Kristina Raave, Juan Carlos Ramos Martinez, Tina Malti, Sayan Mukherjee 0001 |
EC-TEL (2) | 7 |
| 2024 | Play My Math: First Development Cycle of an EdTech Tool Supporting the Teaching and Learning of Fractions Through Music in Algebraic Notation
Eric Roldán Roa, Érika Roldán, Doris Kristina Raave, Jo Van Herwegen, Nina Polytimou, Sayan Mukherjee 0001, Tyler Colasante, Tina Malti, Julia Mori, Marcus Specht |
EC-TEL (2) | 6 |
| 2023 | Global optimality of Elman-type RNNs in the mean-field regimeabstractWe analyze Elman-type recurrent neural networks (RNNs) and their training in the mean-field regime. Specifically, we show convergence of gradient descent training dynamics of the RNN to the corresponding mean-field formulation in the large width limit. We also show that the fixed points of the limiting infinite-width dynamics are globally optimal, under some assumptions on the initialization of the weights. Our results establish optimality for feature-learning with wide RNNs in the mean-field regime. Andrea Agazzi, Sayan Mukherjee 0001 |
ICML | 3 |
| 2023 | Asymptotics of Bayesian Uncertainty Estimation in Random Features RegressionabstractIn this paper we compare and contrast the behavior of the posterior predictive distribution to the risk of the
the maximum a posteriori estimator for the random features regression model in the overparameterized regime. We will focus on the variance of the posterior predictive distribution (Bayesian model average) and compare its asymptotics to that of the risk of the MAP estimator. In the regime where the model dimensions grow faster than any constant multiple of the number of samples, asymptotic agreement between these two quantities is governed by the phase transition in the signal-to-noise ratio. They also asymptotically agree with each other when the number of samples grow faster than any constant multiple of model dimensions. Numerical simulations illustrate finer distributional properties of the two quantities for finite dimensions. We conjecture they have Gaussian fluctuations and exhibit similar properties as found by previous authors in a Gaussian sequence model, this is of independent theoretical interest. Youngsoo Baek, Samuel Berchuck, Sayan Mukherjee 0001 |
NeurIPS | 3 |
| 2023 | Ergodic theorems for dynamic imprecise probability kinematics
Michele Caprio, Sayan Mukherjee 0001 |
Int. J. Approx. Reason. | 2 |
| 2022 | Bayesian Multinomial Logistic Normal Models through Marginally Latent Matrix-T ProcessesabstractBayesian multinomial logistic-normal (MLN) models are popular for the analysis of sequence count data (e.g., microbiome or gene expression data) due to their ability to model multivariate count data with complex covariance structure. However, existing implementations of MLN models are limited to small datasets due to the non-conjugacy of the multinomial and logistic-normal distributions. Motivated by the need to develop efficient inference for Bayesian MLN models, we develop two key ideas. First, we develop the class of Marginally Latent Matrix-T Process (Marginally LTP) models. We demonstrate that many popular MLN models, including those with latent linear, non-linear, and dynamic linear structure are special cases of this class. Second, we develop an efficient inference scheme for Marginally LTP models with specific accelerations for the MLN subclass. Through application to MLN models, we demonstrate that our inference scheme are both highly accurate and often 4-5 orders of magnitude faster than MCMC. Justin D. Silverman, Kimberly Roche, Zachary C. Holmes, Lawrence A. David, Sayan Mukherjee 0001 |
J. Mach. Learn. Res. | 5 |
| 2022 | The accuracy of absolute differential abundance analysis from relative count dataabstractConcerns have been raised about the use of relative abundance data derived from next generation sequencing as a proxy for absolute abundances. For example, in the differential abundance setting, compositional effects in relative abundance data may give rise to spurious differences (false positives) when considered from the absolute perspective. In practice however, relative abundances are often transformed by renormalization strategies intended to compensate for these effects and the scope of the practical problem remains unclear. We used simulated data to explore the consistency of differential abundance calling on renormalized relative abundances versus absolute abundances and find that, while overall consistency is high, with a median sensitivity (true positive rates) of 0.91 and specificity (1-false positive rates) of 0.89, consistency can be much lower where there is widespread change in the abundance of features across conditions. We confirm these findings on a large number of real data sets drawn from 16S metabarcoding, expression array, bulk RNA-seq, and single-cell RNA-seq experiments, where data sets with the greatest change between experimental conditions are also those with the highest false positive rates. Finally, we evaluate the predictive utility of summary features of relative abundance data themselves. Estimates of sparsity and the prevalence of feature-level change in relative abundance data give reasonable predictions of discrepancy in differential abundance calling in simulated data and can provide useful bounds for worst-case outcomes in real data. Kimberly Roche, Sayan Mukherjee 0001 |
PLoS Comput. Biol. | 2 |
| 2022 | A topological data analytic approach for discovering biophysical signatures in protein dynamicsabstractIdentifying structural differences among proteins can be a non-trivial task. When contrasting ensembles of protein structures obtained from molecular dynamics simulations, biologically-relevant features can be easily overshadowed by spurious fluctuations. Here, we present SINATRA Pro, a computational pipeline designed to robustly identify topological differences between two sets of protein structures. Algorithmically, SINATRA Pro works by first taking in the 3D atomic coordinates for each protein snapshot and summarizing them according to their underlying topology. Statistically significant topological features are then projected back onto a user-selected representative protein structure, thus facilitating the visual identification of biophysical signatures of different protein ensembles. We assess the ability of SINATRA Pro to detect minute conformational changes in five independent protein systems of varying complexities. In all test cases, SINATRA Pro identifies known structural features that have been validated by previous experimental and computational studies, as well as novel features that are also likely to be biologically-relevant according to the literature. These results highlight SINATRA Pro as a promising method for facilitating the non-trivial task of pattern recognition in trajectories resulting from molecular dynamics simulations, with substantially increased resolution. Wai-Shing Tang, Gabriel Monteiro da Silva, Henry Kirveslahti, Erin Skeens, Bibo Feng, Timothy Sudijono, Kevin K. Yang, Sayan Mukherjee 0001, Brenda M. Rubenstein, Lorin Crawford |
PLoS Comput. Biol. | 8 |
| 2021 | Statistical robustness of Markov chain Monte Carlo acceleratorsabstractStatistical machine learning often uses probabilistic models and algorithms, such as Markov Chain Monte Carlo (MCMC), to solve a wide range of problems. Probabilistic computations, often considered too slow on conventional processors, can be accelerated with specialized hardware by exploiting parallelism and optimizing the design using various approximation techniques. Current methodologies for evaluating correctness of probabilistic accelerators are often incomplete, mostly focusing only on end-point result quality ("accuracy"). It is important for hardware designers and domain experts to look beyond end-point "accuracy" and be aware of how hardware optimizations impact statistical properties. Xiangyu Zhang 0011, Ramin Bashizade, Sayan Mukherjee 0001, Alvin R. Lebeck |
ASPLOS | 4 |
| 2021 | A Bayesian hierarchical model to estimate DNA methylation conservation in colorectal tumorsabstractMOTIVATION: Conservation is broadly used to identify biologically important (epi)genomic regions. In the case of tumor growth, preferential conservation of DNA methylation can be used to identify areas of particular functional importance to the tumor. However, reliable assessment of methylation conservation based on multiple tissue samples per patient requires the decomposition of methylation variation at multiple levels. RESULTS: We developed a Bayesian hierarchical model that allows for variance decomposition of methylation on three levels: between-patient normal tissue variation, between-patient tumor-effect variation and within-patient tumor variation. We then defined a model-based conservation score to identify loci of reduced within-tumor methylation variation relative to between-patient variation. We fit the model to multi-sample methylation array data from 21 colorectal cancer (CRC) patients using a Monte Carlo Markov Chain algorithm (Stan). Sets of genes implicated in CRC tumorigenesis exhibited preferential conservation, demonstrating the model's ability to identify functionally relevant genes based on methylation conservation. A pathway analysis of preferentially conserved genes implicated several CRC relevant pathways and pathways related to neoantigen presentation and immune evasion. Our findings suggest that preferential methylation conservation may be used to identify novel gene targets that are not consistently mutated in CRC. The flexible structure makes the model amenable to the analysis of more complex multi-sample data structures. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in the NCBI GEO Database, under accession code GSE166212. The R analysis code is available at https://github.com/kevin-murgas/DNAmethylation-hierarchicalmodel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kevin A. Murgas, Yanlin Ma, Lidea Shahidi, Sayan Mukherjee 0001, Andrew S. Allen, Darryl Shibata, Marc D. Ryser |
Bioinform. | 4 |
| 2021 | The Geometry of Synchronization Problems and Learning Group Actions
Tingran Gao, Jacek Brodzki, Sayan Mukherjee 0001 |
Discret. Comput. Geom. | 3 |
| 2021 | Subspace Clustering through Sub-ClustersabstractThe problem of dimension reduction is of increasing importance in modern data analysis. In this paper, we consider modeling the collection of points in a high dimensional space as a union of low dimensional subspaces. In particular we propose a highly scalable sampling based algorithm that clusters the entire data via first spectral clustering of a small random sample followed by classifying or labeling the remaining out-of-sample points. The key idea is that this random subset borrows information across the entire dataset and that the problem of clustering points can be replaced with the more efficient problem of "clustering sub-clusters". We provide theoretical guarantees for our procedure. The numerical results indicate that for large datasets the proposed algorithm outperforms other state-of-the-art subspace clustering algorithms with respect to accuracy and speed. Jan Hannig, Sayan Mukherjee 0001 |
J. Mach. Learn. Res. | 3 |
| 2021 | Measuring and mitigating PCR bias in microbiota datasetsabstractPCR amplification plays an integral role in the measurement of mixed microbial communities via high-throughput DNA sequencing of the 16S ribosomal RNA (rRNA) gene. Yet PCR is also known to introduce multiple forms of bias in 16S rRNA studies. Here we present a paired modeling and experimental approach to characterize and mitigate PCR NPM-bias (PCR bias from non-primer-mismatch sources) in microbiota surveys. We use experimental data from mock bacterial communities to validate our approach and human gut microbiota samples to characterize PCR NPM-bias under real-world conditions. Our results suggest that PCR NPM-bias can skew estimates of microbial relative abundances by a factor of 4 or more, but that this bias can be mitigated using log-ratio linear models. Justin D. Silverman, Rachael J. Bloom, Sharon Jiang, Heather K. Durand, Eric Dallow, Sayan Mukherjee 0001, Lawrence A. David |
PLoS Comput. Biol. | 6 |
| 2018 | Scalable Algorithms for Learning High-Dimensional Linear Mixed Models
Zilong Tan, Kimberly Roche, Sayan Mukherjee 0001 |
UAI | 4 |
| 2017 | Partitioned Tensor Factorizations for Learning Mixed Membership ModelsabstractWe present an efficient algorithm for learning mixed membership models when the number of variables p is much larger than the number of hidden components k. This algorithm reduces the computational complexity of state-of-the-art tensor methods, which require decomposing an $O(p^3)$ tensor, to factorizing $O(p/k)$ sub-tensors each of size $O(k^3)$. In addition, we address the issue of negative entries in the empirical method of moments based estimators. We provide sufficient conditions under which our approach has provable guarantees. Our approach obtains competitive empirical results on both simulated and real data. Zilong Tan, Sayan Mukherjee 0001 |
ICML | 2 |
| 2017 | Adaptive Randomized Dimension Reduction on Massive DataabstractThe scalability of statistical estimators is of increasing importance in modern applications. One approach to implementing scalable algorithms is to compress data into a low dimensional latent space using dimension reduction methods. In this paper, we develop an approach for dimension reduction that exploits the assumption of low rank structure in high dimensional data to gain both computational and statistical advantages. We adapt recent randomized low-rank approximation algorithms to provide an efficient solution to principal component analysis (PCA), and we use this efficient solver to improve estimation in large- scale linear mixed models (LMM) for association mapping in statistical genomics. A key observation in this paper is that randomization serves a dual role, improving both computational and statistical performance by implicitly regularizing the covariance matrix estimate of the random effect in an LMM. These statistical and computational advantages are highlighted in our experiments on simulated data and large-scale genomic studies. Gregory Darnell, Stoyan Georgiev, Sayan Mukherjee 0001, Barbara E. Engelhardt |
J. Mach. Learn. Res. | 3 |
| 2016 | Bayesian group factor analysis with structured sparsityabstractLatent factor models are the canonical statistical tool for exploratory analyses of low-dimensional linear structure for a matrix of $p$ features across $n$ samples. We develop a structured Bayesian group factor analysis model that extends the factor model to multiple coupled observation matrices; in the case of two observations, this reduces to a Bayesian model of canonical correlation analysis. Here, we carefully define a structured Bayesian prior that encourages both element-wise and column-wise shrinkage and leads to desirable behavior on high- dimensional data. In particular, our model puts a structured prior on the joint factor loading matrix, regularizing at three levels, which enables element-wise sparsity and unsupervised recovery of latent factors corresponding to structured variance across arbitrary subsets of the observations. In addition, our structured prior allows for both dense and sparse latent factors so that covariation among either all features or only a subset of features can be recovered. We use fast parameter-expanded expectation-maximization for parameter estimation in this model. We validate our method on simulated data with substantial structure. We show results of our method applied to three high- dimensional data sets, comparing results against a number of state-of-the-art approaches. These results illustrate useful properties of our model, including i) recovering sparse signal in the presence of dense effects; ii) the ability to scale naturally to large numbers of observations; iii) flexible observation- and factor-specific regularization to recover factors with a wide variety of sparsity levels and percentage of variance explained; and iv) tractable inference that scales to modern genomic and text data sizes. Shiwen Zhao, Chuan Gao, Sayan Mukherjee 0001, Barbara E. Engelhardt |
J. Mach. Learn. Res. | 3 |
| 2015 | Contour trees of uncertain terrainsabstractWe study contour trees of terrains, which encode the topological changes of the level set of the height value ℓ as we raise ℓ from -∞ to +∞ on the terrains, in the presence of uncertainty in data. We assume that the terrain is represented by a piecewise-linear height function over a planar triangulation M, by specifying the height of each vertex. We study the case when M is fixed and the uncertainty lies in the height of each vertex in the triangulation, which is described by a probability distribution. We present efficient sampling-based Monte Carlo methods for estimating, with high probability, (i) the probability that two points lie on the same edge of the contour tree, within additive error; (ii) the expected distance of two points p, q and the probability that the distance of p, q is at least ℓ on the contour tree, within additive error, where the distance of p, q on a contour tree is defined to be the difference between the maximum height and the minimum height on the unique path from p to q on the contour tree. The main technical contribution of the paper is to prove that a small number of samples are sufficient to estimate these quantities. We present two applications of these algorithms, and also some experimental results to demonstrate the effectiveness of our approach. Wuzhou Zhang, Pankaj K. Agarwal, Sayan Mukherjee 0001 |
SIGSPATIAL/GIS | 3 |
| 2015 | Cumulon: Matrix-Based Data Analytics in the Cloud with Spot InstancesabstractWe describe Cümülön, a system aimed at helping users develop and deploy matrix-based data analysis programs in a public cloud. A key feature of Cümülön is its end-to-end support for the so-called spot instances ---machines whose market price fluctuates over time but is usually much lower than the regular fixed price. A user sets a bid price when acquiring spot instances, and loses them as soon as the market price exceeds the bid price. While spot instances can potentially save cost, they are difficult to use effectively, and run the risk of not finishing work while costing more. Cümülön provides a highly elastic computation and storage engine on top of spot instances, and offers automatic cost-based optimization of execution, deployment, and bidding strategies. Cümülön further quantifies how the uncertainty in the market price translates into the cost uncertainty of its recommendations, and allows users to specify their risk tolerance as an optimization constraint. Botong Huang, Nicholas W. D. Jarrett, Shivnath Babu, Sayan Mukherjee 0001, Jun Yang 0001 |
Proc. VLDB Endow. | 4 |
| 2015 | The Information Geometry of Mirror DescentabstractWe prove the equivalence of two online learning algorithms: 1) mirror descent and 2) natural gradient descent. Both mirror descent and natural gradient descent are generalizations of online gradient descent when the parameter of interest lies on a non-Euclidean manifold. Natural gradient descent selects the steepest descent along a Riemannian manifold by multiplying the standard gradient by the inverse of the metric tensor. Mirror descent induces non-Euclidean structure by solving iterative optimization problems using different proximity functions. In this paper, we prove that mirror descent induced by Bregman divergence proximity functions is equivalent to the natural gradient descent algorithm on the dual Riemannian manifold. We use techniques from convex analysis and connections between Riemannian manifolds, Bregman divergences, and convexity to prove this result. This equivalence between natural gradient descent and mirror descent, implies that: 1) mirror descent is the steepest descent direction along the Riemannian manifold corresponding to the choice of Bregman divergence and 2) mirror descent with log-likelihood loss applied to parameter estimation in exponential families asymptotically achieves the classical Cramér-Rao lower bound. Garvesh Raskutti, Sayan Mukherjee 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2014 | Fréchet Means for Distributions of Persistence Diagrams
Katharine Turner, Yuriy Mileyko, Sayan Mukherjee 0001, John Harer |
Discret. Comput. Geom. | 3 |
| 2013 | Genome-wide identification and predictive modeling of tissue-specific alternative polyadenylationabstractMOTIVATION: Pre-mRNA cleavage and polyadenylation are essential steps for 3'-end maturation and subsequent stability and degradation of mRNAs. This process is highly controlled by cis-regulatory elements surrounding the cleavage/polyadenylation sites (polyA sites), which are frequently constrained by sequence content and position. More than 50% of human transcripts have multiple functional polyA sites, and the specific use of alternative polyA sites (APA) results in isoforms with variable 3'-untranslated regions, thus potentially affecting gene regulation. Elucidating the regulatory mechanisms underlying differential polyA preferences in multiple cell types has been hindered both by the lack of suitable data on the precise location of cleavage sites, as well as of appropriate tests for determining APAs with significant differences across multiple libraries. RESULTS: We applied a tailored paired-end RNA-seq protocol to specifically probe the position of polyA sites in three human adult tissue types. We specified a linear-effects regression model to identify tissue-specific biases indicating regulated APA; the significance of differences between tissue types was assessed by an appropriately designed permutation test. This combination allowed to identify highly specific subsets of APA events in the individual tissue types. Predictive models successfully classified constitutive polyA sites from a biologically relevant background (auROC = 99.6%), as well as tissue-specific regulated sets from each other. We found that the main cis-regulatory elements described for polyadenylation are a strong, and highly informative, hallmark for constitutive sites only. Tissue-specific regulated sites were found to contain other regulatory motifs, with the canonical polyadenylation signal being nearly absent at brain-specific polyA sites. Together, our results contribute to the understanding of the diversity of post-transcriptional gene regulation. AVAILABILITY: Raw data are deposited on SRA, accession numbers: brain SRX208132, kidney SRX208087 and liver SRX208134. Processed datasets as well as model code are published on our website: http://www.genome.duke.edu/labs/ohler/research/UTR/. CONTACT: [email protected]. Dina Hafez, Ting Ni, Sayan Mukherjee 0001, Uwe Ohler |
Bioinform. | 3 |
| 2013 | A comparative study of covariance selection models for the inference of gene regulatory networksabstractMOTIVATION: The inference, or 'reverse-engineering', of gene regulatory networks from expression data and the description of the complex dependency structures among genes are open issues in modern molecular biology. RESULTS: In this paper we compared three regularized methods of covariance selection for the inference of gene regulatory networks, developed to circumvent the problems raising when the number of observations n is smaller than the number of genes p. The examined approaches provided three alternative estimates of the inverse covariance matrix: (a) the 'PINV' method is based on the Moore-Penrose pseudoinverse, (b) the 'RCM' method performs correlation between regression residuals and (c) 'ℓ(2C)' method maximizes a properly regularized log-likelihood function. Our extensive simulation studies showed that ℓ(2C) outperformed the other two methods having the most predictive partial correlation estimates and the highest values of sensitivity to infer conditional dependencies between genes even when a few number of observations was available. The application of this method for inferring gene networks of the isoprenoid biosynthesis pathways in Arabidopsis thaliana allowed to enlighten a negative partial correlation coefficient between the two hubs in the two isoprenoid pathways and, more importantly, provided an evidence of cross-talk between genes in the plastidial and the cytosolic pathways. When applied to gene expression data relative to a signature of HRAS oncogene in human cell cultures, the method revealed 9 genes (p-value<0.0005) directly interacting with HRAS, sharing the same Ras-responsive binding site for the transcription factor RREB1. This result suggests that the transcriptional activation of these genes is mediated by a common transcription factor downstream of Ras signaling. AVAILABILITY: Software implementing the methods in the form of Matlab scripts are available at: http://users.ba.cnr.it/issia/iesina18/CovSelModelsCodes.zip. Patrizia F. Stifanelli, Teresa Maria Creanza, Roberto Anglani, Vania C. Liuzzi, Sayan Mukherjee 0001, Francesco P. Schena, Nicola Ancona |
J. Biomed. Informatics | 5 |
| 2012 | Local homology transfer and stratification learningabstractThe objective of this paper is to show that point cloud data can under certain circumstances be clustered by strata in a plausible way. For our purposes, we consider a stratified space to be a collection of manifolds of different dimensions which are glued together in a locally trivial manner inside some Euclidean space. To adapt this abstract definition to the world of noise, we first define a multi-scale notion of stratified spaces, providing a stratification at different scales which are indexed by a radius parameter. We then use methods derived from kernel and cokernel persistent homology to cluster the data points into different strata. We prove a correctness guarantee for this clustering method under certain topological conditions. We then provide a probabilistic guarantee for the clustering for the point sample setting – we provide bounds on the minimum number of sample points required to state with high probability which points belong to the same strata. Finally, we give an explicit algorithm for the clustering. Paul Bendich, Bei Wang 0001, Sayan Mukherjee 0001 |
SODA | 3 |
| 2011 | Estimating variable structure and dependence in multitask learning via gradientsabstractWe consider the problem of hierarchical or multitask modeling where we simultaneously learn the regression function and the underlying geometry and dependence between variables. We demonstrate how the gradients of the multiple related regression functions over the tasks allow for dimension reduction and inference of dependencies across tasks jointly and for each task individually. We provide Tikhonov regularization algorithms for both classification and regression that are efficient and robust for high-dimensional data, and a mechanism for incorporating a priori knowledge of task (dis)similarity into this framework. The utility of this method is illustrated on simulated and real data. Justin Guinney, Qiang Wu 0003, Sayan Mukherjee 0001 |
Mach. Learn. | 3 |
| 2010 | On the reproducibility of results of pathway analysis in genome-wide expression studies of colorectal cancers
Rosalia Maglietta, Angela Distaso, Ada Piepoli, Orazio Palumbo, Massimo Carella, Annarita D'Addabbo, Sayan Mukherjee 0001, Nicola Ancona |
J. Biomed. Informatics | 7 |
| 2010 | Learning Gradients: Predictive Models that Infer Geometry and Statistical Dependence
Qiang Wu 0003, Justin Guinney, Mauro Maggioni, Sayan Mukherjee 0001 |
J. Mach. Learn. Res. | 4 |
| 2009 | Comparative study of gene set enrichment methodsabstractBACKGROUND: The analysis of high-throughput gene expression data with respect to sets of genes rather than individual genes has many advantages. A variety of methods have been developed for assessing the enrichment of sets of genes with respect to differential expression. In this paper we provide a comparative study of four of these methods: Fisher's exact test, Gene Set Enrichment Analysis (GSEA), Random-Sets (RS), and Gene List Analysis with Prediction Accuracy (GLAPA). The first three methods use associative statistics, while the fourth uses predictive statistics. We first compare all four methods on simulated data sets to verify that Fisher's exact test is markedly worse than the other three approaches. We then validate the other three methods on seven real data sets with known genetic perturbations and then compare the methods on two cancer data sets where our a priori knowledge is limited. RESULTS: The simulation study highlights that none of the three method outperforms all others consistently. GSEA and RS are able to detect weak signals of deregulation and they perform differently when genes in a gene set are both differentially up and down regulated. GLAPA is more conservative and large differences between the two phenotypes are required to allow the method to detect differential deregulation in gene sets. This is due to the fact that the enrichment statistic in GLAPA is prediction error which is a stronger criteria than classical two sample statistic as used in RS and GSEA. This was reflected in the analysis on real data sets as GSEA and RS were seen to be significant for particular gene sets while GLAPA was not, suggesting a small effect size. We find that the rank of gene set enrichment induced by GLAPA is more similar to RS than GSEA. More importantly, the rankings of the three methods share significant overlap. CONCLUSION: The three methods considered in our study recover relevant gene sets known to be deregulated in the experimental conditions and pathologies analyzed. There are differences between the three methods and GSEA seems to be more consistent in finding enriched gene sets, although no method uniformly dominates over all data sets. Our analysis highlights the deep difference existing between associative and predictive methods for detecting enrichment and the use of both to better interpret results of pathway analysis. We close with suggestions for users of gene set methods. Luca Abatangelo, Rosalia Maglietta, Angela Distaso, Annarita D'Addabbo, Teresa Maria Creanza, Sayan Mukherjee 0001, Nicola Ancona |
BMC Bioinform. | 6 |
| 2008 | Statistical Assessment of MSigDB Gene Sets in Colon Cancer
Angela Distaso, Luca Abatangelo, Rosalia Maglietta, Teresa Maria Creanza, Ada Piepoli, Massimo Carella, Annarita D'Addabbo, Sayan Mukherjee 0001, Nicola Ancona |
KES (2) | 8 |
| 2008 | Localized Sliced Inverse RegressionabstractWe developed localized sliced inverse regression for supervised dimension reduction. It has the advantages of preventing degeneracy, increasing estimation accuracy, and automatic subclass discovery in classification problems. A semisupervised version is proposed for the use of unlabeled data. The utility is illustrated on simulated as well as real data sets. Qiang Wu 0003, Sayan Mukherjee 0001 |
NIPS | 2 |
| 2008 | Modeling Cancer Progression via Pathway DependenciesabstractCancer is a heterogeneous disease often requiring a complexity of alterations to drive a normal cell to a malignancy and ultimately to a metastatic state. Certain genetic perturbations have been implicated for initiation and progression. However, to a great extent, underlying mechanisms often remain elusive. These genetic perturbations are most likely reflected by the altered expression of sets of genes or pathways, rather than individual genes, thus creating a need for models of deregulation of pathways to help provide an understanding of the mechanisms of tumorigenesis. We introduce an integrative hierarchical analysis of tumor progression that discovers which a priori defined pathways are relevant either throughout or in particular steps of progression. Pathway interaction networks are inferred for these relevant pathways over the steps in progression. This is followed by the refinement of the relevant pathways to those genes most differentially expressed in particular disease stages. The final analysis infers a gene interaction network for these refined pathways. We apply this approach to model progression in prostate cancer and melanoma, resulting in a deeper understanding of the mechanisms of tumorigenesis. Our analysis supports previous findings for the deregulation of several pathways involved in cell cycle control and proliferation in both cancer types. A novel finding of our analysis is a connection between ErbB4 and primary prostate cancer. Elena J. Edelman, Justin Guinney, Jen-Tsan A. Chi, Phillip G. Febbo, Sayan Mukherjee 0001 |
PLoS Comput. Biol. | 5 |
| 2007 | Decision Fusion of Circulating Markers for Breast Cancer Detection in Premenopausal WomenabstractCurrent mammographic screening for breast cancer is less effective for younger women. To complement mammography for premenopausal women, we investigated the feasibility screening test using 98 blood serum proteins. Because the data set was very noisy and contained only weak features, we used a classifier designed for noisy data: decision fusion. Decision fusion outperformed both a support vector machine (SVM) and linear regression with forward stepwise feature selection on all three two-class classification tasks: normal tissue vs. cancer, normal tissue vs. benign lesions, and benign lesions vs. cancer. Decision fusion detected cancer moderately well (AUC=0.84 on normal vs. cancer), demonstrating promise as a screening tool. Decision fusion also detected benign lesions similarly well (AUC=0.83 on normal vs. benign lesions) and was the only classifier to achieve any success in separating benign from malignant lesions (AUC=0.64 on benign vs. cancer). The classification results suggest that the assayed proteins are more indicative of a secondary effect, such as immune response, rather than specific for breast cancer. In conclusion, the decision fusion classifier demonstrated some promise in detecting breast lesions and outperformed other classifiers, especially for the very noisy classification problem of distinguishing benign from malignant lesions. Jonathan L. Jesneck, Sayan Mukherjee 0001, Loren W. Nolte, Anna E. Lokshin, Jeffrey R. Marks, Joseph Y. Lo |
BIBE | 2 |
| 2007 | Genomic sweeping for hypermethylated genesabstractMOTIVATION: Genes silenced by the aberrent methylation of nearby CpG islands can contribute to the onset or progression of cancer and represent potential biomarkers for diagnosis and prognosis. Relatively few have thus far been validated as hypermethylated in cancer among over 14,000 candidates with promoter region CpG islands. A descriptive set of genes known to be unmethylated in cancer does not exist. This lack of a negative set and a large number of candidates necessitated the development of a new approach to identify novel genes hypermethylated in cancer. RESULTS: We developed a general method, cluster_boost, that in an imbalanced data setting predicts new minority class members given limited known samples and a large set of unlabeled samples. Synthetic datasets modeled after the hypermethylated genes data show that cluster_boost can successfully identify minority samples within unlabeled data. Using genome sequence features, cluster_boost predicted candidate hypermethylated genes among 14,000 genes of unknown status. In primary ovarian cancers, we determined the methylation status for 15 genes with different levels of support for being hypermethlyated. Results indicate cluster_boost can accurately identify novel genes hypermethylated in cancer. AVAILABILITY: Software and datasets are freely available at http://labs.genome.duke.edu/FureyLab/cluster_boost.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liang Goh, Susan K. Murphy, Sayan Mukherjee 0001, Terrence S. Furey |
Bioinform. | 3 |
| 2007 | Characterizing the Function Space for Bayesian Kernel Models
Natesh S. Pillai, Qiang Wu 0003, Sayan Mukherjee 0001, Robert L. Wolpert |
J. Mach. Learn. Res. | 4 |
| 2006 | Estimation of Gradients and Coordinate Covariation in ClassificationabstractWe introduce an algorithm that simultaneously estimates a classification function as well as its gradient in the supervised learning framework. The motivation for the algorithm is to find salient variables and estimate how they covary. An efficient implementation with respect to both memory and time is given. The utility of the algorithm is illustrated on simulated data as well as a gene expression data set. An error analysis is given for the convergence of the estimate of the classification function and its gradient to the true classification function and true gradient. Sayan Mukherjee 0001, Qiang Wu 0003 |
J. Mach. Learn. Res. | 1 |
| 2006 | Learning Coordinate Covariances via GradientsabstractWe introduce an algorithm that learns gradients from samples in the supervised learning framework. An error analysis is given for the convergence of the gradient estimated by the algorithm to the true gradient. The utility of the algorithm for the problem of variable selection as well as determining variable covariance is illustrated on simulated data as well as two gene expression data sets. For square loss we provide a very efficient implementation with respect to both memory and time. Sayan Mukherjee 0001, Ding-Xuan Zhou |
J. Mach. Learn. Res. | 1 |
| 2006 | Evidence of Influence of Genomic DNA Sequence on Human X Chromosome InactivationabstractA significant number of human X-linked genes escape X chromosome inactivation and are thus expressed from both the active and inactive X chromosomes. The basis for escape from inactivation and the potential role of the X chromosome primary DNA sequence in determining a gene's X inactivation status is unclear. Using a combination of the X chromosome sequence and a comprehensive X inactivation profile of more than 600 genes, two independent yet complementary approaches were used to systematically investigate the relationship between X inactivation and DNA sequence features. First, statistical analyses revealed that a number of repeat features, including long interspersed nuclear element (LINE) and mammalian-wide interspersed repeat repetitive elements, are significantly enriched in regions surrounding transcription start sites of genes that are subject to inactivation, while Alu repetitive elements and short motifs containing ACG/CGT are significantly enriched in those that escape inactivation. Second, linear support vector machine classifiers constructed using primary DNA sequence features were used to correctly predict the X inactivation status for >80% of all X-linked genes. We further identified a small set of features that are important for accurate classification, among which LINE-1 and LINE-2 content show the greatest individual discriminatory power. Finally, as few as 12 features can be used for accurate support vector machine classification. Taken together, these results suggest that features of the underlying primary DNA sequence of the human X chromosome may influence the spreading and/or maintenance of X inactivation. Zhong Wang 0003, Huntington F. Willard, Sayan Mukherjee 0001, Terrence S. Furey |
PLoS Comput. Biol. | 3 |
| 2005 | Permutation Tests for Classification
Polina Golland, Sayan Mukherjee 0001, Dmitry Panchenko |
COLT | 3 |
| 2002 | Choosing Multiple Parameters for Support Vector Machines
Olivier Chapelle, Vladimir Vapnik, Olivier Bousquet, Sayan Mukherjee 0001 |
Mach. Learn. | 4 |
| 2001 | Feature Reduction and Hierarchy of Classifiers for Fast Object Detection in Video ImagesabstractWe present a two-step method to speed-up object detection systems in computer vision that use Support Vector Machines (SVMs) as classifiers. In a first step we perform feature reduction by choosing relevant image features according to a measure derived from statistical learning theory. In a second step we build a hierarchy of classifiers. On the bottom level, a simple and fast classifier analyzes the whole image and rejects large parts of the background On the top level, a slower but more accurate classifier performs the final detection. Experiments with a face detection system show that combining feature reduction with hierarchical classification leads to a speed-up by a factor of 170 with similar classification performance. Bernd Heisele, Thomas Serre, Sayan Mukherjee 0001, Tomaso A. Poggio |
CVPR (2) | 3 |
| 2000 | Feature Selection for SVMsabstractWe introduce a method of feature selection for Support Vector Machines. The method is based upon finding those features which minimize bounds on the leave-one-out error. This search can be efficiently performed via gradient descent. The resulting algorithms are shown to be superior to some standard feature selection algorithms on both toy data and real-life problems of face recognition, pedestrian detection and analyzing DNA micro array data. Jason Weston, Sayan Mukherjee 0001, Olivier Chapelle, Massimiliano Pontil, Tomaso A. Poggio, Vladimir Vapnik |
NIPS | 2 |
| 1996 | Automatic generation of RBF networks using waveletsabstractLearning can be viewed as mapping from an input space to an output space. Examples of these mappings are used to construct a continuous function that approximates the given data and generalizes for intermediate instances. Radial-basis function (RBF) networks are used to formulate this approximating function. A novel method is introduced that automatically constructs a generalized radial-basis function (GRBF) network for a given mapping and error bound. This network is shown to be the smallest network within the error bound for the given mapping. The integral wavelet transform is used to determine the parameters of the network. Simple one-dimensional examples are used to demonstrate how the network constructed using the transform is superior to that constructed using standard ad hoc optimization techniques. The paper concludes with the automatic generation of GRBF networks for a multi-dimensional problem, namely, real-time 3D object recognition and pose estimation. The results of this application are favorable. Sayan Mukherjee 0001, Shree K. Nayar |
Pattern Recognit. | 1 |
| 1995 | Automatic Generation of GRBF Networks for Visual LearningabstractLearning can often be viewed as the problem of mapping from an input space to an output space. Examples of these mappings are used to construct a continuous function that approximates given data and generalizes for intermediate instances. Generalized Radial Basis Function (GRBF) networks are used to formulate this approximating function. A novel method is introduced to construct an optimal GRBF network for a given mapping and error bound using the integral wavelet transform. Simple one-dimensional examples are used to demonstrate how the optimal network is superior to one constructed using standard ad hoc optimization techniques. The paper concludes with an application of optimal GRBF networks to object recognition and pose estimation. The results of this application are favorable.> Sayan Mukherjee 0001, Shree K. Nayar |
ICCV | 1 |