Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Snigdhansu Chatterjee

dblp:14/4350 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0002-7986-0470ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 1 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 94% Graph data management · 6%
Artificial intelligence
2 papers
Learning theory · 62% Representation and self-supervised learning · 31% Probabilistic and Bayesian machine learning · 7%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 90% Environmental and earth informatics · 10%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.622020
Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks · IEEE Trans. Knowl. Data Eng. 2020
A Parameter-Free Spatio-Temporal Pattern Mining Model to Catalog Global Ocean Dynamics · ICDM 2013
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.612022
Feature selection using e-values · ICML 2022
Machine learning › Learning theory
model selection
0.612022
Feature selection using e-values · ICML 2022
Machine learning › Learning theory
statistical estimation
0.612022
Feature selection using e-values · ICML 2022
Data mining › pattern mining › graph pattern mining
clique enumeration
0.412020
Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks · IEEE Trans. Knowl. Data Eng. 2020
Data mining › structured data mining
relational data mining
0.312017
Tripoles: A New Class of Relationships in Time Series Data · KDD 2017
Data mining › temporal data mining
time series mining
0.312017
Tripoles: A New Class of Relationships in Time Series Data · KDD 2017
Algorithms and data structures › randomized algorithms › sampling
resampling algorithms
0.212022
Feature selection using e-values · ICML 2022
Data mining
anomaly detection
0.212013
A Parameter-Free Spatio-Temporal Pattern Mining Model to Catalog Global Ocean Dynamics · ICDM 2013
Data mining › spatiotemporal data mining
spatio-temporal pattern mining
0.212013
A Parameter-Free Spatio-Temporal Pattern Mining Model to Catalog Global Ocean Dynamics · ICDM 2013
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.222008
Reply to "Comment on causality and pathway search in microarray time series experiment" · Bioinform. 2008
Causality and pathway search in microarray time series experiment · Bioinform. 2007
Bioinformatics and computational biology › biological network › network biology › network inference › gene regulatory network inference
granger causality
0.222008
Reply to "Comment on causality and pathway search in microarray time series experiment" · Bioinform. 2008
Causality and pathway search in microarray time series experiment · Bioinform. 2007
Data mining
spatiotemporal data mining
0.112012
Testing the significance of spatio-temporal teleconnection patterns · KDD 2012
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway reconstruction
0.112007
Causality and pathway search in microarray time series experiment · Bioinform. 2007
Environmental and earth informatics › climate science
climate data analysis
0.012012
Testing the significance of spatio-temporal teleconnection patterns · KDD 2012
Bioinformatics and computational biology › gene expression analysis › time-series gene expression analysis
time-course microarray analysis
0.012008
Reply to "Comment on causality and pathway search in microarray time series experiment" · Bioinform. 2008

Methods — techniques the papers use, named apart from their topics

resampling · 1.1e-values · 1.1data depth · 1.1clique enumeration · 0.4wild bootstrap · 0.3statistical significance testing · 0.3incomplete information validation · 0.2granger causality · 0.2probabilistic modeling · 0.1vector autoregression · 0.1minimal spanning tree · 0.1false discovery rate · 0.1
YearPublicationVenuePosition
2023 Probabilistic Inverse Modeling: An Application in Hydrology
abstract
Rapid advancement in inverse modeling methods have brought into light their susceptibility to imperfect data. This has made it imperative to obtain more explainable and trustworthy estimates from these models. In hydrology, basin characteristics can be noisy or missing, impacting streamflow prediction. We propose a probabilistic inverse model framework that can reconstruct robust hydrology basin characteristics from dynamic input weather driver and streamflow response data. We address two aspects of building more explainable inverse models, uncertainty estimation (uncertainty due to imperfect data and imperfect model) and robustness. This can help improve the trust of water managers, handling of noisy data and reduce costs. We also propose an uncertainty based loss regularization that offers removal of 17% of temporal artifacts in reconstructions, 36% reduction in uncertainty and 4% higher coverage rate for basin characteristics. The forward model performance (streamflow estimation) is also improved by 6% using these uncertainty learning based reconstructions.
Somya Sharma, Rahul Ghosh, Arvind Renganathan, Snigdhansu Chatterjee, John Nieber, Christopher J. Duffy, Vipin Kumar 0001
SDM5
2022 Feature selection using e-values
abstract
In the context of supervised learning, we introduce the concept of e-value. An e-value is a scalar quantity that represents the proximity of the sampling distribution of parameter estimates in a model trained on a subset of features to that of the model trained on all features (i.e. the full model). Under general conditions, a rank ordering of e-values separates models that contain all essential features from those that do not. For a p-dimensional feature space, this requires fitting only the full model and evaluating p+1 models, as opposed to the traditional requirement of fitting and evaluating 2^p models. The above e-values framework is applicable to a wide range of parametric models. We use data depths and a fast resampling-based algorithm to implement a feature selection procedure, providing consistency results. Through experiments across several model settings and synthetic and real datasets, we establish that the e-values can be a promising general alternative to existing model-specific methods of feature selection.
Subhabrata Majumdar, Snigdhansu Chatterjee
ICML2
2020 Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks
abstract
In many domains, there is significant interest in capturing novel relationships between time series that represent activities recorded at different nodes of a highly complex system. In this paper, we introduce multipoles, a novel class of linear relationships between more than two time series. A multipole is a set of time series that have strong linear dependence among themselves, with the requirement that each time series makes a significant contribution to the linear dependence. We demonstrate that most interesting multipoles can be identified as cliques of negative correlations in a correlation network. Such cliques are typically rare in a real-world correlation network, which allows us to find almost all multipoles efficiently using a clique-enumeration approach. Using our proposed framework, we demonstrate the utility of multipoles in discovering new physical phenomena in two scientific domains: climate science and neuroscience. In particular, we discovered several multipole relationships that are reproducible in multiple other independent datasets and lead to novel domain insights.
Saurabh Agrawal 0002, Michael S. Steinbach, Daniel Boley, Snigdhansu Chatterjee, Gowtham Atluri, Anh The Dang, Stefan Liess, Vipin Kumar 0001
IEEE Trans. Knowl. Data Eng.4
2017 Tripoles: A New Class of Relationships in Time Series Data
abstract
Mining relationships in time series data is of immense interest to several disciplines such as neuroscience, climate science, and transportation. Traditional approaches for mining relationships focus on discovering pair-wise relationships in the data. In this work, we define a novel relationship pattern involving three interacting time series, which we refer to as a tripole. We show that tripoles capture interesting relationship patterns in the data that are not possible to be captured using traditionally studied pair-wise relationships. We demonstrate the utility of tripoles in multiple real-world datasets from various domains including climate science and neuroscience. In particular, our approach is able to discover tripoles that are statistically significant, reproducible across multiple independent data sets, and lead to novel domain insights.
Saurabh Agrawal 0002, Gowtham Atluri, Anuj Karpatne, William Haltom, Stefan Liess, Snigdhansu Chatterjee, Vipin Kumar 0001
KDD6
2016 A general framework to increase the robustness of model-based change point detection algorithms to outliers and noise
abstract
The autonomous identification of time-steps where the behavior of a time-series significantly deviates from a predefined model, or time-series change point detection, is an active field of research with notable applications in finance, health, and advertising. One family of time-series change detection algorithms, referred to as “model-based methods”, although useful for many applications, performs poor when the data are noisy and have outliers. We introduce a new framework that enables existing model-based methods to be more robust to these data challenges. We demonstrate the effectiveness of our approach on remote sensing and mobile health data. Our method introduces two new concepts: (i) a random sampling procedure allows us to overcome outliers, and (ii) a matrix-based representation of anomaly scores provides a flexible and intuitive way to identify multiple types of changes and test their significance. We show that our method performs better than several baseline methods, including application-specific algorithms, and provide all data and open-source code.
Xi Chen 0120, Yuanshun Yao, Sichao Shi, Snigdhansu Chatterjee, Vipin Kumar 0001, James H. Faghmous
SDM4
2014 Fast algorithm for computing weighted projection quantiles and data depth for high-dimensional large data clouds
abstract
In this paper we present a new algorithm based on a weighted projection quantiles for fast and frugal real time quantile estimation of large sized high dimensional data clouds. We present a projection quantile regression algorithm for high dimensional data. Second, we present a fast algorithm for computing the depth of a point or a new observation in relation to any high-dimensional data cloud, and propose a ranking system for multivariate data. Third, we briefly describe a real time rapid monitoring scheme similar to statistical process monitoring, for actionable analytics with big data. We believe these algorithms would be very useful for real time analysis of high dimensional `big data' sets including streaming data sets. The proposed algorithms would be of immense use in several application areas such as real time financial market analysis, real time remote health monitoring of patients using body area networked devices and real time pricing and inventory decisions in retail and manufacturing sector.
Ujjal Kumar Mukherjee, Snigdhansu Chatterjee
IEEE BigData2
2013 A Parameter-Free Spatio-Temporal Pattern Mining Model to Catalog Global Ocean Dynamics
abstract
As spatio-temporal data have become ubiquitous, an increasing challenge facing computer scientists is that of identifying discrete patterns in continuous spatio-temporal fields. In this paper, we introduce a parameter-free pattern mining application that is able to identify dynamic anomalies in ocean data, known as ocean eddies. Despite ocean eddy monitoring being an active field of research, we provide one of the first quantitative analyses of the performance of the most used monitoring algorithms. We present an incomplete information validation technique, that uses the performance of two methods to construct an imperfect ground truth to test the significance of patterns discovered as well as the relative performance of pattern mining algorithms. These methods, in addition to the validation schemes discussed provide researchers new directions in analyzing large unlabeled climate datasets.
James H. Faghmous, Matt Le 0001, Muhammed Uluyol, Vipin Kumar 0001, Snigdhansu Chatterjee
ICDM5
2013 Contextual Time Series Change Detection
abstract
Time series data are common in a variety of fields ranging from economics to medicine and manufacturing.As a result, time series analysis and modeling has become an active research area in statistics and data mining.In this paper, we focus on a type of change we call contextual time series change (CTC) and propose a novel two-stage algorithm to address it.In contrast to traditional change detection methods, which consider each time series separately, CTC is defined as a change relative to the behavior of a group of related time series.As a result, our proposed method is able to identify novel types of changes not found by other algorithms.We demonstrate the unique capabilities of our approach with several case studies on real-world datasets from the financial and Earth science domains.
Xi Chen 0120, Karsten Steinhaeuser, Shyam Boriah, Snigdhansu Chatterjee, Vipin Kumar 0001
SDM4
2012 Spatially penalized regression for dependence analysis of rare events: A study in precipitation extremes
abstract
Discovery of dependence structure between precipitation extremes and other climate variables (covariates) within a smaller spatial and temporal neighborhood is an important step in better understanding the drivers of this complex phenomenon as well as short-term prediction of extremes occurrence. Apart from the inherent spatio-temporal variability of the dependence, it is further complicated by the availability of the covariates at different vertical levels. The above problem can be split into three different sub-problems. Firstly, a spatio-temporal neighborhood of influence has to be discovered, which can be different for different locations. Secondly, the dependence structure between the precipitation extremes and the covariates has to be discovered within this neighborhood and thirdly, it has to be investigated whether this dependence structure can be exploited for any predictive power. Climate scientists have already discovered some physics-based relations between some of the covariates (e.g. temperature, relative humidity, precipitable water etc.) and precipitation extremes. We are exploring data-dependent alternatives for these problems and any possibility of incorporating the physics-based relations into the resulting data model. In particular, we used elastic net-based sparse optimization technique which solves all three problems of neighborhood discovery, covariate dependence discovery and predictive modeling and at the same time maintains the interpretability of the resulting model. Preliminary results look promising and show potential for some interesting knowledge discovery. We are currently exploring non-linear correlations and the alternatives to combine the physics-based relationships into the data model.
Debasish Das, Auroop R. Ganguly, Snigdhansu Chatterjee, Vipin Kumar 0001, Zoran Obradovic
IGARSS3
2012 Testing the significance of spatio-temporal teleconnection patterns
abstract
Dipoles represent long distance connections between the pressure anomalies of two distant regions that are negatively correlated with each other. Such dipoles have proven important for understanding and explaining the variability in climate in many regions of the world, e.g., the El Nino climate phenomenon is known to be responsible for precipitation and temperature anomalies over large parts of the world. Systematic approaches for dipole detection generate a large number of candidate dipoles, but there exists no method to evaluate the significance of the candidate teleconnections. In this paper, we present a novel method for testing the statistical significance of the class of spatio-temporal teleconnection patterns called as dipoles. One of the most important challenges in addressing significance testing in a spatio-temporal context is how to address the spatial and temporal dependencies that show up as high autocorrelation. We present a novel approach that uses the wild bootstrap to capture the spatio-temporal dependencies, in the special use case of teleconnections in climate data. Our approach to find the statistical significance takes into account the autocorrelation, the seasonality and the trend in the time series over a period of time. This framework is applicable to other problems in spatio-temporal data mining to assess the significance of the patterns.
Jaya Kawale, Snigdhansu Chatterjee, Dominick Ormsby, Karsten Steinhaeuser, Stefan Liess, Vipin Kumar 0001
KDD2
2012 Sparse Group Lasso: Consistency and Climate Applications
abstract
The design of statistical predictive models for climate data gives rise to some unique challenges due to the high dimensionality and spatio-temporal nature of the datasets, which dictate that models should exhibit parsimony in variable selection. Recently, a class of methods which promote structured sparsity in the model have been developed, which is suitable for this task. In this paper, we prove theoretical statistical consistency of estimators with tree-structured norm regularizers. We consider one particular model, the Sparse Group Lasso (SGL), to construct predictors of land climate using ocean climate variables. Our experimental results demonstrate that the SGL model provides better predictive performance than the current state-of-the-art, remains climatologically interpretable, and is robust in its variable selection.
Soumyadeep Chatterjee, Karsten Steinhaeuser, Arindam Banerjee 0001, Snigdhansu Chatterjee, Auroop R. Ganguly
SDM4
2011 Probabilistic Matrix Addition
Amrudin Agovic, Arindam Banerjee 0001, Snigdhansu Chatterjee
ICML3
2008 Reply to "Comment on causality and pathway search in microarray time series experiment"
abstract
We thank Professors Nagarajan and Upreti for their interest in our paper, Mukhopadhyay and Chatterjee (2007). There, we propose using Granger causality-based pathway detection in an acyclic, homoscedastic framework for microarray time-series expressions; which are generally short-duration time series involving very large number of genes. Professors Nagarajan and Upreti point out that in the presence of heteroscedasticity, and a cycle like ‘gene x regulates the expression of gene y and simultaneously gene y regulates the expression of gene x’, Granger causality tests may not be informative. Here, we adopt the term ‘heteroscedasticity’ (‘homoscedasticity’) to mean the unconditional variance of the white noise, represented as a bivariate vector in the Euclidean co-ordinate system, is different (same) in different co-ordinate directions. Thus, in essence, if the assumptions about the acyclic and homoscedastic nature of the time series are violated, tests for causality detection may fail. This is an important point, since when a contemporaneous cyclic relationship is present, the notion of causality makes little sense. In the context of economics, Eichler (2007) present a treatment of contemporaneous correlation as well as Granger causality. Extreme heteroscedasticity may be indicative of improper normalization of gene expressions. At the end of their letter, Dr Nagarajan and Dr Upreti mention the normalization step. Proper normalization should remove wide discrepancy in noise variance, hence nowadays microarray datasets are typically available in de facto normalized version. The data used in Mukhopadhyay and Chatterjee (2007) is also normalized. However, difference in technical variance, as indicated by Professors Nagarajan and Upreti, may still be present. And that will violate the assumption of our method (as well as many other statistical comparison methods relying on common unknown variance). Professor Nagarajan, in review, kindly suggested references for two-gene systems whose time-profile may not fit into to a homoscedastic, cause-effect framework. Thus, a full vector autoregression structure may be needed to capture their mutual dependence at various lags (including lag zero). It can be guessed that multi-gene systems exist whose temporal co-dependency nature is extremely complex. Although current knowledge about gene regulatory networks is limited, some biology experts we consulted believe that cyclical patterns may be found in large multi-gene networks as a part of a feedback procedure, if they are studied over long enough time spans. A proper approach to elicit such patterns would be to conduct multivariate, possibly non-stationary, time-series analysis with all the genes over a long time horizon. This is not feasible currently, since present state-of-the-art microarray time series experiments are of short duration and typically involve very large number of genes. Hence, restricting the network to acyclic ones is, in our opinion, a small price to pay to produce informative analysis. Future microarray experiments over longer duration, along with discoveries of biological and chemical properties relating to gene and protein interactions, will no doubt lead to better understanding of gene networks. We would like to point out in Model 1 (Equation 2), α12, α21, and need to be known constants for the mathematical displays (4)–(7) to hold. As they stand, displays (4)–(7) are missing the O (n−1) terms with each estimated parameters if some (or all) of are estimated from data, where n is the length of the time series data. Also, the equation for s1 does not account for the fact that as univariate time series, both xt and yt are AR(2) (autoregressive of order 2) process and not AR(1). Similar comments hold for Model 2 (Equation 11). The difficulty of modeling microarray time-series can be appreciated from the fact that in the human cell cycle data considered in Mukhopadhyay and Chatterjee (2007), n was 12 in one experiment, while the time-series itself was 802 dimensional.
Nitai D. Mukhopadhyay, Snigdhansu Chatterjee
Bioinform.2
2007 Causality and pathway search in microarray time series experiment
abstract
MOTIVATION: Interaction among time series can be explored in many ways. All the approach has the usual problem of low power and high dimensional model. Here we attempted to build a causality network among a set of time series. The causality has been established by Granger causality, and then constructing the pathway has been implemented by finding the Minimal Spanning Tree within each connected component of the inferred network. False discovery rate measurement has been used to identify the most significant causalities. RESULTS: Simulation shows good convergence and accuracy of the algorithm. Robustness of the procedure has been demonstrated by applying the algorithm in a non-stationary time series setup. Application of the algorithm in a real dataset identified many causalities, with some overlap with previously known ones. Assembled network of the genes reveals features of the network that are common wisdom about naturally occurring networks.
Nitai D. Mukhopadhyay, Snigdhansu Chatterjee
Bioinform.2