VLDB 2026 Research / reviewers in the wild / expert
Gowtham Atluri
dblp:27/5925
· DBLP profile ↗
8ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0001-5619-6688ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-authorArtificial intelligence and machine learning · 2Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data mining · 80% Machine learning and data management · 13% Graph data management · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 63% Computational science and engineering · 37% | |
| Human-computer interaction and pervasive computing
1 paper |
Health and well-being technologies · 62% Ubiquitous computing and smart environments · 19% Wearable and physiological sensing · 19% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.5 | 2 | 2020 | Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks · IEEE Trans. Knowl. Data Eng. 2020 An association analysis approach to biclustering · KDD 2009 |
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis |
0.4 | 1 | 2020 | Resolving single-cell heterogeneity from hundreds of thousands of cells through sequential hybrid clustering and NMF · Bioinform. 2020 |
Data mining › pattern mining › graph pattern mining
clique enumeration |
0.4 | 1 | 2020 | Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks · IEEE Trans. Knowl. Data Eng. 2020 |
Data mining › structured data mining
relational data mining |
0.3 | 1 | 2017 | Tripoles: A New Class of Relationships in Time Series Data · KDD 2017 |
Data mining › temporal data mining
time series mining |
0.3 | 1 | 2017 | Tripoles: A New Class of Relationships in Time Series Data · KDD 2017 |
Health and well-being technologies › behavior change
smoking cessation |
0.2 | 1 | 2016 | mCrave: continuous estimation of craving during smoking cessation · UbiComp 2016 |
Data mining › pattern mining
association analysis |
0.1 | 1 | 2009 | An association analysis approach to biclustering · KDD 2009 |
Data mining › clustering
co-clustering |
0.1 | 1 | 2009 | An association analysis approach to biclustering · KDD 2009 |
Ubiquitous computing and smart environments › mobile sensing
mobile sensor data |
0.1 | 1 | 2016 | mCrave: continuous estimation of craving during smoking cessation · UbiComp 2016 |
Wearable and physiological sensing › cognitive state monitoring
stress detection |
0.1 | 1 | 2016 | mCrave: continuous estimation of craving during smoking cessation · UbiComp 2016 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2009 | An association analysis approach to biclustering · KDD 2009 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.0 | 1 | 2009 | An association analysis approach to biclustering · KDD 2009 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.4sparse non-negative matrix factorization · 0.4pagerank · 0.4clique enumeration · 0.4HOPACH · 0.4feature engineering · 0.2conditional random field · 0.2range support patterns · 0.2biclustering · 0.2association analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Fingerprinting encrypted voice traffic on smart speakers with deep learningabstractThis paper investigates the privacy leakage of smart speakers under an encrypted traffic analysis attack, referred to as voice command fingerprinting. In this attack, an adversary can eavesdrop both outgoing and incoming encrypted voice traffic of a smart speaker, and infers which voice command a user says over encrypted traffic. We first built an automatic voice traffic collection tool and collected two large-scale datasets on two smart speakers, Amazon Echo and Google Home. Then, we implemented proof-of-concept attacks by leveraging deep learning. Our experimental results over the two datasets indicate disturbing privacy concerns. Specifically, compared to 1% accuracy with random guess, our attacks can correctly infer voice commands over encrypted traffic with 92.89% accuracy on Amazon Echo. Sean Kennedy, King Hudson, Gowtham Atluri, Xuetao Wei, Wenhai Sun, Boyang Wang 0001 |
WISEC | 5 |
| 2020 | Resolving single-cell heterogeneity from hundreds of thousands of cells through sequential hybrid clustering and NMFabstractMOTIVATION: The rapid proliferation of single-cell RNA-sequencing (scRNA-Seq) technologies has spurred the development of diverse computational approaches to detect transcriptionally coherent populations. While the complexity of the algorithms for detecting heterogeneity has increased, most require significant user-tuning, are heavily reliant on dimension reduction techniques and are not scalable to ultra-large datasets. We previously described a multi-step algorithm, Iterative Clustering and Guide-gene Selection (ICGS), which applies intra-gene correlation and hybrid clustering to uniquely resolve novel transcriptionally coherent cell populations from an intuitive graphical user interface. RESULTS: We describe a new iteration of ICGS that outperforms state-of-the-art scRNA-Seq detection workflows when applied to well-established benchmarks. This approach combines multiple complementary subtype detection methods (HOPACH, sparse non-negative matrix factorization, cluster 'fitness', support vector machine) to resolve rare and common cell-states, while minimizing differences due to donor or batch effects. Using data from multiple cell atlases, we show that the PageRank algorithm effectively downsamples ultra-large scRNA-Seq datasets, without losing extremely rare or transcriptionally similar yet distinct cell types and while recovering novel transcriptionally distinct cell populations. We believe this new approach holds tremendous promise in reproducibly resolving hidden cell populations in complex datasets. AVAILABILITY AND IMPLEMENTATION: ICGS2 is implemented in Python. The source code and documentation are available at http://altanalyze.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Meenakshi Venkatasubramanian, Kashish Chetal, Daniel J. Schnell, Gowtham Atluri, Nathan Salomonis |
Bioinform. | 4 |
| 2020 | Mining Novel Multivariate Relationships in Time Series Data Using Correlation NetworksabstractIn many domains, there is significant interest in capturing novel relationships between time series that represent activities recorded at different nodes of a highly complex system. In this paper, we introduce multipoles, a novel class of linear relationships between more than two time series. A multipole is a set of time series that have strong linear dependence among themselves, with the requirement that each time series makes a significant contribution to the linear dependence. We demonstrate that most interesting multipoles can be identified as cliques of negative correlations in a correlation network. Such cliques are typically rare in a real-world correlation network, which allows us to find almost all multipoles efficiently using a clique-enumeration approach. Using our proposed framework, we demonstrate the utility of multipoles in discovering new physical phenomena in two scientific domains: climate science and neuroscience. In particular, we discovered several multipole relationships that are reproducible in multiple other independent datasets and lead to novel domain insights. Saurabh Agrawal 0002, Michael S. Steinbach, Daniel Boley, Snigdhansu Chatterjee, Gowtham Atluri, Anh The Dang, Stefan Liess, Vipin Kumar 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2017 | Tripoles: A New Class of Relationships in Time Series DataabstractMining relationships in time series data is of immense interest to several disciplines such as neuroscience, climate science, and transportation. Traditional approaches for mining relationships focus on discovering pair-wise relationships in the data. In this work, we define a novel relationship pattern involving three interacting time series, which we refer to as a tripole. We show that tripoles capture interesting relationship patterns in the data that are not possible to be captured using traditionally studied pair-wise relationships. We demonstrate the utility of tripoles in multiple real-world datasets from various domains including climate science and neuroscience. In particular, our approach is able to discover tripoles that are statistically significant, reproducible across multiple independent data sets, and lead to novel domain insights. Saurabh Agrawal 0002, Gowtham Atluri, Anuj Karpatne, William Haltom, Stefan Liess, Snigdhansu Chatterjee, Vipin Kumar 0001 |
KDD | 2 |
| 2017 | Theory-Guided Data Science: A New Paradigm for Scientific Discovery from DataabstractData science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of data science models in enabling scientific discovery. The overarching vision of TGDS is to introduce scientific consistency as an essential component for learning generalizable models. Further, by producing scientifically interpretable models, TGDS aims to advance our scientific understanding by discovering novel domain insights. Indeed, the paradigm of TGDS has started to gain prominence in a number of scientific disciplines such as turbulence modeling, material discovery, quantum chemistry, bio-medical science, bio-marker discovery, climate science, and hydrology. In this paper, we formally conceptualize the paradigm of TGDS and present a taxonomy of research themes in TGDS. We describe several approaches for integrating domain knowledge in different research themes using illustrative examples from different disciplines. We also highlight some of the promising avenues of novel research for realizing the full potential of theory-guided data science. Anuj Karpatne, Gowtham Atluri, James H. Faghmous, Michael S. Steinbach, Arindam Banerjee 0001, Auroop R. Ganguly, Shashi Shekhar 0001, Nagiza F. Samatova, Vipin Kumar 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | mCrave: continuous estimation of craving during smoking cessationabstractCraving usually precedes a lapse for impulsive behaviors such as overeating, drinking, smoking, and drug use. Passive estimation of craving from sensor data in the natural environment can be used to assist users in coping with craving. In this paper, we take the first steps towards developing a computational model to estimate cigarette craving (during smoking abstinence) at the minute-level using mobile sensor data. We use 2,012 hours of sensor data and 1,812 craving self-reports from 61 participants in a smoking cessation study. To estimate craving, we first obtain a continuous measure of stress from sensor data. We find that during hours of day when craving is high, stress associated with self-reported high craving is greater than stress associated with low craving. We use this and other insights to develop feature functions, and encode them as pattern detectors in a Conditional Random Field (CRF) based model to infer craving probabilities. Soujanya Chatterjee, Karen Hovsepian, Hillol Sarker, Nazir Saleheen, Mustafa al'Absi, Gowtham Atluri, Emre Ertin, Cho Lam, Andrine Lemieux, Motohiro Nakajima, Bonnie Spring, David W. Wetter, Santosh Kumar 0001 |
UbiComp | 6 |
| 2014 | Discovering Groups of Time Series with Similar Behavior in Multiple Small Intervals of TimeabstractThe focus of this paper is to address the problem of discovering groups of time series that share similar behavior in multiple small intervals of time. This problem has two characteristics: i) There are exponentially many combinations of time series that needs to be explored to find these groups, ii) The groups of time series of interest need to have similar behavior only in some subsets of the time dimension. We present an Apriori based approach to address this problem. We evaluate it on a synthetic dataset and demonstrate that our approach can directly find all groups of intermittently correlated time series without finding spurious groups unlike other alternative approaches that find many spurious groups. We also demonstrate, using a neuroimaging dataset, that groups of intermittently coherent time series discovered by our approach are reproducible on independent sets of time series data. In addition, we demonstrate the utility of our approach on an S&P 500 stocks data set. Gowtham Atluri, Michael S. Steinbach, Kelvin O. Lim, Angus W. MacDonald III, Vipin Kumar 0001 |
SDM | 1 |
| 2009 | An association analysis approach to biclusteringabstractThe discovery of biclusters, which denote groups of items that show coherent values across a subset of all the transactions in a data set, is an important type of analysis performed on real-valued data sets in various domains, such as biology. Several algorithms have been proposed to find different types of biclusters in such data sets. However, these algorithms are unable to search the space of all possible biclusters exhaustively. Pattern mining algorithms in association analysis also essentially produce biclusters as their result, since the patterns consist of items that are supported by a subset of all the transactions. However, a major limitation of the numerous techniques developed in association analysis is that they are only able to analyze data sets with binary and/or categorical variables, and their application to real-valued data sets often involves some lossy transformation such as discretization or binarization of the attributes. In this paper, we propose a novel association analysis framework for exhaustively and efficiently mining "range support" patterns from such a data set. On one hand, this framework reduces the loss of information incurred by the binarization- and discretization-based approaches, and on the other, it enables the exhaustive discovery of coherent biclusters. We compared the performance of our framework with two standard biclustering algorithms through the evaluation of the similarity of the cellular functions of the genes constituting the patterns/biclusters derived by these algorithms from microarray data. These experiments show that the real-valued patterns discovered by our framework are better enriched by small biologically interesting functional classes. Also, through specific examples, we demonstrate the ability of the RAP framework to discover functionally enriched patterns that are not found by the commonly used biclustering algorithm ISA. The source code and data sets used in this paper, as well as the supplementary material, are available at http://www.cs.umn.edu/vk/gaurav/rap. Gaurav Pandey 0002, Gowtham Atluri, Michael S. Steinbach, Chad L. Myers, Vipin Kumar 0001 |
KDD | 2 |