Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gowtham Atluri

dblp:27/5925 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0001-5619-6688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 1 first-authorArtificial intelligence and machine learning · 2Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 80% Machine learning and data management · 13% Graph data management · 6%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 63% Computational science and engineering · 37%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 62% Ubiquitous computing and smart environments · 19% Wearable and physiological sensing · 19%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.522020
Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks · IEEE Trans. Knowl. Data Eng. 2020
An association analysis approach to biclustering · KDD 2009
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis
0.412020
Resolving single-cell heterogeneity from hundreds of thousands of cells through sequential hybrid clustering and NMF · Bioinform. 2020
Data mining › pattern mining › graph pattern mining
clique enumeration
0.412020
Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks · IEEE Trans. Knowl. Data Eng. 2020
Data mining › structured data mining
relational data mining
0.312017
Tripoles: A New Class of Relationships in Time Series Data · KDD 2017
Data mining › temporal data mining
time series mining
0.312017
Tripoles: A New Class of Relationships in Time Series Data · KDD 2017
Health and well-being technologies › behavior change
smoking cessation
0.212016
mCrave: continuous estimation of craving during smoking cessation · UbiComp 2016
Data mining › pattern mining
association analysis
0.112009
An association analysis approach to biclustering · KDD 2009
Data mining › clustering
co-clustering
0.112009
An association analysis approach to biclustering · KDD 2009
Ubiquitous computing and smart environments › mobile sensing
mobile sensor data
0.112016
mCrave: continuous estimation of craving during smoking cessation · UbiComp 2016
Wearable and physiological sensing › cognitive state monitoring
stress detection
0.112016
mCrave: continuous estimation of craving during smoking cessation · UbiComp 2016
Bioinformatics and computational biology
gene expression analysis
0.012009
An association analysis approach to biclustering · KDD 2009
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.012009
An association analysis approach to biclustering · KDD 2009

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.4sparse non-negative matrix factorization · 0.4pagerank · 0.4clique enumeration · 0.4HOPACH · 0.4feature engineering · 0.2conditional random field · 0.2range support patterns · 0.2biclustering · 0.2association analysis · 0.2
YearPublicationVenuePosition
2020 Fingerprinting encrypted voice traffic on smart speakers with deep learning
abstract
This paper investigates the privacy leakage of smart speakers under an encrypted traffic analysis attack, referred to as voice command fingerprinting. In this attack, an adversary can eavesdrop both outgoing and incoming encrypted voice traffic of a smart speaker, and infers which voice command a user says over encrypted traffic. We first built an automatic voice traffic collection tool and collected two large-scale datasets on two smart speakers, Amazon Echo and Google Home. Then, we implemented proof-of-concept attacks by leveraging deep learning. Our experimental results over the two datasets indicate disturbing privacy concerns. Specifically, compared to 1% accuracy with random guess, our attacks can correctly infer voice commands over encrypted traffic with 92.89% accuracy on Amazon Echo.
Sean Kennedy, King Hudson, Gowtham Atluri, Xuetao Wei, Wenhai Sun, Boyang Wang 0001
WISEC5
2020 Resolving single-cell heterogeneity from hundreds of thousands of cells through sequential hybrid clustering and NMF
abstract
MOTIVATION: The rapid proliferation of single-cell RNA-sequencing (scRNA-Seq) technologies has spurred the development of diverse computational approaches to detect transcriptionally coherent populations. While the complexity of the algorithms for detecting heterogeneity has increased, most require significant user-tuning, are heavily reliant on dimension reduction techniques and are not scalable to ultra-large datasets. We previously described a multi-step algorithm, Iterative Clustering and Guide-gene Selection (ICGS), which applies intra-gene correlation and hybrid clustering to uniquely resolve novel transcriptionally coherent cell populations from an intuitive graphical user interface. RESULTS: We describe a new iteration of ICGS that outperforms state-of-the-art scRNA-Seq detection workflows when applied to well-established benchmarks. This approach combines multiple complementary subtype detection methods (HOPACH, sparse non-negative matrix factorization, cluster 'fitness', support vector machine) to resolve rare and common cell-states, while minimizing differences due to donor or batch effects. Using data from multiple cell atlases, we show that the PageRank algorithm effectively downsamples ultra-large scRNA-Seq datasets, without losing extremely rare or transcriptionally similar yet distinct cell types and while recovering novel transcriptionally distinct cell populations. We believe this new approach holds tremendous promise in reproducibly resolving hidden cell populations in complex datasets. AVAILABILITY AND IMPLEMENTATION: ICGS2 is implemented in Python. The source code and documentation are available at http://altanalyze.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meenakshi Venkatasubramanian, Kashish Chetal, Daniel J. Schnell, Gowtham Atluri, Nathan Salomonis
Bioinform.4
2020 Mining Novel Multivariate Relationships in Time Series Data Using Correlation Networks
abstract
In many domains, there is significant interest in capturing novel relationships between time series that represent activities recorded at different nodes of a highly complex system. In this paper, we introduce multipoles, a novel class of linear relationships between more than two time series. A multipole is a set of time series that have strong linear dependence among themselves, with the requirement that each time series makes a significant contribution to the linear dependence. We demonstrate that most interesting multipoles can be identified as cliques of negative correlations in a correlation network. Such cliques are typically rare in a real-world correlation network, which allows us to find almost all multipoles efficiently using a clique-enumeration approach. Using our proposed framework, we demonstrate the utility of multipoles in discovering new physical phenomena in two scientific domains: climate science and neuroscience. In particular, we discovered several multipole relationships that are reproducible in multiple other independent datasets and lead to novel domain insights.
Saurabh Agrawal 0002, Michael S. Steinbach, Daniel Boley, Snigdhansu Chatterjee, Gowtham Atluri, Anh The Dang, Stefan Liess, Vipin Kumar 0001
IEEE Trans. Knowl. Data Eng.5
2017 Tripoles: A New Class of Relationships in Time Series Data
abstract
Mining relationships in time series data is of immense interest to several disciplines such as neuroscience, climate science, and transportation. Traditional approaches for mining relationships focus on discovering pair-wise relationships in the data. In this work, we define a novel relationship pattern involving three interacting time series, which we refer to as a tripole. We show that tripoles capture interesting relationship patterns in the data that are not possible to be captured using traditionally studied pair-wise relationships. We demonstrate the utility of tripoles in multiple real-world datasets from various domains including climate science and neuroscience. In particular, our approach is able to discover tripoles that are statistically significant, reproducible across multiple independent data sets, and lead to novel domain insights.
Saurabh Agrawal 0002, Gowtham Atluri, Anuj Karpatne, William Haltom, Stefan Liess, Snigdhansu Chatterjee, Vipin Kumar 0001
KDD2
2017 Theory-Guided Data Science: A New Paradigm for Scientific Discovery from Data
abstract
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of data science models in enabling scientific discovery. The overarching vision of TGDS is to introduce scientific consistency as an essential component for learning generalizable models. Further, by producing scientifically interpretable models, TGDS aims to advance our scientific understanding by discovering novel domain insights. Indeed, the paradigm of TGDS has started to gain prominence in a number of scientific disciplines such as turbulence modeling, material discovery, quantum chemistry, bio-medical science, bio-marker discovery, climate science, and hydrology. In this paper, we formally conceptualize the paradigm of TGDS and present a taxonomy of research themes in TGDS. We describe several approaches for integrating domain knowledge in different research themes using illustrative examples from different disciplines. We also highlight some of the promising avenues of novel research for realizing the full potential of theory-guided data science.
Anuj Karpatne, Gowtham Atluri, James H. Faghmous, Michael S. Steinbach, Arindam Banerjee 0001, Auroop R. Ganguly, Shashi Shekhar 0001, Nagiza F. Samatova, Vipin Kumar 0001
IEEE Trans. Knowl. Data Eng.2
2016 mCrave: continuous estimation of craving during smoking cessation
abstract
Craving usually precedes a lapse for impulsive behaviors such as overeating, drinking, smoking, and drug use. Passive estimation of craving from sensor data in the natural environment can be used to assist users in coping with craving. In this paper, we take the first steps towards developing a computational model to estimate cigarette craving (during smoking abstinence) at the minute-level using mobile sensor data. We use 2,012 hours of sensor data and 1,812 craving self-reports from 61 participants in a smoking cessation study. To estimate craving, we first obtain a continuous measure of stress from sensor data. We find that during hours of day when craving is high, stress associated with self-reported high craving is greater than stress associated with low craving. We use this and other insights to develop feature functions, and encode them as pattern detectors in a Conditional Random Field (CRF) based model to infer craving probabilities.
Soujanya Chatterjee, Karen Hovsepian, Hillol Sarker, Nazir Saleheen, Mustafa al'Absi, Gowtham Atluri, Emre Ertin, Cho Lam, Andrine Lemieux, Motohiro Nakajima, Bonnie Spring, David W. Wetter, Santosh Kumar 0001
UbiComp6
2014 Discovering Groups of Time Series with Similar Behavior in Multiple Small Intervals of Time
abstract
The focus of this paper is to address the problem of discovering groups of time series that share similar behavior in multiple small intervals of time. This problem has two characteristics: i) There are exponentially many combinations of time series that needs to be explored to find these groups, ii) The groups of time series of interest need to have similar behavior only in some subsets of the time dimension. We present an Apriori based approach to address this problem. We evaluate it on a synthetic dataset and demonstrate that our approach can directly find all groups of intermittently correlated time series without finding spurious groups unlike other alternative approaches that find many spurious groups. We also demonstrate, using a neuroimaging dataset, that groups of intermittently coherent time series discovered by our approach are reproducible on independent sets of time series data. In addition, we demonstrate the utility of our approach on an S&P 500 stocks data set.
Gowtham Atluri, Michael S. Steinbach, Kelvin O. Lim, Angus W. MacDonald III, Vipin Kumar 0001
SDM1
2009 An association analysis approach to biclustering
abstract
The discovery of biclusters, which denote groups of items that show coherent values across a subset of all the transactions in a data set, is an important type of analysis performed on real-valued data sets in various domains, such as biology. Several algorithms have been proposed to find different types of biclusters in such data sets. However, these algorithms are unable to search the space of all possible biclusters exhaustively. Pattern mining algorithms in association analysis also essentially produce biclusters as their result, since the patterns consist of items that are supported by a subset of all the transactions. However, a major limitation of the numerous techniques developed in association analysis is that they are only able to analyze data sets with binary and/or categorical variables, and their application to real-valued data sets often involves some lossy transformation such as discretization or binarization of the attributes. In this paper, we propose a novel association analysis framework for exhaustively and efficiently mining "range support" patterns from such a data set. On one hand, this framework reduces the loss of information incurred by the binarization- and discretization-based approaches, and on the other, it enables the exhaustive discovery of coherent biclusters. We compared the performance of our framework with two standard biclustering algorithms through the evaluation of the similarity of the cellular functions of the genes constituting the patterns/biclusters derived by these algorithms from microarray data. These experiments show that the real-valued patterns discovered by our framework are better enriched by small biologically interesting functional classes. Also, through specific examples, we demonstrate the ability of the RAP framework to discover functionally enriched patterns that are not found by the commonly used biclustering algorithm ISA. The source code and data sets used in this paper, as well as the supplementary material, are available at http://www.cs.umn.edu/vk/gaurav/rap.
Gaurav Pandey 0002, Gowtham Atluri, Michael S. Steinbach, Chad L. Myers, Vipin Kumar 0001
KDD2