Arjun Dasgupta

dblp:68/4911 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 6 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 39% Data integration and cleaning · 27% Data mining · 16%
Software engineering, system software, and programming languages
1 paper
Program analysis · 100%
Network and information security
2 papers
Privacy and data protection · 53% Systems and software security · 16% Web and mobile security · 16%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
0.112010
Unbiased estimation of size and other aggregates over hidden web databases · SIGMOD Conference 2010
Query processing and optimization
aggregate query processing
0.112009
HDSampler: revealing data behind web form interfaces · SIGMOD Conference 2009
Data mining
sampling
0.112009
Leveraging COUNT Information in Sampling Hidden Databases · ICDE 2009
Program analysis › static analysis › domain-specific static analysis
database application analysis
0.112009
A Static Analysis Framework for Database Applications · ICDE 2009
Program analysis
static analysis
0.112009
A Static Analysis Framework for Database Applications · ICDE 2009
Web and social media mining › social network sampling
random walk sampling
0.112007
A random walk approach to sampling hidden databases · SIGMOD Conference 2007
Data models and query languages
query interface
0.012009
Leveraging COUNT Information in Sampling Hidden Databases · ICDE 2009
Network security › intrusion detection and prevention
bot detection
0.012009
Privacy preservation of aggregates in hidden databases: why and how? · SIGMOD Conference 2009
Web and mobile security › web security › web vulnerability detection
SQL injection detection
0.012009
A Static Analysis Framework for Database Applications · ICDE 2009
Systems and software security
vulnerability discovery
0.012009
A Static Analysis Framework for Database Applications · ICDE 2009

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 0.2static analysis · 0.2sampling thwarting · 0.2variance reduction · 0.1unbiased estimation · 0.1sampling · 0.1count-based sampling · 0.1random walk · 0.1probabilistic rejection sampling · 0.1
YearPublicationVenuePosition
2010 Turbo-charging hidden database samplers with overflowing queries and skew reduction
abstract
Recently, there has been growing interest in random sampling from online hidden databases. These databases reside behind form-like web interfaces which allow users to execute search queries by specifying the desired values for certain attributes, and the system responds by returning a few (e.g., top-k) tuples that satisfy the selection conditions, sorted by a suitable scoring function. In this paper, we consider the problem of uniform random sampling over such hidden databases. A key challenge is to eliminate the skew of samples incurred by the selective return of highly ranked tuples. To address this challenge, all state-of-the-art samplers share a common approach: they do not use overflowing queries. This is done in order to avoid favoring highly ranked tuples and thus incurring high skew in the retrieved samples. However, not considering overflowing queries substantially impacts sampling efficiency.
Arjun Dasgupta, Nan Zhang 0004, Gautam Das 0001
EDBT1
2010 Unbiased estimation of size and other aggregates over hidden web databases
abstract
Many websites provide restrictive form-like interfaces which allow users to execute search queries on the underlying hidden databases. In this paper, we consider the problem of estimating the size of a hidden database through its web interface. We propose novel techniques which use a small number of queries to produce unbiased estimates with small variance. These techniques can also be used for approximate query processing over hidden databases. We present theoretical analysis and extensive experiments to illustrate the effectiveness of our approach.
Arjun Dasgupta, Bradley Jewell, Nan Zhang 0004, Gautam Das 0001
SIGMOD Conference1
2009 A Static Analysis Framework for Database Applications
abstract
Database developers today use data access APIs such as ADO.NET to execute SQL queries from their application. These applications often have security problems such as SQL injection vulnerabilities and performance problems such as poorly written SQL queries. However today's compilers have little or no understanding of data access APIs or DBMS, and hence the above problems can go undetected until much later in the application lifecycle. We present a framework that adapts traditional program analysis by leveraging understanding of data access APIs in order to identify such problems early on during application development. Our framework can analyze database application binaries that use ADO.NET data access APIs. We show how our framework can be used for a variety of analysis tasks such as SQL injection detection, workload extraction, identifying performance problems, and verifying data integrity constraints in the application.
Arjun Dasgupta, Vivek R. Narasayya, Manoj Syamala
ICDE1
2009 Leveraging COUNT Information in Sampling Hidden Databases
abstract
A large number of online databases are hidden behind form-like interfaces which allow users to execute search queries by specifying selection conditions in the interface. Most of these interfaces return restricted answers (e.g., only top-k of the selected tuples), while many of them also accompany each answer with the COUNT of the selected tuples. In this paper, we propose techniques which leverage the COUNT information to efficiently acquire unbiased samples of the hidden database. We also discuss variants for interfaces which do not provide COUNT information. We conduct extensive experiments to illustrate the efficiency and accuracy of our techniques.
Arjun Dasgupta, Nan Zhang 0004, Gautam Das 0001
ICDE1
2009 Privacy preservation of aggregates in hidden databases: why and how?
abstract
Many websites provide form-like interfaces which allow users to execute search queries on the underlying hidden databases. In this paper, we explain the importance of protecting sensitive aggregate information of hidden databases from being disclosed through individual tuples returned by the search queries. This stands in contrast to the traditional privacy problem where individual tuples must be protected while ensuring access to aggregating information. We propose techniques to thwart bots from sampling the hidden database to infer aggregate information. We present theoretical analysis and extensive experiments to illustrate the effectiveness of our approach.
Arjun Dasgupta, Nan Zhang 0004, Gautam Das 0001, Surajit Chaudhuri
SIGMOD Conference1
2009 HDSampler: revealing data behind web form interfaces
abstract
A large number of online databases are hidden behind the web. Users to these systems can form queries through web forms to retrieve a small sample of the database. Sampling such hidden databases is widely desired for understanding the nature and quality of data stored in them. We have developed HDSampler, which to the best of our knowledge is the first practical system for sampling structured hidden web databases. It enables efficient sampling of the databases and accurate answering of aggregate queries, to provide analysts with valuable information for data analytics, as well as help power a multitude of third-party applications such as web-mashups and meta-search engines. For the purpose of this demo, we present an instance of HDSampler on Google Base - a content-rich hidden web database maintained by Google. By using HDSampler, the demo reveals a snapshot of the marginal distribution of various attributes of Google Base in a matter of minutes.
Anirban Maiti, Arjun Dasgupta, Nan Zhang 0004, Gautam Das 0001
SIGMOD Conference2
2007 A random walk approach to sampling hidden databases
abstract
A large part of the data on the World Wide Web is hidden behind form-like interfaces. These interfaces interact with a hidden back-end database to provide answers to user queries. Generating a uniform random sample of this hidden database by using only the publicly available interface gives us access to the underlying data distribution. In this paper, we propose a random walk scheme over the query space provided by the interface to sample such databases. We discuss variants where the query space is visualized as a fixed and random ordering of attributes. We also propose techniques to further improve the sample quality by using a probabilistic rejection based approach. We conduct extensive experiments to illustrate the accuracy and efficiency of our techniques.
Arjun Dasgupta, Gautam Das 0001, Heikki Mannila
SIGMOD Conference1