Abhijit Pol

dblp:96/4590 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 73% Data stream processing · 17% Information retrieval · 4%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
0.352008
Scalable approximate query processing with the DBO engine · ACM Trans. Database Syst. 2008
Scalable approximate query processing with the DBO engine · SIGMOD Conference 2007
The Sort-Merge-Shrink join · ACM Trans. Database Syst. 2006
Query processing and optimization › aggregation
online aggregation
0.122008
Scalable approximate query processing with the DBO engine · ACM Trans. Database Syst. 2008
The Sort-Merge-Shrink join · ACM Trans. Database Syst. 2006
Query processing and optimization
aggregate query processing
0.112008
Scalable approximate query processing with the DBO engine · ACM Trans. Database Syst. 2008
Data stream processing
random sampling
0.112008
Maintaining very large random samples using the geometric file · VLDB J. 2008
Data stream processing › stream sampling
reservoir sampling
0.112008
Maintaining very large random samples using the geometric file · VLDB J. 2008
Query processing and optimization
long-running queries
0.112007
Stop-and-Restart Style Execution for Long Running Decision Support Queries · VLDB 2007
Query processing and optimization
join processing
0.112006
The Sort-Merge-Shrink join · ACM Trans. Database Syst. 2006
Query processing and optimization › join processing › join algorithms
sort-merge join
0.112006
The Sort-Merge-Shrink join · ACM Trans. Database Syst. 2006
Query processing and optimization
cardinality estimation
0.112005
Online Estimation For Subset-Based SQL Queries · VLDB 2005
Information retrieval › evaluation
confidence interval
0.112005
Relational Confidence Bounds Are Easy With The Bootstrap · SIGMOD Conference 2005
Data mining › statistical analysis
statistical estimation
0.112005
Relational Confidence Bounds Are Easy With The Bootstrap · SIGMOD Conference 2005
Data stream processing
stream sampling
0.012004
Online Maintenance of Very Large Random Samples · SIGMOD Conference 2004
Storage systems › file systems
file organization
0.012008
Maintaining very large random samples using the geometric file · VLDB J. 2008
Query processing and optimization
analytical query processing
0.012007
Scalable approximate query processing with the DBO engine · SIGMOD Conference 2007
Query processing and optimization › analytical query processing
decision support query processing
0.012007
Stop-and-Restart Style Execution for Long Running Decision Support Queries · VLDB 2007
Indexing and storage engines › external memory data structure
disk-based index
0.012004
Online Maintenance of Very Large Random Samples · SIGMOD Conference 2004

Methods — techniques the papers use, named apart from their topics

statistical estimation · 0.2geometric file · 0.2statistical inference · 0.1bootstrap · 0.1
YearPublicationVenuePosition
2008 Scalable approximate query processing with the DBO engine
abstract
This article describes query processing in the DBO database system. Like other database systems designed for ad hoc analytic processing, DBO is able to compute the exact answers to queries over a large relational database in a scalable fashion. Unlike any other system designed for analytic processing, DBO can constantly maintain a guess as to the final answer to an aggregate query throughout execution, along with statistically meaningful bounds for the guess's accuracy. As DBO gathers more and more information, the guess gets more and more accurate, until it is 100% accurate as the query is completed. This allows users to stop the execution as soon as they are happy with the query accuracy, and thus encourages exploratory data analysis.
Chris Jermaine, Subramanian Arumugam 0002, Abhijit Pol, Alin Dobra
ACM Trans. Database Syst.3
2008 Maintaining very large random samples using the geometric file
Abhijit Pol, Chris Jermaine, Subramanian Arumugam 0002
VLDB J.1
2007 Scalable approximate query processing with the DBO engine
abstract
This paper describes query processing in the DBO database system. Like other database systems designed for ad-hoc, analytic processing, DBO is able to compute the exact answer to queries over a large relational database in a scalable fashion. Unlike any other system designed for analytic processing, DBO can constantly maintain a guess as to the final answer to an aggregate query throughout execution, along with statistically meaningful bounds for the guess's accuracy. As DBO gathers more and more information, the guess gets more and more accurate, until it is 100% accurate as the query is completed. This allows users to stop the execution at any time that they are happy with the query accuracy, and encourages exploratory data analysis.
Chris Jermaine, Subramanian Arumugam 0002, Abhijit Pol, Alin Dobra
SIGMOD Conference3
2007 Stop-and-Restart Style Execution for Long Running Decision Support Queries
Surajit Chaudhuri, Raghav Kaushik, Ravishankar Ramamurthy, Abhijit Pol
VLDB4
2006 The Sort-Merge-Shrink join
abstract
One of the most common operations in analytic query processing is the application of an aggregate function to the result of a relational join. We describe an algorithm called the Sort-Merge-Shrink (SMS) Join for computing the answer to such a query over large, disk-based input tables. The key innovation of the SMS join is that if the input data are clustered in a statistically random fashion on disk, then at all times, the join provides an online, statistical estimator for the eventual answer to the query as well as probabilistic confidence bounds. Thus, a user can monitor the progress of the join throughout its execution and stop the join when satisfied with the estimate's accuracy or run the algorithm to completion with a total time requirement that is not much longer than that of other common join algorithms. This contrasts with other online join algorithms, which either do not offer such statistical guarantees or can only offer guarantees so long as the input data can fit into main memory.
Chris Jermaine, Alin Dobra, Subramanian Arumugam 0002, Shantanu Joshi 0001, Abhijit Pol
ACM Trans. Database Syst.5
2005 Highly Scalable and Accurate Seeds for Subsequence Alignment
abstract
We propose a method for finding seeds for the local alignment of two nucleotide sequences. Our method uses randomized algorithms to find approximate seeds. We present a dynamic index to store the fingerprints of k-grams and a highly scalable and accurate (HSA) algorithm to incorporate randomization into process of seed generation. Experimental results show that our method produces better quality seeds with improved running time and memory usage compared to traditional non-spaced and spaced seeds. The presented algorithm scales very well with higher seed lengths while maintaining the quality and performance.
Abhijit Pol, Tamer Kahveci
BIBE1
2005 A Disk-Based Join With Probabilistic Guarantees
abstract
One of the most common operations in analytic query processing is the application of an aggregate function to the result of a relational join. We describe an algorithm for computing the answer to such a query over large, disk-based input tables. The key innovation of our algorithm is that at all times, it provides an online, statistical estimator for the eventual answer to the query, as well as probabilistic confidence bounds. Thus, a user can monitor the progress of the join throughout its execution and stop the join when satisfied with the estimate's accuracy, or run the algorithm to completion with a total time requirement that is not much longer than other common join algorithms. This contrasts with other online join algorithms, which either do not offer such statistical guarantees or can only offer guarantees so long as the input data can fit into core memory.
Chris Jermaine, Alin Dobra, Subramanian Arumugam 0002, Shantanu Joshi 0001, Abhijit Pol
SIGMOD Conference5
2005 Relational Confidence Bounds Are Easy With The Bootstrap
abstract
Statistical estimation and approximate query processing have become increasingly prevalent applications for database systems. However, approximation is usually of little use without some sort of guarantee on estimation accuracy, or "confidence bound." Analytically deriving probabilistic guarantees for database queries over sampled data is a daunting task, not suitable for the faint of heart, and certainly beyond the expertise of the typical database system end-user. This paper considers the problem of incorporating into a database system a powerful "plug-in" method for computing confidence bounds on the answer to relational database queries over sampled or incomplete data. This statistical tool, called the bootstrap, is simple enough that it can be used by a data-base programmer with a rudimentary mathematical background, but general enough that it can be applied to almost any statistical inference problem. Given the power and ease-of-use of the bootstrap, we argue that the algorithms presented for supporting the bootstrap should be incorporated into any database system which is intended to support analytic processing.
Abhijit Pol, Chris Jermaine
SIGMOD Conference1
2005 Online Estimation For Subset-Based SQL Queries
Chris Jermaine, Alin Dobra, Abhijit Pol, Shantanu Joshi 0001
VLDB3
2004 Online Maintenance of Very Large Random Samples
abstract
Random sampling is one of the most fundamental data management tools available. However, most current research involving sampling considers the problem of how to use a sample, and not how to compute one. The implicit assumption is that a "sample" is a small data structure that is easily maintained as new data are encountered, even though simple statistical arguments demonstrate that very large samples of gigabytes or terabytes in size can be necessary to provide high accuracy. No existing work tackles the problem of maintaining very large, disk-based samples from a data management perspective, and no techniques now exist for maintaining very large samples in an online manner from streaming data. In this paper, we present online algorithms for maintaining on-disk samples that are gigabytes or terabytes in size. The algorithms are designed for streaming data, or for any environment where a large sample must be maintained online in a single pass through a data set. The algorithms meet the strict requirement that the sample always be a true, statistically random sample (without replacement) of all of the data processed thus far. Our algorithms are also suitable for biased or unequal probability sampling.
Chris Jermaine, Abhijit Pol, Subramanian Arumugam 0002
SIGMOD Conference2