Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chaitanya Mishra

dblp:24/2530 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 7 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Query processing and optimization · 80% Distributed and cloud data management · 14% Database system architecture and tuning · 6%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
cardinality estimation
0.452009
The design of a query monitoring system · ACM Trans. Database Syst. 2009
Join Reordering by Join Simulation · ICDE 2009
Generating targeted queries for database testing · SIGMOD Conference 2008
Query processing and optimization
multi-query optimization
0.322014
Sharing across Multiple MapReduce Jobs · ACM Trans. Database Syst. 2014
MRShare: Sharing Across Multiple Queries in MapReduce · Proc. VLDB Endow. 2010
Query processing and optimization › query optimization › predicate optimization
filter ordering
0.212014
Sharing across Multiple MapReduce Jobs · ACM Trans. Database Syst. 2014
Distributed and cloud data management
mapreduce
0.212014
Sharing across Multiple MapReduce Jobs · ACM Trans. Database Syst. 2014
Query processing and optimization › multi-query optimization
query sharing
0.112010
MRShare: Sharing Across Multiple Queries in MapReduce · Proc. VLDB Endow. 2010
Query processing and optimization › query optimization
join ordering
0.112009
Join Reordering by Join Simulation · ICDE 2009
Database system architecture and tuning › database performance management
query monitoring
0.112009
The design of a query monitoring system · ACM Trans. Database Syst. 2009
Software testing
database testing
0.112008
Generating targeted queries for database testing · SIGMOD Conference 2008
Query processing and optimization › query execution
query progress estimation
0.112007
ConEx: a system for monitoring queries · SIGMOD Conference 2007
Query processing and optimization › query execution
query progress indicator
0.112007
A Lightweight Online Framework For Query Progress Indicators · ICDE 2007
Distributed and cloud data management › mapreduce
mapreduce query processing
0.012010
MRShare: Sharing Across Multiple Queries in MapReduce · Proc. VLDB Endow. 2010
Query processing and optimization
adaptive query processing
0.012009
The design of a query monitoring system · ACM Trans. Database Syst. 2009

Methods — techniques the papers use, named apart from their topics

cost model · 0.3adaptive filter ordering · 0.2job grouping · 0.1statistical summaries · 0.1random sampling · 0.1lightweight online estimators · 0.1join simulation · 0.1sampling · 0.1online estimation · 0.1
YearPublicationVenuePosition
2014 Sharing across Multiple MapReduce Jobs
abstract
Large-scale data analysis lies in the core of modern enterprises and scientific research. With the emergence of cloud computing, the use of an analytical query processing infrastructure can be directly associated with monetary cost. MapReduce has been a popular framework in the context of cloud computing, designed to serve long-running queries (jobs) which can be processed in batch mode. Taking into account that different jobs often perform similar work, there are many opportunities for sharing. In principle, sharing similar work reduces the overall amount of work, which can lead to reducing monetary charges for utilizing the processing infrastructure. In this article we present a sharing framework tailored to MapReduce, namely, MRShare. Our framework, MRShare, transforms a batch of queries into a new batch that will be executed more efficiently, by merging jobs into groups and evaluating each group as a single query. Based on our cost model for MapReduce, we define an optimization problem and we provide a solution that derives the optimal grouping of queries. Given the query grouping, we merge jobs appropriately and submit them to MapReduce for processing. A key property of MRShare is that it is independent of the MapReduce implementation. Experiments with our prototype, built on top of Hadoop, demonstrate the overall effectiveness of our approach. MRShare is primarily designed for handling I/O-intensive queries. However, with the development of high-level languages operating on top of MapReduce, user queries executed in this model become more complex and CPU intensive. Commonly, executed queries can be modeled as evaluating pipelines of CPU-expensive filters over the input stream. Examples of such filters include, but are not limited to, index probes, or certain types of joins. In this article we adapt some of the standard techniques for filter ordering used in relational and stream databases, propose their extensions, and implement them through MRAdaptiveFilter, an extension of MRShare for expensive filter ordering tailored to MapReduce, which allows one to handle both single- and batch-query execution modes. We present an experimental evaluation that demonstrates additional benefits of MRAdaptiveFilter, when executing CPU-intensive queries in MRShare.
Tomasz Nykiel, Michalis Potamias, Chaitanya Mishra, George Kollios, Nick Koudas
ACM Trans. Database Syst.3
2010 MRShare: Sharing Across Multiple Queries in MapReduce
abstract
Large-scale data analysis lies in the core of modern enterprises and scientific research. With the emergence of cloud computing, the use of an analytical query processing infrastructure (e.g., Amazon EC2) can be directly mapped to monetary value. MapReduce has been a popular framework in the context of cloud computing, designed to serve long running queries (jobs) which can be processed in batch mode. Taking into account that different jobs often perform similar work, there are many opportunities for sharing. In principle, sharing similar work reduces the overall amount of work, which can lead to reducing monetary charges incurred while utilizing the processing infrastructure. In this paper we propose a sharing framework tailored to MapReduce. Our framework, MRShare, transforms a batch of queries into a new batch that will be executed more efficiently, by merging jobs into groups and evaluating each group as a single query. Based on our cost model for MapReduce, we define an optimization problem and we provide a solution that derives the optimal grouping of queries. Experiments in our prototype, built on top of Hadoop, demonstrate the overall effectiveness of our approach and substantial savings.
Tomasz Nykiel, Michalis Potamias, Chaitanya Mishra, George Kollios, Nick Koudas
Proc. VLDB Endow.3
2009 Interactive query refinement
abstract
We investigate the problem of refining SQL queries to satisfy cardinality constraints on the query result. This has applications to the many/few answers problems often faced by database users. We formalize the problem of query refinement and propose a framework to support it in a database system. We introduce an interactive model of refinement that incorporates user feedback to best capture user preferences. Our techniques are designed to handle queries having range and equality predicates on numerical and categorical attributes. We present an experimental evaluation of our framework implemented in an open source data manager and demonstrate the feasibility and practical utility of our approach.
Chaitanya Mishra, Nick Koudas
EDBT1
2009 Join Reordering by Join Simulation
abstract
We introduce a framework for reordering join pipelines at runtime in a database system. This framework incorporates novel techniques for simulating the execution of a join pipeline using random samples and statistical summaries. Our simulation techniques provide accurate runtime cardinality estimates along all alternative execution paths of a join pipeline. These estimates are then utilized to compare costs of alternative execution paths in a dynamic fashion, and reorder the pipeline if a better alternative path is found. We describe simulation techniques for pipelines of different kinds of join operators. We present an experimental evaluation of a prototype implementation of our framework in an open source data manager. The results demonstrate the feasibility and utility of the approach presented herein.
Chaitanya Mishra, Nick Koudas
ICDE1
2009 The design of a query monitoring system
abstract
Query monitoring refers to the problem of observing and predicting various parameters related to the execution of a query in a database system. In addition to being a useful tool for database users and administrators, it can also serve as an information collection service for resource allocation and adaptive query processing techniques. In this article, we present a query monitoring system from the ground up, describing various new techniques for query monitoring, their implementation inside a real database system, and a novel interface that presents the observed and predicted information in an accessible manner. To enable this system, we introduce several lightweight online techniques for progressively estimating and refining the cardinality of different relational operators using information collected at query execution time. These include binary and multiway joins as well as typical grouping operations and combinations thereof. We describe the various algorithms used to efficiently implement estimators and present the results of an evaluation of a prototype implementation of our framework in an open-source data management system. Our results demonstrate the feasibility and practical utility of the approach presented herein.
Chaitanya Mishra, Nick Koudas
ACM Trans. Database Syst.1
2008 Stretch 'n' shrink: resizing queries to user preferences
abstract
We present Stretch 'n' Shrink, a query design framework that explicitly takes into account user preferences about the desired answer size, and subsequently modifies the query with user feedback to meet this target. Our system has been prototyped inside an open source data manager, and requires minimal modifications to the database engine.
Chaitanya Mishra, Nick Koudas
SIGMOD Conference1
2008 Generating targeted queries for database testing
abstract
Tools for generating test queries for databases do not explicitly take into account the actual data in the database. As a consequence, such tools cannot guarantee suitable coverage of test cases commonly required for database testing. In this paper, we investigate the problem of generating queries that satisfy cardinality constraints on intermediate subexpressions when executed on a given test database. Such queries are required to test the performance of a database system under different operating conditions.
Chaitanya Mishra, Nick Koudas, Calisto Zuzarte
SIGMOD Conference1
2007 A Lightweight Online Framework For Query Progress Indicators
abstract
Recently there has been increasing interest in the development of progress indicators for SQL queries. In this paper we present a lightweight online framework for this problem. Our framework is online, in the sense that it refines its estimate of query progress based on feedback received during query execution. It is lightweight, since our techniques are designed to impose minimal overhead on query execution without sacrificing accuracy of estimates. Our framework can estimate progressively the output size of various relational operators and pipelines. These include binary and multiway joins as well as typical grouping operations and combinations thereof. We describe the various algorithms used to efficiently implement the estimators and present the results of a thorough evaluation of a prototype implementation of our framework in an open source data manager. Our results demonstrate the feasibility and practical utility of the approach presented herein.
Chaitanya Mishra, Nick Koudas
ICDE1
2007 ConEx: a system for monitoring queries
abstract
We present a system, ConEx, for monitoring query execution in a relational database management system. ConEx offers a unified view of query execution, providing continuous visual feedback on the progress of the query, and the status of operators in the query evaluation plan. It incorporates novel techniques to dynamically estimate important parameters affecting query progress efficiently. We describe the design and features of ConEx, and discuss its technology.
Chaitanya Mishra, Maksims Volkovs
SIGMOD Conference1