Mitch Cherniack

dblp:c/MitchCherniack · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
0since 2021 · last 2016
0000-0001-7461-891XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 6 first-authorArtificial intelligence and machine learning · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
17 papers
Query processing and optimization · 42% Data stream processing · 27% Database system architecture and tuning · 21%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 20 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query optimization
0.222014
Devel-op: An optimizer development environment · ICDE 2014
Inferring Function Semantics to Optimize Queries · VLDB 1998
Query processing and optimization › query planning
query plan generation
0.212014
Devel-op: An optimizer development environment · ICDE 2014
Data stream processing
stream processing systems
0.132006
Revision Processing in a Stream Processing Engine: A High-Level Design · ICDE 2006
Aurora: a new model and architecture for data stream management · VLDB J. 2003
Aurora: A Data Stream Management System · SIGMOD Conference 2003
Data stream processing › continuous query processing
streaming SQL
0.112008
Towards a streaming SQL standard · Proc. VLDB Endow. 2008
Storage systems › data management › database storage
columnar storage
0.112005
C-Store: A Column-oriented DBMS · VLDB 2005
Query processing and optimization
approximate query processing
0.012003
Load Shedding in a Data Stream Manager · VLDB 2003
Indexing and storage engines › caching
cache management
0.012003
Profile-Driven Cache Management · ICDE 2003
Data stream processing
load shedding
0.012003
Load Shedding in a Data Stream Manager · VLDB 2003
Data stream processing
operator scheduling
0.012003
Operator Scheduling in a Data Stream Manager · VLDB 2003
Information retrieval
query processing
0.012003
Avoiding Ordering and Grouping In Query Processing · VLDB 2003
Query processing and optimization › runtime optimization › prefetching
query result prefetching
0.012003
Profile-Driven Cache Management · ICDE 2003
Query processing and optimization › query rewriting › query transformation
query transformation rules
0.021998
Changing the Rules: Transformations for Rule-Based Optimizers · SIGMOD Conference 1998
Rule Languages and Internal Algebras for Rule-Based Optimizers · SIGMOD Conference 1996
Query processing and optimization › query optimization › transformation-based optimization
rule-based optimization
0.021998
Changing the Rules: Transformations for Rule-Based Optimizers · SIGMOD Conference 1998
Rule Languages and Internal Algebras for Rule-Based Optimizers · SIGMOD Conference 1996
Data stream processing
stream monitoring
0.012002
Monitoring Streams - A New Class of Data Management Applications · VLDB 2002
Database system architecture and tuning › active database
production rules
0.011998
Changing the Rules: Transformations for Rule-Based Optimizers · SIGMOD Conference 1998
Query processing and optimization › semantic query processing
semantic query optimization
0.011998
Inferring Function Semantics to Optimize Queries · VLDB 1998
Data integration and cleaning
data quality
0.012006
Revision Processing in a Stream Processing Engine: A High-Level Design · ICDE 2006
Data models and query languages
query algebra
0.011996
Rule Languages and Internal Algebras for Rule-Based Optimizers · SIGMOD Conference 1996
Data stream processing
continuous query processing
0.012003
Aurora: A Data Stream Management System · SIGMOD Conference 2003
Programming languages and type systems
term rewriting
0.012002
Visual COKO: a debugger for query optimizer development · SIGMOD Conference 2002

Methods — techniques the papers use, named apart from their topics

profiling · 0.2declarative specification · 0.2benchmarking · 0.2randomized approach · 0.1profile language · 0.1partial order semantics · 0.1greedy algorithm · 0.1batching operator · 0.1revision tuples · 0.1data stream processing · 0.0term rewriting · 0.0automated theorem proving · 0.0data management · 0.0context awareness · 0.0
YearPublicationVenuePosition
2016 OptMark: A Toolkit for Benchmarking Query Optimizers
abstract
Query optimizers have long been considered as among the most complex components of a database engine, while the assessment of an optimizer's quality remains a challenging task. Indeed, existing performance benchmarks for database engines (like TPC benchmarks) produce a performance assessment of the query runtime system rather than its query optimizer. To address this challenge, this paper introduces OptMark, a toolkit for evaluating the quality of a query optimizer. OptMark is designed to offer a number of desirable properties. First, it decouples the quality of an optimizer from the quality of its underlying execution engine. Second it evaluates independently both the effectiveness of an optimizer (i.e., quality of the chosen plans) and its efficiency (i.e., optimization time). OptMark includes also a generic benchmarking toolkit that is minimum invasive to the DBMS that wishes to use it. Any DBMS can provide a system-specific implementation of a simple API that allows OptMark to run and generate benchmark scores for the specific system. This paper discusses the metrics we propose for evaluating an optimizer's quality, the benchmark's design and the toolkit's API and functionality. We have implemented OptMark on the open-source MySQL engine as well as two commercial database systems. Using these implementations we are able to assess the quality of the optimizers on these three systems based on the TPC-DS benchmark queries.
Olga Papaemmanouil, Mitch Cherniack
CIKM3
2014 Devel-op: An optimizer development environment
abstract
Recent advances in the underlying architectures of database management systems (DBMS) have motivated the redesign of key DBMS components such as the query optimizer. Optimizers are inherently difficult to build and maintain, and yet there exists no software engineering tools to facilitate their development. In this paper, we introduce a [Devel]opment Environment for Query [Op]timizers (Devel-Op) designed to facilitate the rapid prototyping, profiling and benchmarking of optimizers. Our current version of the tool permits declarative specification and generation of two key optimizer components (the logical plan enumerator and physical plan generator) as well as debugging and visualization tools for profiling generated components.
Zhibo Peng, Mitch Cherniack, Olga Papaemmanouil
ICDE2
2013 Query Steering for Interactive Data Exploration
Ugur Çetintemel, Mitch Cherniack, Justin A. DeBrabant, Yanlei Diao, Kyriaki Dimitriadou, Alexander Kalinin 0001, Olga Papaemmanouil, Stanley B. Zdonik
CIDR2
2013 Data Curation at Scale: The Data Tamer System
Michael Stonebraker, Daniel Bruckner, Ihab F. Ilyas, George Beskales, Mitch Cherniack, Stanley B. Zdonik, Alexander Pagan
CIDR5
2008 Towards a streaming SQL standard
abstract
This paper describes a unification of two different SQL extensions for streams and its associated semantics. We use the data models from Oracle and StreamBase as our examples. Oracle uses a time-based execution model while StreamBase uses a tuple-based execution model. Time-based execution provides a way to model simultaneity while tuple-based execution provides a way to react to primitive events as soon as they are seen by the system. The result is a new model that gives the user control over the granularity at which one can express simultaneity. Of course, it is possible to ignore simultaneity altogether. The proposed model captures ordering and simultaneity through partial orders on batches of tuples. The batching and the ordering are encapsulated in and can be modified by means of a powerful new operator that we call SPREAD. This paper describes the semantics of SPREAD and gives several examples of its use.
Namit Jain, Shailendra Mishra, Anand Srinivasan, Johannes Gehrke, Jennifer Widom, Hari Balakrishnan, Ugur Çetintemel, Mitch Cherniack, Richard Tibbetts, Stanley B. Zdonik
Proc. VLDB Endow.8
2007 One Size Fits All? Part 2: Benchmarking Studies
Michael Stonebraker, Chuck Bear, Ugur Çetintemel, Mitch Cherniack, Tingjian Ge, Nabil Hachem, Stavros Harizopoulos, John Lifter, Jennie Rogers, Stanley B. Zdonik
CIDR4
2006 Improving query I/O performance by permuting and refining block request sequences
abstract
The I/O performance of query processing can be improved using two complementary approaches: improve the buffer and the file system management policies of the DB buffer manager and the OS file system manager (e.g. page replacement), or improve the sequence of requests that are submitted to a file system manager and that lead to actual I/O's (block request sequences). This paper takes the latter approach. Exploiting common file system practices as found in Linux, we propose four techniques for permuting and refining block request sequences: Block-Level I/O Grouping, File-Level I/O Grouping, I/O Ordering, and Block Recycling. To manifest these techniques, we create two new plan operations, MMS and SHJ, each of which adopts some of the block request refinement techniques above. We implement the new plan operations on top of Postgres running on Linux, and show experimental results that demonstrate up to a factor of 4 performance benefit from the use of these techniques.
Mitch Cherniack
CIKM2
2006 Revision Processing in a Stream Processing Engine: A High-Level Design
abstract
Data stream processing systems have become ubiquitous in academic [1, 2, 5, 6] and commercial [11] sectors, with application areas that include financial services, network traffic analysis, battlefield monitoring and traffic control [3]. The append-only model of streams implies that input data is immutable and therefore always correct. But in practice, streaming data sources often contend with noise (e.g., embedded sensors) or data entry errors (e.g., financial data feeds) resulting in erroneous inputs and therefore, erroneous query results. Many data stream sources (e.g., commercial ticker feeds) issue "revision tuples" (revisions) that amend previously issued tuples (e.g. erroneous share prices). Ideally, any stream processing engine should process revision inputs by generating revision outputs that correct previous query results. We know of no stream processing system that presently has this capability.
Esther Ryvkina, Anurag Maskey, Mitch Cherniack, Stanley B. Zdonik
ICDE3
2005 The Design of the Borealis Stream Processing Engine
Daniel J. Abadi, Yanif Ahmad, Magdalena Balazinska, Ugur Çetintemel, Mitch Cherniack, Jeong-Hyon Hwang, Wolfgang Lindner 0001, Anurag Maskey, Alexander Rasin, Esther Ryvkina, Nesime Tatbul, Stanley B. Zdonik
CIDR5
2005 C-Store: A Column-oriented DBMS
Michael Stonebraker, Daniel J. Abadi, Adam Batkin, Xuedong Chen, Mitch Cherniack, Miguel Ferreira, Edmond Lau, Amerson Lin, Samuel Madden 0001, Elizabeth J. O'Neil, Patrick E. O'Neil, Alexander Rasin, Nga Tran 0001, Stanley B. Zdonik
VLDB5
2004 Linear Road: A Stream Data Management Benchmark
Arvind Arasu, Mitch Cherniack, Eduardo F. Galvez, David Maier 0001, Anurag Maskey, Esther Ryvkina, Michael Stonebraker, Richard Tibbetts
VLDB2
2004 Retrospective on Aurora
Hari Balakrishnan, Magdalena Balazinska, Donald Carney, Ugur Çetintemel, Mitch Cherniack, Christian Convey, Eduardo F. Galvez, Jon Salz, Michael Stonebraker, Nesime Tatbul, Richard Tibbetts, Stanley B. Zdonik
VLDB J.5
2003 Scalable Distributed Stream Processing
Mitch Cherniack, Hari Balakrishnan, Magdalena Balazinska, Donald Carney, Ugur Çetintemel, Stanley B. Zdonik
CIDR1
2003 Profile-Driven Cache Management
abstract
Modern distributed information systems cope with disconnection and limited bandwidth by using caches. In communication-constrained situations, traditional demand-driven approaches are inadequate. Instead, caches must be preloaded in order to mitigate the absence of connectivity or the paucity of bandwidth. We propose to use application-level knowledge expressed as profiles to manage the contents of caches. We propose a simple, but rich profile language that permits high-level expression of a user's data needs for the purpose of expressing desirable contents of a cache. We consider techniques for prefetching a cache on the basis of profiles expressed in our framework, both for basic and preemptive prefetching, the latter referring to the case where staging a cache can be interrupted at any point without prior warning. We examine the effectiveness of three profile processing techniques, and show that the rich expressivity of our profile language does not prevent a fairly simple greedy algorithm from being an effective processing technique. We also show that for a large shared cache, multiple clients' profiles can be combined into a single superprofile that is representative of them all, but that when the number of clients with profiles is significantly large, a randomized approach is more scalable than a greedy approach. We believe that profiles, as described, are an enabling technology that could spawn a rich new area of research beyond cache management into network data management in general.
Mitch Cherniack, Eduardo F. Galvez, Michael J. Franklin, Stanley B. Zdonik
ICDE1
2003 Aurora: A Data Stream Management System
abstract
No abstract available.
Daniel J. Abadi, Donald Carney, Ugur Çetintemel, Mitch Cherniack, Christian Convey, C. Erwin, Eduardo F. Galvez, M. Hatoun, Anurag Maskey, Alexander Rasin, A. Singer, Michael Stonebraker, Nesime Tatbul, R. Yan, Stanley B. Zdonik
SIGMOD Conference4
2003 Operator Scheduling in a Data Stream Manager
Donald Carney, Ugur Çetintemel, Alexander Rasin, Stanley B. Zdonik, Mitch Cherniack, Michael Stonebraker
VLDB5
2003 Load Shedding in a Data Stream Manager
Nesime Tatbul, Ugur Çetintemel, Stanley B. Zdonik, Mitch Cherniack, Michael Stonebraker
VLDB4
2003 Avoiding Ordering and Grouping In Query Processing
Mitch Cherniack
VLDB2
2003 Aurora: a new model and architecture for data stream management
Daniel J. Abadi, Donald Carney, Ugur Çetintemel, Mitch Cherniack, Christian Convey, Sangdon Lee, Michael Stonebraker, Nesime Tatbul, Stanley B. Zdonik
VLDB J.4
2002 Visual COKO: a debugger for query optimizer development
abstract
Query optimization generates plans to retrieve data requested by queries. Query rewriting, which is the first step of this process, rewrites a query expression into an equivalent form to prepare it for plan generation. COKO-KOLA introduced a new approach to query rewriting that enables query rewrites to be formally verified using an automated theorem prover [1]. KOLA is a language for expressing term rewriting rules that can be fired on query expressions. COKO is a language for expressing query rewriting transformations that are too complex to express with simple KOLA rules [2].COKO is a programming language designed for query optimizer development. Programming languages require debuggers, and in this demonstration, we illustrate our COKO debugger: Visual COKO. Visual COKO enables a query optimization developer to visually trace the execution of a COKO transformation. At every step of the transformation, the developer can view a tree-display that illustrates how the original query expression has evolved.
Daniel J. Abadi, Mitch Cherniack
SIGMOD Conference2
2002 Monitoring Streams - A New Class of Data Management Applications
Donald Carney, Ugur Çetintemel, Mitch Cherniack, Christian Convey, Sangdon Lee, Greg Seidman, Michael Stonebraker, Nesime Tatbul, Stanley B. Zdonik
VLDB3
2001 Data Management for Pervasive Computing
Mitch Cherniack, Michael J. Franklin, Stanley B. Zdonik
VLDB1
1998 Changing the Rules: Transformations for Rule-Based Optimizers
abstract
Rule-based optimizers are extensible because they consist of modifiable sets of rules. For modification to be straightforward, rules must be easily reasoned about (i.e., understood and verified). At the same time, rules must be expressive and efficient (to fire) for rule-based optimizers to be practical. Production-style rules (as in [15]) are expressed with code and are hard to reason about. Pure rewrite rules (as in [1]) lack code, but cannot atomically express complex transformations (e.g., normalizations). Some systems allow rules to be grouped, but sacrifice efficiency by providing limited control over their firing. Therefore, none of these approaches succeeds in making rules expressive, efficient and understandable.
Mitch Cherniack, Stanley B. Zdonik
SIGMOD Conference1
1998 Inferring Function Semantics to Optimize Queries
Mitch Cherniack, Stanley B. Zdonik
VLDB1
1996 Rule Languages and Internal Algebras for Rule-Based Optimizers
abstract
Rule-based optimizers and optimizer generators use rules to specify query transformations. Rules act directly on query representations, which typically are based on query algebras. But most algebras complicate rule formulation, and rules over these algebras must often resort to calling to externally defined bodies of code. Code makes rules difficult to formulate, prove correct and reason about, and therefore compromises the effectiveness of rule-based systems.In this paper we present KOLA: a combinator-based algebra designed to simplify rule formulation. KOLA is not a user language, and KOLA's variable-free queries are difficult for humans to read. But KOLA is an effective internal algebra because its combinator-style makes queries manipulable and structurally revealing. As a result, rules over KOLA queries are easily expressed without the need for supplemental code. We illustrate this point, first by showing some transformations that despite their simplicity, require head and body routines when expressed over algebras that include variables. We show that these transformations are expressible without supplemental routines in KOLA. We then show complex transformations of a class of nested queries expressed over KOLA. Nested query optimization, while having been studied before, have seriously challenged the rule-based paradigm.
Mitch Cherniack, Stanley B. Zdonik
SIGMOD Conference1