Dachuan Qu

dblp:266/6083 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 96% Data integration and cleaning · 4%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 50% Parallel and multicore computing · 50%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
incremental computation
1.122023
Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › incremental computation
incremental query processing
0.922020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › query optimization
cost-based optimization
0.712023
Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version) · VLDB J. 2023
Query processing and optimization
query optimization
0.412020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.412020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Parallel and multicore computing › parallel scheduling
resource-aware scheduling
0.412020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Data integration and cleaning
data warehouse
0.112020
Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing · SIGMOD Conference 2020
Query processing and optimization › view maintenance
incremental view maintenance
0.112020
Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing · Proc. VLDB Endow. 2020

Methods — techniques the papers use, named apart from their topics

incremental batch processing · 0.9cost-based optimizer framework · 0.7time-varying relations · 0.4rewrite rules · 0.4plan space exploration · 0.4
YearPublicationVenuePosition
2023 Tempura: a general cost-based optimizer framework for incremental data processing (Journal Version)
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
VLDB J.9
2020 Grosbeak: A Data Warehouse Supporting Resource-Aware Incremental Computing
abstract
As the primary approach to deriving decision-support insights, automated recurring routine analytic jobs account for a major part of cluster resource usages in modern enterprise data warehouses. These recurring routine jobs usually have stringent schedule and deadline determined by external business logic, and thus cause dreadful resource skew and severe resource over-provision in the cluster. In this paper, we present Grosbeak, a novel data warehouse that supports resource-aware incremental computing to process recurring routine jobs, smooths the resource skew, and optimizes the resource usage. Unlike batch processing in traditional data warehouses, Grosbeak leverages the fact that data is continuously ingested. It breaks an analysis job into small batches that incrementally process the progressively available data, and schedules these small-batch jobs intelligently when the cluster has free resources. In this demonstration, we showcase Grosbeak using real-world analysis pipelines. Users can interact with the data warehouse by registering recurring queries and observing the incremental scheduling behavior and smoothed resource usage pattern.
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
SIGMOD Conference9
2020 Tempura: A General Cost-Based Optimizer Framework for Incremental Data Processing
abstract
Incremental processing is widely-adopted in many applications, ranging from incremental view maintenance, stream computing, to recently emerging progressive data warehouse and intermittent query processing. Despite many algorithms developed on this topic, none of them can produce an incremental plan that always achieves the best performance, since the optimal plan is data dependent. In this paper, we develop a novel cost-based optimizer framework, called Tempura, for optimizing incremental data processing. We propose an incremental query planning model called TIP based on the concept of time-varying relations, which can formally model incremental processing in its most general form. We give a full specification of Tempura, which can not only unify various existing techniques to generate an optimal incremental plan, but also allow the developer to add their rewrite rules. We study how to explore the plan space and search for an optimal incremental plan. We evaluate Tempura in various incremental processing scenarios to show its effectiveness and efficiency.
Zuozhi Wang, Kai Zeng 0002, Botong Huang, Wei Chen 0133, Xiaozong Cui, Liya Fan, Dachuan Qu, Chen Li 0001, Jingren Zhou 0001
Proc. VLDB Endow.9