Xing Niu 0002

dblp:87/9555-2 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
2since 2021 · last 2026
0000-0003-3006-9546ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 4 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 53% Data models and query languages · 22% Data integration and cleaning · 12%

Topics — the 5 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › runtime optimization
data skipping
0.512021
Provenance-based Data Skipping · Proc. VLDB Endow. 2021
Data integration and cleaning
data provenance
0.422019
Debugging Transactions and Tracking their Provenance with Reenactment · Proc. VLDB Endow. 2017
Heuristic and Cost-Based Optimization for Diverse Provenance Tasks · IEEE Trans. Knowl. Data Eng. 2019
Data models and query languages
temporal data model
0.412019
Snapshot Semantics for Temporal Multiset Relations · Proc. VLDB Endow. 2019
Query processing and optimization › query optimization
cost-based optimization
0.312017
Provenance-Aware Query Optimization · ICDE 2017
Indexing and storage engines › synopsis structure
zone maps
0.112021
Provenance-based Data Skipping · Proc. VLDB Endow. 2021

Methods — techniques the papers use, named apart from their topics

heuristic optimization · 0.7algebraic equivalences · 0.7provenance sketches · 0.5middleware implementation · 0.4k-relations · 0.4cost-based optimization · 0.4reenactment · 0.3declarative replay · 0.3
YearPublicationVenuePosition
2026 In-memory Incremental Maintenance of Provenance Sketches
Pengyuan Li 0007, Boris Glavic, Dieter Gawlick, Vasudha Krishnaswamy, Zhen Hua Liu, Danica Porobic, Xing Niu 0002
EDBT7
2021 Provenance-based Data Skipping
abstract
Database systems use static analysis to determine upfront which data is needed for answering a query and use indexes and other physical design techniques to speed-up access to that data. However, for important classes of queries, e.g., HAVING and top-k queries, it is impossible to determine up-front what data is relevant. To overcome this limitation, we develop provenance-based data skipping (PBDS), a novel approach that generates provenance sketches to concisely encode what data is relevant for a query. Once a provenance sketch has been captured it is used to speed up subsequent queries. PBDS can exploit physical design artifacts such as indexes and zone maps.
Xing Niu 0002, Boris Glavic, Pengyuan Li 0007, Dieter Gawlick, Vasudha Krishnaswamy, Zhen Hua Liu, Danica Porobic
Proc. VLDB Endow.1
2019 Snapshot Semantics for Temporal Multiset Relations
abstract
Snapshot semantics is widely used for evaluating queries over temporal data: temporal relations are seen as sequences of snapshot relations, and queries are evaluated at each snapshot. In this work, we demonstrate that current approaches for snapshot semantics over interval-timestamped multiset relations are subject to two bugs regarding snapshot aggregation and bag difference. We introduce a novel temporal data model based on K -relations that overcomes these bugs and prove it to correctly encode snapshot semantics. Furthermore, we present an efficient implementation of our model as a database middleware and demonstrate experimentally that our approach is competitive with native implementations.
Anton Dignös, Boris Glavic, Xing Niu 0002, Johann Gamper, Michael H. Böhlen
Proc. VLDB Endow.3
2019 Heuristic and Cost-Based Optimization for Diverse Provenance Tasks
abstract
A well-established technique for capturing database provenance as annotations on data is to instrument queries to propagate such annotations. However, even sophisticated query optimizers often fail to produce efficient execution plans for instrumented queries. We develop provenance-aware optimization techniques to address this problem. Specifically, we study algebraic equivalences targeted at instrumented queries and alternative ways of instrumenting queries for provenance capture. Furthermore, we present an extensible heuristic and cost-based optimization framework utilizing these optimizations. Our experiments confirm that these optimizations are highly effective, improving performance by several orders of magnitude for diverse provenance tasks.
Xing Niu 0002, Raghav Kapoor, Boris Glavic, Dieter Gawlick, Zhen Hua Liu, Vasudha Krishnaswamy, Venkatesh Radhakrishnan
IEEE Trans. Knowl. Data Eng.1
2017 Adaptive Schema Databases
William Spoth, Bahareh Arab, Eric S. Chan, Dieter Gawlick, Adel Ghoneimy, Boris Glavic, Beda Christoph Hammerschmidt, Oliver Kennedy, Seokki Lee, Zhen Hua Liu, Xing Niu 0002, Ying Yang 0005
CIDR11
2017 Provenance-Aware Query Optimization
abstract
Data provenance is essential for debugging query results, auditing data in cloud environments, and explaining outputs of Big Data analytics. A well-established technique is to represent provenance as annotations on data and to instrument queries to propagate these annotations to produce results annotated with provenance. However, even sophisticated optimizers are often incapable of producing efficient execution plans for instrumented queries, because of their inherent complexity and unusual structure. Thus, while instrumentation enables provenance support for databases without requiring any modification to the DBMS, the performance of this approach is far from optimal. In this work, we develop provenancespecific optimizations to address this problem. Specifically, we introduce algebraic equivalences targeted at instrumented queries and discuss alternative, equivalent ways of instrumenting a query for provenance capture. Furthermore, we present an extensible heuristic and cost-based optimization (CBO) framework that governs the application of these optimizations and implement this framework in our GProM provenance system. Our CBO is agnostic to the plan space shape, uses a DBMS for cost estimation, and enables retrofitting of optimization choices into existing code by adding a few LOC. Our experiments confirm that these optimizations are highly effective, often improving performance by several orders of magnitude for diverse provenance tasks.
Xing Niu 0002, Raghav Kapoor, Boris Glavic, Dieter Gawlick, Zhen Hua Liu, Venkatesh Radhakrishnan
ICDE1
2017 Debugging Transactions and Tracking their Provenance with Reenactment
abstract
Debugging transactions and understanding their execution are of immense importance for developing OLAP applications, to trace causes of errors in production systems, and to audit the operations of a database. However, debugging transactions is hard for several reasons: 1) after the execution of a transaction, its input is no longer available for debugging, 2) internal states of a transaction are typically not accessible, and 3) the execution of a transaction may be affected by concurrently running transactions. We present a debugger for transactions that enables non-invasive, postmortem debugging of transactions with provenance tracking and supports what-if scenarios (changes to transaction code or data). Using reenactment , a declarative replay technique we have developed, a transaction is replayed over the state of the DB seen by its original execution including all its interactions with concurrently executed transactions from the history. Importantly, our approach uses the temporal database and audit logging capabilities available in many DBMS and does not require any modifications to the underlying database system nor transactional workload.
Xing Niu 0002, Bahareh Arab, Seokki Lee, Su Feng, Xun Zou, Dieter Gawlick, Vasudha Krishnaswamy, Zhen Hua Liu, Boris Glavic
Proc. VLDB Endow.1