Xi Liang 0002

dblp:117/1897-2 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0002-7580-5933ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2023 JanusAQP: Efficient Partition Tree Maintenance for Dynamic Approximate Query Processing
abstract
Approximate query processing over dynamic databases, i.e., under insertions/deletions, has applications ranging from high-frequency trading to internet-of-things analytics. We present JanusAQP, a new dynamic AQP system, which supports SUM, COUNT, AVG, MIN, and MAX queries under insertions and deletions to the dataset. JanusAQP extends static partition tree synopses, which are hierarchical aggregations of datasets, into the dynamic setting. This paper contributes new methods for: (1) efficient initialization of the data synopsis in the presence of incoming data, (2) maintenance of the data synopsis under insertions/deletions, and (3) re-optimization of the partitioning to reduce the approximation error. JanusAQP reduces the error of a state-of-the-art baseline by more than 60% using only 10% storage cost. JanusAQP can process more than 100K updates per second in a single node setting and keep the query latency at a millisecond level.
Xi Liang 0002, Stavros Sintos, Sanjay Krishnan
ICDE1
2021 CIAO: An Optimization Framework for Client-Assisted Data Loading
abstract
Data loading has been one of the most common performance bottlenecks for many big data applications, especially when they are running on inefficient human-readable formats, such as JSON or CSV. Parsing, validating, integrity checking and data structure maintenance are all computationally expensive steps in loading these formats. Regardless of these costs, many records may be filtered later during query evaluation due to highly selective predicates - resulting in wasted computation. Meanwhile, the computing power of client ends is typically not exploited. Here, we explore investing limited cycles of clients on prefiltering to accelerate data loading and enable data skipping for query execution. In this paper, we present CIAO, a tunable system to enable client cooperation with the server to enable efficient partial loading and data skipping for a given workload. We proposed an efficient algorithm that would select a near-optimal predicate set to push down within a given budget. Our experiments show that CIAO can significantly accelerate the data loading processing and query execution.
Cong Ding 0002, Dixin Tang, Xi Liang 0002, Aaron J. Elmore, Sanjay Krishnan
ICDE3
2021 Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing
abstract
Sample-based approximate query processing (AQP) suffers from many pitfalls such as the inability to answer very selective queries and unreliable confidence intervals when sample sizes are small. Recent research presented an intriguing solution of combining materialized, pre-computed aggregates with sampling for accurate and more reliable AQP. We explore this solution in detail in this work and propose an AQP physical design called PASS, or Precomputation-Assisted Stratified Sampling. PASS builds a tree of partial aggregates that cover different partitions of the dataset. The leaf nodes of this tree form the strata for stratified samples. Aggregate queries whose predicates align with the partitions (or unions of partitions) are exactly answered with a depth-first search, and any partial overlaps are approximated with the stratified samples. We propose an algorithm for optimally partitioning the data into such a data structure with various practical approximation techniques.
Xi Liang 0002, Stavros Sintos, Zechao Shang, Sanjay Krishnan
SIGMOD Conference1
2020 CrocodileDB: Efficient Database Execution through Intelligent Deferment
Zechao Shang, Xi Liang 0002, Dixin Tang, Cong Ding 0002, Aaron J. Elmore, Sanjay Krishnan, Michael J. Franklin
CIDR2
2020 Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints
abstract
Today, data analysts largely rely on intuition to determine whether missing or withheld rows of a dataset significantly affect their analyses. We propose a framework that can produce automatic contingency analysis, i.e., the range of values an aggregate SQL query could take, under formal constraints describing the variation and frequency of missing data tuples. We describe how to process SUM, COUNT, AVG, MIN, and MAX queries in these conditions resulting in hard error bounds with testable constraints. We propose an optimization algorithm based on an integer program that reconciles a set of such constraints, even if they are overlapping, conflicting, or unsatisfiable, into such bounds. Our experiments on real-world datasets against several statistical imputation and inference baselines show that statistical techniques can have a deceptively high error rate that is often unpredictable. In contrast, our framework offers hard bounds that are guaranteed to hold if the constraints are not violated. In spite of these hard bounds, we show competitive accuracy to statistical baselines.
Xi Liang 0002, Zechao Shang, Sanjay Krishnan, Aaron J. Elmore, Michael J. Franklin
SIGMOD Conference1
2017 Cross-layer refresh mitigation for efficient and reliable DRAM systems: A comparative study
abstract
DRAM is a crucial component in computing systems, and is expected to be even more important as data-intensive applications become more prominent. A key challenge in advancing DRAM technology is the growing cost of refresh operations, which can impose a large impact on the energy efficiency of DRAM modules. Existing refresh mitigation techniques all require hardware modifications, which may be undesirable. In this paper, we make two major contributions. First, we present a new cross-layer refresh mitigation approach that takes into account both DRAM retention time characteristics and application behaviors. The main ideas are: (1) uniformly lower refresh rate to improve DRAM energy efficiency, (2) utilize a combination of system-level memory tests and ECC/scrubbing to detect and correct errors resulting from the reduced refresh rate, and (3) perform software-based memory repair so that DRAM cells in which errors occur because of the reduced refresh rate are not mapped to application space. Our approach reduces refresh power by 98.6% and improves overall DRAM energy efficiency by up to 37.7% without sacrificing reliability, at the low price of a small reduction in main memory capacity. Second, we perform thorough experiments and analysis to compare our cross-layer approach with prior work. In general, our approach achieves better or similar power benefits, but it does not require any modifications to the hardware.
Xiaoan Ding, Xi Liang 0002, Yanjing Li
ITC2