Gaozhong Liang

dblp:325/9463 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Database system architecture and tuning · 56% Query processing and optimization · 24% Data mining · 20%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
database tuning
0.712023
Active Sampling for Sparse Table by Bayesian Optimization with Adaptive Resolution · ICDE 2023
Data mining › data reduction › data summarization
histogram construction
0.712023
Active Sampling for Sparse Table by Bayesian Optimization with Adaptive Resolution · ICDE 2023
Database system architecture and tuning › self-managing database systems
database diagnosis
0.612022
PinSQL: Pinpoint Root Cause SQLs to Resolve Performance Issues in Cloud Databases · ICDE 2022
Database system architecture and tuning › self-managing database systems › database diagnosis
root cause analysis
0.612022
PinSQL: Pinpoint Root Cause SQLs to Resolve Performance Issues in Cloud Databases · ICDE 2022
Query processing and optimization
cardinality estimation
0.212023
Active Sampling for Sparse Table by Bayesian Optimization with Adaptive Resolution · ICDE 2023
Cloud and datacenter computing
database-as-a-service
0.212022
PinSQL: Pinpoint Root Cause SQLs to Resolve Performance Issues in Cloud Databases · ICDE 2022

Methods — techniques the papers use, named apart from their topics

propagation chain tracking · 1.1anomaly detection · 1.1gaussian process regression · 0.7bayesian optimization · 0.7active sampling · 0.7
YearPublicationVenuePosition
2023 Active Sampling for Sparse Table by Bayesian Optimization with Adaptive Resolution
abstract
Open-source relational database systems have become increasingly popular in the cloud era. However, practitioners are often beset with query performance issues. Thus a general-purpose database performance tuning tool independent of the various DBMS kernels becomes desired to lower the bar of using these systems. The first mandatory step in developing such a tool is to design an effective sampling method that collects representative records from different tables. Although one could leverage standard SQL statements and indexes to achieve this, sampling performance and statistical efficiency are not guaranteed when the underlying tables are frequently updated, especially for Sparse Tables where the range of index values is significantly greater than the table size.To this end, we propose a novel Active Sampling algorithm that queries regions more likely to contain data records from Sparse Tables. It relies on Gaussian process regression to characterize the probability density of whether a data record is non-null at a given index value. With the help of this estimated density function, the proposed method achieves efficient sampling by actively querying records with adaptive resolutions of interval lengths and provides an unbiased estimator for histogram construction. Comprehensive experiments on synthetic and real-world datasets demonstrate that the proposed Active Sampling method can effectively improve the estimation accuracy and use less query cost than other commonly used sampling methods.
Xiao He 0008, Jian Tan 0001, Bin Wu 0003, Feifei Li 0001, Gaozhong Liang, Jinfeng Xu 0001
ICDE6
2022 PinSQL: Pinpoint Root Cause SQLs to Resolve Performance Issues in Cloud Databases
abstract
Deploying database services on cloud systems has gained increasing popularity and has become a common practice in the industry. However, the complicated cloud environments make performance issues inevitable, which could violate the service level guarantee if not addressed in a timely manner. Among the various problems, anomalies in SQL queries are the most commonly reported sources that cause performance issues in database applications. These anomalous queries can be divided into High-impact SQLs (H-SQLs) and Root Cause SQLs (R-SQLs), representing the related SQLs that are correlated with the anomalies and the ones that are the root causes of the performance issue, respectively. In the presence of a large number of queries, to pinpoint the R-SQLs is far more difficult than to identify the H-SQLs. To address this challenge, we aim at automatically pinpointing the R-SQLs to resolve performance issues in cloud databases. This paper introduces PinSQL, an autonomous diagnosing system for Alibaba Cloud, which has four modules that are executed sequentially, including data collection and pre-processing, anomaly detection, root cause analysis, and repairing actions. First, the related performance metrics and query logs from monitored cloud database instances are collected and aggregated as the data sources. Then, based on these inputs, efficient anomaly detection is conducted in real-time. Upon the detection of an anomaly, the root cause SQLs are pinpointed through tracking the propagation chain of the involved SQLs. Finally, repairing actions are suggested and then executed on R-SQLs to address the anomalies. Extensive experiments on an Alibaba production system show that PinSQL can achieve an 80% accuracy for pinpointing the top-1 R-SQLs and successfully resolve the database performance issues resultantly.
Xiaoze Liu, Zheng Yin, Congcong Ge, Lu Chen 0001, Yunjun Gao, Dimeng Li, Ziting Wang, Gaozhong Liang, Jian Tan 0001, Feifei Li 0001
ICDE9