EDBT 2026 Demo / reviewers in the wild / expert
Gaoxiang Xu
dblp:224/6709
· DBLP profile ↗
8ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Database system architecture and tuning · 97% Query processing and optimization · 3% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 68% Storage systems · 13% Energy-efficient computing · 12% |
Topics — the 12 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Database system architecture and tuning
index tuning |
1.2 | 2 | 2025 | Understanding and Detecting Query Performance Regression in Practical Index Tuning: [Experiments & Analysis] · Proc. ACM Manag. Data 2025 Automatically Indexing Millions of Databases in Microsoft Azure SQL Database · SIGMOD Conference 2019 |
Database system architecture and tuning › database performance management
query performance regression |
0.9 | 1 | 2025 | Understanding and Detecting Query Performance Regression in Practical Index Tuning: [Experiments & Analysis] · Proc. ACM Manag. Data 2025 |
Database system architecture and tuning › database performance management › query performance analysis
query plan analysis |
0.9 | 1 | 2025 | Understanding and Detecting Query Performance Regression in Practical Index Tuning: [Experiments & Analysis] · Proc. ACM Manag. Data 2025 |
Memory systems
non-volatile memory |
0.7 | 2 | 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main Memory · DAC 2019 CACF: A Novel Circuit Architecture Co-optimization Framework for Improving Performance, Reliability and Energy of ReRAM-based Main Memory System · ACM Trans. Archit. Code Optim. 2018 |
Database system architecture and tuning
index recommendation |
0.4 | 1 | 2019 | Automatically Indexing Millions of Databases in Microsoft Azure SQL Database · SIGMOD Conference 2019 |
Memory systems › non-volatile memory
bit-flip reduction |
0.4 | 1 | 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main Memory · DAC 2019 |
Storage systems › data representation
data encoding |
0.4 | 1 | 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main Memory · DAC 2019 |
Memory systems › non-volatile memory › write reliability
write endurance |
0.4 | 1 | 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main Memory · DAC 2019 |
Energy-efficient computing
circuit-architecture co-optimization |
0.3 | 1 | 2018 | CACF: A Novel Circuit Architecture Co-optimization Framework for Improving Performance, Reliability and Energy of ReRAM-based Main Memory System · ACM Trans. Archit. Code Optim. 2018 |
Memory systems › processing-in-memory
ReRAM crossbar |
0.3 | 1 | 2018 | CACF: A Novel Circuit Architecture Co-optimization Framework for Improving Performance, Reliability and Energy of ReRAM-based Main Memory System · ACM Trans. Archit. Code Optim. 2018 |
Memory systems
cache |
0.1 | 1 | 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main Memory · DAC 2019 |
Cloud and datacenter computing
database-as-a-service |
0.1 | 1 | 2019 | Automatically Indexing Millions of Databases in Microsoft Azure SQL Database · SIGMOD Conference 2019 |
Methods — techniques the papers use, named apart from their topics
pattern detection · 0.9machine learning · 0.9production experimentation · 0.8index recommendation · 0.8tag-bit sharing · 0.4adaptive granularity encoding · 0.4region partition with address remapping · 0.3double-sided write driver · 0.3RESET disturbance detection · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding and Detecting Query Performance Regression in Practical Index Tuning: [Experiments & Analysis]abstractExisting index tuners typically rely on the ''what if'' API provided by the query optimizer to estimate the execution cost of a query on top of an index configuration. Such cost estimates can be inaccurate and may therefore lead to significant query performance regression (QPR) once the recommended indexes are materialized. This becomes a serious problem for cloud database providers, such as Microsoft's Azure SQL Database, that offer index tuning as an automated service (a.k.a. ''auto-indexing''). Previous work has explored use of supervised machine learning (ML) to reduce the likelihood of QPR. However, the trained ML models have limited generalization capability when applied to new databases and workloads. We propose an alternative approach where we analyze the query plans with significant QPRs and look for structural changes due to the new index configuration that could explain the QPR. We perform such study for index tuning data across many benchmark and real-world database workloads, for multiple realistic index tuning scenarios. Our study reveals that most of the significant QPRs can be attributed to a small number of common ''regression patterns'' characterizing the structural plan changes, and we further propose a pattern-based QPR detector accordingly. Our experimental evaluation shows that the pattern-based QPR detector can significantly outperform existing ML-based QPR detectors. Wentao Wu 0001, Anshuman Dutt, Gaoxiang Xu, Vivek R. Narasayya, Surajit Chaudhuri |
Proc. ACM Manag. Data | 3 |
| 2021 | MORE2: Morphable Encryption and Encoding for Secure NVMabstractMemory encryption can enhance the security of Non-volatile memories (NVMs), but it significantly increases the data bits written to NVMs and leads to severe lifetime and performance degradation. Current encryption techniques aim to reduce the re-encryption to many existing clean words, which unfortunately suffer from high encryption overheads (i.e. latency and energy) and many unnecessary writes. In the meantime, compression techniques can reduce the writes of encrypted NVM. However, we find that they may destroy the data patterns and increase the modified words, resulting in many encryptions in secure NVM. In this paper, we propose the MORphable Encryption and Encoding (MORE2) scheme to address these problems. Our MORphable Encryption (MORE) technique aims to reduce the full-line re-encryption and avoid clean line encryption. Besides, MORE proposes a prediction-based write scheme to avoid the encryption of clean lines, and pre-encrypt the lines that are predicted as dirty. Therefore, MORE can remove the encryption from the critical path of NVM. Furthermore, MORE2proposes the Morphable Selective Encoding (MSE) scheme to compress the modified words while preserving clean words. MORE2encrypts all metadata with the line counter to guarantee high security. Experimental results show that MORE2reduces the bit flips of encrypted NVM by 53.5 %, decreases the access latency by 27.32%, improves the IPC performance by 12.1 %, and reduces the write energy by 29.1 % compared with the state-of-the-art design. Wei Zhao 0034, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Jie Xu 0013, Gaoxiang Xu, Yiran Chen 0001 |
ICCAD | 8 |
| 2019 | Adaptive Granularity Encoding for Energy-efficient Non-Volatile Main MemoryabstractData encoding methods have been proposed to alleviate the high write energy and limited write endurance disadvantages of Non-Volatile Memories (NVMs). Encoding methods are proved to be effective through theoretical analysis. Under the data patterns of workloads, existing encoding methods could become inefficient. We observe that the new cache line and the old cache line have many redundant (or unmodified) words. This makes the utilization ratio of the tag bits of data encoding methods become very low, and the efficiency of data encoding method decreases. To fully exploit the tag bits to reduce the bit flips of NVMs, we propose REdundant word Aware Data encoding (READ). The key idea of READ is to share the tag bits among all the words of the cache line and dynamically assign the tag bits to the modified words. The high utilization ratio of the tag bits in READ leads to heavy bit flips of the tag bits. To reduce the bit flips of the tag bits in READ, we further propose Sequential flips Aware Encoding (SAE). SAE is designed based on the observation that many sequential bits of the new data and the old data are opposite. For those writes, the bit flips of the tag bits will increase with the number of tag bits. SAE dynamically selects the encoding granularity which causes the minimum bit flips instead of using the minimum encoding granularity. Experimental results show that our schemes can reduce the energy consumption by 20.3%, decrease the bit flips by 25.0%, and improve the lifetime by 52.1%. Jie Xu 0013, Dan Feng 0001, Yu Hua 0001, Wei Tong 0001, Jingning Liu, Gaoxiang Xu, Yiran Chen 0001 |
DAC | 7 |
| 2019 | RFPL: A Recovery Friendly Parity Logging Scheme for Reducing Small Write Penalty of SSD RAIDabstractParity based RAID suffers from poor small write performance due to heavy parity update overhead. The recently proposed method EPLOG constructs a new stripe with updated data chunks without updating old parity chunks. However, due to skewness of data accesses, old versions of updated data chunks often need to be kept to protect other data chunks of the same stripe. This seriously hurts the efficiency of recovering system from device failures due to the need of reconstructing the preserved old data chunks on failed devices. Gaoxiang Xu, Dan Feng 0001, Jie Xu 0013, Xi Shu |
ICPP | 1 |
| 2019 | Automatically Indexing Millions of Databases in Microsoft Azure SQL DatabaseabstractAn appropriate set of indexes can result in orders of magnitude better query performance. Index management is a challenging task even for expert human administrators. Fully automating this process is of significant value. We describe the challenges, architecture, design choices, implementation, and learnings from building an industrial-strength auto-indexing service for Microsoft Azure SQL Database, a relational database service. Our service has been generally available for more than two years, generating index recommendations for every database in Azure SQL Database, automatically implementing them for a large fraction, and significantly improving performance of hundreds of thousands of databases. We also share our experience from experimentation at scale with production databases which gives us confidence in our index recommendation quality for complex real applications. Sudipto Das, Miroslav Grbic, Igor Ilic, Isidora Jovandic, Andrija Jovanovic, Vivek R. Narasayya, Miodrag Radulovic, Maja Stikic, Gaoxiang Xu, Surajit Chaudhuri |
SIGMOD Conference | 9 |
| 2019 | FvRS: Efficiently identifying performance-critical data for improving performance of big data processing
Gaoxiang Xu, Dan Feng 0001, Laurence T. Yang, Wei Zhou 0013, Yang Zhang 0051, Jie Xu 0013 |
Future Gener. Comput. Syst. | 1 |
| 2018 | Cap: Exploiting Data Correlations to Improve the Performance and Endurance of SSD RAIDabstractParity-based RAID provides system-level fault tolerance. However, parity updates caused by small writes introduce lots of extra I/Os, degrading I/O performance and wearing SSDs out. It has been proposed to use Non-Volatile Memory (NVM) as a parity cache on an SSD RAID to postpone parity updates until the whole stripe has been updated. However, this often fails because of skewed distribution of hot data chunks within a stripe. In real workloads, it is often difficult to achieve a full-stripe update even after a long delay. In this paper, we propose a Correlation aware parity caching scheme, called Cap, for SSD-based RAIDs. The key idea behind Cap is to periodically reconstruct correlated hot data chunks into a new stripe. Since these data chunks have a strong correlation, they tend to be updated together within a short time span. This co-update within a stripe more efficiently utilizes the parity cache to convert partial-stripe updates into a full-stripe update. We have implemented Cap on a RAID-5 SSD array in Linux Kernel 4.3. Experimental results show that Cap improves the I/O bandwidth by 54%~145% compared with the Linux software RAID. Compared with the state-of-the-art parity caching scheme PPC, Cap improves the I/O bandwidth by 14%~31%. Gaoxiang Xu, Dan Feng 0001, Jie Xu 0013 |
ICCD | 1 |
| 2018 | CACF: A Novel Circuit Architecture Co-optimization Framework for Improving Performance, Reliability and Energy of ReRAM-based Main Memory SystemabstractEmerging Resistive Random Access Memory (ReRAM) is a promising candidate as the replacement for DRAM due to its low standby power, high density, high scalability, and nonvolatility. By employing the unique crossbar structure, ReRAM can be constructed with extremely high density. However, the crossbar ReRAM faces some serious challenges in terms of performance, reliability, and energy consumption. First, ReRAM’s crossbar structure causes an IR drop problem due to wire resistance and sneak currents, which results in nonuniform access latency in ReRAM banks and reduces its reliability. Second, without access transistors in the crossbar structure, write disturbance results in serious data reliability problem. Third, the access latency, reliability, and energy use of ReRAM arrays are significantly influenced by the data patterns involved in a write operation. To overcome the challenges of the crossbar ReRAM, we propose a novel circuit architecture co-optimization framework for improving the performance, reliability, and energy use of ReRAM-based main memory system, called CACF. The proposed CACF consists of three levels, including the circuit level, circuit architecture level, and architecture level. At the circuit level, to reduce the IR drops along bitlines, we propose a double-sided write driver design by applying write drivers along both sides of bitlines and selectively activating the write drivers. At the circuit architecture level, to address the write disturbance with low overheads, we propose a RESET disturbance detection scheme by adding disturbance reference cells and conditionally performing refresh operations. At the architecture level, a region partition with address remapping method is proposed to leverage the nonuniform access latency in ReRAM banks, and two flip schemes are proposed in different regions to optimize the data patterns involved in a write operation. The experimental results show that CACF improves system performance by 26.1%, decreases memory access latency by 22.4%, shortens running time by 20.1%, and reduces energy consumption by 21.6% on average over an aggressive baseline. Meanwhile, CACF significantly improves the reliability of ReRAM-based memory systems. Yang Zhang 0051, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Chengning Wang, Bing Wu 0001, Zheng Li 0005, Gaoxiang Xu |
ACM Trans. Archit. Code Optim. | 10 |