Guanduo Chen

dblp:352/6148 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2024
0009-0001-6882-1417ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Indexing and storage engines · 57% Data models and query languages · 43%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Indexing and storage engines
bitmap index
0.812024
Oasis: An Optimal Disjoint Segmented Learned Range Filter · Proc. VLDB Endow. 2024
Indexing and storage engines › filter data structures
range filter
0.812024
Oasis: An Optimal Disjoint Segmented Learned Range Filter · Proc. VLDB Endow. 2024
Data models and query languages › natural language interface
natural language interface to database
0.712023
Gar: A Generate-and-Rank Approach for Natural Language to SQL Translation · ICDE 2023
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.712023
Gar: A Generate-and-Rank Approach for Natural Language to SQL Translation · ICDE 2023
Indexing and storage engines
key-value store
0.212024
Oasis: An Optimal Disjoint Segmented Learned Range Filter · Proc. VLDB Endow. 2024
Natural language and speech › Language models and text generation
text generation
0.212023
Gar: A Generate-and-Rank Approach for Natural Language to SQL Translation · ICDE 2023

Methods — techniques the papers use, named apart from their topics

rule-based translation · 1.3learning to rank · 1.3generate-and-rank · 1.3theoretical analysis · 0.8piecewise linear function · 0.8
YearPublicationVenuePosition
2024 Oasis: An Optimal Disjoint Segmented Learned Range Filter
abstract
The learning-enhanced data structure has inspired the development of the range filter, bringing significantly better false positive rate (FPR) than traditional non-learned range filters. Its core idea is to employ piece-wise linear functions that uniformly map the entire key space into a bitmap sequentially. Nonetheless, such uniform mapping can be space-ineffective, impacting FPRs. This paper introduces Oasis, a novel learned range filter that divides the key space into disjointed intervals by excluding large empty ranges explicitly and optimally maps those unpruned intervals into a compressed bitmap. The configuration optimality in Oasis is guaranteed by a careful theoretical analysis. To enhance the versatility of Oasis, we further propose Oasis+, which integrates the design space of both learned and non-learned filters, delivering robust performance across a wide range of workloads. We evaluate the performance of both Oasis and Oasis+ when integrated into the key-value system RocksDB, using a diverse set of real-world and synthetic datasets and workloads. In RocksDB, Oasis and Oasis+ improve the performance by up to 1.4× and 6.2× when compared to state-of-the-art learned and non-learned range filters.
Guanduo Chen, Meng Li 0010, Siqiang Luo, Zhenying He
Proc. VLDB Endow.1
2023 Gar: A Generate-and-Rank Approach for Natural Language to SQL Translation
abstract
A Natural Language (NL) Interface to Databases (NLIDB) aims to help end-users access databases. State-of-the-art approaches primarily construct language translation models to convert NL queries to SQL queries. While these models exhibit good performance on NLIDB benchmarks, the translation accuracy seems to have stalled at between 70%-75%, and most erroneous translations happen with complex queries that require an understanding of the structure and semantics specific to a database. This paper proposes a Generate-And-Rank approach called Gar. Gar assumes that a set of sample SQL queries is given to represent the possible user-intended queries to the database. In order to provide a broad coverage, akin to avoiding over-fitting, Gar extracts the basic components from the sample set to form the basic building blocks to generate a set of generalized SQL queries. By leveraging a simple rule-based SQL to NL technique, a less natural NL expression called a dialect expression for each sample and generalized SQL query is obtained. Finally, a learning-to-rank method is used for a given NL query to retrieve the best dialect expression and hence the resulting SQL query. Extensive experiments are performed to study Gar in comparison with other approaches. The results show that Gar achieves better performance on the NLIDB benchmarks, including in particular a 78.5% translation accuracy on the popular Spider benchmark, outperforming the best reported accuracy in the literature. An extension to Gar, called Gar-j, is further introduced to aid the translation by annotating join semantics in the sample queries. The experimental results show that Gar-j can further improve translation accuracy on queries with joins. Code for Gar can be found at https://github.com/Kaimary/GAR.
Yuankai Fan, Zhenying He, Tonghui Ren, Dianjun Guo, Ruisi Zhu, Guanduo Chen, Yinan Jing, Kai Zhang 0006, Xiaoyang Sean Wang
ICDE7