Yuxuan Zhu 0003

dblp:146/0939-3 · DBLP profile ↗
← Back
6ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0009-0001-6738-8930ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 6 (2 first)
YearPublicationVenuePosition
2026 Text-to-SQL Benchmarks are Broken: An In-Depth Analysis of Annotation Errors
Tengjun Jin, Yoojin Choi, Yuxuan Zhu 0003, Daniel Kang 0001
CIDR3
2026 Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
Tengjun Jin, Yoojin Choi, Yuxuan Zhu 0003, Daniel Kang 0001
Proc. VLDB Endow.3
2025 Efficient Approximate Query Processing with Block Sampling
Yuxuan Zhu 0003, Daniel Kang 0001
CIDR1
2025 PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees
abstract
After decades of research in approximate query processing (AQP), its adoption in the industry remains limited. Existing methods struggle to simultaneously provide user-specified error guarantees, eliminate maintenance overheads, and avoid modifications to database management systems. To address these challenges, we introduce two novel techniques, TAQA and BSAP. TAQA is a two-stage online AQP algorithm that achieves all three properties for arbitrary queries. However, it can be slower than exact queries if we use standard row-level sampling. BSAP resolves this by enabling block-level sampling with statistical guarantees in TAQA. We implement TAQA and BSAP in a prototype middleware system, PilotDB, that is compatible with all DBMSs supporting efficient block-level sampling. We evaluate PilotDB on PostgreSQL, SQL Server, and DuckDB over real-world benchmarks, demonstrating up to 126X speedups when running with a 5% guaranteed error.
Yuxuan Zhu 0003, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang, Charith Mendis, Daniel Kang 0001
Proc. ACM Manag. Data1
2025 ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
Tengjun Jin, Yuxuan Zhu 0003, Daniel Kang 0001
Proc. VLDB Endow.2
2023 SlabCity: Whole-Query Optimization using Program Synthesis
abstract
Query rewriting is often a prerequisite for effective query optimization, particularly for poorly-written queries. Prior work on query rewriting has relied on a set of "rules" based on syntactic pattern-matching. Whether relying on manual rules or auto-generated ones, rule-based query rewriters are inherently limited in their ability to handle new query patterns. Their success is limited by the quality and quantity of the rules provided to them. To our knowledge, we present the first synthesis-based query rewriting technique, SlabCity, capable of whole-query optimization without relying on any rewrite rules. SlabCity directly searches the space of SQL queries using a novel query synthesis algorithm that leverages a new concept called query dataflows. We evaluate SlabCity on four workloads, including a newly curated benchmark with more than 1000 real-life queries. We show that not only can SlabCity optimize more queries than state-of-the-art query rewriting techniques, but interestingly, it also leads to queries that are significantly faster than those generated by rule-based systems.
Rui Dong 0006, Jie Liu 0048, Yuxuan Zhu 0003, Cong Yan, Barzan Mozafari, Xinyu Wang 0006
Proc. VLDB Endow.3