Xiaozheng Zhang 0005

dblp:281/3909-5 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0007-4339-3317ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CAMEL Hash Table: Striking a Balance Between CPU and Memory Efficiency in Main-Memory Hash Join
Sudip Chatterjee 0002, Xiaozheng Zhang 0005, Suprio Ray, Ian Finlay, Calisto Zuzarte, Mark Stoodley
EDBT2
2024 Scalable Big Spatial Data Processing with SQL Query Compilation and Distributed Morsel-driven Parallelism
abstract
The rapid rise in spatial data volumes from diverse sources necessitate efficient spatial data processing capability. Although most relational databases support spatial extensions of SQL query features, they offer limited scalability. Traditional relational database query processing follows a pull-based (or tuple-at-a-time) model of query processing. This is not efficient for processing large volumes of data. A number of specialized spatial data processing systems were developed that extend cluster computing frameworks, such as Spark and Hadoop. However, these systems are characterized by limited or no support for spatial SQL query execution. The few systems that support SQL querying, suffer from the overheads of the pull-based model.We present a compilation-based distributed SQL query processing system. It follows a data-centric query compilation approach that takes a SQL query and generates distributed C++ (UPC++) based physical query plans. The generated code is compiled and executed on a distributed in-memory high performance framework based on the Partitioned Global Address Space (PGAS) paradigm. We also introduce morsel-driven parallelism for scalable spatial query execution in a distributed runtime. We conduct experimental evaluation of our system with two real-world datasets on a number of spatial query workloads. Experimental results demonstrate that our system performs significantly better than a leading spatial big data system Apache Sedona and distributed parallel relational database Citus.
Rahul Sahni, Xiaozheng Zhang 0005, Sudip Chatterjee 0002, Suprio Ray
IEEE Big Data2
2024 Query Compilation based Distributed Morsel-driven Parallel Spatial Query Processing
abstract
Driven by the need to support spatial data applications, most relational databases offer spatial SQL query features. However, traditional relational databases are not scalable, and their query processing follows a pull-based tuple-at-a-time model, which is not efficient for large data volumes. Although several specialized spatial data processing systems were developed by extending frameworks, such as Spark and Hadoop, these systems offer limited or no support for spatial SQL query execution. The few systems that support SQL querying, suffer from the overheads of the pull-based model.
Rahul Sahni, Xiaozheng Zhang 0005, Sudip Chatterjee 0002, Suprio Ray
SIGSPATIAL/GIS2