Weiwei Gong

dblp:30/841 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Rethinking Analytical Processing in the GPU Era
Bobbi W. Yogatama, Kevin Kristensen, Devesh Sarda, Abigale Kim, Adrian Cockcroft, Yu Teng, Joshua Patterson, Gregory Kimball, Wes McKinney, Weiwei Gong, Xiangyao Yu
CIDR11
2024 Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs
abstract
GPU-accelerated databases have been gaining popularity in recent years due to their massive parallelism and high memory bandwidth. The limited GPU memory capacity, however, is still a major bottleneck for GPU databases. Existing approaches have attempted to address this limitation by using (1) hybrid CPU-GPU DBMS or (2) multi-GPU DBMS. We aim to improve prior solutions further by leveraging both hybrid CPU-GPU DBMS and multi-GPU DBMS at the same time. In particular, we explore the design space and optimize the data placement and query execution in hybrid CPU and multi-GPU DBMS. To improve data placement, we introduce the cache-aware replication policy which takes into account the cost of shuffle when replicating data and could coordinate both caching and replication decisions for the best performance. To improve query execution, we extend the existing hybrid CPU-GPU query execution strategy with distributed query processing techniques to support multiple GPUs. We build a system called Lancelot , a hybrid CPU and Multi-GPU data analytics engine with all the optimizations integrated. Our evaluation shows that the cache-aware replication outperforms other policies by up to 2.5× and Lancelot outperforms existing GPU DBMSes by at least 2× on Star Schema Benchmark and 12× on TPC-H Benchmark.
Bobbi W. Yogatama, Weiwei Gong, Xiangyao Yu
Proc. VLDB Endow.2
2022 Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS
abstract
There has been a growing interest in using GPU to accelerate data analytics due to its massive parallelism and high memory bandwidth. The main constraint of using GPU for data analytics is the limited capacity of GPU memory. Heterogeneous CPU-GPU query execution is a compelling approach to mitigate the limited GPU memory capacity and PCIe bandwidth. However, the design space of heterogeneous CPU-GPU query execution has not been fully explored. We aim to improve state-of-the-art CPU-GPU data analytics engine by optimizing data placement and heterogeneous query execution. First, we introduce a semantic-aware fine-grained caching policy which takes into account various aspects of the workload such as query semantics, data correlation, and query frequency when determining data placement between CPU and GPU. Second, we introduce a heterogeneous query executor which can fully exploit data in both CPU and GPU and coordinate query execution at a fine granularity. We integrate both solutions in Mordred, our novel hybrid CPU-GPU data analytics engine. Evaluation on the Star Schema Benchmark shows that the semantic-aware caching policy can outperform the best traditional caching policy by up to 3x. Compared to existing GPU DBMSs, Mordred can outperform by an order of magnitude.
Bobbi W. Yogatama, Weiwei Gong, Xiangyao Yu
Proc. VLDB Endow.2
2019 A Morsel-Driven Query Execution Engine for Heterogeneous Multi-Cores
abstract
Currently, we face the next major shift in processor designs that arose from the physical limitations known as the "dark silicon effect". Due to thermal limitations and shrinking transistor sizes, multi-core scaling is coming to an end. A major new direction that hardware vendors are currently investigating involves specialized and energy-efficient hardware accelerators (e.g., ASICs) placed on the same die as the normal CPU cores. In this paper, we present a novel query processing engine called SiliconDB that targets such heterogeneous processor environments. We leverage the Sparc M7 platform to develop and test our ideas. Based on the SSB benchmarks, as well as other micro benchmarks, we compare the efficiency of SiliconDB with existing execution strategies that make use of co-processors (e.g., FPGAs, GPUs) and demonstrate speed-up improvements of up to 2x.
Kayhan Dursun, Carsten Binnig, Ugur Çetintemel, Garret Swart, Weiwei Gong
Proc. VLDB Endow.5
2014 Improving MMDB distributed transactional concurrency
abstract
Main Memory Database Systems (MMDBs) have been studied since the 80s [3,4], when memory was quite costly ($1500 per MByte in 1984). We can now buy memory for about $10 per GByte. An advantage of MMDBs is that serial execution of a non-distributed transaction on a uniprocessor from start to finish saves the work of disk I/O, locking, latching and deadlock handling [7]. The 2013 Bulletin on Data Engineering [11] had eight articles on recent MMDBs and only three mentioned distributed transactions. Implementing fast, serializable, distributed transactions on an MMDB is difficult, since communication delays typically leave some CPUs idle and reduce total throughput.
Weiwei Gong, Patrick E. O'Neil, Elizabeth J. O'Neil
IDEAS1