Yihua Wei

dblp:72/10325 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-9091-2446ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 DCSM: Enabling Inter-Batch Parallelism for Continuous Subgraph Matching on GPU
abstract
Continuous subgraph matching (CSM) is a fundamental building block in many real-world applications. While prior studies have explored executing CSM on heterogeneous systems with GPUs, they only exploit intra-batch parallelism and cannot process multiple batches concurrently—a capability essential for handling real-time requests. In this work, we propose a GPU-based system to accelerate CSM in practical, real-world settings. We adopt an algorithm-system co-design approach to unlock inter-batch parallelism. We introduce several key components, including a warp-specialized execution model and a multi-version graph data structure, along with version control logic for CSM tasks. Additionally, we propose optimizations such as warp-level parallel execution for data copying and incremental matching. Experimental results show that our system demonstrates optimal throughput and response time on GPU platforms across various update arrival rates.
Yihua Wei, Peng Jiang 0004
ICS1
2025 Matcha: A Language and Compiler for Backtracking-Based Subgraph Matching
abstract
Subgraph matching is one of the most fundamental tasks in graph analytics. Numerous algorithms and systems have been proposed for the task. However, due to the diverse optimizations proposed in previous work and their targeting of different hardware, comparing and integrating existing techniques has become increasingly challenging. In this work, we propose Matcha, a domain-specific language for implementing subgraph matching algorithms. Compared to previous systems, Matcha provides a lower-level programming interface that allows users to express a wider variety of subgraph matching algorithms. This simplifies the comparisons of existing techniques and facilitates the development of new algorithms. We implement a compiler that translates and optimizes Matcha programs into C++/CUDA code for CPU and GPU execution. Our experiments show that Matcha can readily reproduce the performance of state-of-the-art subgraph matching systems. By incorporating additional optimizations, Matcha can achieve speedups up to 60x against the existing systems.
Yihua Wei, Lihan Hu, Peng Jiang 0004
IPDPS1
2024 GCSM: GPU-Accelerated Continuous Subgraph Matching for Large Graphs
abstract
Continuous subgraph matching (CSM) is a key building block in many graph mining applications. Previous research has primarily focused on CSM algorithms on CPU. However, due to the dynamic nature of input graphs and the data movement bottlenecks, it has been challenging to efficiently run CSM on heterogeneous systems with GPUs. This work proposes the first system that exploits GPU to accelerate CSM. The main idea of our system is to cache the frequently accessed graph data in GPU memory so that most of the data communication between CPU and GPU can be avoided during the matching process. To identify the frequent data, we propose an efficient frequency estimation technique based on random walks on the input graph. We also provide an end-to-end design that achieves efficient dynamic graph updates on the CPU and efficient incremental matching on the GPU. The experiments show that our system significantly improves the performance of CSM compared to existing CPU solutions and supports CSM on extremely large graphs.
Yihua Wei, Peng Jiang 0004
IPDPS1
2022 SampleMine: A Framework for Applying Random Sampling to Subgraph Pattern Mining through Loop Perforation
abstract
Subgraph Pattern Mining (SPM) is an important class of graph applications that aim to discover structural patterns in a graph. Due to the enormous exploration space, SPM is in general computationally challenging. To accelerate SPM, many random sampling techniques have been proposed. While the existing sampling techniques are effective for conventional SPM tasks such as motif counting and frequent subgraph mining, they cannot be easily adapted for new applications.
Peng Jiang 0004, Yihua Wei, Jiya Su, Rujia Wang, Bo Wu 0002
PACT2
2022 STMatch: Accelerating Graph Pattern Matching on GPU with Stack-Based Loop Optimizations
abstract
Graph pattern matching is a fundamental task in many graph analytics and graph mining applications. As an NP-hard problem, it is often a performance bottleneck in these applications. Previous work has proposed to use GPU to accelerate the computation. However, we find that the existing GPU solutions fail to show a performance advantage over the state-of-the-art CPU implementation due to their subgraph-centric design. This work proposes a novel stack-based graph pattern matching system on GPU that avoids the synchronization and memory consumption issues of the previous subgraph-centric systems. We also propose a two-level work-stealing and a loop-unrolling technique to improve the inter-warp and intra-warp GPU resource utilization of our system. The experiments show that our system significantly advances the state-of-the-art for graph pattern matching on GPU.
Yihua Wei, Peng Jiang 0004
SC1