EDBT 2026 Demo / reviewers in the wild / expert
Junyi Mei
dblp:279/2556
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0008-1956-1242ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DGS: A GPU-based Adaptive Graph Sampling FrameworkabstractGraph sampling plays a critical role in graph learning applications, notably within Graph Neural Networks (GNNs). Typically, the performance of GPU-based graph sampling is determined by the efficiency of sampling kernels. Different sampling methods excel under different conditions, and no single method consistently outperforms others in all scenarios. As sampling applications become increasingly complex, graph-related sparse operations can dominate the computational workload, with performance heavily influenced by storage formats. In this article, we propose DGS, a GPU-based graph sampling framework that can detach the kernel implementation from computation logic. In addition to sampling kernels, DGS jointly optimizes sparse graph kernels. It can adaptively switch between different execution strategies based on various inputs. Experiments show that DGS outperforms current state-of-the-art GPU sampling frameworks, achieving speedups ranging from 1.1× to 92.0×. This adaptability and performance improvement establish DGS as a highly effective and efficient solution for diverse graph sampling scenarios. Junyi Mei, Shixuan Sun, Chao Li 0009, Xinkai Wang 0003, Xiaofeng Hou, Minyi Guo, Yongchao Liu 0004, Chuntao Hong |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | Enhancing High-Throughput GPU Random Walks Through Multi-Task Concurrency OrchestrationabstractRandom walk is a powerful tool for large-scale graph learning, but its high computational demand presents a challenge. While GPUs can accelerate random walk tasks, current frameworks fail to fully utilize GPU parallelism due to memory-to-compute bandwidth imbalance. In this article, CoWalker, an efficient GPU framework, is proposed to facilitate concurrent execution of random walks for high overall throughput. CoWalker features three novel designs. First, it incorporates a multi-level execution model that effectively orchestrates diverse walk tasks and reduces GPU stalls based on multiple graph characteristics. Second, it collaboratively manages graph data and streaming multiprocessors to minimize memory access interference and maximize core utilization under concurrent tasks. Finally, a multi-dimensional scheduler selects compatible random walk task combinations based on memory footprints to achieve maximum throughput. CoWalker significantly improves throughput over state-of-the-art baselines by mitigating concurrency overheads and effectively harnessing GPU parallelism. Our extensive evaluations on real-world workloads demonstrate that CoWalker achieves 2.75× higher overall system throughput compared with commercial tools and 1.56× over the SOTA academic system. Chao Li 0009, Xiaofeng Hou, Junyi Mei, Jing Wang 0055, Pengyu Wang 0003, Shixuan Sun, Minyi Guo, Baoping Hao |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk FrameworkabstractDynamic graph random walk (DGRW) emerges as a practical tool for capturing structural relations within a graph. Effectively executing DGRW on GPU presents certain challenges. First, existing sampling methods demand a pre-processing buffer, causing substantial space complexity. Moreover, the power-law distribution of graph vertex degrees introduces workload imbalance issues, rendering DGRW embarrassed to parallelize. In this paper, we propose FlowWalker, a GPU-based dynamic graph random walk framework. FlowWalker implements an efficient parallel sampling method to fully exploit the GPU parallelism and reduce space complexity. Moreover, it employs a sampler-centric paradigm alongside a dynamic scheduling strategy to handle the huge amounts of walking queries. FlowWalker stands as a memory-efficient framework that requires no auxiliary data structures in GPU global memory. We examine the performance of FlowWalker extensively on ten datasets, and experiment results show that FlowWalker achieves up to 752.2×, 72.1×, and 16.4× speedup compared with existing CPU, GPU, and FPGA random walk frameworks, respectively. Case study shows that FlowWalker diminishes random walk time from 35% to 3% in a pipeline of ByteDance friend recommendation GNN training. Junyi Mei, Shixuan Sun, Chao Li 0009, Cheng Chen 0008, Jing Wang 0055, Cheng Zhao 0001, Xiaofeng Hou, Minyi Guo, Bingsheng He, Xiaoliang Cong |
Proc. VLDB Endow. | 1 |
| 2023 | Fargraph+: Excavating the parallelism of graph processing workload on RDMA-based far memory system
Jing Wang 0055, Chao Li 0009, Taolei Wang, Junyi Mei, Lu Zhang 0049, Pengyu Wang 0003, Minyi Guo |
J. Parallel Distributed Comput. | 5 |
| 2022 | HyFarM: Task Orchestration on Hybrid Far Memory for High Performance Per BitabstractTapping into secondary memory resources, i.e., far memory (FM), has shown huge potential to improve the cost-efficiency of data centers. Recent advances in both storage-based vertical FM and network-based horizontal FM have raised new questions about leveraging hybrid FM tiers to achieve the best performance per bit of memory. It is still unclear how to efficiently place tasks when far memory access is enabled.In this work, we propose HyFarM, a novel task management strategy for hybrid FM clusters. We analyze FM sensitivity and cooperatively co-locate tasks to enable high utilization and scalability. Further, by tapping into dynamic memory adaption within and across servers, our strategy allows one to consistently deliver high performance on memory-intensive tasks. We evaluate our design with a heavily instrumented testbench. Compared with the state-of-the-art designs, HyFarM respectively improves memory utilization and the overall performance per bit (PPB) by up to 17.6% and 20.5%, with minor overhead. Jing Wang 0055, Chao Li 0009, Junyi Mei, Taolei Wang, Pengyu Wang 0003, Lu Zhang 0049, Minyi Guo, Dongbai Chen, Xiangwen Liu |
ICCD | 3 |
| 2022 | Excavating the Potential of Graph Workload on RDMA-based Far Memory ArchitectureabstractDisaggregated architecture brings new opportunities to memory -consuming applications like graph processing. It allows one to outspread memory access pressure from local to far memory, providing an attractive alternative to disk-based processing. Although existing works on general-purpose far mem-ory platforms show great potentials for application expansion, it is unclear how graph processing applications could benefit from disaggregated architecture, and how different optimization methods influence the overall performance. In this paper, we take the first step to analyze the impact of graph processing workload on disaggregated architecture by extending the GridGraph framework on top of the RDMA-based far memory system. We design Fargraph, a far memory coordi-nation strategy for enhancing graph processing workload. Specif-ically, Fargraph reduces the overall data movement through a well-crafted, graph-aware data segment offloading mechanism. In addition, we use optimal data segment splitting and asynchronous data buffering to achieve graph iteration-friendly far memory access. We show that Fargraph achieves near-oracle performance for typical in-local-memory graph processing systems. Fargraph shows up to 8.3 x speedup compared to Fastswap, the state-of-the-art, general-purpose far memory platform. Jing Wang 0055, Chao Li 0009, Taolei Wang, Lu Zhang 0049, Pengyu Wang 0003, Junyi Mei, Minyi Guo |
IPDPS | 6 |