VLDB 2026 Research / reviewers in the wild / expert
Zhe Pan 0001
dblp:232/1625-1
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-0355-8099ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpiderFlow: Efficient Topology-Aware Scheduling for LLM Training Across Decentralized GPU ClustersabstractIn response to increasing demands for largescale machine learning training jobs, many organizations have deployed GPU clusters across geographically distributed regions.However, existing Integer Linear Programming (ILP)-or genetic-based cross-cluster training approaches largely overlook the topology of decentralized clusters, lacking both topology-aware task scheduling mechanisms and automated model parallelization strategies.As a result, naively applying these optimization-based methods in cross-cluster settings leads to prohibitive scheduling overhead, due to the drastically enlarged search space induced by complex inter-cluster topologies.To address these challenges, we propose SpiderFlow, a topologyaware scheduling system specifically designed for decentralized GPU clusters.We formulate cross-cluster task scheduling as a graph optimization problem and introduce SpinSearch, a low-overhead topology-aware scheduling algorithm.In addition, for automated model parallelization, we propose Topology-aware Parallelism Automation (TPA), a two-level scheduling framework that combines heuristic methods at the inter-cluster level with ILP-based optimization within clusters, effectively reducing the search space while maintaining high training throughput with substantially lower scheduling overhead.We evaluate SpiderFlow on a physical platform comprising 8 decentralized clusters, as well as on a simulation platform with up to 64 decentralized clusters.Experimental results demonstrate that SpiderFlow reduces job completion time (JCT) by 1.2-1.3×,improves throughput by 1.12-1.25×,and reduces scheduling overhead by 20-90× on average compared to state-of-the-art scheduling systems. Zihan Chang, Shuibing He, Sheng Xiao, Siling Yang, Rui Wang 0076, Zhe Pan 0001 |
ACL (1) | 7 |
| 2026 | Look Before You Leap : Precision Instruction Supply via SmartScoutabstractModern high-performance processors extensively employ Fetch-Directed Instruction Prefetching (FDIP) to mitigate instruction supply bottlenecks. However, the efficacy of FDIP is fundamentally constrained by the accuracy of the Branch Prediction Unit (BPU). As the critical component within the BPU, the Branch Target Buffer (BTB) faces severe capacity bottlenecks. While pre-decoding-based prefetching offers a remedy, existing approaches suffer from two critical impediments: (1) The Noise Dilemma: Suboptimal trade-off between coverage and accuracy. (2) Inefficient Miss Resolution: Current designs rely on reactive recovery or stalls, failing to leverage available front-end slack for proactive correction. Peng Qu 0001, Tingji Zhang, Fang Su, Zhe Pan 0001, Youhui Zhang |
ICS | 5 |
| 2026 | VDHA: Vector-Driven Hash Aggregation for Sparse Matrix-Sparse Vector Multiplication on GPUsabstractSparse matrix-sparse vector multiplication (SpMSpV) is a core primitive in graph analytics and scientific computing, also arising in spiking neural networks for event-driven spike propagation. On GPUs, the performance of the prevalent and efficient SpMSpV paradigm is often bottlenecked by the write-back phase of accumulating non-zero multiply–accumulate results; its many-to-one index scatter pattern causes severe conflicts and poor bandwidth utilization on GPUs. We present VDHA, a GPU-based weighted SpMSpV kernel that leverages block-private hash tables for local aggregation, substantially reducing write conflicts and improving memory coalescing. To further amplify this benefit, we incorporate column splitting with lightweight reordering to expose more locality, and employ a fetch–compute–writeback pipeline to overlap hash computation with memory accesses. Extensive evaluation on over 300 matrices with more than 5 million nonzeros, including web-scale graphs (Konect/LAW) and scientific workloads (SuiteSparse), shows that VDHA consistently outperforms state-of-the-art baselines. On web graphs, it achieves a 1.41× geometric-mean speedup (up to 3.42×), while on SuiteSparse it delivers 1.13× (up to 2.55×). We also provide a lightweight predictive model that identifies matrices favorable to VDHA with 91.3% accuracy. Zhe Pan 0001, Peng Qu 0001, Youhui Zhang |
PPoPP | 2 |
| 2026 | Root-Down Exposure for Maximal Clique Enumeration on GPUsabstractMaximal clique enumeration (MCE) in large-scale graphs is critical across various application domains, including social network analysis, bioinformatics, and computer vision. However, existing GPU-based MCE solutions suffer from inefficient load-balancing mechanisms. These mechanisms force busy workers to pause and hand over workloads to idle workers, which introduces significant synchronization overhead and increases memory usage. To address these limitations, we introduce a root-down exposure mechanism where busy workers dynamically expose their current root, enabling idle workers to pull workloads from the exposed node directly without synchronization. We then propose a bitmap-centric MCE scheme and an aggressive node generation rule to further simplify the memory layout and accelerate enumeration. We combine them into RDMCE, a Root-Down MCE solution on GPUs. Across large real-world graphs with up to 146 billion maximal cliques, RDMCE is 1.25-5.38× faster than any next-best state-of-the-art GPU-based solutions and is the only one that completes enumeration on every test dataset, offering a more efficient and scalable MCE solution. Zhe Pan 0001, Peng Qu 0001, Youhui Zhang |
PPoPP | 1 |
| 2025 | Advanced Maximal Biclique Enumeration on GPUs Using BitmapsabstractMaximal biclique enumeration (MBE) in bipartite graphs is an important problem in data mining with many real-world applications. Parallel MBE algorithms for GPUs are needed for MBE acceleration leveraging its many computing cores. However, enumerating maximal bicliques using GPUs has three main challenges including large memory requirement, thread divergence, and load imbalance. In this paper, we propose GMBE+, an advanced GPU solution for the MBE problem. To overcome the challenges, we design (1) a node-reuse approach to reduce GPU memory usage with advanced node pruning, (2) a bitmap-based set intersection approach to minimize thread divergence, and (3) a load-aware task scheduling framework to achieve load balance among threads within GPU warps, facilitated by a novel set union approach. Our experiments reveal that GMBE+ is 1.2× faster than the latest GPU-based MBE algorithm GMBE on average when running on the same NVIDIA A100 GPU. Zhe Pan 0001, Shuibing He, Xu Li 0026, Xuechen Zhang 0001, Rui Wang 0076, Yanlong Yin, Gang Chen 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Enumeration of Billions of Maximal Bicliques in Bipartite Graphs without Using GPUsabstractMaximal biclique enumeration (MBE) is crucial in bipartite graph analysis. Recent studies rely on extensive set intersections on static bipartite graphs to solve the MBE problem. However, the computational subgraphs dynamically change during enumeration, leading to redundant memory accesses and degraded set intersection performance. To overcome this limitation, we propose an AdaMBE algorithm. First, we redesign its core operations using local neighborhood information derived from computational subgraphs to minimize redundant memory accesses. Second, we dynamically create computational subgraphs using bitmaps leveraging its fast bitwise operations to accelerate set intersections. Finally, we integrate them in AdaMBE. Our experimental results show that AdaMBE is $1.6 \times-49.7 \times$ faster than its closest CPU-based competitor and successfully enumerates all 19 billion maximal bicliques on the TVTropes dataset, a large task beyond the capabilities of existing algorithms. Notably, on certain datasets, our parallel version, ParAdaMBE, on CPUs even outperforms GMBE on GPUs by up to $5.07 \times$. Zhe Pan 0001, Shuibing He, Xu Li 0026, Xuechen Zhang 0001, Yanlong Yin, Rui Wang 0076, Lidan Shou, Mingli Song, Xian-He Sun, Gang Chen 0001 |
SC | 1 |
| 2024 | AMBEA: Aggressive Maximal Biclique Enumeration in Large Bipartite Graph ComputingabstractMaximal biclique enumeration (MBE) in bipartite graphs is a fundamental problem in data mining with widespread applications. Many recent works solve this problem based on the set-enumeration (SE) tree, which sequentially traverses vertices to generate the enumeration tree nodes representing distinct bicliques, then checks whether these bicliques are maximal or not. However, existing MBE algorithms only expand bicliques with untraversed vertices to ensure distinction, which often necessitate extensive node checks to eliminate non-maximal bicliques, resulting in significant computational overhead during the enumeration process. To address this issue, we propose an aggressive set-enumeration (ASE) tree that aggressively expands all bicliques to their maximal form, thus avoiding costly node checks on non-maximal bicliques. This aggressive enumeration may produce multiple duplicate maximal bicliques, but we efficiently eliminate these duplicates by leveraging the connection between parent and child nodes and conducting low-cost node checking. Additionally, we introduce an aggressive merge-based pruning (AMP) approach that aggressively merges vertices sharing the same local neighbors. This helps prune numerous duplicate node generations caused by subsets of merged vertices. We integrate the AMP approach into the ASE tree, and present the Aggressive Maximal Biclique Enumeration Algorithm (AMBEA). Experimental results show that AMBEA is 1.15$\times$to 5.32$\times$faster than its closest competitor and exhibits better scalability and parallelization capabilities on larger bipartite graphs. Zhe Pan 0001, Xu Li 0026, Shuibing He, Xuechen Zhang 0001, Rui Wang 0076, Yunjun Gao, Gang Chen 0001, Xian-He Sun |
IEEE Trans. Computers | 1 |
| 2023 | Efficient Maximal Biclique Enumeration on GPUsabstractMaximal biclique enumeration (MBE) in bipartite graphs is an important problem in data mining with many real-world applications. All existing solutions for MBE are designed for CPUs. Parallel MBE algorithms for GPUs are needed for MBE acceleration leveraging its many computing cores. However, enumerating maximal bicliques using GPUs has three main challenges including large memory requirement, thread divergence, and load imbalance. In this paper, we propose GMBE, the first highly-efficient GPU solution for the MBE problem. To overcome the challenges, we design a node-reuse approach to reduce GPU memory usage, a pro-active pruning method using the vertex's local neighborhood size to alleviate thread divergence, and a load-aware task scheduling framework to achieve load balance among threads within GPU warps and blocks. Our experimental results show that GMBE on an NVIDIA A100 GPU can achieve 70.6× speedup over the state-of-the-art parallel MBE algorithm ParMBE on a 96-core CPU machine. Zhe Pan 0001, Shuibing He, Xu Li 0026, Xuechen Zhang 0001, Rui Wang 0076, Gang Chen 0001 |
SC | 1 |
| 2020 | An FPGA-Optimized Architecture of Real-time Farneback Optical FlowabstractOptical flow estimation is a fundamental tool for computer vision applications. As a classical optical flow algorithm, Farneback version was a good blend of accuracy and runtime performance for a long time. This work presents a dataflow-based architecture of Farneback optical flow with high level synthesis (HLS) tools. Multi-dimension separable convolution, Block RAM array and deep pipeline optimization techniques are applied to reduce both space and time. The system is implemented on XC7K325T FPGA with an image size of $640 \times 480.$ It supports complete Farneback algorithm for any user-defined image pyramid levels and iteration times. It can process 31 fps, 8x faster than OpenCV version with 3 pyramid levels and 1 iteration time. Zhe Pan 0001, Yuruo Jin, Xiaohong Jiang 0002 |
FCCM | 1 |
| 2019 | Hybrid XML Parser Based on Software and Hardware Co-designabstractExtensible Markup Language (XML) is widely used in web services. However, the task of XML parsing is always the bottleneck which consumes a lot of time and resources. In this work, we present a hybrid XML parser based on software and hardware co-design. We place hardware acceleration into a software-driven context. Our parser is based on document object model (DOM). It is capable of well-formed checking and tree construction at throughput of 1 cycle per byte (CPB). We implement the design on a Xilinx Kintex-7 FPGA with 0.8Gbps parsing throughput. Zhe Pan 0001, Xiaohong Jiang 0002, Xiang Li 0017 |
FCCM | 1 |