EDBT 2026 Demo / reviewers in the wild / expert
Shenghao Qiu
dblp:305/4387
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-6345-0306ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Tensor-Train Decomposition on Graph Neural NetworksabstractMemory footprint is a major concern when training graph neural networks (GNNs) on large graph data. Tensor-train decomposition (TTD) offers a potential solution by representing high-dimensional tensors with a set of smaller tensors, reducing memory overhead. However, existing TTD-based solutions for GNNs fail to reuse intermediate computation results and minimize memory data transfers to improve GNN performance. We introduce FALCON, a software framework to accelerate TTDbased GNN training. FALCON leverages the observation that a small subset of graph nodes with high edge degrees are frequently accessed, enabling the caching of intermediate results to reduce redundant computation and data transfers. Additionally, it incorporates multi-level graph partitioning and kernel optimization techniques to boost computational efficiency. We evaluated FALCON using three real-world datasets on three GPU platforms-NVIDIA 3090, 4090, and A100. Experimental results show that FALCON outperforms previous TTD-based frameworks, delivering a 1.3 to$8.17 \times$improvement in throughput while maintaining comparable or better efficiencies in memory footprint and model accuracy. Shenghao Qiu, Chunwei Xia, Zheng Wang 0001 |
IPDPS | 1 |
| 2024 | GraphCube: Interconnection Hierarchy-aware Graph ProcessingabstractProcessing large-scale graphs with billions to trillions of edges requires efficiently utilizing parallel systems. However, current graph processing engines do not scale well beyond a few tens of computing nodes because they are oblivious to the communication cost variations across the interconnection hierarchy. We introduce GraphCube, a better approach to optimizing graph processing on large-scale parallel systems with complex interconnections. GraphCube features a new graph partitioning approach to achieve better load balancing and minimize communication overhead across multiple levels of the interconnection hierarchy. We evaluate GraphCube by applying it to fundamental graph operations performed on synthetic and real-world graph datasets. Our evaluation used up to 79,024 computing nodes and 1.2+ million processor cores. Our large-scale experiments show that GraphCube outperforms state-of-the-art parallel graph processing methods in throughput and scalability. Furthermore, GraphCube outperformed the top-ranked systems on the Graph 500 list. Xinbiao Gan, Shenghao Qiu, Jiaqi Si, Jianbin Fang, Dezun Dong, Chunye Gong, Zheng Wang 0001 |
PPoPP | 3 |
| 2022 | AIACC-Training: Optimizing Distributed Deep Learning Training through Multi-streamed and Concurrent Gradient CommunicationsabstractThere is a growing interest in training deep neural networks (DNNs) in a GPU cloud environment. This is typically achieved by running parallel training workers on multiple GPUs across computing nodes. Under such a setup, the communication overhead is often responsible for long training time and poor scalability. This paper presents AIACC-Training, a unified communication framework designed for the distributed training of DNNs in a GPU cloud environment. AIACC-Training permits a training worker to participate in multiple gradient communication operations simultaneously to improve network bandwidth utilization and reduce communication latency. It employs auto-tuning techniques to dynamically determine the right communication parameters based on the input DNN workloads and the underlying network infrastructure. AIACC-Training has been deployed to production at Alibaba GPU Cloud with 3000+ GPUs executing AIACC-Training optimized code at any time. Experiments performed on representative DNN workloads show that AIACC-Training outperforms existing solutions, improving the training throughput and scalability by a large margin. Lixiang Lin, Shenghao Qiu, Liang You, Long Xin, Jie Xu 0007, Zheng Wang 0001 |
ICDCS | 2 |
| 2022 | STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model TrainingabstractDeep neural networks (DNNs) with billion-scale parameters have demonstrated impressive performance in solving many tasks. Unfortunately, training a billion-scale DNN is out of the reach of many data scientists because it requires high-performance GPU servers that are too expensive to purchase and maintain. We present STRONGHOLD, a novel approach for enabling large DNN model training with no change to the user code. STRONGHOLD scales up the largest trainable model size by dynamically offloading data to the CPU RAM and enabling the use of secondary storage. It automatically determines the minimum amount of data to be kept in the GPU memory to minimize GPU memory usage. Compared to state-of-the-art offloading-based solutions, STRONGHOLD improves the trainable model size by 1.9x~6. Sx on a 32GB V100 GPU, with 1.2x~3.7x improvement on the training throughput. It has been deployed into production to successfully support large-scale DNN training. Wei Wang 0225, Shenghao Qiu, Renyu Yang, Songfang Huang, Jie Xu 0007, Zheng Wang 0001 |
SC | 3 |