EDBT 2026 Demo / reviewers in the wild / expert
Hao Qi 0008
dblp:24/8493-8
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0007-8795-5262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | eGPU: Production-Scale Elastic Sharing Over 10,000 GPUsabstractAs the cost of GPUs continues to rise, GPU-sharing solutions have become increasingly important for improving efficiency and maximizing resource utilization. At the same time, large-scale operational deployments of such solutions remain relatively less explored, especially in heterogeneous production environments where workload dynamics and orchestration complexity introduce new practical considerations. In this paper, we introduce eGPU, an elastic, efficient, and scalable GPU-sharing framework tailored for production-scale concurrent machine learning (ML) training and inference. eGPU enables fine-grained, runtime-adjustable sharing of GPUs across multiple jobs, while preserving high resource utilization and fault isolation. To address communication bottlenecks, eGPU supports native NVLink/NCCL-based communication between shared GPU instances, capabilities that are limited or unavailable in many existing designs. Built with production deployment in mind, eGPU integrates with Kubernetes (K8s) to support large-scale orchestration. It has been deployed and running stably in production clusters with over$\text{1 0, 0 0 0 ~ G P U s}$for five years. Our evaluation results show that eGPU achieves elastic and precise control over instance sizes, improves job efficiency by 21 % to 31% than SOTA sharing solutions, saves the number of GPUs required by up to$8 \times$, and improves cluster GPU utilization by more than$\mathrm{3} \times$. Xiaochuan Tang, Hao Qi 0008, Jianbo Dong, Yinghao Yu, Zhennan Xue, Daocheng Ying, Zheng Cao 0003, Xiaoyi Lu 0001 |
HPCA | 2 |
| 2025 | SBMGT: Scaling Bayesian Multinomial Group TestingabstractGroup testing is a widely used binary classification method that efficiently distinguishes between samples with and without a binary-classifiable attribute by pooling and testing subsets of a group. Bayesian Group Testing (BGT) is the state-of-the-art approach, which integrates prior risk information into a Bayesian Boolean Lattice framework to minimize test counts and reduce false classifications. However, BGT, like other existing group testing techniques, struggles with multinomial group testing, where samples have multiple binary-classifiable attributes that can be individually distinguished simultaneously. We address this need by proposing Bayesian Multinomial Group Testing (BMGT), which includes a new Bayesian-based model and supporting theorems for an efficient and precise multinomial pooling strategy. We further design and develop SBMGT, a high-performance and scalable framework to tackle BMGT's computational challenges by proposing three key innovations: 1) a parallel binary-encoded product lattice model with up to 99.8% efficiency; 2) the Bayesian Balanced Partitioning Algorithm (BBPA), a multinomial pooling strategy optimized for parallel computation with up to 97.7% scaling efficiency on 4096 cores; and 3) a scalable multinomial group testing analytics framework, demonstrated in a real-world disease surveillance case study using AIDS and STDs datasets from Uganda, where SBMGT reduced tests by up to 54% and lowered false classification rates by 92% compared to BGT. Weicong Chen 0002, Hao Qi 0008, Curtis Tatsuoka, Xiaoyi Lu 0001 |
PPoPP | 2 |
| 2025 | DPAR: High-Performance, Secure, and Scalable Differential Privacy-based AllReduceabstractSecure, efficient, and scalable AllReduce-based data aggregation is essential for Artificial Intelligence (AI) and scientific applications on modern High-Performance Computing (HPC) and cloud infrastructures. As AllReduce is increasingly used across these distributed infrastructures, privacy has become a critical concern. State-of-the-art (SOTA) Homomorphic Encryption (HE)-based AllReduce solutions introduce high overhead, require secure key exchanges, and remain vulnerable to collusion. We propose DPAR, the first differentially private, collusion-resistant AllReduce framework optimized for large-scale HPC and AI workloads. DPAR introduces three key innovations: integrating Differential Privacy (DP) to eliminate collusion risks without key exchanges, scalable noise growth to preserve accuracy, and performance optimizations using a noise pooling mechanism. DPAR is a drop-in Message Passing Interface (MPI) AllReduce replacement, providing strong privacy with minimal performance cost. Evaluated on Delta and Frontier supercomputers with up to 8192 cores, DPAR outperforms the SOTA HE solution by up to 34.7% in modern AI workloads. Hao Qi 0008, Weicong Chen 0002, Chenghong Wang, Xiaoyi Lu 0001 |
SC | 1 |
| 2025 | HPC-R1: Characterizing R1-like Large Reasoning Models on HPCabstractLarge Reasoning Models (LRMs) are becoming increasingly popular as they offer advanced capabilities in logical inference, mathematical reasoning, and knowledge synthesis, even beyond those of standard language models. However, their complex training workflows present significant challenges in reproducibility, efficiency, and system-level optimization. This paper introduces HPC-R1, a comprehensive characterization of LRM training on the NERSC Perlmutter supercomputer, representing behavior on a Top500-ranked system. We analyze all major stages, including supervised fine-tuning (SFT), Group Relative Policy Optimization (GRPO)-based reinforcement learning (RL), autoregressive generation, and distillation using customized state-of-the-art frameworks. Our detailed performance analysis reveals key system inefficiencies and scaling behaviors. Through our in-depth analysis, we present 19 key observations across all stages, including 4 for SFT, 7 for GRPO-based RL, 6 for generation, and 2 for distillation. Based on these findings, we present several key recommendations to guide future HPC-AI system design. Adam Weingram, Zhonghao Chen, Hao Qi 0008, Xiaoyi Lu 0001 |
SC | 4 |
| 2024 | Kspeed: Beating I/O Bottlenecks of Data Provisioning for RDMA Training ClustersabstractThe rapidly-increasing computing power of GPUs has rendered the I/O subsystem a bottleneck for distributed deep learning (DL) training. Currently, substantial data preprocessing work (e.g., decoding) has to be conducted on CPUs for a wide range of training scenarios such as computer vision (CV) and audio. Unfortunately, the involvement of training nodes' host memory and/or CPUs on the critical path of loading data to GPUs incurs significant GPU stalls in modern RDMA training clusters, because CPUs are much slower than GPUs and the connection from PCIe switches to host memory tends to suffer from incast problems. Moreover, this also incurs high CPU usage and resource contention, which consequently causes data loading performance variation and stragglers. This paper presents KSpeed, a novel data provisioning framework for large-scale RDMA training clusters. As many data preprocessing tasks need to be done by CPUs, KSpeed organizes host memory and CPU resources in the cluster to build a disaggregated memory/CPU pool, where the nodes can read raw input data from backend storage to their host memory, preprocess the data by their CPUs if necessary, and write cached/preprocessed data (on demand) directly to the training workers' GPU memory to minimize GPU stalls. KSpeed leverages the multi-rail RDMA network to eliminate unnecessary memory copies, interference, and congestion. Evaluation on a 96-GPU cluster shows that KSpeed delivers$5.4 \times \sim 100 \times$higher data loading performance over the state-of-the-art designs (DPP and Alluxio). KSpeed achieves near-linear scalability as the GPU number increases from 8 to 512. Jianbo Dong, Hao Qi 0008, Tianjing Xu, Xiaoli Liu 0002, Rongyao Wang, Xiaoyi Lu 0001, Zheng Cao 0003, Binzhang Fu |
ICNP | 2 |
| 2024 | On the Feasibility and Benefits of Extensive EvaluationabstractBenchmark and system parameters often have a significant impact on performance evaluation, which raises a long-lasting question about which settings we should use. This paper studies the feasibility and benefits of extensive evaluation. A full extensive evaluation, which tests all possible settings, is usually too expensive. This work investigates whether it is possible to sample a subset of the settings and, upon them, generate observations that match those from a full extensive evaluation. Towards this goal, we have explored the incremental sampling approach, which starts by measuring a small subset of random settings, builds a prediction model on these samples using the popular ANOVA approach, adds more samples if the model is not accurate enough, and terminates otherwise. To summarize our findings: 1) Enhancing a research prototype to support extensive evaluation mostly involves changing hard-coded configurations, which does not take much effort. 2) Some systems are highly predictable, which means that they can achieve accurate predictions with a low sampling rate, but some systems are less predictable. 3) We have not found a method that can consistently outperform random sampling + ANOVA. Based on these findings, we provide recommendations to improve artifact predictability and strategies for selecting parameter values during evaluation. Yujie Hui, Miao Yu 0023, Hao Qi 0008, Yifan Gan, Tianxi Li, Yuke Li 0003, Xueyuan Ren, Sixiang Ma, Xiaoyi Lu 0001, Yang Wang 0009 |
Proc. ACM Manag. Data | 3 |
| 2023 | Performance Characterization of Large Language Models on High-Speed InterconnectsabstractLarge Language Models (LLMs) have recently gained significant popularity due to their ability to generate human-like text and perform a wide range of natural language processing tasks. Training these models usually requires a large amount of computational resources and is often done in a distributed manner. The use of high-speed interconnects can significantly influence the efficiency of distributed training. Therefore, there poses a need for systematic studies to explore the distributed training characteristics of these models on high-speed interconnects. This paper presents a comprehensive performance characterization of representative large language models: GPT, BERT, and T5. We evaluate their training performance in terms of iteration time, interconnect utilization, and scalability, over different high-speed interconnects and communication protocols, including TCP/IP, IPoIB, and RDMA. We observe that interconnects play a vital role in LLM training. Specifically, RDMA-100 Gbps outperforms IPoIB-100 Gbps and TCP/IP-10 Gbps by an average of 2.51x and 4.79x regarding training iteration time, and scores the highest interconnect utilization (up to 60 Gbps) in both strong and weak scaling, compared to IPoIB with up to 20 Gbps and TCP/IP with up to 9 Gbps, leading to the shortest training time. We also observe that larger models tend to have higher requirements for communication bandwidth, especially for AllReduce during backward propagation, which can take up to 91.12% of training time. Through our evaluation, we envision opportunities to improve the communication time for better training performance of LLMs. We extensively explore and summarize the role communication plays in distributed LLM training. Hao Qi 0008, Liuyao Dai, Weicong Chen 0002, Zhen Jia 0001, Xiaoyi Lu 0001 |
HOTI | 1 |
| 2023 | SBGT: Scaling Bayesian-based Group Testing for Disease SurveillanceabstractThe COVID-19 pandemic underscored the necessity for disease surveillance using group testing. Novel Bayesian methods using lattice models were proposed, which offer substantial improvements in group testing efficiency by precisely quantifying uncertainty in diagnoses, acknowledging varying individual risk and dilution effects, and guiding optimally convergent sequential pooled test selections using a Bayesian Halving Algorithm. Computationally, however, Bayesian group testing poses considerable challenges as computational complexity grows exponentially with sample size. This can lead to shortcomings in reaching a desirable scale without practical limitations. We propose a new framework for scaling Bayesian group testing based on Spark: SBGT. We show that SBGT is lightning fast and highly scalable. In particular, SBGT is up to 376x, 1733x, and 1523x faster than the state-of-the-art framework in manipulating lattice models, performing test selections, and conducting statistical analyses, respectively, while achieving up to 97.9% scaling efficiency up to 4096 CPU cores. More importantly, SBGT fulfills our mission towards reaching applicable scale for guiding pooling decisions in wide-scale disease surveillance, and other large scale group testing applications. Weicong Chen 0002, Hao Qi 0008, Xiaoyi Lu 0001, Curtis Tatsuoka |
IPDPS | 2 |
| 2023 | xCCL: A Survey of Industry-Led Collective Communication Libraries for Deep Learning
Adam Weingram, Yuke Li 0003, Hao Qi 0008, Darren Ng, Liuyao Dai, Xiaoyi Lu 0001 |
J. Comput. Sci. Technol. | 3 |