EDBT 2026 Demo / reviewers in the wild / expert
Hazem Zaky
dblp:435/3954
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 62% GPUs and heterogeneous computing · 38% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
performance variability |
1.0 | 1 | 2026 | Quantifying Performance Variability in GPU Clusters · IEEE Trans. Parallel Distributed Syst. 2026 |
GPUs and heterogeneous computing
GPU-accelerated scientific computing |
0.3 | 1 | 2026 | Quantifying Performance Variability in GPU Clusters · IEEE Trans. Parallel Distributed Syst. 2026 |
GPUs and heterogeneous computing
GPU performance analysis |
0.3 | 1 | 2026 | Quantifying Performance Variability in GPU Clusters · IEEE Trans. Parallel Distributed Syst. 2026 |
Methods — techniques the papers use, named apart from their topics
microbenchmarking · 1.0characterization study · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantifying Performance Variability in GPU ClustersabstractModern supercomputers are equipped with massive amounts of Graphics Processing Units (GPUs) to meet the growing demands of scientific computing and machine learning. However, GPUs, even of the same type, exhibit performance variability, leading to resource under-utilization and prolonged execution times. In this work, we perform a characterization study on the performance variability of NVIDIA A100 GPUs and GH200 superchips using the GEMM and STREAM micro-benchmarks and seven real-world applications across two supercomputing systems: TACC Vista (GH200) and NERSC Perlmutter (A100). Our results reveal GEMM performance variability ranging from 0.1% to 8.8%, though outliers can push deviations significantly higher. For scientific applications, the variability ranges from 1.0% to 12.2% on Vista. The variability of single-GPU GPT 4.8B and Llama 3B training is 10.4% and 8.9%, respectively. We further extend this study by training a GPT 19B model using eight GH200s in the 3D parallel configuration, observing up to 3.6% throughput variability and 8.5% slowdown in the worst case. By comparing the variability across GPU types, compute paths, data precisions, and applications, we observe that applications that use Tensor Cores and FP64 on CUDA cores show higher variability than other applications. Supercomputer users can leverage the observed variability to diagnose performance issues and enhance execution efficiency. Computing centers can design variability-aware scheduling approaches to achieve higher machine utilization without sacrificing individual application performance. Michael Mogilevsky, Hazem Zaky, Mingkai Zheng, Lishan Yang 0001, Zhao Zhang 0007 |
IEEE Trans. Parallel Distributed Syst. | 2 |