EDBT 2026 Demo / reviewers in the wild / expert
Ziheng Chen 0006
dblp:246/8429-6 · also Ziheng (Jack) Chen
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0006-2335-8656ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware reliability and fault tolerance · 61% GPUs and heterogeneous computing · 30% High-performance computing · 9% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware reliability and fault tolerance
error recovery |
0.9 | 1 | 2025 | Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs · SC 2025 |
GPUs and heterogeneous computing
GPU reliability |
0.9 | 1 | 2025 | Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs · SC 2025 |
Hardware reliability and fault tolerance › memory reliability
memory error characterization |
0.9 | 1 | 2025 | Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs · SC 2025 |
Methods — techniques the papers use, named apart from their topics
operational data analysis · 0.9failure projection · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsabstractThis study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures. Shengkun Cui, Archit Patke, Aditya Ranjan, Ziheng Chen 0006, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandrasekhar Narayanaswami 0001, Daby M. Sow, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 5 |
| 2024 | iPrism: Characterize and Mitigate Risk by Quantifying Change in Escape Routes
Shengkun Cui, Saurabh Jha, Ziheng Chen 0006, Zbigniew T. Kalbarczvk, Ravishankar K. Iyer |
DSN | 3 |