Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ziheng Chen 0006

dblp:246/8429-6 · also Ziheng (Jack) Chen · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0006-2335-8656ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware reliability and fault tolerance · 61% GPUs and heterogeneous computing · 30% High-performance computing · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance
error recovery
0.912025
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs · SC 2025
GPUs and heterogeneous computing
GPU reliability
0.912025
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs · SC 2025
Hardware reliability and fault tolerance › memory reliability
memory error characterization
0.912025
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs · SC 2025

Methods — techniques the papers use, named apart from their topics

operational data analysis · 0.9failure projection · 0.9
YearPublicationVenuePosition
2025 Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
abstract
This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures.
Shengkun Cui, Archit Patke, Aditya Ranjan, Ziheng Chen 0006, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandrasekhar Narayanaswami 0001, Daby M. Sow, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SC5
2024 iPrism: Characterize and Mitigate Risk by Quantifying Change in Escape Routes
Shengkun Cui, Saurabh Jha, Ziheng Chen 0006, Zbigniew T. Kalbarczvk, Ravishankar K. Iyer
DSN3