Jiahao Lu 0003

dblp:237/8931-3 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0003-3790-6519ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware reliability and fault tolerance · 47% Memory systems · 35% Distributed systems · 17%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
DRAM
1.122026
Removing Obstacles before Breaking Through the Memory Wall: A Close Look at HBM Errors in the Field · USENIX ATC 2024
Looking Back to Move Forward: Unveiling the Mysteries of HBM Errors to Predict Future Failures · ACM Trans. Storage 2026
Distributed systems › fault tolerance › proactive fault tolerance
failure prediction
1.012026
Looking Back to Move Forward: Unveiling the Mysteries of HBM Errors to Predict Future Failures · ACM Trans. Storage 2026
Memory systems › DRAM › DRAM architecture
high bandwidth memory
1.012026
Looking Back to Move Forward: Unveiling the Mysteries of HBM Errors to Predict Future Failures · ACM Trans. Storage 2026
Hardware reliability and fault tolerance › memory reliability
memory error characterization
1.012026
Looking Back to Move Forward: Unveiling the Mysteries of HBM Errors to Predict Future Failures · ACM Trans. Storage 2026
Hardware reliability and fault tolerance › memory reliability
memory failure prediction
1.012026
Looking Back to Move Forward: Unveiling the Mysteries of HBM Errors to Predict Future Failures · ACM Trans. Storage 2026
Hardware reliability and fault tolerance › memory reliability
memory errors
0.812024
Removing Obstacles before Breaking Through the Memory Wall: A Close Look at HBM Errors in the Field · USENIX ATC 2024

Methods — techniques the papers use, named apart from their topics

empirical prediction models · 1.0clustering · 1.0
YearPublicationVenuePosition
2026 Looking Back to Move Forward: Unveiling the Mysteries of HBM Errors to Predict Future Failures
abstract
High-bandwidth memory (HBM) is regarded as a promising technology for fundamentally overcoming the memory wall. It stacks up multiple DRAM dies vertically to dramatically improve the memory access bandwidth. However, this architecture also comes with more severe reliability issues, since HBM not only inherits error patterns of the conventional DRAM, but also introduces new error causes. In this article, we conduct the first systematical study on HBM errors, which cover over 460 million error events collected from 19 data centers and span over two years of deployment under a variety of services. Through error analyses and methodology validations, we confirm that the HBM exhibits different error patterns from conventional DRAM, in terms of spatial locality, temporal correlation, and sensor metrics which make empirical prediction models for DRAM error prediction ineffective for HBM. We design and implement Calchas , a hierarchical failure prediction framework for HBM based on our findings, which integrate spatial, temporal, and sensor information from various device levels to predict upcoming failures. The results demonstrate the feasibility of failure prediction across hierarchical levels.
Shuyue Zhou, Xinbin Hu, Ronglong Wu, Jiahao Lu 0003, Zhirong Shen, Yue Yu 0001, Yuze Jiang, Jiwu Shu, Feilong Lin, Yiming Zhang 0003
ACM Trans. Storage4
2024 Removing Obstacles before Breaking Through the Memory Wall: A Close Look at HBM Errors in the Field
Ronglong Wu, Shuyue Zhou, Jiahao Lu 0003, Zhirong Shen, Jiwu Shu, Feilong Lin, Yiming Zhang 0003
USENIX ATC3