Haishan Zhu

dblp:00/7084 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 44% Cloud and datacenter computing · 24% Hardware accelerators and domain-specific architectures · 19%
Artificial intelligence
2 papers
Efficient and distributed learning · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.822020
Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020
Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019
Cloud and datacenter computing
quality of service
0.622019
Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019
Dirigent: Enforcing QoS for Latency-Critical Tasks on Shared Multicore Systems · ASPLOS 2016
Machine learning › Efficient and distributed learning
model compression
0.412020
Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020
Cloud and datacenter computing › resource management
datacenter resource management
0.412019
Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019
Memory systems › memory interference
memory contention
0.412019
Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019
Memory systems
cache design
0.312018
SIPT: Speculatively Indexed, Physically Tagged Caches · HPCA 2018
Memory systems › cache design
cache indexing
0.312018
SIPT: Speculatively Indexed, Physically Tagged Caches · HPCA 2018
Memory systems › cache › CPU cache
l1 data cache
0.312018
SIPT: Speculatively Indexed, Physically Tagged Caches · HPCA 2018
Memory systems › cache design
virtually-indexed physically-tagged cache
0.312018
SIPT: Speculatively Indexed, Physically Tagged Caches · HPCA 2018
Parallel and multicore computing › task scheduling
contention-aware scheduling
0.212016
Dirigent: Enforcing QoS for Latency-Critical Tasks on Shared Multicore Systems · ASPLOS 2016
Embedded and real-time systems › real-time scheduling
multicore scheduling
0.212016
Dirigent: Enforcing QoS for Latency-Critical Tasks on Shared Multicore Systems · ASPLOS 2016
Machine learning › Efficient and distributed learning
distributed training
0.112019
Kelp: QoS for Accelerated Machine Learning Systems · HPCA 2019
Memory systems
cache
0.112018
SIPT: Speculatively Indexed, Physically Tagged Caches · HPCA 2018
Memory systems › memory access latency
cache access latency
0.112018
SIPT: Speculatively Indexed, Physically Tagged Caches · HPCA 2018
Processor architecture and microarchitecture
resource contention
0.112016
Dirigent: Enforcing QoS for Latency-Critical Tasks on Shared Multicore Systems · ASPLOS 2016

Methods — techniques the papers use, named apart from their topics

quantization · 0.9microsoft floating point · 0.9runtime isolation · 0.8microarchitectural simulation · 0.3runtime performance management · 0.2fine-grained qos control · 0.2
YearPublicationVenuePosition
2020 Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point
abstract
In this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bfloat16 and INT8 at 3x and 4x lower cost, respectively. MSFP incurs negligible impact to accuracy (<1%), requires no changes to the model topology, and is integrated with a mature cloud production pipeline. MSFP supports various classes of deep learning models including CNNs, RNNs, and Transformers without modification. Finally, we characterize the accuracy and implementation of MSFP and demonstrate its efficacy on a number of production scenarios, including models that power major online scenarios such as web search, question-answering, and image classification.
Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, Alessandro Forin, Haishan Zhu, Taesik Na, Prerak Patel, Shuai Che, Lok Chand Koppaka, Subhojit Som, Kaustav Das, Saurabh Tiwary, Steven K. Reinhardt, Sitaram Lanka, Eric S. Chung, Doug Burger
NeurIPS12
2019 Kelp: QoS for Accelerated Machine Learning Systems
abstract
Development and deployment of machine learning (ML) accelerators in Warehouse Scale Computers (WSCs) demand significant capital investments and engineering efforts. However, even though heavy computation can be offloaded to the accelerators, applications often depend on the host system for various supporting tasks. As a result, contention on host resources, such as memory bandwidth, can significantly discount the performance and efficiency gains of accelerators. The impact of performance interference is further amplified in distributed learning, which has become increasingly common as model sizes continue to grow. In this work, we study the performance of four production machine learning workloads on three accelerator platforms. Our experiments show that these workloads are highly sensitive to host memory bandwidth contention, which can cause 40% average performance degradation when left unmanaged. To tackle this problem, we design and implement Kelp, a software runtime that isolates high priority accelerated ML tasks from memory resource interference. We evaluate Kelp with both production and artificial aggressor workloads, and compare its effectiveness with previously proposed solutions. Our evaluation shows that Kelp is effective in mitigating performance degradation of the accelerated tasks, and improves performance by 24% on average. Compared to previous work, Kelp reduces performance degradation of ML tasks by 7% and improves system efficiency by 17%. Our results further expose opportunities in future architecture designs.
Haishan Zhu, David Lo 0003, Liqun Cheng, Rama Govindaraju, Parthasarathy Ranganathan, Mattan Erez
HPCA1
2018 SIPT: Speculatively Indexed, Physically Tagged Caches
abstract
First-level (L1) data cache access latency is critical to performance because it services the vast majority of loads and stores. To keep L1 latency low while ensuring low-complexity and simple-to-verify operation, current processors most-typically utilize a virtually-indexed physically-tagged (VIPT) cache architecture. While VIPT caches decrease latency by proceeding with cache access and address translation concurrently, each cache way is constrained by the size of a virtual page. Thus, larger L1 caches are highly-associative, which degrades their access latency and energy. We propose speculatively-indexed physically-tagged (SIPT) caches to enable simultaneously larger, faster, and more efficient L1 caches. A SIPT cache speculates on the value of a few address bits beyond the page offset concurrently with address translation, maintaining the overall safe and reliable architecture of a VIPT cache while eliminating the VIPT design constraints. SIPT is a purely microarchitectural approach that can be used with any software and for all accesses. We evaluate SIPT with simulations of applications under standard Linux. SIPT improves performance by 8.1% on average and reduces total cache-hierarchy energy by 15.6%.
Tianhao Zheng, Haishan Zhu, Mattan Erez
HPCA2
2016 Dirigent: Enforcing QoS for Latency-Critical Tasks on Shared Multicore Systems
abstract
Latency-critical applications suffer from both average performance degradation and reduced completion time predictability when collocated with batch tasks. Such variation forces the system to overprovision resources to ensure Quality of Service (QoS) for latency-critical tasks, degrading overall system throughput. We explore the causes of this variation and exploit the opportunities of mitigating variation directly to simultaneously improve both QoS and utilization. We develop, implement, and evaluate Dirigent, a lightweight performance-management runtime system that accurately controls the QoS of latency-critical applications at fine time scales, leveraging existing architecture mechanisms. We evaluate Dirigent on a real machine and show that it is significantly more effective than configurations representative of prior schemes.
Haishan Zhu, Mattan Erez
ASPLOS1