Vrushabh Sanghavi

dblp:320/3273 · also Vrushabh H. Sanghavi · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0001-9886-7419ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 56% Processor architecture and microarchitecture · 22% Hardware accelerators and domain-specific architectures · 22%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
CPU optimization
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems
memory access latency
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model inference
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems
software prefetching
0.712023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems › cache
cache performance
0.212023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023
Memory systems › memory access patterns
irregular memory access
0.212023
Optimizing CPU Performance for Recommendation Systems At-Scale · ISCA 2023

Methods — techniques the papers use, named apart from their topics

software prefetching · 0.7hyperthreading · 0.7
YearPublicationVenuePosition
2023 Optimizing CPU Performance for Recommendation Systems At-Scale
abstract
Deep Learning Recommendation Models (DLRMs) are very popular in personalized recommendation systems and are a major contributor to the data-center AI cycles. Due to the high computational and memory bandwidth needs of DLRMs, specifically the embedding stage in DLRM inferences, both CPUs and GPUs are used for hosting such workloads. This is primarily because of the heavy irregular memory accesses in the embedding stage of computation that leads to significant stalls in the CPU pipeline. As the model and parameter sizes keep increasing with newer recommendation models, the computational dominance of the embedding stage also grows, thereby, bringing into question the suitability of CPUs for inference. In this paper, we first quantify the cause of irregular accesses and their impact on caches and observe that off-chip memory access is the main contributor to high latency. Therefore, we exploit two well-known techniques: (1) Software prefetching, to hide the memory access latency suffered by the demand loads and (2) Overlapping computation and memory accesses, to reduce CPU stalls via hyperthreading to minimize the overall execution time. We evaluate our work on a single-core and 24-core configuration with the latest recommendation models and recently released production traces. Our integrated techniques speed up the inference by up to 1.59x, and on average by 1.4x.
Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi, Samvit Kaul, Meena Arunachalam, Kiwan Maeng, Adwait Jog, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das
ISCA4