Javin Pombra

dblp:293/6682 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 69% Memory systems · 21% Parallel and multicore computing · 10%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.512021
RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance · MICRO 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
recommendation model accelerator
0.512021
RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance · MICRO 2021
Memory systems
cache
0.112021
RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance · MICRO 2021
Memory systems › cache
embedding cache
0.112021
RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance · MICRO 2021
Parallel and multicore computing › parallel scheduling
heterogeneous multiprocessor scheduling
0.112021
RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance · MICRO 2021

Methods — techniques the papers use, named apart from their topics

top-k filtering · 0.5sub-batch processing · 0.5static and dynamic caching · 0.5model pipelining · 0.5
YearPublicationVenuePosition
2021 RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance
abstract
Deep learning recommendation systems must provide high quality, personalized content under strict tail-latency targets and high system loads. This paper presents RecPipe, a system to jointly optimize recommendation quality and inference performance. Central to RecPipe is decomposing recommendation models into multi-stage pipelines to maintain quality while reducing compute complexity and exposing distinct parallelism opportunities. RecPipe implements an inference scheduler to map multi-stage recommendation engines onto commodity, heterogeneous platforms (e.g., CPUs, GPUs). While the hardware-aware scheduling improves ranking efficiency, the commodity platforms suffer from many limitations requiring specialized hardware. Thus, we design RecPipeAccel (RPAccel), a custom accelerator that jointly optimizes quality, tail-latency, and system throughput. RPAccel is designed specifically to exploit the distinct design space opened via RecPipe. In particular, RPAccel processes queries in sub-batches to pipeline recommendation stages, implements dual static and dynamic embedding caches, a set of top-k filtering units, and a reconfigurable systolic array. Compared to previously proposed specialized recommendation accelerators and at iso-quality, we demonstrate that RPAccel improves latency and throughput by 3 × and 6 ×.
Udit Gupta 0001, Samuel Hsia, Jeff Zhang 0001, Mark Wilkening, Javin Pombra, Hsien-Hsin S. Lee, Gu-Yeon Wei, Carole-Jean Wu, David Brooks 0001
MICRO5