Priyank Faldu

dblp:208/0336 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
1since 2021 · last 2022
0000-0002-1772-2048ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%
Databases, data mining, and information retrieval
1 paper
Graph data management · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.412020
Domain-Specialized Cache Management for Graph Analytics · HPCA 2020
Memory systems › cache › cache performance
cache thrashing
0.412020
Domain-Specialized Cache Management for Graph Analytics · HPCA 2020
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.412020
Domain-Specialized Cache Management for Graph Analytics · HPCA 2020
Graph data management
graph analytics
0.112020
Domain-Specialized Cache Management for Graph Analytics · HPCA 2020

Methods — techniques the papers use, named apart from their topics

software-assisted hot vertex identification · 0.9
YearPublicationVenuePosition
2022 Reconsidering OS memory optimizations in the presence of disaggregated memory
abstract
Tiered memory systems introduce an additional memory level with higher-than-local-DRAM access latency and require sophisticated memory management mechanisms to achieve cost-efficiency and high performance. Recent works focus on byte-addressable tiered memory architectures which offer better performance than pure swap-based systems. We observe that adding disaggregation to a byte-addressable tiered memory architecture requires important design changes that deviate from the common techniques that target lower-latency non-volatile memory systems. Our comprehensive analysis of real workloads shows that the high access latency to disaggregated memory undermines the utility of well-established memory management optimizations Based on these insights, we develop HotBox – a disaggregated memory management subsystem for Linux that strives to maximize the local memory hit rate with low memory management overhead. HotBox introduces only minor changes to the Linux kernel while outperforming state-of-the-art systems on memory-intensive benchmarks by up to 2.25×.
Shai Bergman, Priyank Faldu, Boris Grot, Lluís Vilanova, Mark Silberstein
ISMM2
2020 Domain-Specialized Cache Management for Graph Analytics
abstract
Graph analytics power a range of applications in areas as diverse as finance, networking and business logistics. A common property of graphs used in the domain of graph analytics is a power-law distribution of vertex connectivity, wherein a small number of vertices are responsible for a high fraction of all connections in the graph. These richly-connected, hot, vertices inherently exhibit high reuse. However, this work finds that state-of-the-art hardware cache management schemes struggle in capitalizing on their reuse due to highly irregular access patterns of graph analytics. In response, we propose GRASP, domain-specialized cache management at the last-level cache for graph analytics. GRASP augments existing cache policies to maximize reuse of hot vertices by protecting them against cache thrashing, while maintaining sufficient flexibility to capture the reuse of other vertices as needed. GRASP keeps hardware cost negligible by leveraging lightweight software support to pinpoint hot vertices, thus eliding the need for storage-intensive prediction mechanisms employed by state-of-the-art cache management schemes. On a set of diverse graph-analytic applications with large high-skew graph datasets, GRASP outperforms prior domain-agnostic schemes on all datapoints, yielding an average speed-up of 4.2% (max 9.4%) over the best-performing prior scheme. GRASP remains robust on low-/no-skew datasets, whereas prior schemes consistently cause a slowdown.
Priyank Faldu, Jeff Diamond, Boris Grot
HPCA1
2019 POSTER: Domain-Specialized Cache Management for Graph Analytics
abstract
In the domain of graph analytics, power-law graphs are prevalent. In such graphs, a small fraction of vertices are responsible for a large share of all graph connections. These richly-connected (hot) vertices inherently exhibit high reuse. However, this work finds that the state-of-the-art hardware cache management schemes struggle in capitalizing on their reuse due to highly irregular access patterns of graph analytics. In response, we argue in favor of leveraging software knowledge of graph data structures to accurately pinpoint hot vertices in hardware. To that end, we propose GRASP, a domain-specialized LLC management scheme that enables high cache efficiency for graph analytics with minimal modifications to existing cache structures.
Priyank Faldu, Jeff Diamond, Boris Grot
PACT1
2017 Leeway: Addressing Variability in Dead-Block Prediction for Last-Level Caches
abstract
The looming breakdown of Moore's Law and the end of voltage scaling are ushering a new era where neither transistors nor the energy to operate them is free. This calls for a new regime in computer systems, one in which every transistor counts. Caches are essential for processor performance and represent the bulk of modern processor's transistor budget. To get more performance out of the cache hierarchy, future processors will rely on effective cache management policies.This paper identifies variability in generational behavior of cache blocks as a key challenge for cache management policies that aim to identify dead blocks as early and as accurately as possible to maximize cache efficiency. We show that existing management policies are limited by the metrics they use to identify dead blocks, leading to low coverage and/or low accuracy in the face of variability. In response, we introduce a new metric - Live Distance - that uses the stack distance to learn the temporal reuse characteristics of cache blocks, thus enabling a dead block predictor that is robust to variability in generational behavior. Based on the reuse characteristics of an application's cache blocks, our predictor - Leeway - classifies application's behavior as streaming-oriented or reuse-oriented and dynamically selects an appropriate cache management policy. By leveraging live distance for LLC management, Leeway outperforms state-of-the-art approaches on single- and multi-core SPEC and manycore CloudSuite workloads.
Priyank Faldu, Boris Grot
PACT1