Krishnendra Nathella

dblp:242/9022 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0002-8091-2389ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
YearPublicationVenuePosition
2023 A Characterization of the Effects of Software Instruction Prefetching on an Aggressive Front-end
abstract
Growing application sizes continue to strain the memory system. As more complex applications are developed, and the instruction memory footprint increases, the cache hierarchy cannot contain relevant instructions causing the front end to become idle as it awaits fetched instructions. Hardware instruction prefetchers can alleviate this problem by learning instruction stream behavior and prefetching instructions into the cache before use. Fetch Directed Prefetching (FDP) is a ubiquitous form of hardware instruction prefetching that uses branch predictor to predict future instruction cache references. Modern processors generally implement aggressive, deep FDP to decouple the front-end from the rest of the machine. Instruction accesses, however, provide a small amount of information over long periods due to instruction stream variability and lack of information regarding the context of instruction accesses. Capturing the instruction stream’s context requires significant storage overhead to correlate instruction accesses and ensure timely accesses. Software prefetching techniques overcome this problem by profiling and statically analyzing an application’s behavior. Prior work demonstrates software instruction prefetching’s potential performance benefit but does not evaluate performance in the context of aggressive, decoupled front-ends. While software prefetching provides ~ 20% improvement in conservative front-ends, we find that it does not yield performance benefit when modeling a baseline with an aggressive FDP, in some cases hurting performance. Our analysis finds that using software instruction prefetching negatively impacts an aggressive front-end’s behavior. We investigate this finding and characterize the different states a front-end can be in and how introducing instructions into an application can change the front-end’s behavior resulting in destructive interference with the software prefetcher.
Gino Chacon, Nathan Gober, Krishnendra Nathella, Paul Gratz, Daniel A. Jiménez
ISPASS3
2022 Whisper: Profile-Guided Branch Misprediction Elimination for Data Center Applications
abstract
Modern data center applications experience frequent branch mispredictions– degrading performance, increasing cost, and reducing energy efficiency in data centers. Even the state-of the-art branch predictor, TAGE-SC-L, suffers from an average branch Mispredictions Per Kilo Instructions (branch-MPKI) of 3.0 (0.5-7.2) for these applications since their large code footprints exhaust TAGE-SC-L’s intended capacity. In this work, we propose Whisper, a novel profile-guided mechanism to avoid branch mispredictions. Whisper investigates the in-production profile of data center applications to identify precise program contexts that lead to branch mispredictions. Corresponding prediction hints are then inserted into code to strategically avoid those mispredictions during program execution. Whisper presents three novel profile-guided techniques: (1) hashed history correlation which efficiently encodes hard-to-predict correlations in branch history using lightweight Boolean formulas, (2) randomized formula testing which selects a locally-optimal Boolean formula from a randomly selected subset of possible formulas to predict a branch, and (3) the extension of Read-Once Monotone Boolean Formulas with Implication and Converse Non-Implication to improve the branch history coverage of these formulas with minimal overhead. We evaluate Whisper on 12 widely-used data center applications and demonstrate that Whisper enables traditional branch predictors to achieve a speedup close to that of an ideal branch predictor. Specifically, Whisper achieves an average speedup of 2.8% (0.4%-4.6%) by reducing 16.8% (1.7%-32.4%) of branch mispredictions over TAGE-SC-L and outperforms the state-of the-art profile-guided branch prediction mechanisms by 7.9% on average.
Tanvir Ahmed Khan 0001, Muhammed Ugur, Krishnendra Nathella, Dam Sunwoo, Heiner Litz, Daniel A. Jiménez, Baris Kasikci
MICRO3
2022 Practical Temporal Prefetching With Compressed On-Chip Metadata
abstract
Temporal prefetchers are powerful because they can prefetch irregular sequences of memory accesses, but temporal prefetchers are commercially infeasible because they store large amounts of metadata in DRAM. This article presents Triage, the first temporal data prefetcher that does not require off-chip metadata. Triage builds on two insights: (1) Metadata are not equally useful, so the less useful metadata need not be saved, and (2) for irregular workloads, it is more profitable to use portions of the LLC to store metadata than data. We also introduce novel schemes to identify useful metadata, to compress metadata, and to determine the fraction of the LLC to dedicate for metadata. Using an industrial-strength simulator running irregular workloads on a single-core system, we show that at a prefetch degree of 4, Triage improves performance by 41.1 percent compared to a baseline with no prefetching, whereas BO, a state-of-the-art prefetcher that uses only on-chip metadata, sees only 10.9 percent improvement. Compared with MISB, a temporal prefetcher that uses off-chip metadata, Triage provides a design alternative that reduces memory traffic by an order of magnitude (260.8 percent extra traffic for MISB at degree 1 versus 56.9 percent for Triage), while reducing coverage by 20 percent.
Krishnendra Nathella, Matthew Pabst, Dam Sunwoo, Akanksha Jain, Calvin Lin
IEEE Trans. Computers2
2021 Re-establishing Fetch-Directed Instruction Prefetching: An Industry Perspective
abstract
Instruction prefetching can play a pivotal role in improving the performance of workloads with large instruction footprints and frequent, costly frontend stalls. In particular, Fetch Directed Prefetching (FDP) is an effective technique to mitigate frontend stalls since it leverages existing branch prediction resources in a processor and incurs very little hardware overhead. Modern processors have been trending towards provisioning more frontend resources, which bodes well for FDP as it requires these resources to be effective. However, recent academic research has been using outdated and less than optimal frontend baselines that employ smaller structures, resulting in equivocal outcomes. This paper presents a detailed FDP microarchitecture and evaluates two improvements, better branch history management and post-fetch correction. Our mechanism provides a 41.0% speedup over the baseline (no prefetching, no FDP) with only 195 bytes of hardware overhead and outperforms the 1st Instruction Prefetching Championship (IPC-1) winners that had a 128KB storage budget. We believe that our FDP-based frontend design can serve as a new reference baseline for instruction prefetching research to bridge the gap between academia and industry.
Yasuo Ishii, Jaekyu Lee, Krishnendra Nathella, Dam Sunwoo
ISPASS3
2019 Efficient metadata management for irregular data prefetching
abstract
Temporal prefetchers have the potential to prefetch arbitrary memory access patterns, but they require large amounts of metadata that must typically be stored in DRAM. In 2013, the Irregular Stream Buffer (ISB), showed how this metadata could be cached on chip and managed implicitly by synchronizing its contents with that of the TLB. This paper reveals the inefficiency of that approach and presents a new metadata management scheme that uses a simple metadata prefetcher to feed the metadata cache. The result is the Managed ISB (MISB), a temporal prefetcher that significantly advances the state-of-the-art in terms of both traffic overhead and IPC.
Krishnendra Nathella, Dam Sunwoo, Akanksha Jain, Calvin Lin
ISCA2
2019 Temporal Prefetching Without the Off-Chip Metadata
abstract
Temporal prefetching offers great potential, but this potential is difficult to achieve because of the need to store large amounts of prefetcher metadata off chip. To reduce the latency and traffic of off-chip metadata accesses, recent advances in temporal prefetching have proposed increasingly complex mechanisms that cache and prefetch this off-chip metadata. This paper suggests a return to simplicity: We present a temporal prefetcher whose metadata resides entirely on chip. The key insights are (1) only a small portion of prefetcher metadata is important, and (2) for most workloads with irregular accesses, the benefits of an effective prefetcher outweigh the marginal benefits of a larger data cache. Thus, our solution, the Triage prefetcher, identifies important metadata and uses a portion of the LLC to store this metadata, and it dynamically partitions the LLC between data and metadata.
Krishnendra Nathella, Joseph Pusdesris, Dam Sunwoo, Akanksha Jain, Calvin Lin
MICRO2