VLDB 2026 Research / reviewers in the wild / expert
Nayana P. Nagendra
dblp:217/1291 · also Nayana Prasad Nagendra
· DBLP profile ↗
4ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0003-3972-0497ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Processor architecture and microarchitecture · 39% Memory systems · 31% Parallel and multicore computing · 22% | |
| Network and information security
1 paper |
Hardware security and side channels · 77% Systems and software security · 23% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 54% Concurrent programming · 46% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache management
cache replacement |
0.7 | 1 | 2023 | EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction Caching · ISCA 2023 |
Hardware security and side channels › hardware security primitives
hardware root of trust |
0.4 | 1 | 2019 | Architectural Support for Containment-based Security · ASPLOS 2019 |
Processor architecture and microarchitecture › front-end
front-end stalls |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Memory systems › cache › CPU cache
instruction cache |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Memory systems › cache
instruction cache miss |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Processor architecture and microarchitecture › hardware-assisted security
secure architecture |
0.4 | 1 | 2019 | Architectural Support for Containment-based Security · ASPLOS 2019 |
Cloud and datacenter computing › datacenter architecture
warehouse-scale computer |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.3 | 1 | 2018 | Hardware Multithreaded Transactions · ASPLOS 2018 |
Parallel and multicore computing
speculative parallelization |
0.3 | 1 | 2018 | Hardware Multithreaded Transactions · ASPLOS 2018 |
Parallel and multicore computing
transactional memory |
0.3 | 1 | 2018 | Hardware Multithreaded Transactions · ASPLOS 2018 |
Processor architecture and microarchitecture › instruction fetch › instruction prefetching
fetch directed instruction prefetching |
0.2 | 1 | 2023 | EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction Caching · ISCA 2023 |
Processor architecture and microarchitecture › instruction fetch
instruction prefetching |
0.2 | 1 | 2023 | EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction Caching · ISCA 2023 |
Systems and software security › trusted computing
trusted execution |
0.1 | 1 | 2019 | Architectural Support for Containment-based Security · ASPLOS 2019 |
Compilers and program optimization
code layout optimization |
0.1 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Methods — techniques the papers use, named apart from their topics
supply chain diversification · 0.8hardware performance monitoring · 0.8formal verification · 0.8control flow graph analysis · 0.8trace-driven simulation · 0.7hardware transactional memory · 0.7cost-aware replacement · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction CachingabstractFor decades, architects have designed cache replacement policies to reduce cache misses. Since not all cache misses affect processor performance equally, researchers have also proposed cache replacement policies focused on reducing the total miss cost rather than the total miss count. However, all prior cost-aware replacement policies have been proposed specifically for data caching and are either inappropriate or unnecessarily complex for instruction caching. This paper presents EMISSARY, the first cost-aware cache replacement family of policies specifically designed for instruction caching. Observing that modern architectures entirely tolerate many instruction cache misses, EMISSARY resists evicting those cache lines whose misses cause costly decode starvations. In the context of a modern processor with fetch-directed instruction prefetching and other aggressive front-end features, EMISSARY applied to L2 cache instructions delivers an impressive 3.24% geomean speedup (up to 23.7%) and a geomean energy savings of 2.1% (up to 17.7%) when evaluated on widely used server applications with large code footprints. This speedup is 21.6% of the total speedup obtained by an unrealizable L2 cache with a zero-cycle miss latency for all capacity and conflict instruction misses. Nayana P. Nagendra, Bhargav Reddy Godala, Ishita Chaturvedi, Atmn Patel, Svilen Kanev, Tipp Moseley, Jared Stark, Gilles Pokam, Simone Campanoni, David I. August |
ISCA | 1 |
| 2019 | Architectural Support for Containment-based SecurityabstractSoftware security techniques rely on correct execution by the hardware. Securing hardware components has been challenging due to their complexity and the proportionate attack surface they present during their design, manufacture, deployment, and operation. Recognizing that external communication represents one of the greatest threats to a system's security, this paper introduces the TrustGuard containment architecture. TrustGuard contains malicious and erroneous behavior using a relatively simple and pluggable gatekeeping hardware component called the Sentry. The Sentry bridges a physical gap between the untrusted system and its external interfaces. TrustGuard allows only communication that results from the correct execution of trusted software, thereby preventing the ill effects of actions by malicious hardware or software from leaving the system. The simplicity and pluggability of the Sentry, which is implemented in less than half the lines of code of a simple in-order processor, enables additional measures to secure this root of trust, including formal verification, supervised manufacture, and supply chain diversification with less than a 15% impact on performance. Hansen Zhang, Soumyadeep Ghosh, Jordan Fix, Sotiris Apostolakis, Stephen R. Beard, Nayana P. Nagendra, Taewook Oh, David I. August |
ASPLOS | 6 |
| 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computersabstractThe large instruction working sets of private and public cloud workloads lead to frequent instruction cache misses and costs in the millions of dollars. While prior work has identified the growing importance of this problem, to date, there has been little analysis of where the misses come from, and what the opportunities are to improve them. To address this challenge, this paper makes three contributions. First, we present the design and deployment of a new, always-on, fleet-wide monitoring system, AsmDB, that tracks front-end bottlenecks. AsmDB uses hardware support to collect bursty execution traces, fleet-wide temporal and spatial sampling, and sophisticated offline post-processing to construct full-program dynamic control-flow graphs. Second, based on a longitudinal analysis of AsmDB data from real-world online services, we present two detailed insights on the sources of front-end stalls: (1) cold code that is brought in along with hot code leads to significant cache fragmentation and a corresponding large number of instruction cache misses; (2) distant branches and calls that are not amenable to traditional cache locality or next-line prefetching strategies account for a large fraction of cache misses. Third, we prototype two optimizations that target these insights. For misses caused by fragmentation, we focus on memcmp, one of the hottest functions contributing to cache misses, and show how fine-grained layout optimizations lead to significant benefits. For misses at the targets of distant jumps, we propose new hardware support for software code prefetching and prototype a new feedback-directed compiler optimization that combines static program flow analysis with dynamic miss profiles to demonstrate significant benefits for several large warehouse-scale workloads. Improving upon prior work, our proposal avoids invasive hardware modifications by prefetching via software in an efficient and scalable way. Simulation results show that such an approach can eliminate up to 96% of instruction cache misses with negligible overheads. Grant Ayers, Nayana P. Nagendra, David I. August, Hyoun Kyu Cho, Svilen Kanev, Christoforos E. Kozyrakis, Trivikram Krishnamurthy, Heiner Litz, Tipp Moseley, Parthasarathy Ranganathan |
ISCA | 2 |
| 2018 | Hardware Multithreaded TransactionsabstractSpeculation with transactional memory systems helps pro- grammers and compilers produce profitable thread-level parallel programs. Prior work shows that supporting transactions that can span multiple threads, rather than requiring transactions be contained within a single thread, enables new types of speculative parallelization techniques for both programmers and parallelizing compilers. Unfortunately, software support for multi-threaded transactions (MTXs) comes with significant additional inter-thread communication overhead for speculation validation. This overhead can make otherwise good parallelization unprofitable for programs with sizeable read and write sets. Some programs using these prior software MTXs overcame this problem through significant efforts by expert programmers to minimize these sets and optimize communication, capabilities which compiler technology has been unable to equivalently achieve. Instead, this paper makes speculative parallelization less laborious and more feasible through low-overhead speculation validation, presenting the first complete design, implementation, and evaluation of hardware MTXs. Even with maximal speculation validation of every load and store inside transactions of tens to hundreds of millions of instructions, profitable parallelization of complex programs can be achieved. Across 8 benchmarks, this system achieves a geomean speedup of 99% over sequential execution on a multicore machine with 4 cores. Jordan Fix, Nayana P. Nagendra, Sotiris Apostolakis, Hansen Zhang, Sophie Qiu, David I. August |
ASPLOS | 2 |