Krishnam Tibrewala

dblp:385/3375 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0002-9309-4810ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Correct Wrong Path Simulation
abstract
Modern OoO CPUs employ deep pipelines with high branch misprediction recovery penalties. Instructions speculatively executed along mispredicted paths can significantly alter microarchitectural state. During design space exploration, architects often rely on trace-driven simulators, which are significantly faster than execution-driven models but trade accuracy for speed. Despite this benefit, trace-driven simulation often fails to adequately model the effects of wrong-path execution because traces are typically collected only from the correct-path. While prior work can accurately model wrong-path effects on the instruction stream, it often makes unrealistic assumptions when modeling the impact on the data stream. In this work, we examine the effects of wrong-path execution and present an infrastructure for enabling its modeling in a tracedriven simulator. Our analysis shows that wrong-path execution extensively affects structures on both the instruction and data sides, yielding performance variations ranging from $-3.6 \%$ to $85.7 \%$ compared to a baseline that ignores these effects. To benefit the research community and enhance the accuracy of simulators, we provide our traces and tracing utility. We aim for this to encourage industry to provide wrong-path traces generated by internal simulators, enabling fast academic research without exposing industry proprietary IP.
Chrysanthos Pepi, Krishnam Tibrewala, Bhargav Reddy Godala, Sankara Prasad Ramesh, Alberto Ros 0001, Daniel A. Jiménez, Gilles Pokam, Paul Gratz
ISPASS2
2025 Skia: Exposing Shadow Branches
abstract
Modern processors implement a decoupled front-end, often using a form of Fetch Directed Instruction Prefetching (FDIP), to avoid front-end stalls. FDIP is driven by the Branch Prediction Unit (BPU), relying on the BPU's accuracy and branch target tracking structures to speculatively fetch instructions into the Instruction Cache (L1-I cache). As contemporary data center applications become more complex, their code footprints also grow, resulting in a high number of Branch Target Buffer (BTB) misses. These BTB missing branches typically have previously been decoded and placed in the BTB, but have since been evicted, leading to BTB misses now. FDIP can alleviate L1-I cache misses, but its reliance on the BPU's tracking structures means that when it encounters a BTB miss, the BPU may not identify the current instruction as a branch to FDIP. This can prevent FDIP from prefetching or cause it to speculate down the wrong path, further polluting the L1-I cache.
Chrysanthos Pepi, Bhargav Reddy Godala, Krishnam Tibrewala, Gino Chacon, Paul Gratz, Daniel A. Jiménez, Gilles Pokam, David I. August
ASPLOS (2)3
2025 Light-weight Cache Replacement for Instruction Heavy Workloads
abstract
The last-level cache (LLC) is the last chance for memory accesses from the processor to avoid the costly latency of accessing the main memory.In recent years, an increasing number of instruction heavy workloads have put pressure on the last-level cache.We find that, for instruction heavy workloads, a simple replacement policy with minimal overhead provides at least the same benefit as a stateof-the-art, high-overhead replacement policy in the presence of aggressive prefetching.Our proposal is based on specifying insertion and promotion vectors (IPVs) as a generalization of re-reference interval prediction (RRIP) in such a way that the space of feasible policies may be searched exhaustively to find the best policy for the training set of workloads.The policies are formulated to deliver the best performance taking into account demand and prefetch accesses.We show that our technique, Prefetch Aware Coarse-grained Insertion and Promotion Vectors (PACIPV), improves performance over a state-of-the-art LLC replacement policy (Mockingjay) for instruction heavy workloads, and remains competitive for data heavy workloads with significantly less hardware overhead.We show that RRIP-based IPVs are very easy to implement but outperform far more complex replacement policies.PACIPV achieves a speedup of 3.3% over the baseline of LRU, outperforming SRRIP by 1.1% and the much more hardware intensive Mockingjay by 0.1%.
Saba Mostofi, Setu Gupta, Ahmad Hassani, Krishnam Tibrewala, Elvira Teran, Paul Gratz, Daniel A. Jiménez
ISCA4