Anirudh Seshadri

dblp:222/1598 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 62% Reconfigurable computing and FPGAs · 24% Memory systems · 14%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
branch prediction
1.022025
Delinquent Loop Pre-execution Using Predicated Helper Threads · HPCA 2025
Post-Fabrication Microarchitecture · MICRO 2021
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable logic
0.512021
Post-Fabrication Microarchitecture · MICRO 2021
Processor architecture and microarchitecture › multithreading
helper threads
0.312025
Delinquent Loop Pre-execution Using Predicated Helper Threads · HPCA 2025
Memory systems
cache
0.112021
Post-Fabrication Microarchitecture · MICRO 2021
Memory systems › cache › prefetching
data prefetching
0.112021
Post-Fabrication Microarchitecture · MICRO 2021

Methods — techniques the papers use, named apart from their topics

store-load forwarding · 0.9predicated execution · 0.9reconfigurable logic fabric · 0.5IPC analysis · 0.5
YearPublicationVenuePosition
2025 Delinquent Loop Pre-execution Using Predicated Helper Threads
abstract
Branch pre-execution targets delinquent branches that are not predictable by conventional branch predictors. Helper threads attempt to resolve branches ahead of the main thread. Pre-executed branch outcomes are communicated to the main thread’s fetch unit via a global branch queue or local branch queues (one per branch PC). Two key challenges are discussed in this paper. 1) Handling a delinquent branch b2 that is control-dependent on another delinquent branch b1. Prior works that include both branches resort to branch prediction of b1 in the helper thread to determine whether or not to pre-execute b2. But b1 is hard-to-predict and the misprediction bottleneck merely shifts from the main thread to the helper thread. 2) Handling a store instruction that both influences a delinquent branch and is control-dependent on it. We propose predicated helper threads (Phelps) to address these challenges. Phelps constructs a helper thread for each inner loop containing delinquent branches. All delinquent branches, even control-dependent ones (b2), are unconditionally pre-executed in each loop iteration. Per-branch queues are managed in lockstep based on loop iterations, allowing the helper thread to deposit outcomes for both b1 and b2 each iteration and the main thread to consume or ignore b2 outcomes in the correct sequence dictated by b1. The helper thread also retains influential stores for dynamic disambiguation and store-load forwarding. Any such store that is control-dependent on a delinquent branch is predicated on the branch’s outcome, which is necessary because the helper thread no longer has control-flow (except for the loop branch). Phelps also features dual decoupled helper threads for outer-inner loop pairs, for effective branch pre-execution when the inner loop has a short and unpredictable trip count.
Anirudh Seshadri, Eric Rotenberg
HPCA1
2021 Post-Fabrication Microarchitecture
abstract
Microarchitectural enhancements that improve performance generally, across many workloads, are favored in superscalar processor design. Targeting general performance is necessary but it also constrains some microarchitecture innovation. We explore relieving this constraint, via a new paradigm called Post-Fabrication Microarchitecture (PFM). A high-performance superscalar core is coupled with a reconfigurable logic fabric, RF. A programmable interface, or Agent, allows for RF to observe and microarchitecturally intervene at key pipeline stages of the superscalar core. New microarchitectural components, specific to applications, are synthesized on-demand to RF. All instructions still flow through the superscalar pipeline, as usual, but their execution is streamlined (better instructions per cycle (IPC)) through microarchitectural intervention by RF. Our research shows that one can achieve large speedups of individual applications, by analyzing their bottlenecks and providing customized microarchitectural solutions to target these bottlenecks. Examples of PFM use-cases explored in this paper include custom branch predictors and data prefetchers.
Chanchal Kumar, Anirudh Seshadri, Aayush Chaudhary, Shubham Bhawalkar, Eric Rotenberg
MICRO2