VLDB 2026 Research / reviewers in the wild / expert
Anirudh Seshadri
dblp:222/1598
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Processor architecture and microarchitecture · 62% Reconfigurable computing and FPGAs · 24% Memory systems · 14% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
branch prediction |
1.0 | 2 | 2025 | Delinquent Loop Pre-execution Using Predicated Helper Threads · HPCA 2025 Post-Fabrication Microarchitecture · MICRO 2021 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable logic |
0.5 | 1 | 2021 | Post-Fabrication Microarchitecture · MICRO 2021 |
Processor architecture and microarchitecture › multithreading
helper threads |
0.3 | 1 | 2025 | Delinquent Loop Pre-execution Using Predicated Helper Threads · HPCA 2025 |
Memory systems
cache |
0.1 | 1 | 2021 | Post-Fabrication Microarchitecture · MICRO 2021 |
Memory systems › cache › prefetching
data prefetching |
0.1 | 1 | 2021 | Post-Fabrication Microarchitecture · MICRO 2021 |
Methods — techniques the papers use, named apart from their topics
store-load forwarding · 0.9predicated execution · 0.9reconfigurable logic fabric · 0.5IPC analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Delinquent Loop Pre-execution Using Predicated Helper ThreadsabstractBranch pre-execution targets delinquent branches that are not predictable by conventional branch predictors. Helper threads attempt to resolve branches ahead of the main thread. Pre-executed branch outcomes are communicated to the main thread’s fetch unit via a global branch queue or local branch queues (one per branch PC). Two key challenges are discussed in this paper. 1) Handling a delinquent branch b2 that is control-dependent on another delinquent branch b1. Prior works that include both branches resort to branch prediction of b1 in the helper thread to determine whether or not to pre-execute b2. But b1 is hard-to-predict and the misprediction bottleneck merely shifts from the main thread to the helper thread. 2) Handling a store instruction that both influences a delinquent branch and is control-dependent on it. We propose predicated helper threads (Phelps) to address these challenges. Phelps constructs a helper thread for each inner loop containing delinquent branches. All delinquent branches, even control-dependent ones (b2), are unconditionally pre-executed in each loop iteration. Per-branch queues are managed in lockstep based on loop iterations, allowing the helper thread to deposit outcomes for both b1 and b2 each iteration and the main thread to consume or ignore b2 outcomes in the correct sequence dictated by b1. The helper thread also retains influential stores for dynamic disambiguation and store-load forwarding. Any such store that is control-dependent on a delinquent branch is predicated on the branch’s outcome, which is necessary because the helper thread no longer has control-flow (except for the loop branch). Phelps also features dual decoupled helper threads for outer-inner loop pairs, for effective branch pre-execution when the inner loop has a short and unpredictable trip count. Anirudh Seshadri, Eric Rotenberg |
HPCA | 1 |
| 2021 | Post-Fabrication MicroarchitectureabstractMicroarchitectural enhancements that improve performance generally, across many workloads, are favored in superscalar processor design. Targeting general performance is necessary but it also constrains some microarchitecture innovation. We explore relieving this constraint, via a new paradigm called Post-Fabrication Microarchitecture (PFM). A high-performance superscalar core is coupled with a reconfigurable logic fabric, RF. A programmable interface, or Agent, allows for RF to observe and microarchitecturally intervene at key pipeline stages of the superscalar core. New microarchitectural components, specific to applications, are synthesized on-demand to RF. All instructions still flow through the superscalar pipeline, as usual, but their execution is streamlined (better instructions per cycle (IPC)) through microarchitectural intervention by RF. Our research shows that one can achieve large speedups of individual applications, by analyzing their bottlenecks and providing customized microarchitectural solutions to target these bottlenecks. Examples of PFM use-cases explored in this paper include custom branch predictors and data prefetchers. Chanchal Kumar, Anirudh Seshadri, Aayush Chaudhary, Shubham Bhawalkar, Eric Rotenberg |
MICRO | 2 |