VLDB 2026 Research / reviewers in the wild / expert
Aniket Deshmukh 0002
dblp:375/1776-2
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0007-6393-6139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enabling Ahead Prediction with Practical Energy ConstraintsabstractAccurate branch predictors require multiple cycles to produce a prediction, and that latency hurts processor performance."Ahead prediction" solves the performance problem by starting the prediction early.Unfortunately, this means making the prediction without the directions of the N branches between when the prediction starts and when the relevant branch's prediction is needed.The energy required to consider all 2 𝑛 possible cases increases by 14.6x, making ahead prediction not viable.This paper shows that most of the intermediate branch directions never materialize, reducing the number of observed missing history patterns significantly (usually, only one or two).We modified the TAGE predictor to eliminate those branches from having to be considered.The result, our ahead predictor can produce a performance benefit of 4.4%, while causing an increase in energy consumption of only 1.5x, far less than the 14.6x that was thought to be necessary, and very much viable. Lingzhe Chester Cai, Aniket Deshmukh 0002, Yale N. Patt |
ISCA | 2 |
| 2024 | Alternate Path FetchabstractModern out-of-order cores rely on a large instruction supply from the processor frontend to achieve high performance. This requires building wider pipelines with more accurate branch predictors. However, scaling the pipeline width is becoming more challenging due to limitations on the number of instructions that can be renamed and branches that can be predicted in a single cycle. Moreover, mispredictions reduce the useful fetch bandwidth that can be extracted from a wider frontend. Our work, Alternate Path Fetch (APF), effectively uses a wide frontend by dividing the pipeline into two parallel sections. One processes regular instructions, and the other uses a separate pipeline to Branch Predict, Fetch, Decode, and partially Rename instructions on the alternate path of hard-to-predict (H2P) branches. The pipelines operate simultaneously using a Parallel Fetch scheme we developed. This allows APF to more efficiently utilize the bandwidth of a wider frontend without the overhead associated with building a monolithic, wider pipeline. APF improves performance by reducing the pipeline re-fill delay on branch mispredictions. Unlike other solutions that fully rename and execute instructions on both sides of a branch, we show that stopping after partial Renaming on the alternate path provides better performance through improved coverage and avoids the complexity associated with further processing. APF provides a $5 \%$ geomean speedup over an aggressive 8 -wide out-of-order core. Aniket Deshmukh 0002, Lingzhe Chester Cai, Yale N. Patt |
ISCA | 1 |
| 2024 | Timely, Efficient, and Accurate Branch PrecomputationabstractOut-of-order cores rely on high-accuracy branch predictors to supply useful instructions to the processor backend. However, there remains a large fraction of mispredictions caused by hard-to-predict (H2P) branches that modern predictors have been unable to improve. Precomputation is an alternative to prediction that speculatively executes the dependence chain of a branch earlier to override the branch predictor at Fetch time. However, prior work sacrifices H2P branch coverage and precomputation accuracy to produce timely results. Our work relaxes this timeliness constraint by using precomputation results to issue early misprediction flushes instead of overriding the branch predictor. This allows us to construct a highly accurate precomputation thread with good misprediction coverage, without sacrificing timeliness. The thread is efficient as it utilizes on-core execution resources and re-uses existing hardware for issuing early flushes. Using our Timely, Efficient, and Accurate thread for precomputation yields a 10.1% improvement in performance over an aggressive baseline OoO core. Aniket Deshmukh 0002, Lingzhe Chester Cai, Yale N. Patt |
MICRO | 1 |
| 2021 | Performance Characterization of .NET BenchmarksabstractManaged language frameworks are pervasive today, especially in modern datacenters. .NET is one such framework that is used widely in Microsoft Azure but has not been well-studied. Applications built on these frameworks have different characteristics compared to traditional SPEC-like programs due to the presence of a managed runtime. This affects the tradeoffs associated with designing hardware for such applications. Our goal is to study hardware performance bottlenecks in .NET applications. To find suitable benchmarks, we use Principal Component Analysis (PCA) to find redundancies in a set of open-source .NET and ASP.NET benchmarks and use hierarchical clustering to create representative subsets. We perform microarchitecture and application-level characterization of these subsets and show that they are significantly different from SPEC CPU17 benchmarks in branch and memory behavior, and hence merit consideration in architecture research. In-depth analysis using the Top-Down methodology reveals that .NET benchmarks are significantly more frontend bound. We also analyze the effect of managed runtime events such as JIT (Just-in-Time) compilation and GC (Garbage Collection). Among other findings, GC improves cache performance significantly and JITing could benefit from aggressive prefetching and transformation of hardware microarchitectural state to prevent frequent cold starts. As computing increasingly moves to the cloud and managed languages grow even more in popularity, it is important to consider .NET-like benchmarks in architecture studies. Aniket Deshmukh 0002, Ruihao Li 0002, Rathijit Sen, Monica Beckwith, Gagan Gupta 0003 |
ISPASS | 1 |
| 2021 | Criticality Driven FetchabstractModern OoO cores achieve high levels of performance using large instruction windows. Scaling the window size improves performance by making visible more of the parallelism present in programs. However, this leads to an exponential increase in area and power. We specify Criticality Driven Fetch (CDF), a new execution paradigm that preferentially fetches, allocates, and executes instructions on the critical path of the program. By skipping over non-critical instructions, critical instructions in the ROB can span a sequential instruction window larger than the size of the ROB. This increases the amount of parallelism that can be extracted from critical instructions, thereby improving performance. Aniket Deshmukh 0002, Yale N. Patt |
MICRO | 1 |