Akshay Bhosale

dblp:254/0921 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-1264-5274ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 The Future of Instruction-Level Parallelism (ILP)
abstract
High-performance processors have long used instruction-level parallelism (ILP) to achieve performance, but in the past decade processor vendors have dramatically increased their reliance upon this technique. We therefore take another look at the theoretical limits of ILP, in order to evaluate challenges and opportunities for processor architectures. Using the dynamic dependency graph of general-purpose workloads, we find that the upper bound on ILP is surprisingly close to the IPC capabilities of current state-of-the-art cores. Our results suggest that there may be as little as a decade of further scaling on current trends before hardware capabilities exceed the ILP bound.
Alexandra W. Chadwick, Márton Erdos, Utpal Bora 0003, Akshay Bhosale, Bob Lytton, Giacomo Gabrielli, Timothy M. Jones 0001
ISPASS4
2025 LoopFrog: In-Core Hint-Based Loop Parallelization
abstract
To scale ILP, designers build deeper and wider out-of-order superscalar CPUs.However, this approach incurs quadratic scaling complexity, area, and energy costs with each generation.While small loops may benefit from increased instruction-window sizes and large loops may see speedups via thread-level parallelism across cores, there remains unexploited medium-granularity parallelism.We propose LoopFrog to tap into this potential by bringing thread-level speculation schemes into the modern era.LoopFrog runs multiple loop iterations from a single thread in parallel within the microarchitecture.The core can spawn future loop iterations as new microarchitectural threadlets based on compiler-inserted hints, which can leapfrog execution beyond the parent thread's instruction window, exposing a new, medium-grained parallelism, orthogonal to traditional ILP and TLP.LoopFrog monitors data dependencies between executing threadlets, forwards data for true dependencies and squashes speculative threadlets on ordering violations.Using an LLVM-based compiler to insert hints, we achieve a geometric mean loop speedup of 43%, translating to whole-program speedups of 9.2% on SPEC CPU 2006 and 9.5% on SPEC CPU 2017 benchmarks, with only modest area and power overheads.
Márton Erdos, Utpal Bora 0001, Akshay Bhosale, Bob Lytton, Ali Mustafa Zaidi, Alexandra W. Chadwick, Giacomo Gabrielli, Timothy M. Jones 0001
MICRO3
2025 Ghost Threading: Helper-Thread Prefetching for Real Systems
Akshay Bhosale, Utpal Bora 0001, Alexandra W. Chadwick, Márton Erdos, Giacomo Gabrielli, Timothy M. Jones 0001
MICRO2
2024 Recurrence Analysis for Automatic Parallelization of Subscripted Subscripts
abstract
Introducing correct and optimal OpenMP parallelization directives in applications is a challenge. To parallelize a loop in an input application code automatically, parallelizing compilers need to disprove dependences with respect to variables across iterations of the loop. Performing such dependence analysis in the presence of index arrays or subscripted subscripts - a[b[i]] - has long been a challenge for automatic parallelizers. Loops with subscripted subscripts can be parallelized if the subscript array is known to possess a property such as monotonicity. This paper presents a compiletime algorithm that can analyze complex recurrence relations and determine irregular or intermittent monotonicity of one-dimensional and monotonicity of multi-dimensional subscript arrays. The new algorithm builds on a prior approach that is capable of analyzing simple recurrence relations and determining monotonic one-dimensional subscript arrays. Experimental results show that automatic parallelizers equipped with our new analysis techniques can substantially improve the performance of ten out of twelve or 83.33% of the benchmarks evaluated, 25--33.33% more than possible with state-of-the-art compile-time automatic parallelization techniques.
Akshay Bhosale, Rudolf Eigenmann
PPoPP1
2021 On the automatic parallelization of subscripted subscript patterns using array property analysis
abstract
Parallelizing loops with subscripted subscript patterns at compile-time has long been a challenge for automatic parallelizers. In the class of irregular applications that we have analyzed, the presence of subscripted subscript patterns was one of the primary reasons why a significant number of loops could not be automatically parallelized. Loops with such patterns can be parallelized, if the subscript array or the expression in which the subscript array appears possess certain properties, such as monotonicity. The information required to prove the existence of these properties is often present in the application code itself. This suggests that their automatic detection may be feasible. In this paper, we present an algebra for representing and reasoning about subscript array properties, and we discuss a compile-time algorithm, based on symbolic range aggregation, that can prove monotonicity and parallelize key loops. We show that this algorithm can produce significant performance gains, not only in the parallelized loops, but also in the overall applications.
Akshay Bhosale, Rudolf Eigenmann
ICS1