Vinayaka Bandishti

dblp:121/2114 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › parallelization
automatic parallelization
0.312017
Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations · IEEE Trans. Parallel Distributed Syst. 2017
Compilers and program optimization › loop optimization
stencil computation optimization
0.312017
Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations · IEEE Trans. Parallel Distributed Syst. 2017
Parallel and multicore computing › parallel program transformation
tiling
0.312017
Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations · IEEE Trans. Parallel Distributed Syst. 2017
Compilers and program optimization › loop optimization
loop tiling
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Compilers and program optimization
stencil computation
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Parallel and multicore computing
load balancing
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Parallel and multicore computing
parallel programming models
0.112012
Tiling stencil computations to maximize parallelism · SC 2012

Methods — techniques the papers use, named apart from their topics

polyhedral compilation · 0.6affine scheduling · 0.6
YearPublicationVenuePosition
2017 Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations
abstract
Most stencil computations allow tile-wise concurrent start, i.e., there always exists a face of the iteration space and a set of tiling directions such that all tiles along that face can be started concurrently. This provides load balance and maximizes parallelism. However, existing automatic tiling frameworks often choose hyperplanes that lead to pipelined start-up and load imbalance. We address this issue with a new tiling technique, called diamond tiling, that ensures concurrent start-up as well as perfect load-balance whenever possible. We first provide necessary and sufficient conditions for a set of tiling hyperplanes to allow concurrent start for programs with affine data accesses. We then provide an approach to automatically find such hyperplanes. Experimental evaluation on a 12-core Intel Westmere shows that diamond tiled code is able to outperform a tuned domain-specific stencil code generator by 10 to 40 percent, and previous compiler techniques by a factor of 1.3x to 10.1x.
Uday Bondhugula, Vinayaka Bandishti, Irshad Pananilath
IEEE Trans. Parallel Distributed Syst.2
2014 Tiling and optimizing time-iterated computations on periodic domains
abstract
This paper deals with optimizing time-iterated computations on periodic data domains. These computations are prevalent in computational sciences, particularly in partial differential equation solvers. We propose a fully automatic technique suitable for implementation in a compiler or in a domain-specific code generator for such computations. Dependence patterns on periodic data domains prevent existing algorithms from finding tiling opportunities. Our approach augments a state-of-the-art parallelization and locality-enhancing algorithm from the polyhedral framework to allow time-tiling of stencil computations on periodic domains. Experimental results on the swim SPEC CPU2000fp benchmark show a speedup of 5× and 4.2× over the highest SPEC performance achieved by native compilers on Intel Xeon and AMD Opteron multicore SMP systems, respectively. On other representative stencil computations, our scheme provides performance similar to that achieved with no periodicity, and a very high speedup is obtained over the native compiler. We also report a mean speedup of about 1.5× over a domain-specific stencil compiler supporting limited cases of periodic boundary conditions. To the best of our knowledge, it has been infeasible to manually reproduce such optimizations on swim or any other periodic stencil, especially on a data grid of two-dimensions or higher.
Uday Bondhugula, Vinayaka Bandishti, Albert Cohen 0001, Guillain Potron, Nicolas Vasilache
PACT2
2012 Tiling stencil computations to maximize parallelism
abstract
Most stencil computations allow tile-wise concurrent start, i.e., there always exists a face of the iteration space and a set of tiling hyperplanes such that all tiles along that face can be started concurrently. This provides load balance and maximizes parallelism. However, existing automatic tiling frameworks often choose hyperplanes that lead to pipelined start-up and load imbalance. We address this issue with a new tiling technique that ensures concurrent start-up as well as perfect load-balance whenever possible. We first provide necessary and sufficient conditions on tiling hyperplanes to enable concurrent start for programs with affine data accesses. We then provide an approach to find such hyperplanes. Experimental evaluation on a 12-core Intel Westmere shows that our code is able to outperform a tuned domain-specific stencil code generator by 4% to 27%, and previous compiler techniques by a factor of 2× to 10.14×.
Vinayaka Bandishti, Irshad Pananilath, Uday Bondhugula
SC1