Irshad Pananilath

dblp:121/2283 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 53% High-performance computing · 40% Performance modeling and evaluation · 6%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › parallelization
automatic parallelization
0.312017
Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations · IEEE Trans. Parallel Distributed Syst. 2017
Compilers and program optimization › loop optimization
stencil computation optimization
0.312017
Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations · IEEE Trans. Parallel Distributed Syst. 2017
Parallel and multicore computing › parallel program transformation
tiling
0.312017
Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations · IEEE Trans. Parallel Distributed Syst. 2017
Compilers and program optimization
code generation
0.212015
An Optimizing Code Generator for a Class of Lattice-Boltzmann Computations · ACM Trans. Archit. Code Optim. 2015
Compilers and program optimization › loop transformation
polyhedral compilation
0.212015
An Optimizing Code Generator for a Class of Lattice-Boltzmann Computations · ACM Trans. Archit. Code Optim. 2015
High-performance computing › scientific computing systems › computational fluid dynamics
lattice boltzmann method
0.212015
An Optimizing Code Generator for a Class of Lattice-Boltzmann Computations · ACM Trans. Archit. Code Optim. 2015
High-performance computing
scientific computing
0.212015
An Optimizing Code Generator for a Class of Lattice-Boltzmann Computations · ACM Trans. Archit. Code Optim. 2015
Compilers and program optimization › loop optimization
loop tiling
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Compilers and program optimization
stencil computation
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Parallel and multicore computing
load balancing
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Parallel and multicore computing
parallel programming models
0.112012
Tiling stencil computations to maximize parallelism · SC 2012
Performance modeling and evaluation › analytical modeling
roofline model
0.112015
An Optimizing Code Generator for a Class of Lattice-Boltzmann Computations · ACM Trans. Archit. Code Optim. 2015

Methods — techniques the papers use, named apart from their topics

polyhedral compilation · 0.6affine scheduling · 0.6time tiling · 0.4polyhedral optimization · 0.4
YearPublicationVenuePosition
2017 Diamond Tiling: Tiling Techniques to Maximize Parallelism for Stencil Computations
abstract
Most stencil computations allow tile-wise concurrent start, i.e., there always exists a face of the iteration space and a set of tiling directions such that all tiles along that face can be started concurrently. This provides load balance and maximizes parallelism. However, existing automatic tiling frameworks often choose hyperplanes that lead to pipelined start-up and load imbalance. We address this issue with a new tiling technique, called diamond tiling, that ensures concurrent start-up as well as perfect load-balance whenever possible. We first provide necessary and sufficient conditions for a set of tiling hyperplanes to allow concurrent start for programs with affine data accesses. We then provide an approach to automatically find such hyperplanes. Experimental evaluation on a 12-core Intel Westmere shows that diamond tiled code is able to outperform a tuned domain-specific stencil code generator by 10 to 40 percent, and previous compiler techniques by a factor of 1.3x to 10.1x.
Uday Bondhugula, Vinayaka Bandishti, Irshad Pananilath
IEEE Trans. Parallel Distributed Syst.3
2015 An Optimizing Code Generator for a Class of Lattice-Boltzmann Computations
abstract
The Lattice-Boltzmann method (LBM), a promising new particle-based simulation technique for complex and multiscale fluid flows, has seen tremendous adoption in recent years in computational fluid dynamics. Even with a state-of-the-art LBM solver such as Palabos, a user has to still manually write the program using library-supplied primitives. We propose an automated code generator for a class of LBM computations with the objective to achieve high performance on modern architectures. Few studies have looked at time tiling for LBM codes. We exploit a key similarity between stencils and LBM to enable polyhedral optimizations and in turn time tiling for LBM. We also characterize the performance of LBM with the Roofline performance model. Experimental results for standard LBM simulations like Lid Driven Cavity, Flow Past Cylinder, and Poiseuille Flow show that our scheme consistently outperforms Palabos—on average by up to 3× while running on 16 cores of an Intel Xeon (Sandybridge). We also obtain an improvement of 2.47× on the SPEC LBM benchmark.
Irshad Pananilath, Aravind Acharya, Vinay Vasista, Uday Bondhugula
ACM Trans. Archit. Code Optim.1
2012 Tiling stencil computations to maximize parallelism
abstract
Most stencil computations allow tile-wise concurrent start, i.e., there always exists a face of the iteration space and a set of tiling hyperplanes such that all tiles along that face can be started concurrently. This provides load balance and maximizes parallelism. However, existing automatic tiling frameworks often choose hyperplanes that lead to pipelined start-up and load imbalance. We address this issue with a new tiling technique that ensures concurrent start-up as well as perfect load-balance whenever possible. We first provide necessary and sufficient conditions on tiling hyperplanes to enable concurrent start for programs with affine data accesses. We then provide an approach to find such hyperplanes. Experimental evaluation on a 12-core Intel Westmere shows that our code is able to outperform a tuned domain-specific stencil code generator by 4% to 27%, and previous compiler techniques by a factor of 2× to 10.14×.
Vinayaka Bandishti, Irshad Pananilath, Uday Bondhugula
SC2