Ganesh Bikshandi

dblp:68/2188 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 54% GPUs and heterogeneous computing · 26% High-performance computing · 11%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 72% Programming languages and type systems · 28%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
attention acceleration
0.812024
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision · NeurIPS 2024
GPUs and heterogeneous computing
GPU kernel optimization
0.812024
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision · NeurIPS 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision · NeurIPS 2024
Machine learning › Deep learning architectures and training
transformer
0.212024
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision · NeurIPS 2024
High-performance computing › parallel numerical algorithms
communication-avoiding algorithms
0.212013
Tera-scale 1D FFT with low-communication algorithm and Intel® Xeon Phi™ coprocessors · SC 2013
Compilers and program optimization › loop transformation
tiling
0.122008
Programming with tiles · PPoPP 2008
Programming for parallelism and locality with hierarchically tiled arrays · PPoPP 2006
Parallel and multicore computing › parallel programming models › distributed memory programming models
partitioned global address space
0.112009
Efficient, portable implementation of asynchronous multi-place programs · PPoPP 2009
Compilers and program optimization › memory optimization
data layout transformation
0.112008
Programming with tiles · PPoPP 2008
Parallel and multicore computing
parallel programming models
0.112008
Programming with tiles · PPoPP 2008
Compilers and program optimization › memory optimization
data locality optimization
0.112006
Programming for parallelism and locality with hierarchically tiled arrays · PPoPP 2006
Parallel and multicore computing
data-parallel programming
0.112006
Programming for parallelism and locality with hierarchically tiled arrays · PPoPP 2006
Hardware accelerators and domain-specific architectures › many-core accelerator
many-core coprocessor
0.012013
Tera-scale 1D FFT with low-communication algorithm and Intel® Xeon Phi™ coprocessors · SC 2013
Memory systems › data locality
data locality exploitation
0.012008
Programming with tiles · PPoPP 2008

Methods — techniques the papers use, named apart from their topics

warp specialization · 1.5asynchronous execution · 1.5FP8 quantization · 1.5PGAS · 0.2performance modeling · 0.2cache-aware data movement optimization · 0.2array operation overloading · 0.1
YearPublicationVenuePosition
2024 FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
abstract
Attention, as a core layer of the ubiquitous Transformer architecture, is the bottleneck for large language models and long-context applications. elaborated an approach to speed up attention on GPUs through minimizing memory reads/writes. However, it has yet to take advantage of new capabilities present in recent hardware, with FlashAttention-2 achieving only 35% utilization on the H100 GPU. We develop three main techniques to speed up attention on Hopper GPUs: exploiting asynchrony of the Tensor Cores and TMA to (1) overlap overall computation and data movement via warp-specialization and (2) interleave block-wise matmul and softmax operations, and (3) block quantization and incoherent processing that leverages hardware support for FP8 low-precision. We demonstrate that our method, FlashAttention-3, achieves speedup on H100 GPUs by 1.5-2.0$\times$ with BF16 reaching up to 840 TFLOPs/s (85\% utilization), and with FP8 reaching 1.3 PFLOPs/s. We validate that FP8 FlashAttention-3 achieves 2.6$\times$ lower numerical error than a baseline FP8 attention.
Jay Shah, Ganesh Bikshandi, Vijay Thakkar, Pradeep Ramani, Tri Dao
NeurIPS2
2013 Tera-scale 1D FFT with low-communication algorithm and Intel® Xeon Phi™ coprocessors
abstract
This paper demonstrates the first tera-scale performance of Intel® Xeon Phi™ coprocessors on 1D FFT computations. Applying a disciplined performance programming methodology of sound algorithm choice, valid performance model, and well-executed optimizations, we break the tera-flop mark on a mere 64 nodes of Xeon Phi and reach 6.7 TFLOPS with 512 nodes, which is 1.5x than achievable on a same number of Intel® Xeon® nodes. It is a challenge to fully utilize the compute capability presented by many-core wide-vector processors for bandwidth-bound FFT computation. We leverage a new algorithm, Segment-of-Interest FFT, with low inter-node communication cost, and aggressively optimize data movements in node-local computations, exploiting caches. Our coordination of low communication algorithm and massively parallel architecture for scalable performance is not limited to running FFT on Xeon Phi; it can serve as a reference for other bandwidth-bound computations and for emerging HPC systems that are increasingly communication limited.
Jongsoo Park, Ganesh Bikshandi, Karthikeyan Vaidyanathan, Ping Tak Peter Tang, Pradeep Dubey, Daehyun Kim 0001
SC2
2012 Optimization techniques for efficient HTA programs
Basilio B. Fraguela, Ganesh Bikshandi, María Jesús Garzarán, David A. Padua, Christoph von Praun
Parallel Comput.2
2009 Efficient, portable implementation of asynchronous multi-place programs
abstract
The X10 programming language is organized around the notion of places (an encapsulation of data and activities operating on the data), partitioned global address space (PGAS), and asynchronous computation and communication.
Ganesh Bikshandi, José G. Castaños, Sreedhar B. Kodali, V. Krishna Nandivada, Igor Peshansky, Vijay A. Saraswat, Sayantan Sur, Pradeep Varma, Tong Wen
PPoPP1
2009 Writing productive stencil codes with overlapped tiling
abstract
Abstract Stencil computations constitute the kernel of many scientific applications. Tiling is often used to improve the performance of stencil codes for data locality and parallelism. However, tiled stencil codes typically require shadow regions, whose management becomes a burden to programmers. In fact, it is often the case that the code required to manage these regions, and in particular their updates, is much longer than the computational kernel of the stencil. As a result, shadow regions usually impact programmers' productivity negatively. In this paper, we describeoverlapped tiling, a construct that supports shadow regions in a convenient, flexible and efficient manner in the context of the hierarchically tiled array (HTA) data type. The HTA is a class designed to express algorithms with a high degree of parallelism and/or locality as naturally as possible in terms of tiles. We discuss the syntax and implementation of overlapped HTAs as well as our experience in rewriting parallel and sequential codes using them. The results have been satisfactory in terms of both productivity and performance. For example, overlapped HTAs reduced the number of communication statements in non‐trivial codes by 78% on average while speeding them up. We also examine different implementation options and compare overlapped HTAs with previous approaches. Copyright © 2008 John Wiley & Sons, Ltd.
Ganesh Bikshandi, Basilio B. Fraguela, David A. Padua
Concurr. Comput. Pract. Exp.2
2008 Programming with tiles
abstract
The importance of tiles or blocks in scientific computing cannot be overstated. Many algorithms, both iterative and recursive, can be expressed naturally if tiles are represented explicitly. From the point of view of performance, tiling, either as a code or a data layout transformation, is one of the most effective ways to exploit locality, which is a must to achieve good performance in current computers because of the significant difference in speed between processor and memory. Furthermore, tiles are also useful to express data distribution in parallel computations. However, despite the importance of tiles, most languages do not support them directly. This gives place to bloated programs populated with numerous subscript expressions which make the code difficult to read and coding mistakes more likely.
Ganesh Bikshandi, Basilio B. Fraguela, María Jesús Garzarán, David A. Padua
PPoPP2
2006 Hierarchically tiled arrays for parallelism and locality
abstract
Parallel programming is facilitated by constructs which, unlike the widely used SPMD paradigm, provide programmers with a global view of the code and data structures. These constructs could be compiler directives containing information about data and task distribution, language extensions specifically designed for parallel computation, or classes that encapsulate parallelism. In this paper, we describe a class developed at Illinois and its Matlab implementation. This class can be used to conveniently express both parallelism and locality. A C++ implementation is now underway. Its characteristics will be reported in a future paper. We have implemented most of the NAS benchmarks using our HTA Matlab extensions and found during that HTAs enable the fast prototyping of parallel algorithms and produce programs that are easy to understand and maintain.
Ganesh Bikshandi, Daniel Hoeflinger, Gheorghe Almási 0001, Basilio B. Fraguela, María Jesús Garzarán, David A. Padua, Christoph von Praun
IPDPS2
2006 Programming for parallelism and locality with hierarchically tiled arrays
abstract
Tiling has proven to be an effective mechanism to develop high performance implementations of algorithms. Tiling can be used to organize computations so that communication costs in parallel programs are reduced and locality in sequential codes or sequential components of parallel programs is enhanced.In this paper, a data type - Hierarchically Tiled Arrays or HTAs - that facilitates the direct manipulation of tiles is introduced. HTA operations are overloaded array operations. We argue that the implementation of HTAs in sequential OO languages transforms these languages into powerful tools for the development of high-performance parallel codes and codes with high degree of locality. To support this claim, we discuss our experiences with the implementation of HTAs for MATLAB and C++ and the rewriting of the NAS benchmarks and a few other programs into HTA-based parallel form.
Ganesh Bikshandi, Daniel Hoeflinger, Gheorghe Almási 0001, Basilio B. Fraguela, María Jesús Garzarán, David A. Padua, Christoph von Praun
PPoPP1