Raghu Prabhakar

dblp:13/11061 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0003-0230-4377ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Invited: SambaNova SN40L: Unleashing Agentic AI with Dataflow
abstract
AI inference clouds are increasingly being tasked with running a diverse set of models to support interactive, agentic workloads such as LLM-powered assistants, chatbots, and autonomous agents. Efficient AI cloud inference requires producing tokens during the memory-bound decode phase at peak performance, as well as economically hosting and rapidly switching between a vast array of models. We describe the SambaNova SN40L Reconfigurable Dataflow Unit (RDU) that combines dataflow with a three-tier memory system with SRAM, HBM, and DDR. Dataflow enables peak token generation performance with aggressive fusion of large compute graphs into a single kernel. HBM and high-capacity DDR drastically lower hardware footprint to host and serve trillions of parameters at scale. The SN40L RDU produce tokens over $\mathbf{3} \times$ faster, consume $3 \times$ less energy, and lower model hosting costs by up to $19 \times$ over a DGX H100.
Raghu Prabhakar, Pushkar Nandkar, Darshan Gandhi, Nasim Farahini, Håkan Zeffer
DAC1
2024 SambaNova SN40L RDU: Breaking the Barrier of Trillion+ Parameter Scale Gen AI Computing
abstract
•Configurable as a systolic array or a SIMD vector unit with M lanes •BF16, FP32, INT32, and INT8 compute data types, configurable storage data types •Arithmetic, Logical, and Bitwise operations •A cross-lane reduction tree (blue) to reduce along the vectorized dimension •Tail stage provides transcendental functions, casting, and stochastic rounding capabilities
Raghu Prabhakar
HCS1
2024 Revet: A Language and Compiler for Dataflow Threads
abstract
Spatial dataflow architectures such as reconfigurable dataflow accelerators (RDA) can provide much higher performance and efficiency than CPUs and GPUs. In particular, vectorized reconfigurable dataflow accelerators (vRDA) in recent literature represent a design point that enhances the efficiency of dataflow architectures with vectorization. Today, vRDAs can be exploited using either hard-coded kernels or MapReduce languages like Spatial, which cannot vectorize data-dependent control flow. In contrast, CPUs and GPUs can be programmed using general-purpose threaded abstractions. The ideal combination would be the generality of a threaded programming model coupled with the efficient execution model of a vRDA. We introduce Revet: a programming model, compiler, and execution model that lets threaded applications run efficiently on vRDAs. The Revet programming language uses threads to support a broader range of applications than prior parallel-patterns approaches, and our MLIR-based compiler lowers this language to a generic dataflow backend that operates on streaming tensors. Finally, we show that mapping threads to dataflow out-performs GPUs, the current state-of-the-art for threaded accelerators, by 3.8×.
Alexander Rucker, Shiv Sundram, Coleman Smith, Matthew Vilim, Raghu Prabhakar, Fredrik Kjolstad, Kunle Olukotun
HPCA5
2024 SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
abstract
Monolithic large language models (LLMs) like GPT-4 have paved the way for modern generative AI applications. Training, serving, and maintaining monolithic LLMs at scale, however, remains prohibitively expensive and challenging. The disproportionate increase in compute-to-memory ratio of modern AI accelerators have created a memory wall, necessitating new methods to deploy AI. Recent research has shown that a composition of many smaller expert models, each with several orders of magnitude fewer parameters, can match or exceed the capabilities of monolithic LLMs. Composition of Experts (CoE) is a modular approach that lowers the cost and complexity of training and serving. However, this approach presents two key challenges when using conventional hardware: (1) without fused operations, smaller models have lower operational intensity, which makes high utilization more challenging to achieve; and (2) hosting a large number of models can be either prohibitively expensive or slow when dynamically switching between them. In this paper, we describe how combining CoE, streaming dataflow, and a three-tier memory system scales the AI memory wall. We describe Samba-CoE, a CoE system with 150 experts and a trillion total parameters. We deploy Samba-CoE on the SambaNova SN40L Reconfigurable Dataflow Unit (RDU) -a commercial dataflow accelerator architecture that has been codesigned for enterprise inference and training applications. The chip introduces a new three-tier memory system with on-chip distributed SRAM, on-package HBM, and off-package DDR DRAM. A dedicated inter-RDU network enables scaling up and out over multiple sockets. We demonstrate speedups ranging from 2× to 13× on various benchmarks running on eight RDU sockets compared with an unfused baseline. We show that for CoE inference deployments, the 8-socket RDU Node reduces machine footprint by up to 19 ×, speeds up model switching time by 15× to 31×, and achieves an overall speedup of 3.7× over a DGX H100 and 6.6× over a DGX A100.
Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi, Mingran Wang, Kejie Zhang, Tianren Gao, Angela Wang, Yongning Sheng, Joshua Brot, Denis Sokolov, Apurv Vivek, Calvin Leung, Arjun Sabnis, Jiayu Bai, Tuowen Zhao, Mark Gottscho, Mark Luttrell, Manish K. Shah, Zhengyu Chen 0002, Kaizhao Liang, Swayambhoo Jain, Urmish Thakker, Dawei Huang, Sumti Jairath, Kevin J. Brown, Kunle Olukotun
MICRO1
2021 SambaNova SN10 RDU: Accelerating Software 2.0 with Dataflow
abstract
The following is intended to outline our general product direction at this time. There is no obligation to update this presentation and the Company’s products and direction are always subject to change. This presentation is intended for information purposes only and may not be relied upon for any purchasing, partnership, or other decisions.
Raghu Prabhakar, Sumti Jairath
HCS1
2021 Capstan: A Vector RDA for Sparsity
abstract
This paper proposes Capstan: a scalable, parallel-patterns-based, reconfigurable dataflow accelerator (RDA) for sparse and dense tensor applications. Instead of designing for one application, we start with common sparse data formats, each of which supports multiple applications. Using a declarative programming model, Capstan supports application-independent sparse iteration and memory primitives that can be mapped to vectorized, high-performance hardware. We optimize random-access sparse memories with configurable out-of-order execution to increase SRAM random-access throughput from 32% to 80%.
Alexander Rucker, Matthew Vilim, Tian Zhao 0001, Yaqi Zhang 0001, Raghu Prabhakar, Kunle Olukotun
MICRO5
2019 Scalable interconnects for reconfigurable spatial architectures
abstract
Recent years have seen the increased adoption of Coarse-Grained Reconfigurable Architectures (CGRAs) as flexible, energy-efficient compute accelerators. Obtaining performance using spatial architectures while supporting diverse applications requires a flexible, high-bandwidth interconnect. Because modern CGRAs support vector units with wide datapaths, designing an interconnect that balances dynamism, communication granularity, and programmability is a challenging task.
Yaqi Zhang 0001, Alexander Rucker, Matthew Vilim, Raghu Prabhakar, William Hwang, Kunle Olukotun
ISCA4
2018 Spatial: a language and compiler for application accelerators
abstract
Industry is increasingly turning to reconfigurable architectures like FPGAs and CGRAs for improved performance and energy efficiency. Unfortunately, adoption of these architectures has been limited by their programming models. HDLs lack abstractions for productivity and are difficult to target from higher level languages. HLS tools are more productive, but offer an ad-hoc mix of software and hardware abstractions which make performance optimizations difficult.
David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang 0001, Stefan Hadjis, Ruben Fiszel, Tian Zhao 0001, Luigi Nardi, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
PLDI3
2017 Plasticine: A Reconfigurable Architecture For Parallel Paterns
abstract
Reconfigurable architectures have gained popularity in recent years as they allow the design of energy-efficient accelerators. Fine-grain fabrics (e.g. FPGAs) have traditionally suffered from performance and power inefficiencies due to bit-level reconfigurable abstractions. Both fine-grain and coarse-grain architectures (e.g. CGRAs) traditionally require low level programming and suffer from long compilation times. We address both challenges with Plasticine, a new spatially reconfigurable architecture designed to efficiently execute applications composed of parallel patterns. Parallel patterns have emerged from recent research on parallel programming as powerful, high-level abstractions that can elegantly capture data locality, memory access patterns, and parallelism across a wide range of dense and sparse applications.
Raghu Prabhakar, Yaqi Zhang 0001, David Koeplinger, Matthew Feldman, Tian Zhao 0001, Stefan Hadjis, Ardavan Pedram, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA1
2016 Generating Configurable Hardware from Parallel Patterns
abstract
In recent years the computing landscape has seen an increasing shift towards specialized accelerators. Field programmable gate arrays (FPGAs) are particularly promising for the implementation of these accelerators, as they offer significant performance and energy improvements over CPUs for a wide class of applications and are far more flexible than fixed-function ASICs. However, FPGAs are difficult to program. Traditional programming models for reconfigurable logic use low-level hardware description languages like Verilog and VHDL, which have none of the productivity features of modern software languages but produce very efficient designs, and low-level software languages like C and OpenCL coupled with high-level synthesis (HLS) tools that typically produce designs that are far less efficient. Functional languages with parallel patterns are a better fit for hardware generation because they provide high-level abstractions to programmers with little experience in hardware design and avoid many of the problems faced when generating hardware from imperative languages. In this paper, we identify two important optimizations for using parallel patterns to generate efficient hardware: tiling and metapipelining. We present a general representation of tiled parallel patterns, and provide rules for automatically tiling patterns and generating metapipelines. We demonstrate experimentally that these optimizations result in speedups up to 39.4× on a set of benchmarks from the data analytics domain.
Raghu Prabhakar, David Koeplinger, Kevin J. Brown, HyoukJoong Lee, Christopher De Sa, Christoforos E. Kozyrakis, Kunle Olukotun
ASPLOS1
2016 Automatic Generation of Efficient Accelerators for Reconfigurable Hardware
abstract
Acceleration in the form of customized datapaths offer large performance and energy improvements over general purpose processors. Reconfigurable fabrics such as FPGAs are gaining popularity for use in implementing application-specific accelerators, thereby increasing the importance of having good high-level FPGA design tools. However, current tools for targeting FPGAs offer inadequate support for high-level programming, resource estimation, and rapid and automatic design space exploration. We describe a design framework that addresses these challenges. We introduce a new representation of hardware using parameterized templates that captures locality and parallelism information at multiple levels of nesting. This representation is designed to be automatically generated from high-level languages based on parallel patterns. We describe a hybrid area estimation technique which uses template-level models and design-level artificial neural networks to account for effects from hardware place-and-route tools, including routing overheads, register and block RAM duplication, and LUT packing. Our runtime estimation accounts for off-chip memory accesses. We use our estimation capabilities to rapidly explore a large space of designs across tile sizes, parallelization factors, and optional coarse-grained pipelining, all at multiple loop levels. We show that estimates average 4.8% error for logic resources, 6.1% error for runtimes, and are 279 to 6533 times faster than a commercial high-level synthesis tool. We compare the best-performing designs to optimized CPU code running on a server-grade 6 core processor and show speedups of up to 16.7×.
David Koeplinger, Raghu Prabhakar, Yaqi Zhang 0001, Christina Delimitrou, Christoforos E. Kozyrakis, Kunle Olukotun
ISCA2
2012 Compilation and architecture support for customized vector instruction extension
abstract
Vectorization has been commonly employed in modern processors. In this work we identify the opportunities to explore customized vector instructions and build an automatic compilation flow to efficiently identify those instructions. A composable vector unit (CVU) is proposed to support a large number of customized vector instructions with small area overhead. The results show that our approach achieves an average 1.41X speedup over the state-of-art vector ISA. We also observe a large area gain (around 11.6X) over the dedicated ASIC-based design.
Jason Cong, Mohammad Ali Ghodrat, Michael Gill, Hui Huang 0001, Bin Liu 0006, Raghu Prabhakar, Glenn Reinman, Marco Vitanza
ASP-DAC6
2012 CUDA-For-Clusters: A System for Efficient Execution of CUDA Kernels on Multi-core Clusters
Raghu Prabhakar, R. Govindarajan, Matthew J. Thazhuthaveetil
Euro-Par1
2012 Static and dynamic co-optimizations for blocks mapping in hybrid caches
abstract
In this paper, a combined static and dynamic scheme is proposed to optimize the block placement for endurance and energy-efficiency in a hybrid SRAM and STT-RAM cache. With the proposed scheme, STT-RAM endurance is maximized while performance is maintained. We use the compiler to provide static hints to guide initial data placement, and use the hardware to correct the hints based on the run-time cache behavior. Experimental results show that the combined scheme improves the endurance by 23.9x and 5.9x compared to pure static and pure dynamic optimizations respectively. Furthermore, the system energy can be reduced by 17% compared to pure dynamic optimization through minimizing STT-RAM writes.
Yuting Chen 0003, Jason Cong, Hui Huang 0001, Chunyue Liu, Raghu Prabhakar, Glenn Reinman
ISLPED5
2012 Towards layout-friendly high-level synthesis
abstract
There are two prominent problems with technology scaling: increasing design complexity and more challenges with interconnect design, including routability. High-level synthesis has been proposed to solve the complexity problem by raising the abstraction level. In this paper, we share our vision that high-level synthesis can potentially help the routability problem as well. We show that many interconnect problems that occur in layout can be avoided or mitigated by adopting a layout-friendly RTL architecture generated from high-level synthesis. We also evaluate some structural metrics that can be used to estimate the routability impact of design decisions in high-level synthesis. Experimental results have demonstrated correlations between the metrics and the routability of the resulting design.
Jason Cong, Bin Liu 0006, Guojie Luo, Raghu Prabhakar
ISPD4