Vinay Gangadhar

dblp:162/9963 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 39% Processor architecture and microarchitecture · 19% GPUs and heterogeneous computing · 12%

Topics — the 12 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › spatial architecture
dataflow accelerator
0.612022
The Mozart reuse exposed dataflow processor for AI and beyond: industrial product · ISCA 2022
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.612022
The Mozart reuse exposed dataflow processor for AI and beyond: industrial product · ISCA 2022
Processor architecture and microarchitecture
dataflow architecture
0.312017
Stream-Dataflow Acceleration · ISCA 2017
Parallel and multicore computing › parallel programming models › dataflow programming
dataflow execution model
0.212015
Exploring the potential of heterogeneous von neumann/dataflow execution models · ISCA 2015
GPUs and heterogeneous computing
GPU architecture
0.212015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Electronic design automation › hardware verification and test › hardware verification
RTL simulation
0.212015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Performance modeling and evaluation
simulation
0.212015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing
energy-efficient architecture
0.112016
Pushing the limits of accelerator efficiency while retaining programmability · HPCA 2016
Electronic design automation › hardware verification and test
design validation
0.112015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Electronic design automation
hardware verification and test
0.112015
Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015
Energy-efficient computing › low-power design
low-power processor design
0.112015
Exploring the potential of heterogeneous von neumann/dataflow execution models · ISCA 2015
Processor architecture and microarchitecture › speculation
speculation control
0.112015
Exploring the potential of heterogeneous von neumann/dataflow execution models · ISCA 2015

Methods — techniques the papers use, named apart from their topics

dataflow execution · 0.3SIMD · 0.3GPGPU comparison · 0.3scratchpads · 0.2configurable spatial architecture · 0.2DMA · 0.2explicit-dataflow execution · 0.2data-dependence graph extraction · 0.2RTL implementation · 0.2OpenCL · 0.2
YearPublicationVenuePosition
2022 The Mozart reuse exposed dataflow processor for AI and beyond: industrial product
abstract
In this paper we introduce the Mozart Processor, which implements a new processing paradigm called Reuse Exposed Dataflow (RED). RED is a counterpart to existing execution models of Von-Neumann, SIMT, Dataflow, and FPGA. Dataflow and data reuse are the fundamental architecture primitives in RED, implemented with mechanisms for inter-worker communication and synchronization. The paper defines the processor architecture, the details of the microarchitecture, chip implementation, software stack development, and performance results. The architecture's goal is to achieve near-CPU like flexibility while having ASIC-like efficiency for a large-class of data-intensive workloads. An additional goal was software maturity --- have large coverage of applications immediately, avoiding the need for a long-drawn hand-tuning software development phase. The architecture was defined with this software-maturity/compiler friendliness in mind. In short, the goal was to do to GPUs, what GPUs did to CPUs --- i.e. be a better solution for a large range of workloads, while preserving flexibility and programmability. The chip was implemented with HBM and PCIe interfaces and taken to production on a 16nm TSMC FFC process. For ML inference tasks with batch-size=4, Mozart is integer factors better than state-of-the-art GPUs even while being nearly 2 technology nodes behind. We conclude with a set of lessons learned, the unique challenges of a clean-slate architecture in a commercial setting, and pointers for uncovered research problems.
Karthikeyan Sankaralingam, Tony Nowatzki, Vinay Gangadhar, Preyas Shah, William Galliher, Ziliang Guo, Jitu Khare, Deepak Vijay, Poly Palamuttam, Maghawan Punde, Alex Tan, Vijayraghavan Thiruvengadam, Rongyi Wang, Shunmiao Xu
ISCA3
2021 Mozart: Designing for Software Maturity and the Next Paradigm for Chip Architectures
abstract
Where does AI hardware/software stand today? 1. The computational diversity needed to support AI is increasing2. The software user experience expectations is increasing3. GPU software maturity* is unrivalled in completeness and hence allows near complete dominance among AI industry deployment and researchers.4. This support for model diversity is fuelling these trends and increasing GPU adoption!* NVIDIA DL stack - cuDNN, TensorRT, etc.
Karthikeyan Sankaralingam, Tony Nowatzki, Greg Wright, Poly Palamuttam, Jitu Khare, Vinay Gangadhar, Preyas Shah
HCS6
2017 Stream-Dataflow Acceleration
abstract
Demand for low-power data processing hardware continues to rise inexorably. Existing programmable and "general purpose" solutions (eg. SIMD, GPGPUs) are insufficient, as evidenced by the order-of-magnitude improvements and industry adoption of application and domain-specific accelerators in important areas like machine learning, computer vision and big data. The stark tradeoffs between efficiency and generality at these two extremes poses a difficult question: how could domain-specific hardware efficiency be achieved without domain-specific hardware solutions?
Tony Nowatzki, Vinay Gangadhar, Newsha Ardalani, Karthikeyan Sankaralingam
ISCA2
2016 Pushing the limits of accelerator efficiency while retaining programmability
abstract
The waning benefits of device scaling have caused a push towards domain specific accelerators (DSAs), which sacrifice programmability for efficiency. While providing huge benefits, DSAs are prone to obsoletion due to domain volatility, have recurring design and verification costs, and have large area footprints when multiple DSAs are required in a single device. Because of the benefits of generality, this work explores how far a programmable architecture can be pushed, and whether it can come close to the performance, energy, and area efficiency of a DSA-based approach. Our insight is that DSAs employ common specialization principles for concurrency, computation, communication, data-reuse and coordination, and that these same principles can be exploited in a programmable architecture using a composition of known microarchitectural mechanisms. Specifically, we propose and study an architecture called LSSD, which is composed of many low-power and tiny cores, each having a configurable spatial architecture, scratchpads, and DMA. Our results show that a programmable, specialized architecture can indeed be competitive with a domain-specific approach. Compared to four prominent and diverse DSAs, LSSD can match the DSAs' 10× to 150× speedup over an OOO core, with only up to 4x more area and power than a single DSA, while retaining programmability.
Tony Nowatzki, Vinay Gangadhar, Karthikeyan Sankaralingam, Greg Wright
HPCA2
2015 MIAOW: An open source GPGPU
Vinay Gangadhar, Raghuraman Balasubramanian, Mario Drumond, Ziliang Guo, Jai Menon 0003, Cherin Joseph, Robin Prakash, Sharath Prasad, Pradip Valathol, Karthikeyan Sankaralingam
Hot Chips Symposium1
2015 Exploring the potential of heterogeneous von neumann/dataflow execution models
abstract
General purpose processors (GPPs), from small inorder designs to many-issue out-of-order, incur large power overheads which must be addressed for future technology generations. Major sources of overhead include structures which dynamically extract the data-dependence graph or maintain precise state. Considering irregular workloads, current specialization approaches either heavily curtail performance, or provide simply too little benefit. Interestingly, well known explicit-dataflow architectures eliminate these overheads by directly executing the data-dependence graph and eschewing instruction-precise recoverability. However, even after decades of research, dataflow architectures have yet to come into prominence as a solution. We attribute this to a lack of effective control speculation and the latency overhead of explicit communication, which is crippling for certain codes.
Tony Nowatzki, Vinay Gangadhar, Karthikeyan Sankaralingam
ISCA2
2015 Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU
abstract
Graphic processing unit (GPU)-based general-purpose computing is developing as a viable alternative to CPU-based computing in many domains. Today’s tools for GPU analysis include simulators like GPGPU-Sim, Multi2Sim, and Barra. While useful for modeling first-order effects, these tools do not provide a detailed view of GPU microarchitecture and physical design. Further, as GPGPU research evolves, design ideas and modifications demand detailed estimates of impact on overall area and power. Fueled by this need, we introduce MIAOW (Many-core Integrated Accelerator Of Wisconsin), an open-source RTL implementation of the AMD Southern Islands GPGPU ISA, capable of running unmodified OpenCL-based applications. We present our design motivated by our goals to create a realistic, flexible, OpenCL-compatible GPGPU, capable of emulating a full system. We first explore if MIAOW is realistic and then use four case studies to show that MIAOW enables the following: physical design perspective to “traditional” microarchitecture, new types of research exploration, and validation/calibration of simulator-based characterization of hardware. The findings and ideas are contributions in their own right, in addition to MIAOW’s utility as a tool for others’ research.
Raghuraman Balasubramanian, Vinay Gangadhar, Ziliang Guo, Chen-Han Ho, Cherin Joseph, Jai Menon 0003, Mario Drumond, Robin Paul, Sharath Prasad, Pradip Valathol, Karthikeyan Sankaralingam
ACM Trans. Archit. Code Optim.2