VLDB 2026 Research / reviewers in the wild / expert
Vinay Gangadhar
dblp:162/9963
· DBLP profile ↗
7ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 39% Processor architecture and microarchitecture · 19% GPUs and heterogeneous computing · 12% |
Topics — the 12 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › spatial architecture
dataflow accelerator |
0.6 | 1 | 2022 | The Mozart reuse exposed dataflow processor for AI and beyond: industrial product · ISCA 2022 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.6 | 1 | 2022 | The Mozart reuse exposed dataflow processor for AI and beyond: industrial product · ISCA 2022 |
Processor architecture and microarchitecture
dataflow architecture |
0.3 | 1 | 2017 | Stream-Dataflow Acceleration · ISCA 2017 |
Parallel and multicore computing › parallel programming models › dataflow programming
dataflow execution model |
0.2 | 1 | 2015 | Exploring the potential of heterogeneous von neumann/dataflow execution models · ISCA 2015 |
GPUs and heterogeneous computing
GPU architecture |
0.2 | 1 | 2015 | Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015 |
Electronic design automation › hardware verification and test › hardware verification
RTL simulation |
0.2 | 1 | 2015 | Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015 |
Performance modeling and evaluation
simulation |
0.2 | 1 | 2015 | Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015 |
Energy-efficient computing
energy-efficient architecture |
0.1 | 1 | 2016 | Pushing the limits of accelerator efficiency while retaining programmability · HPCA 2016 |
Electronic design automation › hardware verification and test
design validation |
0.1 | 1 | 2015 | Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015 |
Electronic design automation
hardware verification and test |
0.1 | 1 | 2015 | Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPU · ACM Trans. Archit. Code Optim. 2015 |
Energy-efficient computing › low-power design
low-power processor design |
0.1 | 1 | 2015 | Exploring the potential of heterogeneous von neumann/dataflow execution models · ISCA 2015 |
Processor architecture and microarchitecture › speculation
speculation control |
0.1 | 1 | 2015 | Exploring the potential of heterogeneous von neumann/dataflow execution models · ISCA 2015 |
Methods — techniques the papers use, named apart from their topics
dataflow execution · 0.3SIMD · 0.3GPGPU comparison · 0.3scratchpads · 0.2configurable spatial architecture · 0.2DMA · 0.2explicit-dataflow execution · 0.2data-dependence graph extraction · 0.2RTL implementation · 0.2OpenCL · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | The Mozart reuse exposed dataflow processor for AI and beyond: industrial productabstractIn this paper we introduce the Mozart Processor, which implements a new processing paradigm called Reuse Exposed Dataflow (RED). RED is a counterpart to existing execution models of Von-Neumann, SIMT, Dataflow, and FPGA. Dataflow and data reuse are the fundamental architecture primitives in RED, implemented with mechanisms for inter-worker communication and synchronization. The paper defines the processor architecture, the details of the microarchitecture, chip implementation, software stack development, and performance results. The architecture's goal is to achieve near-CPU like flexibility while having ASIC-like efficiency for a large-class of data-intensive workloads. An additional goal was software maturity --- have large coverage of applications immediately, avoiding the need for a long-drawn hand-tuning software development phase. The architecture was defined with this software-maturity/compiler friendliness in mind. In short, the goal was to do to GPUs, what GPUs did to CPUs --- i.e. be a better solution for a large range of workloads, while preserving flexibility and programmability. The chip was implemented with HBM and PCIe interfaces and taken to production on a 16nm TSMC FFC process. For ML inference tasks with batch-size=4, Mozart is integer factors better than state-of-the-art GPUs even while being nearly 2 technology nodes behind. We conclude with a set of lessons learned, the unique challenges of a clean-slate architecture in a commercial setting, and pointers for uncovered research problems. Karthikeyan Sankaralingam, Tony Nowatzki, Vinay Gangadhar, Preyas Shah, William Galliher, Ziliang Guo, Jitu Khare, Deepak Vijay, Poly Palamuttam, Maghawan Punde, Alex Tan, Vijayraghavan Thiruvengadam, Rongyi Wang, Shunmiao Xu |
ISCA | 3 |
| 2021 | Mozart: Designing for Software Maturity and the Next Paradigm for Chip ArchitecturesabstractWhere does AI hardware/software stand today? 1. The computational diversity needed to support AI is increasing2. The software user experience expectations is increasing3. GPU software maturity* is unrivalled in completeness and hence allows near complete dominance among AI industry deployment and researchers.4. This support for model diversity is fuelling these trends and increasing GPU adoption!* NVIDIA DL stack - cuDNN, TensorRT, etc. Karthikeyan Sankaralingam, Tony Nowatzki, Greg Wright, Poly Palamuttam, Jitu Khare, Vinay Gangadhar, Preyas Shah |
HCS | 6 |
| 2017 | Stream-Dataflow AccelerationabstractDemand for low-power data processing hardware continues to rise inexorably. Existing programmable and "general purpose" solutions (eg. SIMD, GPGPUs) are insufficient, as evidenced by the order-of-magnitude improvements and industry adoption of application and domain-specific accelerators in important areas like machine learning, computer vision and big data. The stark tradeoffs between efficiency and generality at these two extremes poses a difficult question: how could domain-specific hardware efficiency be achieved without domain-specific hardware solutions? Tony Nowatzki, Vinay Gangadhar, Newsha Ardalani, Karthikeyan Sankaralingam |
ISCA | 2 |
| 2016 | Pushing the limits of accelerator efficiency while retaining programmabilityabstractThe waning benefits of device scaling have caused a push towards domain specific accelerators (DSAs), which sacrifice programmability for efficiency. While providing huge benefits, DSAs are prone to obsoletion due to domain volatility, have recurring design and verification costs, and have large area footprints when multiple DSAs are required in a single device. Because of the benefits of generality, this work explores how far a programmable architecture can be pushed, and whether it can come close to the performance, energy, and area efficiency of a DSA-based approach. Our insight is that DSAs employ common specialization principles for concurrency, computation, communication, data-reuse and coordination, and that these same principles can be exploited in a programmable architecture using a composition of known microarchitectural mechanisms. Specifically, we propose and study an architecture called LSSD, which is composed of many low-power and tiny cores, each having a configurable spatial architecture, scratchpads, and DMA. Our results show that a programmable, specialized architecture can indeed be competitive with a domain-specific approach. Compared to four prominent and diverse DSAs, LSSD can match the DSAs' 10× to 150× speedup over an OOO core, with only up to 4x more area and power than a single DSA, while retaining programmability. Tony Nowatzki, Vinay Gangadhar, Karthikeyan Sankaralingam, Greg Wright |
HPCA | 2 |
| 2015 | MIAOW: An open source GPGPU
Vinay Gangadhar, Raghuraman Balasubramanian, Mario Drumond, Ziliang Guo, Jai Menon 0003, Cherin Joseph, Robin Prakash, Sharath Prasad, Pradip Valathol, Karthikeyan Sankaralingam |
Hot Chips Symposium | 1 |
| 2015 | Exploring the potential of heterogeneous von neumann/dataflow execution modelsabstractGeneral purpose processors (GPPs), from small inorder designs to many-issue out-of-order, incur large power overheads which must be addressed for future technology generations. Major sources of overhead include structures which dynamically extract the data-dependence graph or maintain precise state. Considering irregular workloads, current specialization approaches either heavily curtail performance, or provide simply too little benefit. Interestingly, well known explicit-dataflow architectures eliminate these overheads by directly executing the data-dependence graph and eschewing instruction-precise recoverability. However, even after decades of research, dataflow architectures have yet to come into prominence as a solution. We attribute this to a lack of effective control speculation and the latency overhead of explicit communication, which is crippling for certain codes. Tony Nowatzki, Vinay Gangadhar, Karthikeyan Sankaralingam |
ISCA | 2 |
| 2015 | Enabling GPGPU Low-Level Hardware Explorations with MIAOW: An Open-Source RTL Implementation of a GPGPUabstractGraphic processing unit (GPU)-based general-purpose computing is developing as a viable alternative to CPU-based computing in many domains. Today’s tools for GPU analysis include simulators like GPGPU-Sim, Multi2Sim, and Barra. While useful for modeling first-order effects, these tools do not provide a detailed view of GPU microarchitecture and physical design. Further, as GPGPU research evolves, design ideas and modifications demand detailed estimates of impact on overall area and power. Fueled by this need, we introduce MIAOW (Many-core Integrated Accelerator Of Wisconsin), an open-source RTL implementation of the AMD Southern Islands GPGPU ISA, capable of running unmodified OpenCL-based applications. We present our design motivated by our goals to create a realistic, flexible, OpenCL-compatible GPGPU, capable of emulating a full system. We first explore if MIAOW is realistic and then use four case studies to show that MIAOW enables the following: physical design perspective to “traditional” microarchitecture, new types of research exploration, and validation/calibration of simulator-based characterization of hardware. The findings and ideas are contributions in their own right, in addition to MIAOW’s utility as a tool for others’ research. Raghuraman Balasubramanian, Vinay Gangadhar, Ziliang Guo, Chen-Han Ho, Cherin Joseph, Jai Menon 0003, Mario Drumond, Robin Paul, Sharath Prasad, Pradip Valathol, Karthikeyan Sankaralingam |
ACM Trans. Archit. Code Optim. | 2 |