Vijay Narayanan

dblp:64/3690 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0008-8433-963XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Assessing the Performance of Analog Training for Transfer Learning
abstract
Analog in-memory computing is a next-generation computing paradigm that promises fast, parallel, and energy-efficient deep learning training and transfer learning (TL). However, achieving this promise has remained elusive due to a lack of suitable training algorithms. Analog memory devices exhibit asymmetric and non-linear switching behavior in addition to device-to-device variation, meaning that most, if not all, of the current off-the-shelf training algorithms cannot achieve good training outcomes. Also, recently introduced algorithms have enjoyed limited attention, as they require bi-directionally switching devices of unrealistically high symmetry and precision and are highly sensitive. A new algorithm chopped TTv2 (c-TTv2), has been introduced, which leverages the chopped technique to address many of the challenges mentioned above. In this paper, we assess the performance of the c-TTv2 algorithm for analog TL using a Swin-ViT model on a subset of the CIFAR100 dataset. We also investigate the robustness of our algorithm to changes in some device specifications, including weight transfer noise, symmetry point skew, and symmetry point variability.
Omobayode Fagbohungbe, Corey Lammie, Malte J. Rasch, Takashi Ando, Tayfun Gokmen, Vijay Narayanan
ISCAS6
2024 Analog AI as a Service: A Cloud Platform for In-Memory Computing
abstract
This paper introduces the Analog AI Cloud Composer platform, a service that allows users to access Analog In-Memory Computing (AIMC) simulation and computing resources over the cloud. We introduce the concept of an Analog AI as a Service (AAaaS). AIMC offers a novel approach for decreasing both the latency and energy usage associated with Deep Neural Network (DNN) inference and training. This platform democratizes access to AIMC computing, making it available to a broader audience, including researchers, developers, and businesses. Emphasizing a user-friendly, no-code approach, AAaaS integrates the Analog Hardware Acceleration Kit (AIHWKit) simulation platform within a fully managed cloud environment. We discuss the architecture of the Analog AI Cloud Composer (AAICC), focusing on its key services such as inference, training, and AIMC hardware access. The platform's design, grounded in cloud services and guidelines, ensures a secure, data-centric user experience with robust control and validation mechanisms.
Kaoutar El Maghraoui, Kim Tran, Kurtis Ruby, Borja Godoy, Jordan Murray, Manuel Le Gallo-Bourdeau, Todd Deshane, Pablo Gonzalez, Diego Moreda, Hadjer Benmeziane, Corey Lammie, Julian Büchel, Malte J. Rasch, Abu Sebastian, Vijay Narayanan
SSE15
2023 Architectures and Circuits for Analog-memory-based Hardware Accelerators for Deep Neural Networks (Invited)
abstract
Analog non-volatile memory (NVM)-based accelerators for Deep Neural Networks (DNNs) can achieve high-throughput and energy-efficient multiply-accumulate (MAC) operations by taking advantage of massively parallelized analog compute, implemented with Ohm's law and Kirchhoff's current law on arrays of resistive memory devices. Competitive end-to-end DNN accuracies can be obtained, provided that weights are accurately programmed onto NVM devices and MAC operations are sufficiently linear. In this paper, we report architectural and circuit advances for such Analog NVM-based accelerators. We describe a highly heterogeneous and programmable accelerator architecture for DNN inference that combines analog NVM memory-array “Tiles” for weight-stationary, energy-efficient MAC operations, together with heterogeneous special-function Compute-Cores for auxiliary digital computation. Massively parallel vectors of neuron-activation data are exchanged over short distances using a dense and efficient circuit-switched 2D mesh, enabling a wide range of DNN workloads, including CNNs, LSTMs, and Transformers. We also show a 14-nm inference chip consisting of multiple$\mathbf{512}\times \mathbf{512}$arrays of Phase Change Memory (PCM) devices which implements multiple DNN benchmarks using such a circuit-switched 2D mesh.
Hsinyu Tsai, Pritish Narayanan, Shubham Jain 0004, Stefano Ambrogio, Kohji Hosokawa, Masatoshi Ishii, Charles Mackin, Ching-Tzu Chen, Atsuya Okazaki, Akiyo Nomura, Irem Boybat, Ramachandran Muralidhar, Martin M. Frank, Takeo Yasuda, Alexander M. Friz, Yasuteru Kohda, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Stanislaw Wozniak, Jose Luquin, Vijay Narayanan, Geoffrey W. Burr
ISCAS22
2023 A Heterogeneous and Programmable Compute-In-Memory Accelerator Architecture for Analog-AI Using Dense 2-D Mesh
abstract
We introduce a highly heterogeneous and programmable compute-in-memory (CIM) accelerator architecture for deep neural network (DNN) inference. This architecture combines spatially distributed CIM memory array “tiles” for weight-stationary, energy-efficient multiply–accumulate (MAC) operations, together with heterogeneous special-function compute cores for auxiliary digital computation. Massively parallel vectors of neuron activation data are exchanged over short distances using a dense and efficient circuit-switched 2-D mesh, offering full end-to-end support for a wide range of DNN workloads, including CNNs, long-short-term-memory (LSTM), and transformers. We discuss the design of the “analog fabric”—the 2-D grid of tiles and compute cores interconnected by the 2-D mesh—and address the efficiency in both mapping of DNNs onto the hardware and in pipelining of various DNN workloads across a range of batch sizes. We show, for the first time, system-level assessments using projected component parameters for a realistic “analog AI” system, based on dense crossbar arrays of low-power nonvolatile analog memory elements, while incorporating a single common analog fabric design that can scale to large networks by introducing data transport between multiple analog AI chips. Our performance estimates for several networks, including large LSTM and bidirectional encoder representations from transformers (BERT), show highly competitive throughput while offering$40\times $–$140\times $higher energy efficiency than NVIDIA A100—thus illustrating the strong promise of analog AI and the proposed architecture for DNN inference applications.
Shubham Jain 0004, Hsinyu Tsai, Ching-Tzu Chen, Ramachandran Muralidhar, Irem Boybat, Martin M. Frank, Stanislaw Wozniak, Milos Stanisavljevic, Praneet Adusumilli, Pritish Narayanan, Kohji Hosokawa, Masatoshi Ishii, Vijay Narayanan, Geoffrey W. Burr
IEEE Trans. Very Large Scale Integr. Syst.14
2020 Hardware and Software Co-optimization for the Initialization Failure of the ReRAM-based Cross-bar Array
abstract
Recent advances in deep neural network demand more than millions of parameters to handle and mandate the high-performance computing resources with improved efficiency. The cross-bar array architecture has been considered as one of the promising deep learning architectures that shows a significant computing gain over the conventional processors. To investigate the feasibility of the architecture, we examine non-idealities and their impact on the performance. Specifically, we study the impact of failed cells due to the initialization process of the resistive memory-based cross-bar array. Unlike the conventional memory array, individual memory elements cannot be rerouted and, thus, may have a critical impact on model accuracy. We categorize the possible failures and propose hardware implementation that minimizes catastrophic failures. Such hardware optimization bounds the possible logical value of the failed cells and allows us to compensate for the loss of accuracy via off-line training. By introducing the random weight defects during the training, we show that the model becomes more resilient on the device initialization failures, therefore, less prone to degrade the inference performance due to the failed devices. Our study sheds light on the hardware and software co-optimization procedure to cope with potentially catastrophic failures in the cross-bar array.
Chun-Chen Yeh, Vijay Narayanan, Jungwook Choi
ACM J. Emerg. Technol. Comput. Syst.4
2010 Characterizing the soft error vulnerability of multicores running multithreaded applications
abstract
Multicores have become the platform of choice across all market segments. Cost-eective protection against soft er-rors is important in these environments, due to the need to move to lower technology generations and the exploding number of transistors on a chip. While multicores oer the exibility of varying the number of application threads and the number of cores on which they run, the reliability im-pact of choosing one conguration over another is unclear. Our study reveals that the reliability costs vary dramatically between congurations and being unaware could lead to a sub-optimal choice.
Niranjan Soundararajan, Anand Sivasubramaniam, Vijay Narayanan
SIGMETRICS3
2009 3D GPU architecture using cache stacking: Performance, cost, power and thermal analysis
abstract
Graphics Processing Units (GPUs) offer tremendous computational and processing power. The architecture requires high communication bandwidth and lower latency between computation units and caches. 3D die-stacking technology is a promising approach to meet such requirements. To the best of our knowledge no other study has investigated the implementation of 3D technology in GPUs. In this paper, we study the impact of stacking caches using the 3D technology on GPU performance. We also investigate the benefits of using 3D stacked MRAM on GPUs. Our work includes cost, power, and thermal analysis of the proposed architectural designs. Our results show a 53% geometric mean performance speedup for iso-cycle time architectures and about 19% for iso-cost architectures.
Ahmed Al-Maashri, Guangyu Sun 0003, Xiangyu Dong 0001, Vijay Narayanan, Yuan Xie 0001
ICCD4