Makesh Chandran

dblp:268/2321 · also Tarun Makesh Chandran · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 80% Hardware accelerators and domain-specific architectures · 20%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › in-memory computing
in-cache acceleration
0.412020
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration · MICRO 2020
Memory systems › processing-in-memory
LUT-based PIM
0.412020
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration · MICRO 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412020
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration · MICRO 2020
Memory systems › processing-in-memory
processing-in-cache
0.412020
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration · MICRO 2020
Memory systems
processing-in-memory
0.412020
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration · MICRO 2020
Machine learning › Deep learning architectures and training
neural network inference
0.112020
Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration · MICRO 2020

Methods — techniques the papers use, named apart from their topics

systolic dataflow · 0.9lookup table · 0.9bitline computing · 0.9
YearPublicationVenuePosition
2020 PSB-RNN: A Processing-in-Memory Systolic Array Architecture using Block Circulant Matrices for Recurrent Neural Networks
abstract
Recurrent Neural Networks (RNNs) are widely used in Natural Language Processing (NLP) applications as they inherently capture contextual information across spatial and temporal dimensions. Compared to other classes of neural networks, RNNs have more weight parameters as they primarily consist of fully connected layers. Recently, several techniques such as weight pruning, zero-skipping, and block circulant compression have been introduced to reduce the storage and access requirements of RNN weight parameters. In this work, we present a ReRAM crossbar based processing-in-memory (PIM) architecture with systolic dataflow incorporating block circulant compression for RNNs. The block circulant compression decomposes the operations in a fully connected layer into a series of Fourier transforms and point-wise operations resulting in reduced space and computational complexity. We formulate the Fourier transform and point-wise operations into in-situ multiply-and-accumulate (MAC) operations mapped to ReRAM crossbars for high energy efficiency and throughput. We also incorporate systolic dataflow for communication within the crossbar arrays, in contrast to broadcast and multicast communications, to further improve energy efficiency. The proposed architecture achieves average improvements in compute efficiency of 44× and 17× over a custom FPGA architecture and conventional crossbar based architecture implementations, respectively.
Nagadastagiri Challapalle, Sahithi Rampalli, Makesh Chandran, Gurpreet S. Kalsi, Sreenivas Subramoney, Jack Sampson, Narayanan Vijaykrishnan
DATE3
2020 Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration
abstract
This paper presents a Look-Up Table (LUT) based Processing-In-Memory (PIM) technique with the potential for running Neural Network inference tasks. We implement a bitline computing free technique to avoid frequent bitline accesses to the cache sub-arrays and thereby considerably reducing the memory access energy overhead. LUT in conjunction with the compute engines enables sub-array level parallelism while executing complex operations through data lookup which otherwise requires multiple cycles. Sub-array level parallelism and systolic input data flow ensure data movement to be confined to the SRAM slice.Our proposed LUT based PIM methodology exploits substantial parallelism using look-up tables, which does not alter the memory structure/organization, that is, preserving the bit-cell and peripherals of the existing SRAM monolithic arrays. Our solution achieves 1.72× higher performance and 3.14x lower energy as compared to a state-of-the-art processing-in-cache solution. Sub-array level design modifications to incorporate LUT along with the compute engines will increase the overall cache area by 5.6%. We achieve 3.97x speedup w.r.t neural network systolic accelerator with a similar area. The re-configurable nature of the compute engines enables various neural network operations and thereby supporting sequential networks (RNNs) and transformer models. Our quantitative analysis demonstrates 101×, 3× faster execution and 91×, 11× energy efficient than CPU and GPU respectively while running the transformer model, BERT-Base.
Akshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Rangachar Srinivasa, Makesh Chandran, Kamlesh R. Pillai, Om Ji Omer, Narayanan Vijaykrishnan, Sreenivas Subramoney
MICRO4