Harideep Nair

dblp:243/2845 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0008-8015-4339ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
abstract
The increasing complexity of deep neural networks (DNNs) poses significant challenges for edge inference deployment due to resource and power constraints of edge devices. Recent works on unary-based matrix multiplication hardware aim to leverage data sparsity and low-precision values to enhance hardware efficiency. However, the adoption and integration of such unary hardware into commercial deep learning accelerators (DLA) remain limited due to processing element (PE) array dataflow differences. This work presents Tempus Core, a convolution core with highly scalable unary-based PE array comprising of tub (temporal-unary-binary) multipliers that seamlessly integrates with the NVDLA (NVIDIA's open-source DLA for accelerating CNNs) while maintaining dataflow compliance and boosting hardware efficiency. Analysis across various datapath granularities shows that for INT8 precision in 45nm CMOS, Tempus Core's PE cell unit (PCU) yields 59.3% and 15.3% reductions in area and power consumption, respectively, over NVDLA's CMAC unit. Considering a 16x16 PE array in Tempus Core, area and power improves by 75% and 62%, respectively, while delivering 5x and 4x iso-area throughput improvements for INT8 and INT4 precisions. Post-place and route analysis of Tempus Core's PCU shows that the 16x4 PE array for INT4 precision in 45nm CMOS requires only 0.017mm2die area and consumes only 6.2mW of total power. We demonstrate that area-power efficient unary-based hardware can be seamlessly integrated into conventional DLAs, paving the path for efficient unary hardware for edge AI inference.
Prabhu Vellaisamy, Harideep Nair, Thomas Kang, Yichen Ni, Haoyang Fan, Jeff Chen, R. D. (Shawn) Blanton, John Paul Shen
DATE2
2024 Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
abstract
General Matrix Multiply (GEMM) units, consisting of multiply-accumulate (MAC) arrays, perform bulk of the computation in deep learning (DL). Recent work has proposed a novel MAC design, Bit-Pragmatic (PRA), capable of dynamically exploiting bit sparsity. This work presents OzMAC (Omit-zero-MAC), a modified re-implementation of PRA, but extends beyond earlier works by performing rigorous post-synthesis evaluation against binary MAC design across multiple bitwidths and clock frequencies using TSMC N5 process node to assess commercial implementation potential. We demonstrate the existence of high bit sparsity in eight pretrained INT8 DL workloads and show that 8-bit OzMAC improves all three metrics of area, power, and energy significantly by 21%, 70%, and 28%, respectively. Similar improvements are achieved when scaling data precisions (4, 8, 16 bits) and clock frequencies (0.5 GHz, 1 GHz, 1.5 GHz). For the 8-bit OzMAC, scaling its frequency to normalize the throughput, it still achieves 30% improvement on both power and energy.
Harideep Nair, Prabhu Vellaisamy, Tsung-Han Lin, Perry H. Wang, R. D. (Shawn) Blanton, John Paul Shen
VLSI-SoC1
2023 tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
abstract
General matrix multiplication (GEMM) is a ubiqui-tous computing kernel/algorithm for data processing in diverse applications, including artificial intelligence (AI) and deep learning (DL). Recent shift towards edge computing has inspired GEMM architectures based on unary computing, which are predominantly stochastic and rate-coded systems. This paper proposes a novel GEMM architecture based on temporal-coding, called tuGEMM, that performs exact computation. We introduce two variants of tuGEMM, serial and parallel, with distinct area/power-latency trade-offs. Post-synthesis Power-Performance-Area (PPA) in 45 nm CMOS are reported for 2-bit, 4-bit, and 8-bit computations. The designs illustrate significant advantages in area-power efficiency over state-of-the-art stochastic unary systems especially at low precisions, e.g. incurring just 0.03 mm2and 9 mW for 4 bits, and 0.01 mm2and 4 mW for 2 bits. This makes tuGEMM ideal for power constrained mobile and edge devices performing always-on real-time sensory processing.
Harideep Nair, Prabhu Vellaisamy, Albert Chen 0002, Joseph Finn, Anna Li, Manav Trivedi, John Paul Shen
ISCAS1
2021 Unsupervised Clustering of Time Series Signals Using Neuromorphic Energy-Efficient Temporal Neural Networks
abstract
Unsupervised time series clustering is a challenging problem with diverse industrial applications such as anomaly detection, bio-wearables, etc. These applications typically involve small, low-power devices on the edge that collect and process real-time sensory signals. State-of-the-art time-series clustering methods perform some form of loss minimization that is extremely computationally intensive from the perspective of edge devices. In this work, we propose a neuromorphic approach to unsupervised time series clustering based on Temporal Neural Networks that is capable of ultra low-power, continuous online learning. We demonstrate its clustering performance on a subset of UCR Time Series Archive datasets. Our results show that the proposed approach either outperforms or performs similarly to most of the existing algorithms while being far more amenable for efficient hardware implementation. Our hardware assessment analysis shows that in 7 nm CMOS the proposed architecture, on average, consumes only about 0.005 mm2die area and 22 μW power and can process each signal with about 5 ns latency.
Shreyas Chaudhari, Harideep Nair, José M. F. Moura, John Paul Shen
ICASSP2
2019 Freeflow Core: Enhancing Performance of In-Order Cores with Energy Efficiency
abstract
The Out-of-Order (OoO) superscalar core design has been widely adopted for high performance computing. It exploits both instruction level parallelism (ILP) and memory level parallelism (MLP) to speedup the program's execution. However, due to unresolved data dependencies among instructions, exploiting ILP becomes at times difficult and make the OoO core idle, thus reducing its energy-efficiency. Our study focuses on selective exploitation of inherent ILP present in the program. Our proposed architecture, Freeflow Core (FC), focuses on discovering the selective opportunities for conversion of instruction guided execution to data guided execution. Such selective mechanism improves the performance without incurring substantial energy budget overheads. Memory-related instructions are handled in FC by memory-access pipeline and compute instructions are handled by compute pipeline. Giving priority to load/store instructions has been one of the known techniques to improve performance. However, less importance has been given to non-ready instructions which may block the head of the compute pipeline. We observe that such instructions are mainly those whose producers are unresolved in the memory-access pipeline. Hence, FC detects instructions that are dependent on unresolved memory instructions and guides them to a dedicated independent execution path. This segregation enables the younger ready instructions to free flow through the compute pipeline. Our evaluations show that FC outperforms InO and state-of-the-art Load Slice Core (LSC) both in performance and energy efficiency metrics.
Raj Kumar Choudhary, Newton, Harideep Nair, Rishabh Rawat, Virendra Singh
ICCD3