Gian Singh

dblp:17/7491 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-6649-8487ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Compact, Low Power Transprecision ALU for Smart Edge Devices
abstract
Transprecision computing (TC) is a promising approach for energy-efficient machine learning (ML) computation on resource-constrained platforms. This work presents a novel ASIC design of a Transprecision Arithmetic and Logic Unit (TALU) that can support multiple number formats: Posit, Floating Point (FP), and Integer (INT) data with variable bitwidth of 8, 16, and 32 bits. Additionally, TALU can be reconfigured in runtime to support TC without overprovisioning the hardware. Posit is a new number format, gaining traction for ML computations, producing similar accuracy in lower bitwidth than FP representation. This paper thus proposes a novel algorithm for decoding Posit for energy-efficient computation. TALU implementation achieves a 54.6× reduction in power consumption and 19.8× reduction in the area as compared to a state-of-the-art unified MAC unit (UMAC) [1] for Posit and FP computation. Experimental results on an ML compute kernel executed on a Vector Processor of TALUs integrated with a RISC-V processor achieves about 2× improvement in energy efficiency and similar throughput as compared to a state-of-the-art TC-based vector processor.
Ayushi Dube, Gian Singh, Sarma B. K. Vrudhula
ISLPED2
2024 A DRAM-based Near-Memory Architecture for Accelerated and Energy-Efficient Execution of Transformers
abstract
Transformers-based language models have achieved remarkable accuracy in various NLP tasks, employing self-attention mechanisms primarily based on matrix multiplication. However, their significant size leads to data movement issues, causing latency and energy efficiency challenges in conventional Von-Neumann systems. To mitigate these issues, several in-memory and near-memory architectures have been proposed. This paper introduces PACT-3D, a near-memory architecture featuring novel computing units integrated with DRAM banks. PACT-3D significantly reduces latency by 1.7 × and improves energy efficiency by 18.7 × compared to state-of-the-art near-memory architectures.
Gian Singh, Sarma B. K. Vrudhula
ACM Great Lakes Symposium on VLSI1
2024 Hardware-Software Co-Design for Path Planning by Drones
abstract
This work consists of two main components: designing a hardware-software co-design, MT+, for adapting the Mikami-Tabuchi algorithm for on-board path planning by drones in a 3D environment; and development of a specialized custom hardware accelerator CDU, as a part of MT+, for parallel collision detection. Collision detection is a performance bottleneck in path planning. MT+reduces the delay in path planning without using any heuristic. A comparative analysis between the state-of-the-art path planning algorithm A* and Mikami-Tabuchi is performed to show that Mikami-Tabuchi is faster than A* in typical real-world environments. In custom-generated environments, path planning using Mikami-Tabuchi shows a latency improvement of 1.7× across varying average sizes of obstacles and 2.7× across varying obstacle density over state-of-the-art path planning algorithm, A*. Further, the experiments show that the co-design achieves speedups over a full software implementation on CPU, averaging between 10% to 60% across different densities and sizes of obstacles. CDU area and power overheads are negligible against a conventional single-core processor.
Ayushi Dube, Omkar Patil, Gian Singh, Nakul Gopalan, Sarma B. K. Vrudhula
IROS3
2024 A High Throughput, Energy-Efficient Architecture for Variable Precision Computing in DRAM
abstract
DRAM-based near-memory architectures are recognized for their ability to deliver substantial energy efficiency and throughput to execute data-intensive tasks. However, the inherent limitations regarding area, power, and timing within DRAM allow the integration of only primitive processing elements with limited operations and application support. This paper introduces a near-memory processing architecture based on DRAM featuring a novel computing unit termed the neuron processing element (NPE). NPEs are capable of performing multiple arithmetic, logical, and predicate operations. With a well-defined instruction set, the NPEs can be programmed to support standard data formats for floating point and fixed point precision used in different AI/ML and signal processing applications. They can be dynamically reconfigured to switch operations during run-time without increasing overall latency or power consumption. The NPEs have a small area and power footprint compared to conventional MAC units and other functionally equivalent implementations, making them suitable for integration with DRAM without compromising its organization or timing constraints. Furthermore, this paper shows a substantial improvement in latency and energy consumption compared to prior in-memory architectures and demonstrates the efficacy of the proposed architecture for the acceleration of neural network inference.
Gian Singh, Ayushi Dube, Sarma B. K. Vrudhula
VLSI-SoC1
2024 An ASIC Accelerator for QNN With Variable Precision and Tunable Energy Efficiency
abstract
This paper presents TULIP, a new architecture for a variable precision Quantized Neural Network (QNN) inference. It is designed with the goal of maximizing energy efficiency per classification. TULIP is constructed by arranging a collection of unique processing elements (TULIP-PEs) in a single instruction multiple data (SIMD) fashion. Each TULIP-PE contains binary neurons that are interconnected using multiplexers. Each neuron also has a small dedicated local register connected to it. The binary neurons are implemented as standard cells and used for implementing threshold functions, i.e., an inner-product and thresholding operation on its binary inputs. The neurons can be reconfigured with a single change in the control signals to implement all the standard operations used in a QNN. This paper presents novel algorithms for implementing the operations of a QNN on the TULIP-PEs in the form of a schedule of threshold functions. TULIP was implemented as an ASIC in TSMC 40nm-LP technology. A QNN accelerator that employs a conventional MAC-based arithmetic processor was also implemented in the same technology to provide a fair comparison. The results show that TULIP is 30-50X more energy-efficient than an equivalent design, without any penalty in performance, area, or accuracy. Furthermore, TULIP achieves these improvements without using traditional techniques such as voltage scaling or approximate computing. Finally, the paper also demonstrates how the run-time trade-off between accuracy and energy efficiency is done on the TULIP architecture.
Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 PARAG: PIM Architecture for Real-Time Acceleration of GCNs
abstract
Graph Convolutional Networks (GCNs) have successfully incorporated deep learning to graph structures for social network analysis, bio-informatics, etc. The execution pattern of GCNs is a hybrid of graph processing and neural networks which poses unique and significant challenges for hardware implementation. Graph processing involves a large amount of irregular memory access with little computation whereas processing of neural networks involves a large number of operations with regular memory access. Existing graph processing and neural network accelerators are therefore inefficient for computing GCNs. This paper presents Parag, processing in memory (PIM) architecture for GCN computation. It consists of customized logic with minuscule computing units called Neural Processing Elements (NPEs) interfaced to each bank of the DRAM to support parallel graph processing and neural network computation. It utilizes the massive internal parallelism of DRAM to accelerate the GCN execution with high energy efficiency. Simulation results for inference of GCN over standard datasets show a latency and energy reduction by three orders of magnitude over a CPU implementation. When compared to a state-of-the-art PIM architecture, PARAG achieves on an average 4x reduction in latency and 4.23x reduction in the energy-delay-product (EDP).
Gian Singh, Sanmukh R. Kuppannagari, Sarma B. K. Vrudhula
HiPC1
2022 Tunable Precision Control for Approximate Image Filtering in an In-Memory Architecture with Embedded Neurons
abstract
This paper presents a novel hardware-software co-design consisting of a Processing in-Memory (PiM) architecture with embedded neural processing elements (NPE) that are highly reconfigurable. The PiM platform and proposed approximation strategies are employed for various image filtering applications while providing the user with fine-grain dynamic control over energy efficiency, precision, and throughput (EPT). The proposed co-design can change the Peak Signal to Noise Ratio (PSNR, output quality metric for image filtering applications) from 25dB to 50dB (acceptable PSNR range for image filtering applications) without incurring any extra cost in terms of energy or latency. While switching from accurate to approximate mode of computation in the proposed co-design, the maximum improvement in energy efficiency and throughput is 2X. However, the gains in energy efficiency against a MAC-based PE array with the proposed memory platform are 3X-6X. The corresponding improvements in throughput are 2.26X-4.52X, respectively.
Ayushi Dube, Ankit Wagle, Gian Singh, Sarma B. K. Vrudhula
ICCAD3
2022 A Novel ASIC Design Flow Using Weight-Tunable Binary Neurons as Standard Cells
abstract
In this paper, we describe a design of a mixed-signal circuit for an binary neuron (a.k.a perceptron, threshold logic gate) and a methodology for automatically embedding such cells in ASICs. The binary neuron, referred to as an FTL (flash threshold logic) uses floating gate or flash transistors whose threshold voltages serve as a proxy for the weights of the neuron. Algorithms for mapping the weights to the flash transistor threshold voltages are presented. The threshold voltages are determined to maximize both the robustness of the cell and its speed. The performance, power, and area of a single FTL cell are shown to be significantly smaller (79.4%), consume less power (61.6%), and operate faster (40.3%) compared to conventional CMOS logic equivalents. Also included are the architecture and the algorithms to program the flash devices of an FTL. The FTL cells are implemented as standard cells, and are designed to allow commercial synthesis and P&R tools to automatically use them in synthesis of ASICs. Substantial reductions in area and power without sacrificing performance are demonstrated on several ASIC benchmarks by the automatic embedding of FTL cells. The paper also demonstrates how FTL cells can be used for fixing timing errors after fabrication
Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 CIDAN: Computing in DRAM with Artificial Neurons
abstract
Numerous applications such as graph processing, cryptography, databases, bioinformatics, etc., involve the repeated evaluation of Boolean functions on large bit vectors. In-memory architectures which perform processing in memory (PIM) are tailored for such applications. This paper describes a different architecture for in-memory computation called CIDAN, that achieves a 3X improvement in performance and a 2X improvement in energy for a representative set of algorithms over the state-of-the-art in-memory architectures. CIDAN uses a new basic processing element called a TLPE, which comprises a threshold logic gate (TLG) (a.k.a artificial neuron or perceptron). The implementation of a TLG within a TLPE is equivalent to a multi-input, edge-triggered flipflop that computes a subset of threshold functions of its inputs. The specific threshold function is selected on each cycle by enabling/disabling a subset of the weights associated with the threshold function, by using logic signals. In addition to the TLG, a TLPE realizes some non-threshold functions by a sequence of TLG evaluations. An equivalent CMOS implementation of a TLPE requires a substantially higher area and power. CIDAN has an array of TLPE(s) that is integrated with a DRAM, to allow fast evaluation of any one of its set of functions on large bit vectors. Results of running several common in-memory applications in graph processing and cryptography are presented.
Gian Singh, Ankit Wagle, Sarma B. K. Vrudhula, Sunil P. Khatri
ICCD1
2019 Threshold Logic in a Flash
abstract
This paper describes a novel design of a threshold logic gate (a binary perceptron) and its implementation as a standard cell. This new cell structure, referred to as flash threshold logic (FTL), uses floating gate (flash) transistors to realize the weights associated with a threshold function. The threshold voltages of the flash transistors serve as proxy for the weights. An FTL cell can be equivalently viewed as a multi-input, edge-triggered flipflop which computes a threshold function on a clock edge. Consequently it can used in automatic synthesis of ASICs. The use of flash transistors in the FTL cell allows programming of the weights after fabrication, thereby preventing discovery of its function by a foundry or by reverse engineering. This paper focuses on the design and characteristics of the FTL cell. We present a novel method for programming the weights of an FTL cell for a specified threshold function using a modified perceptron learning algorithm. The algorithm is further extended to select weights to maximize the robustness of the design in the presence of process variations. The FTL circuit was designed in 40nm technology and simulations with layout-extracted parasitics included, demonstrate significant improvements in area (79.7%), power (61.1%), and performance (42.5%) when compared to the equivalent implementations of the same function in conventional static CMOS design. Weight selection targeting robustness is demonstrated using Monte Carlo simulations. The paper also shows how FTL cells can be used for fixing timing errors after fabrication.
Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula
ICCD2