Ankit Wagle

dblp:231/4785 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 first-author · 6 since 2021
YearPublicationVenuePosition
2024 An ASIC Accelerator for QNN With Variable Precision and Tunable Energy Efficiency
abstract
This paper presents TULIP, a new architecture for a variable precision Quantized Neural Network (QNN) inference. It is designed with the goal of maximizing energy efficiency per classification. TULIP is constructed by arranging a collection of unique processing elements (TULIP-PEs) in a single instruction multiple data (SIMD) fashion. Each TULIP-PE contains binary neurons that are interconnected using multiplexers. Each neuron also has a small dedicated local register connected to it. The binary neurons are implemented as standard cells and used for implementing threshold functions, i.e., an inner-product and thresholding operation on its binary inputs. The neurons can be reconfigured with a single change in the control signals to implement all the standard operations used in a QNN. This paper presents novel algorithms for implementing the operations of a QNN on the TULIP-PEs in the form of a schedule of threshold functions. TULIP was implemented as an ASIC in TSMC 40nm-LP technology. A QNN accelerator that employs a conventional MAC-based arithmetic processor was also implemented in the same technology to provide a fair comparison. The results show that TULIP is 30-50X more energy-efficient than an equivalent design, without any penalty in performance, area, or accuracy. Furthermore, TULIP achieves these improvements without using traditional techniques such as voltage scaling or approximate computing. Finally, the paper also demonstrates how the run-time trade-off between accuracy and energy efficiency is done on the TULIP architecture.
Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A New Approach to Clock Skewing for Area and Power Optimization of ASICs Using Differential Flipflops and Local Clocking
abstract
A new design methodology for reducing the area and power of standard cell ASICs that uses a combination of differential flipflops and a method of deliberate clock-skewing, called local clocking (LC), is described. LC introduces clock skew without the use of extra buffers in the clock network. This is done by having some flipflops, called sources, generate clock signals for other flipflops, called targets. The method involves two key features: 1) the design of a new differential flipflop, referred to as KVFF, that is functionally identical to a double-latch edge-triggered$D$flipflop, but in addition, produces a completion signal that is a skewed version of its input clock, which is used to clock other flipflops and 2) an efficient algorithm that identifies the sources and targets involved in the new clocking scheme, with the objective of reducing area and power. These are reduced because deliberate skew introduces extra slack on the logic cones that feed the target flipflops, which is exploited by synthesis tools to reduce area and power. Furthermore, the area and power overhead of conventional methods of introducing skew, e.g., buffers, is eliminated. LC is shown to result in significant improvements in area, power, and wirelength for several, publicly available, benchmark circuits for 65 nm bulk CMOS and 28 nm FDSOI technologies. For 65 nm, the average improvement in area, power and wirelength were 27.7%, 13.4%, and 21.0%, respectively. For 28 nm FDSOI the average improvement in area, power, and wirelength were 20.0%, 10.5%, and 30.5%, respectively. In addition, this article demonstrates how LC can be used to eliminate hold time violations.
Ankit Wagle, Niranjan Kulkarni, Sarma B. K. Vrudhula
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Tunable Precision Control for Approximate Image Filtering in an In-Memory Architecture with Embedded Neurons
abstract
This paper presents a novel hardware-software co-design consisting of a Processing in-Memory (PiM) architecture with embedded neural processing elements (NPE) that are highly reconfigurable. The PiM platform and proposed approximation strategies are employed for various image filtering applications while providing the user with fine-grain dynamic control over energy efficiency, precision, and throughput (EPT). The proposed co-design can change the Peak Signal to Noise Ratio (PSNR, output quality metric for image filtering applications) from 25dB to 50dB (acceptable PSNR range for image filtering applications) without incurring any extra cost in terms of energy or latency. While switching from accurate to approximate mode of computation in the proposed co-design, the maximum improvement in energy efficiency and throughput is 2X. However, the gains in energy efficiency against a MAC-based PE array with the proposed memory platform are 3X-6X. The corresponding improvements in throughput are 2.26X-4.52X, respectively.
Ayushi Dube, Ankit Wagle, Gian Singh, Sarma B. K. Vrudhula
ICCAD2
2022 Heterogeneous FPGA Architecture Using Threshold Logic Gates for Improved Area, Power, and Performance
abstract
The flexibility of field-programmable gate arrays (FPGAs) is attributed to the reconfigurability of their basic logic elements (BLEs). Traditionally, the BLEs are comprised of one or more lookup tables (LUTs) of$n$inputs, that are designed to implement Boolean functions of$n$or fewer inputs. In an attempt to reduce the area and power consumption that comes from using LUTs, a number of complex LUT architectures have been reported. Although most of the proposed complex LUT architectures have resulted in reduced area and power, this has always been at cost of the decreased performance. This article proposes a new FPGA architecture, called threshold logic FPGA (TLFPGA), which results in significant improvement in performance, power, and area (PPA). TLFPGA is comprised of a combination of LUTs and a new type of BLE referred to as a threshold logic cell (TLC) (Muroga, 1987). Although TLCs implement a relatively small subset of Boolean functions known as threshold functions (Muroga, 1987), they require far fewer registers and multiplexers than an LUT, and are also significantly faster. This article describes the architecture of the TLFPGA and a technology mapping algorithm tailored for a TLFGA. On average, TLFPGA designs use 18% fewer registers and multiplexers, which improves the collective area of BLEs by approximately 16%, power by 14%, and performance by 5%. The improvements have been demonstrated in both 40 and 28 nm technologies for ISCAS-85 circuits as well as practical circuits, using industry-standard flows. The improvements were also demonstrated using a layout of the architecture.
Ankit Wagle, Sarma B. K. Vrudhula
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 A Novel ASIC Design Flow Using Weight-Tunable Binary Neurons as Standard Cells
abstract
In this paper, we describe a design of a mixed-signal circuit for an binary neuron (a.k.a perceptron, threshold logic gate) and a methodology for automatically embedding such cells in ASICs. The binary neuron, referred to as an FTL (flash threshold logic) uses floating gate or flash transistors whose threshold voltages serve as a proxy for the weights of the neuron. Algorithms for mapping the weights to the flash transistor threshold voltages are presented. The threshold voltages are determined to maximize both the robustness of the cell and its speed. The performance, power, and area of a single FTL cell are shown to be significantly smaller (79.4%), consume less power (61.6%), and operate faster (40.3%) compared to conventional CMOS logic equivalents. Also included are the architecture and the algorithms to program the flash devices of an FTL. The FTL cells are implemented as standard cells, and are designed to allow commercial synthesis and P&R tools to automatically use them in synthesis of ASICs. Substantial reductions in area and power without sacrificing performance are demonstrated on several ASIC benchmarks by the automatic embedding of FTL cells. The paper also demonstrates how FTL cells can be used for fixing timing errors after fabrication
Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 CIDAN: Computing in DRAM with Artificial Neurons
abstract
Numerous applications such as graph processing, cryptography, databases, bioinformatics, etc., involve the repeated evaluation of Boolean functions on large bit vectors. In-memory architectures which perform processing in memory (PIM) are tailored for such applications. This paper describes a different architecture for in-memory computation called CIDAN, that achieves a 3X improvement in performance and a 2X improvement in energy for a representative set of algorithms over the state-of-the-art in-memory architectures. CIDAN uses a new basic processing element called a TLPE, which comprises a threshold logic gate (TLG) (a.k.a artificial neuron or perceptron). The implementation of a TLG within a TLPE is equivalent to a multi-input, edge-triggered flipflop that computes a subset of threshold functions of its inputs. The specific threshold function is selected on each cycle by enabling/disabling a subset of the weights associated with the threshold function, by using logic signals. In addition to the TLG, a TLPE realizes some non-threshold functions by a sequence of TLG evaluations. An equivalent CMOS implementation of a TLPE requires a substantially higher area and power. CIDAN has an array of TLPE(s) that is integrated with a DRAM, to allow fast evaluation of any one of its set of functions on large bit vectors. Results of running several common in-memory applications in graph processing and cryptography are presented.
Gian Singh, Ankit Wagle, Sarma B. K. Vrudhula, Sunil P. Khatri
ICCD2
2020 A Configurable BNN ASIC using a Network of Programmable Threshold Logic Standard Cells
abstract
This paper presents Tulip, a new architecture for a binary neural network (BNN) that uses an optimal schedule for executing the operations of an arbitrary BNN. It was constructed with the goal of maximizing energy efficiency per classification. At the top-level, Tulip consists of a collection of unique processing elements (TULIP-PEs) that are organized in a SIMD fashion. Each Tulip- Peconsists of a small network of binary neurons, and a small amount of local memory per neuron. The unique aspect of the binary neuron is that it is implemented as a mixed-signal circuit that natively performs the inner-product and thresholding operation of an artificial binary neuron. Moreover, the binary neuron, which is implemented as a single CMOS standard cell, is reconfigurable, and with a change in a single parameter, can implement all standard operations involved in a BNN. We present novel algorithms for mapping arbitrary nodes of a BNN onto the TULIP-PEs. Tulip was implemented as an ASIC in TSMC 40nm-LP technology. To provide a fair comparison, a recently reported BNN that employs a conventional MAC-based arithmetic processor was also implemented in the same technology. The results show that Tulip is consistently 3X more energy-efficient than the conventional design, without any penalty in performance, area, or accuracy.
Ankit Wagle, Sunil P. Khatri, Sarma B. K. Vrudhula
ICCD1
2019 Embedding Binary Perceptrons in FPGA to improve Area, Power and Performance
abstract
For the flexibility of implementing any given Boolean function(s), the FPGA uses re-configurable building blocks called LUTs. The price for this reconfigurability is a large number of registers and multiplexers required to construct the FPGA. While researchers have been working on complex LUT structures to reduce the area and power for several years, most of these implementations come at the cost of performance penalty. This paper demonstrates simultaneous improvement in area, power, and performance in an FPGA by using special logic cells called Threshold Logic Cells (TLCs) (also known as binary perceptrons). The TLCs are capable of implementing a complex threshold function, which if implemented using conventional gates would require several levels of logic gates. The TLCs only require 7 SRAM cells and are significantly faster than the conventional LUTs. The implementation of the proposed FPGA architecture has been done using 28nm FDSOI standard cells and has been evaluated using ISCAS-85, ISCAS-89, and a few large industrial designs. Experiments demonstrate that the proposed architecture can be used to get an average reduction of 18.1% in configuration registers, 18.1% reduction in multiplexer count, 12.3% in Basic Logic Element (BLE) area, 16.3% in BLE power, 5.9% improvement in operating frequency, with a slight reduction in track count, routing area and routing power. The improvements are also demonstrated on the physically designed version of the architecture.
Ankit Wagle, Elham Azari, Sarma B. K. Vrudhula
ICCAD1
2019 Threshold Logic in a Flash
abstract
This paper describes a novel design of a threshold logic gate (a binary perceptron) and its implementation as a standard cell. This new cell structure, referred to as flash threshold logic (FTL), uses floating gate (flash) transistors to realize the weights associated with a threshold function. The threshold voltages of the flash transistors serve as proxy for the weights. An FTL cell can be equivalently viewed as a multi-input, edge-triggered flipflop which computes a threshold function on a clock edge. Consequently it can used in automatic synthesis of ASICs. The use of flash transistors in the FTL cell allows programming of the weights after fabrication, thereby preventing discovery of its function by a foundry or by reverse engineering. This paper focuses on the design and characteristics of the FTL cell. We present a novel method for programming the weights of an FTL cell for a specified threshold function using a modified perceptron learning algorithm. The algorithm is further extended to select weights to maximize the robustness of the design in the presence of process variations. The FTL circuit was designed in 40nm technology and simulations with layout-extracted parasitics included, demonstrate significant improvements in area (79.7%), power (61.1%), and performance (42.5%) when compared to the equivalent implementations of the same function in conventional static CMOS design. Weight selection targeting robustness is demonstrated using Monte Carlo simulations. The paper also shows how FTL cells can be used for fixing timing errors after fabrication.
Ankit Wagle, Gian Singh, Sunil P. Khatri, Sarma B. K. Vrudhula
ICCD1
2018 FPGAs with Reconfigurable Threshold Logic Gates for Improved Performance, Power and Area
abstract
This paper proposes an alternative FPGA tile structure that consists of three traditional LUTs combined with a new reconfigurable threshold logic cell (TLC). The TLC requires only 7 SRAM cells and can be configured to implement one of several threshold functions. The proposed architecture is implemented in a 28nm FDSOI process, and is evaluated on standard benchmark circuits and several large complex function blocks. The results demonstrate an average reduction of 8.9% in register count, 15.4% in multiplexer count, 7% average reduction in Basic Logic Element (BLE) area, and 8.2% average reduction in BLE power, with a maximum decrease in register count up to 64%, BLE multiplexer count up to 68%, BLE Area up to 51.6% and BLE power up to 61.6% without loss in performance. We also show a reduction of 21% in the area of a tile.
Ankit Wagle, Aykut Dengi, Sarma B. K. Vrudhula
FPL1