EDBT 2026 Demo / reviewers in the wild / expert
Aishwarya Natarajan
dblp:201/3929
· DBLP profile ↗
12ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0001-5260-4111ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analog In-Memory Computing Enhanced FPGA for High-Throughput and Energy-Efficient AccelerationabstractThe ever-growing demand for AI computing, coupled with slowing performance gains in chip manufacturing, has heightened the role of FPGA-based accelerators. FPGAs enable the implementation of application-customized parallel dataflows due to their reconfigurability, achieving high energy efficiency. However, the bit-level routing fabric on FPGAs often results in high overheads because large amounts of data must be shuttled between compute blocks and memory blocks on the FPGA. We propose enhancing FPGAs with in-memory computing macros, specifically analog Dot Product Engines based on non-volatile RRAM devices. Using the Verilog to Routing (VTR) framework, we simulate a novel 40 nm, 26.2 mm × 26.2 mm architecture and employ a custom event-driven simulator to evaluate its performance. Our design achieves 25.5 ×103TOPS/W, an average ×31.4 throughput improvement and an average ×9,380 energy efficiency improvement when compared to state-of-the-art FPGA implementations of AI models. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Xia Sheng, Giacomo Pedretti, Aman Arora 0001, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FCCM | 4 |
| 2025 | Enhancing FPGAs with Analog In-Memory Computing MacrosabstractWhile the AI computing needs are ever-increasing and the innovation in models generates tens of new architectures yearly, the performance gain from improvements in chip manufacturing has slowed down. Within this context, FPGA-based accelerators play a fundamental role. FPGAs are the backbone of specialized architectures, their reconfigurability being the key differentiation that enables an effective design space exploration. At the same time, to overcome the limitations induced by the memory bottleneck, the computing architectures community has proposed the in-memory computing paradigm: storage and computations are both performed in non-volatile memory devices. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Rand Jean, Xia Sheng, Giacomo Pedretti, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FPGA | 4 |
| 2025 | RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer AccelerationabstractTransformer models represent the cutting edge of Deep Neural Networks (DNNs) and excel in a wide range of machine learning tasks. However, processing these models demands significant computational resources and results in a substantial memory footprint. While In-memory Computing (IMC) offers promise for accelerating Vector-Matrix Multiplications (VMMs) with high computational parallelism and minimal data movement, employing it for other crucial DNN operators remains a formidable task. This challenge is exacerbated by the extensive use of complex activation functions, Softmax, and data-dependent matrix multiplications (DMMuls) within Transformer models. To address this challenge, we introduce a Reconfigurable Analog Computing Engine (RACE) by enhancing Analog Content Addressable Memories (ACAMs) to support broader operations. Based on the RACE, we propose the RACE-IT accelerator (meaning RACE for In-memory Transformers) to enable efficient analog-domain execution of all core operations of Transformer models. Given the flexibility of our proposed RACE in supporting arbitrary computations, RACE-IT is well-suited for adapting to emerging and non-traditional DNN architectures without requiring hardware modifications. We compare RACE-IT with various accelerators. Results show that RACE-IT increases performance by 453× and 15×, and reduces energy by 354× and 122× over the state-of-the-art GPUs and existing Transformer-specific IMC accelerators, respectively. Aishwarya Natarajan, Luca Buonanno, Archit Gajjar, Ron M. Roth, Sergey Serebryakov, John Moon, Omar Eldash, Jim Ignowski, Giacomo Pedretti |
ICCD | 2 |
| 2024 | Memristive Quaternary Content-Addressable Memories for Implementing Boolean FunctionsabstractIn-memory computing is, in current literature, the most common paradigm used to counteract the Von-Neumann bottleneck, proposing the use of memory elements to define complex input-output relations of the computing kernels. While in classical CMOS computing a similar paradigm can be implemented with look-up tables (LUT), this solution is power and area-hungry. This paper presents the use of Quaternary Content-Addressable Memories (QCAMs), a generalization of the Ternary Content-Addressable Memories (TCAMs), for implementing boolean functions. Content-Addressable Memories can be used as a building block for in-memory processing, using the states of the cells to define a ${\mathbb{B}^{\text{N}}} \to {\mathbb{B}^{\text{M}}}$ function which projects the search word into a new string of bits. The quaternary alphabet allows to represent a more complex function space with respect to the TCAMs while using the same number of cells, enhancing area, power consumption and latency performances achieved when representing arbitrary functions with the CAM hardware. For comparison, it can be demonstrated that QCAMs represent arbitrary Boolean functions with half the number of cells than that would be needed in a standard TCAM implementation, and a ×10 smaller area with respect to SRAM-based LUTs. Along with the table of states and a toy example where the QCAM states are used to define the product among two 2-bit precision real values, this paper presents multiple circuit schemes and encoding schemes for memristor-based QCAMs. Luca Buonanno, Giacomo Pedretti, Aishwarya Natarajan, Todd Richmond, John Moon, Rand Jean, Xia Sheng, Ron M. Roth, Jim Ignowski |
ISCAS | 4 |
| 2021 | Continuous-Time, Configurable Analog Linear System Solutions With Transconductance AmplifiersabstractThis paper addresses and experimentally demonstrates a programmable linear equation solver by analog computation. A set of differential equations using transconductance devices directly translated from circuit theory converges to the linear equation solution. These energy-efficient analog techniques are experimentally demonstrated in a configurable analog platform. The resulting analog linear equation solution circuits are effectively analog filters. The paper analyzes the algorithmic issues and analog numerical analysis issues, including accuracy, convergence time, and the interpretation of condition number for analog solutions. Jennifer Hasler, Aishwarya Natarajan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | Analog Solutions of Systems of Linear Equations on a Configurable PlatformabstractEven though analog computation is better suited for differential equation solutions (ODE, PDE), sometimes it needs to solve systems of linear equations. This discussion focuses on analog solutions of linear equation systems, implemented on a configurable platform. Digital systems rely on solving linear equations as the fundamental numerical computation. Systems of linear equations are used to solve static circuits illustrating that at least a reduced class of analog physical linear system computation should be possible. The analog approaches utilize iterative techniques, setting up a set of ODEs to solve the system of linear equations, rather than relying on matrix decompositions (e.g. LU decomposition). The approach allows for multiple potential configurable circuit approaches. A set of amplifier networks has been designed to demonstrate the solutions for different matrices. These techniques provide energy-efficient continuous-time solutions. The resulting algorithm has been studied considering the analog numerical analysis for the solution and convergence time. Aishwarya Natarajan, Jennifer Hasler |
ISCAS | 1 |
| 2020 | Built-in Self-Test of Vector Matrix Multipliers on a Reconfigurable DeviceabstractAn analog Vector Matrix Multiplier (VMM) along with its interface through Multi-Input Translinear Elements (MITE) is compiled on a reconfigurable Field Programmable Analog Array (FPAA) on a 350nm process. The paper focuses on the tuning algorithm to set the weights on the VMM as well as to produce levels of voltage near the power supply rails, Vdd, to the source-driven VMMs, by application of a wide range of input levels. The constraints during the design and implementation process, accounting for the mismatch in devices, are discussed. A significant reduction in the variation in the levels of the desired voltage has been demonstrated in the paper. Aishwarya Natarajan, Jennifer Hasler |
ISCAS | 1 |
| 2019 | Implementation of Synapses with Hodgkin Huxley Neurons on the FPAAabstractThe neuronal responses through excitatory and inhibitory synapses are implemented on a reconfigurable and programmable platform. The synaptic cleft between the pre synaptic and post synaptic neuron has been depicted as a ramp generator while the post synaptic potentials are observed from the transistor channel neuron model which emulates the ion channels in biological neurons. The synaptic strengths are tuned by modulating the charge on the floating gate devices on the hardware demonstrated through particular voltage levels and time constants. The models have been designed and built in such a way that the tools associated with the chip make it possible to build up and compile a bigger network of neurons and synapses. The experimental measurements are taken from the circuits compiled on a Field Programmable Analog Array fabricated on a 350nm process. Aishwarya Natarajan, Jennifer Hasler |
ISCAS | 1 |
| 2018 | Dynamics of Hodgkin Huxley Neuron across chips implemented on a reconfigurable platformabstractThe spiking dynamics from a Hodgkin Huxley neuron, implemented on a reconfigurable and programmable platform is presented here. The similarity between biology and silicon has been exploited to model the ion channels in the neuron, and their voltages and time constants on hardware. We demonstrate the reproducibility of the results by replicating the dynamics across different chips along with a discussion on the methodology of tuning to reproduce similar responses. The reconfigurability enables one to make use of a single primary design to obtain a variety of results. The measurements are taken from the system compiled on a Field Programmable Analog Array (FPAA) fabricated on a 350nm process. Aishwarya Natarajan, Jennifer Hasler |
ISCAS | 1 |
| 2018 | Temperature Sensitivity and Compensation on a Reconfigurable PlatformabstractThis brief investigates temperature compensation techniques for circuits and systems on a reconfigurable platform. The work demonstrates use of large-scale reconfigurable system-on-chip for reducing the variability of circuits and systems compiled on a floating gate (FG)-based field-programmable analog array (FPAA). The work presents current and voltage reference which could help in reducing the variability caused due to changes in temperature. These references are standard blocks in the Scilab/Xcos environment, which could be easily compiled on the FPAA. An FG-based current reference is then used for biasing a second-order $G_{m}-C$ bandpass filter to demonstrate the compilation and usage of these voltage/current reference in a reconfigurable fabric. The large-scale FG FPAA presented here is fabricated in 350-nm CMOS process. Sahil Shah, Hakan Toreyin, Jennifer Hasler, Aishwarya Natarajan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Using SoC FPAA and integrated simulator for implementation of circuits and systems in educationabstractAn openly available simulator is presented here, which can be used to design, characterize and simulate circuits including floating gate (FG) components and larger systems building upon these elements, while measuring the silicon data too from the Field Programmable Analog Array (FPAA), incorporated within the same tool infrastructure. A single system for simulation as well as obtaining experimental measurements from reconfigurable hardware is significantly useful especially in classroom environments. The efficiency of our simulator is shown through the comparison in performance to that of conventional simulators and its close overlap with the data measured for the same experiments from the FPAA fabricated in a 350nm process. Aishwarya Natarajan, Jennifer Hasler |
ISCAS | 1 |
| 2017 | Quantitative growth analysis of pulp necrotic tooth (post-op) using modified region growing active contour modelabstractIn the field of dentistry, prospective clinical study reports affirm the need for approximate growth analysis of endodontic tooth post treatment. There is no difference in the frequency, appearance or extent of root resorption in the teeth. It is necessary to elucidate the role of endodontic treatment in the root resorption. Differences between the two samples (radiographs), which are taken with specified period of intervals, in terms of the frequency of growth changes in treated teeth is needed to observe accurately. This study mainly aims at this requirement by utilising slightly modified region‐based growing active contour model for quantitative growth analysis of tooth. Recall radiographs of endodontic regeneration involving immature permanent teeth with pulp necrosis have been considered as input in this research. Image enhancing techniques such as dilation and erosion of mathematical morphology are then performed sequentially to emphasise the outlying pixels. Finally, the criterions which include the root length, apical diameter and the dentinal wall thickness are calibrated and given in experimental results. The visual sample results and along with its measurements proved the efficiency of the proposed algorithm. Leninisha Shanmugam, Krithika Gunasekaran, Aishwarya Natarajan, Vani Kaliaperumal |
IET Image Process. | 3 |