Vijaya Lakshmi

dblp:253/9946 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0001-7300-6284ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 In-Memory Implementation of an Approximate Adder With Reduced Latency and Error
abstract
In-memory computing has been a prominent solution to Von Neumann bottleneck that degrades the performance of a computing system. Approximate computing is widely used to improve the performance of multimedia and other applications that are error-tolerant. Approximate adders being the basic units used to design other complex units, get benefited when implemented in-memory by taking the advantages of both in-memory computing and approximate computing. In this work, we have improved the speculative carry select adder to minimize error and critical path delay by eliminating multiplexers. The proposed adder achieves less critical path, area, improved error characteristics such as error rate, normalized mean error distance and mean relative error distance when compared to the state-of-the-art approximate adders. Error rate of the proposed adder is 34.48% less than the best reported 32-bit adder with sub-adder size of 8-bit. When the proposed approximate adders are implemented in-memory using majority logic, they achieve better performance compared to the existing in-memory approximate adders. Latency of the proposed adders is observed to be a constant irrespective of adder size for a fixed sub-adder size.
Vijaya Lakshmi, Vikramkumar Pudi, John Reuben
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 Inner Product Computation In-Memory Using Distributed Arithmetic
abstract
In-memory computing using emerging technologies such as Resistive Random-Access Memory (ReRAM) has been proposed as a promising substitute for future computing applications to address the ‘von Neumann bottleneck’. Multiplication is the key component for inner product computation in every digital signal processing (DSP) application and the complexity of multipliers increases greatly with bit-width. Distributed arithmetic (DA) using look-up tables and adder-shifter module has been proposed for inner product computation to achieve multiplier-less efficient DSP architectures, particularly when one of the vectors is a constant and known in advance. Due to the memory wall, DA can be made furthermore latency and energy-efficient when implemented ‘in memory’. In this work, for the first time, we propose two design techniques to compute inner product completely in memory using DA. This is accomplished by storing the precomputed look-up table contents in a ReRAM array and implementing adder-shifter module also in the same array. The adder-shifter is implemented in memory using majority gates which are in turn realized as READ operations in the memory array. Two methods of mapping: latency-optimized and area-optimized and their comparison in terms of latency and area are presented. The proposed method-1 achieves$\approx 60$% energy savings compared to CMOS and the proposed method-2 achieves 10.59 times higher throughput compared to CMOS.
Vijaya Lakshmi, Vikramkumar Pudi, John Reuben
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 A Novel In-Memory Wallace Tree Multiplier Architecture Using Majority Logic
abstract
In-memory computing using emerging technologies such as resistive random-access memory (ReRAM) addresses the ‘von Neumann bottleneck’ and strengthens the present research impetus to overcome the memory wall. While many methods have been recently proposed to implement Boolean logic in memory, the latency of arithmetic circuits (adders and consequently multipliers) implemented as a sequence of such Boolean operations increases greatly with bit-width. Existing in-memory multipliers require$O(n^{2})$cycles which is inefficient both in terms of latency and energy. In this work, we tackle this exorbitant latency by adopting Wallace Tree multiplier architecture and optimizing the addition operation in each phase of the Wallace Tree. Majority logic primitive was used for addition since it is better than NAND/NOR/IMPLY primitives. Furthermore, high degree of gate-level parallelism is employed at the array level by executing multiple majority gates in the columns of the array. In this manner, an in-memory multiplier of$O(n.log(n))$latency is achieved which outperforms all reported in-memory multipliers. Furthermore, the proposed multiplier can be implemented in a regular transistor-accessed memory array without any major modifications to its peripheral circuitry and is also energy-efficient.
Vijaya Lakshmi, John Reuben, Vikramkumar Pudi
IEEE Trans. Circuits Syst. I Regul. Pap.1