EDBT 2026 Demo / reviewers in the wild / expert
John Reuben
dblp:146/1596
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-7891-4975ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Smart Sensing of Multi-bit Resistive Memory using a Single ReferenceabstractEmerging Non-Volatile Memories (NVMs) are increasingly being researched for multi-bit capabilities which can be exploited not only for storage (increased memory density) but also for novel applications like in-memory computing e.g. matrix vector multiplication. Such applications require readout circuits, i.e. Sense Amplifiers (SAs) which convert the resistance of the memory cell to digital data. In this work, we propose such a Sense Amplifier (SA) to distinguish between the four states of a ReRAM cell. Unlike conventional multi-bit sensing schemes which require three references to sense four states, the proposed sensing technique requires only a single voltage reference. The proposed SA uses current comparison to sense the first bit and based on the value of the sensed bit, the crucial currents of the SA are manipulated to sense the second bit. The proposed SA is very energy efficient when compared to the state-of-the-art (consuming 82.4 fJ in 130 nm process node) while avoiding the need to generate multiple voltage/current references on-chip. Requiring only 25 transistors and a single $V_{R E F}$, the proposed SA is one of the most compact SA among others in the state-of-the-art. John Reuben, Dietmar Fey |
DSD | 1 |
| 2025 | In-Memory Implementation of an Approximate Adder With Reduced Latency and ErrorabstractIn-memory computing has been a prominent solution to Von Neumann bottleneck that degrades the performance of a computing system. Approximate computing is widely used to improve the performance of multimedia and other applications that are error-tolerant. Approximate adders being the basic units used to design other complex units, get benefited when implemented in-memory by taking the advantages of both in-memory computing and approximate computing. In this work, we have improved the speculative carry select adder to minimize error and critical path delay by eliminating multiplexers. The proposed adder achieves less critical path, area, improved error characteristics such as error rate, normalized mean error distance and mean relative error distance when compared to the state-of-the-art approximate adders. Error rate of the proposed adder is 34.48% less than the best reported 32-bit adder with sub-adder size of 8-bit. When the proposed approximate adders are implemented in-memory using majority logic, they achieve better performance compared to the existing in-memory approximate adders. Latency of the proposed adders is observed to be a constant irrespective of adder size for a fixed sub-adder size. Vijaya Lakshmi, Vikramkumar Pudi, John Reuben |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Inner Product Computation In-Memory Using Distributed ArithmeticabstractIn-memory computing using emerging technologies such as Resistive Random-Access Memory (ReRAM) has been proposed as a promising substitute for future computing applications to address the ‘von Neumann bottleneck’. Multiplication is the key component for inner product computation in every digital signal processing (DSP) application and the complexity of multipliers increases greatly with bit-width. Distributed arithmetic (DA) using look-up tables and adder-shifter module has been proposed for inner product computation to achieve multiplier-less efficient DSP architectures, particularly when one of the vectors is a constant and known in advance. Due to the memory wall, DA can be made furthermore latency and energy-efficient when implemented ‘in memory’. In this work, for the first time, we propose two design techniques to compute inner product completely in memory using DA. This is accomplished by storing the precomputed look-up table contents in a ReRAM array and implementing adder-shifter module also in the same array. The adder-shifter is implemented in memory using majority gates which are in turn realized as READ operations in the memory array. Two methods of mapping: latency-optimized and area-optimized and their comparison in terms of latency and area are presented. The proposed method-1 achieves$\approx 60$% energy savings compared to CMOS and the proposed method-2 achieves 10.59 times higher throughput compared to CMOS. Vijaya Lakshmi, Vikramkumar Pudi, John Reuben |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A Novel In-Memory Wallace Tree Multiplier Architecture Using Majority LogicabstractIn-memory computing using emerging technologies such as resistive random-access memory (ReRAM) addresses the ‘von Neumann bottleneck’ and strengthens the present research impetus to overcome the memory wall. While many methods have been recently proposed to implement Boolean logic in memory, the latency of arithmetic circuits (adders and consequently multipliers) implemented as a sequence of such Boolean operations increases greatly with bit-width. Existing in-memory multipliers require$O(n^{2})$cycles which is inefficient both in terms of latency and energy. In this work, we tackle this exorbitant latency by adopting Wallace Tree multiplier architecture and optimizing the addition operation in each phase of the Wallace Tree. Majority logic primitive was used for addition since it is better than NAND/NOR/IMPLY primitives. Furthermore, high degree of gate-level parallelism is employed at the array level by executing multiple majority gates in the columns of the array. In this manner, an in-memory multiplier of$O(n.log(n))$latency is achieved which outperforms all reported in-memory multipliers. Furthermore, the proposed multiplier can be implemented in a regular transistor-accessed memory array without any major modifications to its peripheral circuitry and is also energy-efficient. Vijaya Lakshmi, John Reuben, Vikramkumar Pudi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Accelerated Addition in Resistive RAM Array Using Parallel-Friendly Majority GatesabstractTo overcome the “von Neumann bottleneck,” methods to compute in memory are being researched in many emerging memory technologies, including resistive RAMs (ReRAMs). Majority logic is efficient for synthesizing arithmetic circuits when compared to NAND/NOR/IMPLY logic. In this work, we propose a method to implement a majority gate in a transistor-accessed ReRAM array during the READ operation. Together with NOT gate, which is also implemented in memory, the proposed gate forms a functionally complete Boolean logic, capable of implementing any digital logic. Computing is simplified to a sequence of READ and WRITE operations and does not require any major modifications to the peripheral circuitry of the array. While many methods have been proposed recently to implement the Boolean logic in memory, the latency of in-memory adders implemented as a sequence of such Boolean operations is exorbitant ( O( n)). Parallel-prefix (PP) adders use prefix computation to accelerate addition in conventional CMOS-based adders. By exploiting the parallel-friendly nature of the proposed majority gate and the regular structure of the memory array, it is demonstrated how PP adders can be implemented in memory in O(log( n)) latency. The proposed in-memory addition technique incurs a latency of 4·log( n)+6 for n-bit addition and is energy-efficient due to the absence of sneak currents in 1Transistor-1Resistor configuration. John Reuben, Stefan Pechmann |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | A Parallel-friendly Majority Gate to Accelerate In-memory ComputationabstractEfforts to combat the ‘von Neumann bottleneck’ have been strengthened by Resistive RAMs (RRAMs), which enable computation in the memory array. Majority logic can accelerate computation when compared to NAND/NOR/IMPLY logic due to it’s expressive power. In this work, we propose a method to compute majority while reading from a transistor-accessed RRAM array. The proposed gate was verified by simulations using a physics-based model (for RRAM) and industry standard model (for CMOS sense amplifier) and, found to tolerate reasonable variations in the RRAMs’ resistive states. Together with NOT gate, which is also implemented in-memory, the proposed gate forms a functionally complete Boolean logic, capable of implementing any digital logic. Computing is simplified to a sequence of READ and WRITE operations and does not require any major modifications to the peripheral circuitry of the array. The parallel-friendly nature of the proposed gate is exploited to implement an eight-bit parallel-prefix adder in memory array. The proposed in-memory adder could achieve a latency reduction of 70% and 50% when compared to IMPLY and NAND/NOR logic-based adders, respectively. John Reuben, Stefan Pechmann |
ASAP | 1 |