EDBT 2026 Demo / reviewers in the wild / expert
Vikramkumar Pudi
dblp:168/6457 · also Vikram Kumar Pudi
· DBLP profile ↗
11ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3992-0624ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | In-Memory Implementation of an Approximate Adder With Reduced Latency and ErrorabstractIn-memory computing has been a prominent solution to Von Neumann bottleneck that degrades the performance of a computing system. Approximate computing is widely used to improve the performance of multimedia and other applications that are error-tolerant. Approximate adders being the basic units used to design other complex units, get benefited when implemented in-memory by taking the advantages of both in-memory computing and approximate computing. In this work, we have improved the speculative carry select adder to minimize error and critical path delay by eliminating multiplexers. The proposed adder achieves less critical path, area, improved error characteristics such as error rate, normalized mean error distance and mean relative error distance when compared to the state-of-the-art approximate adders. Error rate of the proposed adder is 34.48% less than the best reported 32-bit adder with sub-adder size of 8-bit. When the proposed approximate adders are implemented in-memory using majority logic, they achieve better performance compared to the existing in-memory approximate adders. Latency of the proposed adders is observed to be a constant irrespective of adder size for a fixed sub-adder size. Vijaya Lakshmi, Vikramkumar Pudi, John Reuben |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Design of Low-Complexity Quantized Compressive Sensing Using Measurement Predictive CodingabstractBlock-based compressive sensing (BCS) has evolved as a promising method for smart devices with limited bandwidth and computing capabilities, striking a balance between image/video quality and transmission efficiency. Despite its advantages, BCS falls short in reducing bitrate compared with traditional acquisition systems, because it increases the number of bits per measurement, which leads to high storage and transmission costs. In this context, we propose a measurement predictive coding (MPC) along with the quantization method in integration with BCS named BCS-MPC; here, we have performed the quantization with bit shifts only instead of binary division. The proposed method reduces the number of bits per compressive sensing (CS) measurement as well as the transmission of the quantization step size. Furthermore, it reduces the latency and hardware resources. The proposed method improved on average +3.44 to +8.28 dB in PSNR over the current works. From the synthesis results, the proposed BCS-MPC method requires 26.11%, 18.89%, and 82.53% less area, power, and delay over the existing work. We have achieved a reduction in delay with bit-shift operations. Lakshmi Bhanuprakash Reddy Konduru, Vikramkumar Pudi, Balasubramanyam Appina |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Redefining Clock Network Construction: The Nested Flex Paradigm for Enhanced PPA DynamicsabstractIn the evolving technological landscape, the pursuit of high-performance systems is paramount. Achieving a balance between Power, Performance and Area (PPA) – the core elements of chip design – becomes increasingly daunting in light of mounting technological intricacies. To address this challenge, this paper introduces the Nested Flex technique, specifically devised to construct resilient clock networks in non-uniform channel partitions by maximizing the shared path in the clock network. This approach elevates clock quality in comparison to conventional clock tree networks. Relative to classic clock tree synthesis, Nested Flex demonstrates 42.85% enhancement in latency and 32.30% skew reduction. When compared with Multi Source Clock Tree Synthesis (MSCTS), it exhibits 20.83% reduction in latency and 43.18% decrease in skew. Impressively, Nested Flex efficiently condenses a 23-level clock depth to a mere 15 levels, addressing the on-chip variation (OCV) dilemma and establishing a new benchmark for reliable clocking solutions. In terms of power metrics, Nested Flex yields a 15% improvement in total power and 15.82% enhancement in clock power over classic CTS. Additionally, when compared to MSCTS, it achieves 8.38% betterment in total power and 7.03% improvement in clock power. Lakshmi Sarvaani P, Subba Ramkumar Reddy Annapalli, Vikramkumar Pudi |
ISCAS | 3 |
| 2023 | A Configurable Multi Source Clock Tree Synthesis For High Frequency Network On ChipsabstractPerformance driven designs have led to an increase in the multitude of transistors on an IC following the Moore law. The goal of design engineers is to achieve high frequency operation within the given power budget. Clock tree synthesis distributes the clock signal to the design sinks and plays a vital role in determining the power, performance and area of a design. The H-tree is effective for square shaped blocks in building a clock with minimum skew and latency but fails in constructing an electrically symmetric clock tree for rectilinear and tubular network on chips. In this paper, we present a configurable multi source clock tree synthesis which minimizes the latency, skew, routing resources, power consumption and on chip variations. On examining various network on chips and performing several iterations the average improvement values are as follows. Latency decreased by 35.32%, skew reduced by 47.51 %, clock power improved by 15.17%, routing resources were saved by 15.41 % and number of hold buffers reduced by 16.53%. Lakshmi Sarvaani P, Subba Ramkumar Reddy Annapalli, Vikramkumar Pudi, Naga Teja Babu M |
ISCAS | 3 |
| 2022 | Inner Product Computation In-Memory Using Distributed ArithmeticabstractIn-memory computing using emerging technologies such as Resistive Random-Access Memory (ReRAM) has been proposed as a promising substitute for future computing applications to address the ‘von Neumann bottleneck’. Multiplication is the key component for inner product computation in every digital signal processing (DSP) application and the complexity of multipliers increases greatly with bit-width. Distributed arithmetic (DA) using look-up tables and adder-shifter module has been proposed for inner product computation to achieve multiplier-less efficient DSP architectures, particularly when one of the vectors is a constant and known in advance. Due to the memory wall, DA can be made furthermore latency and energy-efficient when implemented ‘in memory’. In this work, for the first time, we propose two design techniques to compute inner product completely in memory using DA. This is accomplished by storing the precomputed look-up table contents in a ReRAM array and implementing adder-shifter module also in the same array. The adder-shifter is implemented in memory using majority gates which are in turn realized as READ operations in the memory array. Two methods of mapping: latency-optimized and area-optimized and their comparison in terms of latency and area are presented. The proposed method-1 achieves$\approx 60$% energy savings compared to CMOS and the proposed method-2 achieves 10.59 times higher throughput compared to CMOS. Vijaya Lakshmi, Vikramkumar Pudi, John Reuben |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | A Novel In-Memory Wallace Tree Multiplier Architecture Using Majority LogicabstractIn-memory computing using emerging technologies such as resistive random-access memory (ReRAM) addresses the ‘von Neumann bottleneck’ and strengthens the present research impetus to overcome the memory wall. While many methods have been recently proposed to implement Boolean logic in memory, the latency of arithmetic circuits (adders and consequently multipliers) implemented as a sequence of such Boolean operations increases greatly with bit-width. Existing in-memory multipliers require$O(n^{2})$cycles which is inefficient both in terms of latency and energy. In this work, we tackle this exorbitant latency by adopting Wallace Tree multiplier architecture and optimizing the addition operation in each phase of the Wallace Tree. Majority logic primitive was used for addition since it is better than NAND/NOR/IMPLY primitives. Furthermore, high degree of gate-level parallelism is employed at the array level by executing multiple majority gates in the columns of the array. In this manner, an in-memory multiplier of$O(n.log(n))$latency is achieved which outperforms all reported in-memory multipliers. Furthermore, the proposed multiplier can be implemented in a regular transistor-accessed memory array without any major modifications to its peripheral circuitry and is also energy-efficient. Vijaya Lakshmi, John Reuben, Vikramkumar Pudi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2018 | Efficient and Lightweight Quantized Compressive Sensing using μ-LawabstractIoT devices for video sensing need to operate within the constraints of limited bandwidth and low computing capabilities. To that effect, Compressive Sensing (CS) emerged as a prominent technique for balancing the quality of images/video and the computing/communication overheads. For CS of video data, the Block-based CS (BCS) is typically used due to low complexity. However, while CS reduces the number of samples to be transmitted, the bit-width of each sample increases due to the linear algebraic operations involved in CS, thus making CS less attractive in its pure and straightforward form. To further optimize the use of CS in IoT devices for video sensing, we explore the use of μ-law quantization technique due to its low hardware implementation overhead. We designed and implemented a complete CS platform with the integration of μ-law quantization, and studied the image quality at different compression ratios. The results show that the proposed quantization technique requires only up to 40 additional LUTs compared to the baseline algorithm, while achieving an additional compression of up to 280% in the best case. Vikramkumar Pudi, Anupam Chattopadhyay, Kwok-Yan Lam |
ISCAS | 1 |
| 2018 | Lightweight and High Performance SHA-256 using Architectural Folding and 4-2 Adder CompressorabstractThe modern era of Internet-of-Things (IoT) is naturally imposing a tight area/runtime constraint on the computing kernels. Security kernels, as part of the standardized protocols as well as custom defense techniques, are among the most common tasks executed on every digital device. Therefore, low area cost and high performance implementation of security kernels is an important goal of current system designers. In this paper, we revisit the state-of-the-art implementations of SHA-256, a standardized security primitive for authentication and propose novel optimizations. Our optimizations, based on architectural folding and 4-2 adder compressor, are geared toward both lightweight and high performance implementations. Detailed experiments of our optimized architecture on different FPGA fabrics clearly demonstrate their benefits. Our presented design point successfully attained the highest hardware efficiency (throughput/area) figures among the published literature so far. Ming Ming Wong, Vikramkumar Pudi, Anupam Chattopadhyay |
VLSI-SoC | 2 |
| 2017 | Majority Logic Formulations for Parallel Adder Designs at Reduced Delay and Circuit ComplexityabstractThe design of high-performance adders has experienced a renewed interest in the last few years; among high performance schemes, parallel prefix adders constitute an important class. They require a logarithmic number of stages and are typically realized using AND-OR logic; moreover with the emergence of new device technologies based on majority logic, new and improved adder designs are possible. However, the best existing majority gate-based prefix adder incurs a delay of 2log2(n) - 1 (due to the nth carry); this is only marginally better than a design using only AND-OR gates (the latter design has a 2log2(n) + 1 gate delay). This paper initially shows that this delay is caused by the output carry equation in majority gate-based adders that is still largely defined in terms of AND-OR gates. In this paper, two new majority gate-based recursive techniques are proposed. The first technique relies on a novel formulation of the majority gate-based equations in the used group generate and group propagate hardware; this results in a new definition for the output carry, thus reducing the delay. The second contribution of this manuscript utilizes recursive properties of majority gates (through a novel operator) to reduce the circuit complexity of prefix adder designs. Overall, the proposed techniques result in the calculation of the output carry of an n-bit adder with only a majority gate delay of log2(n) + 1. This leads to a reduction of 40percent in delay and 30percent in circuit complexity (in terms of the number of majority gates) for multi-bit addition in comparison to the best existing designs found in the technical literature. Vikramkumar Pudi, K. Sridharan 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2015 | Very large-scale integration architecture for video stabilisation and implementation on a field programmable gate array-based autonomous vehicleabstractAutonomous vehicles engaged in terrain exploration are typically equipped with a camera. The camera is subjected to vibration as the vehicle moves so that the videos captured require stabilisation to facilitate accurate interpretation by remote operators. Dedicated architectures for video stabilisation that offer high performance while consuming low area and power are desirable for this application. This study presents a pipelined very large‐scale integration architecture. It is based on exploiting the separability property of the two‐dimensional (2‐D) Sobel matrix and the 2‐D Gaussian filtering matrix to obtain an efficient corner point detection architecture. It also employs the coordinate rotation digital computer architecture for global motion vector calculation. The proposed architecture has been coded in Verilog and synthesised for a field programmable gate array (FPGA), which offers massive parallelism at fairly low power. The proposed architecture is shown to be highly area efficient. An FPGA‐based autonomous vehicle has been fabricated, and experiments with a camera mounted on the vehicle are presented and analysed. Tahiyah Nou Shene, Vikramkumar Pudi, K. Sridharan 0001, Vineetha Thomas, J. Arthi |
IET Comput. Vis. | 2 |
| 2015 | A Bit-Serial Pipelined Architecture for High-Performance DHT Computation in Quantum-Dot Cellular AutomataabstractIn this brief, we consider quantum-dot cellular automata (QCA) realization of the discrete Hadamard transform (DHT). An analysis of a full-parallel solution based on efficient multibit addition in QCA is first presented. We show that this leads to large area as well as delay. We then propose a bit-serial pipelined architecture for QCA-based DHT. The proposed architecture is based on a new one-bit adder-subtractor requiring only six majority gates and a feedback latch that requires only one majority gate and limited wiring. The approach leads to a reduction in area-delay-cycle product of 74% and 91% (over a full-parallel solution) for wordlengths of 4 and 8, respectively. Results of simulations in QCADesigner are also presented. Vikramkumar Pudi, K. Sridharan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |