EDBT 2026 Demo / reviewers in the wild / expert
Santosh Kumar Vishvakarma
dblp:57/11307
· DBLP profile ↗
15ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-4223-0077ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gated logic controlled 10T-SRAM for low-power bidirectional ring oscillators
Neha Maheshwari, Ambika Prasad Shah, Santosh Kumar Vishvakarma |
Integr. | 3 |
| 2026 | Adaptive-precision SIMD architecture for high-throughput and resource-efficient DNN acceleration
Vasundhara Trivedi, Harman Singh Bagga, Gopal Raut, Santosh Kumar Vishvakarma |
Integr. | 4 |
| 2026 | CIM-Enabled MBRO PUF: Integrating Multistage BRO and Reconfigurable SRAM for Edge Security ApplicationsabstractThe Internet of Things (IoT) ecosystem and the rapid advancement of consumer devices necessitates solutions that provide both computer efficiency and hardware-level security. Compute-in-memory (CIM) is developing as an effective edge computing approach in IoT systems because it reduces data transit between memory and processors. In this work, a 10T SRAM-based MBRO PUF is designed using a CIM approach to reduce latency and generate double challenge response pairs (CRPs) that improve hardware security. Reconfigurable SRAM cells with tristate inverters allow bidirectional control, and the integration of transmission gates enables the configuration of multistage ring oscillators, proving the potential of the SRAM array for diverse operations. The proposed architecture not only ensures reliable functioning but also allows flexible activation and deactivation of the oscillator and inverter, lowering power consumption during idle periods and improving the stability of oscillation patterns with frequencies of 1.53 GHz, 942.9 MHz and 652.5 MHz, respectively, for 3, 5 and 7 stage MBRO PUF. The area utilized by the 7 stage MBRO is 11.43μm2with low power consumption of 32uW The selection of MBRO PUF based on the number of stages according to the application affirms that the proposed MBRO PUF are highly efficient designs with strong uniqueness and reliability. Its uniqueness is 49.6%, 49.4% and 48% for 3, 5, and 7 stages. By balancing security and power efficiency requirements, making it suitable for integration into resource-constrained environment and embedded systems. Neha Maheshwari, Brij B. Gupta, Santosh Kumar Vishvakarma |
IEEE Internet Things J. | 3 |
| 2026 | ReLANCE: A Resource-Efficient Low-Latency Cortical Neural Acceleration EngineabstractWe present a Cortical Neural Pool (CNP) architecture featuring a high-speed, resource-efficient CORDIC based Hodgkin-Huxley (RCHH) neuron model. Unlike shared CORDIC-based DNN approaches, the proposed neuron leverages modular and performance-optimised CORDIC stages with a latency-area trade-off. We introduce a novel Constraint-Aware Modular Parallelism (CAMP) with Precision & Stability handling to leverage maximum speedup and utilisation of hardware through hardware software co-design. The FPGA implementation of the RCHH neuron shows 24.5% LUT reduction and 35.2% improved speed, compared to SoTA designs, with 70% better normalised root mean square error (NRMSE). Furthermore, the CNP exhibits 2.85x higher throughput (12.69 GOPS) than a functionally equivalent CORDIC-based DNN engine, with only a 0.35% accuracy drop relative to the DNN counterpart on the MNIST dataset. The overall results indicate that the design shows biologically accurate, low-resource spiking neural network implementations for resource-constrained edge AI applications. The reproducibility codes are publicly available at https://github.com/mukullokhande99/CNP RCHH, facilitating rapid integration and further development by researchers. Arjun S. Nair, Bhawna Chaudhary, Mukul Lokhande, Santosh Kumar Vishvakarma |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | LPRE: Logarithmic Posit-enabled Reconfigurable edge-AI EngineabstractEdge-AI applications face huge challenges in resource-constrained environments, particularly in enhancing computational efficiency within bandwidth limitations. This work proposes the Logarithmic-Posit-enabled Reconfigurable edgeAI Engine (LPRE) that enhances hardware efficiency without compromising accuracy. The proposed architecture utilizes time-multiplexed dynamically configurable single-layer hardware to balance resource reuse and bandwidth for multi-layer perceptron and CNN models. Evaluations on LeNet-5 using MNIST demonstrate that LPRE achieves up to 4× throughput enhancement at 8-bit precision with negligible accuracy loss (compared to FP32 baseline), while requiring up to 80% and 50% fewer resources than fixed-point arithmetic and state-of-the-art works, respectively. The design is viable for various edge-AI applications, such as real-time number plate recognition, offering scalable, energy-efficient IoT solutions. Omkar Kokane, Mukul Lokhande, Gopal Raut, Adam Teman, Santosh Kumar Vishvakarma |
ISCAS | 5 |
| 2025 | Flex-PE: Flexible and SIMD Multiprecision Processing Element for AI WorkloadsabstractThe rapid evolution of artificial intelligence (AI) models, from deep neural networks (DNNs) to transformers/large-language models (LLMs), demands flexible hardware solutions to meet diverse execution needs across edge and cloud platforms. Existing accelerators lack unified support for multiprecision arithmetic and runtime-configurable activation functions (AFs). This work proposes Flex-PE, a single instruction, multiple data (SIMD)-enabled multiprecision processing element that efficiently integrates multiply-and-accumulate operations with configurable AFs using unified hardware, including Sigmoid, Tanh, ReLU, and SoftMax. The proposed design achieves throughput improvements of up to$16\times $FxP4,$8\times $FxP8,$4\times $FxP16, and$1\times $FxP32, with maximum hardware efficiency for both iterative and pipelined architectures. An area-efficient iterative Flex-PE-based SIMD systolic array reduces DMA reads by up to$62\times $and$371\times $for input feature maps and weight filters in VGG-16, achieving 8.42 GOPS/W energy efficiency with minimal accuracy loss (<2%). Flex-PE scales from 4-bit edge inference to FxP8/16/32, supporting edge and cloud high-performance computing (HPC) while providing high-performance adaptable AI hardware with optimal precision, throughput, and energy efficiency. Mukul Lokhande, Gopal Raut, Santosh Kumar Vishvakarma |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | An Empirical Approach to Enhance Performance for Scalable CORDIC-Based Deep Neural NetworksabstractPractical implementation of deep neural networks (DNNs) demands significant hardware resources, necessitating high computational power and memory bandwidth. While existing field-programmable gate array (FPGA)–based DNN accelerators are primarily optimized for fast single-task performance, cost, energy efficiency, and overall throughput are crucial considerations for their practical use in various applications. This article proposes a performance-centric pipeline Coordinate Rotation Digital Computer (CORDIC)–based MAC unit and implements a scalable CORDIC-based DNN architecture that is area- and power-efficient and has high throughput. The CORDIC-based neuron engine uses bit-rounding to maintain input-output precision and minimal hardware resource overhead. The results demonstrate the versatility of the proposed pipelined MAC, which operates at 460 MHz and allows for higher network throughput. A software-based implementation platform evaluates the proposed MAC operation’s accuracy for more extensive neural networks and complex datasets. The DNN accelerator with parameterized and modular layer-multiplexed architecture is designed. Empirical evaluation through Pareto analysis is used to improve the efficiency of DNN implementations by fixing the arithmetic precision and optimal pipeline stages. The proposed architecture utilizes layer-multiplexing, a technique that effectively reuses a single DNN layer to enhance efficiency while maintaining modularity and adaptability for integrating various network configurations. The proposed CORDIC MAC-based DNN architecture is scalable for any bit-precision network size, and the DNN accelerator is prototyped using the Xilinx Virtex-7 VC707 FPGA board, operating at 66 MHz. The proposed design does not use any Xilinx macros, making it easily adaptable for ASIC implementation. Compared with state-of-the-art designs, the proposed design reduces resource use by 45% and power consumption by 4× without sacrificing performance. The accelerator is validated using the MNIST dataset, achieving 95.06% accuracy, only 0.35% less than other cutting-edge implementations. Gopal Raut, Saurabh Karkun, Santosh Kumar Vishvakarma |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | Loading Effect Free MOS-only Voltage Reference Ladder for ADC in RRAM-crossbar ArrayabstractIn the analog domain, with the increase in ReRAM m × n crossbar array, the Loading Effect (LE) seems to grow at the input of the comparator stage in analog to digital converter (ADC). The reference voltage generating ladder nodes for ADC are susceptible to design parameters due to small input voltages. We used the PMOS transistor to design this ladder circuitry. Further, sleep mode is applied using the power-gating (PG) technique to lower power dissipation. In this article, a Pareto study has been performed to evaluate robust and stable circuitry with minimum LE in the reference voltage ladder for ADC. An NMOS-based Current mirror is also designed and used with the proposed reference voltage ladder to achieve better stability in terms of power supply and reference voltage variations. Further, we analyzed the Process, Voltage, and Temperature (PVT) variation impact on the proposed circuitry. Finally, the power consumption of the proposed ladder at the 180nm technology node, is 0.7uW. Also, the circuit supports the power-gating technique in sleep mode, saving 43% of total power. Circuit's Monte-Carlo simulation for node voltage variation shows minimum mean and σ deviation. The circuit supports the power-gating technique in sleep mode, saving 43% of total power. Varun Bhatnagar, Gopal Raut, Santosh Kumar Vishvakarma |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Data multiplexed and hardware reused architecture for deep neural network acceleratorabstractDespite many decades of research on high-performance Deep Neural Network (DNN) accelerators, their massive computational demand still requires resource-efficient, optimized and parallel architecture for computational acceleration. Contemporary hardware implementations of DNNs face the burden of excess area requirement due to resource-intensive elements such as multipliers and non-linear Activation Functions (AFs). This paper proposes DNN with reused hardware-costly AF by multiplexing data using shift-register. The on-chip quantized log2 based memory addressing with an optimized technique is used to access input features, weights, and biases. This way the external memory bandwidth requirement is reduced and dynamically adjusted for DNNs. Further, high-throughput and resource-efficient memory elements for sigmoid activation function are extracted using the Taylor series and its order expansion have been tuned for better test accuracy. The performance is validated and compared with previous works for the MNIST dataset. Besides, the digital design of AF is synthesized at 45 nm technology node and physical parameters are compared with previous works. The proposed hardware reused architecture is verified for neural network 16:16:10:4 using 8-bit dynamic fixed-point arithmetic and implemented on Xilinx Zynq xc7z010clg400 SoC using 100 MHz clock. The implemented architecture uses 25% less hardware resources and consumes 12% less power without performance loss, compared to other state-of-the-art implementations, as lower hardware resources and power consumption are especially important for increasingly important edge computing solutions. Gopal Raut, Anton Biasizzo, Narendra Singh Dhakad, Gregor Papa, Santosh Kumar Vishvakarma |
Neurocomputing | 6 |
| 2021 | Voltage Bootstrapped Schmitt Trigger based Radiation Hardened Latch Design for Reliable CircuitsabstractSoft error is one of the major reliability issue with technology scaling. In this work, we propose a radiation hardened voltage bootstrapped schmitt trigger (VB-ST) latch. To evaluate the circuit radiation resilience, we calculated the critical charge under the PVT variations at the most sensitive node and observed that the proposed latch has the highest critical charge and the lowest soft error rate ratio when compared to existing latches. We analyzed the impact of process variations on our design and observed that the VB-ST latch has 0.42x less critical voltage variability as compared to ST latch. Further, dynamic power and propagation delay are examined for various supply voltages, and we observed that the VB-ST latch has the lowest power consumption and delay propagation when compared to the other considered latches. For the validation of the proposed latch, a charge to power-delay-area product ratio (QPAR) is calculated and we clearly observed that the proposed VB-ST based latch significantly outperforms the performance of existing designs. Nikhil Agrawal, Narendra Singh Dhakad, Ambika Prasad Shah, Santosh Kumar Vishvakarma, Patrick Girard 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2020 | Soft Error Hardened Asymmetric 10T SRAM Cell for Aerospace Applications
Ambika Prasad Shah, Santosh Kumar Vishvakarma, Michael Hübner 0001 |
J. Electron. Test. | 2 |
| 2018 | Tunneling Field Effect Transistors for Enhancing Energy Efficiency and Hardware Security of IoT Platforms: Challenges and OpportunitiesabstractTunneling Field-Effect Transistor (TFET) is a leading future transistor option for next generation VLSI applications and Internet of things (IoT). Many have demonstrated the energy efficiency of TFET circuits. In this work, we demonstrate for the first time utilizing TFET ambipolar device characteristics for jitter generation in post-processing circuits of true random number generators (TRNGs) and suitability for enhancing hardware security of IoT platforms. Device and circuit design challenges for TFET based transceiver designs for capacitive coupled interconnect in 3DIC and on-chip low dropout digital voltage regulators (DLDOs) are further explored towards energy efficient IoT platforms. Aditya Japa, T. Nagateja, Santosh Kumar Vishvakarma, Yellappa Palagani, Jun Rim Choi, Ramesh Vaddi |
ISCAS | 3 |
| 2018 | Ultra low power-high stability, positive feedback controlled (PFC) 10T SRAM cell for look up table (LUT) design
Pooran Singh, Bhupendra Singh Reniwal, Vikas Vijayvargiya, Santosh Kumar Vishvakarma |
Integr. | 5 |
| 2016 | A Single-Ended With Dynamic Feedback Control 8T Subthreshold SRAM CellabstractA novel 8-transistor (8T) static random access memory cell with improved data stability in subthreshold operation is designed. The proposed single-ended with dynamic feedback control 8T static RAM (SRAM) cell enhances the static noise margin (SNM) for ultralow power supply. It achieves write SNM of 1.4× and 1.28× as that of isoarea 6T and read-decoupled 8T (RD-8T), respectively, at 300 mV. The standard deviation of write SNM for 8T cell is reduced to 0.4× and 0.56× as that for 6T and RD-8T, respectively. It also possesses another striking feature of high read SNM 2.33×, 1.23×, and 0.89× as that of 5T, 6T, and RD-8T, respectively. The cell has hold SNM of 1.43×, 1.23×, and 1.05× as that of 5T, 6T, and RD-8T, respectively. The write time is 71% lesser than that of single-ended asymmetrical 8T cell. The proposed 8T consumes less write power 0.72×, 0.6×, and 0.85× as that of 5T, 6T, and isoarea RD-8T, respectively. The read power is 0.49× of 5T, 0.48× of 6T, and 0.64× of RD-8T The power/energy consumption of 1-kb 8T SRAM array during read and write operations is 0.43× and 0.34×, respectively, of 1-kb 6T array. These features enable ultralow power applications of 8T. C. B. Kushwah, Santosh Kumar Vishvakarma |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Dataline Isolated Differential Current Feed/Mode Sense Amplifier for Small Icell SRAM Using FinFETabstractThis paper for the first time presents a novel, high-performance and robust current feed sense amplifiers (CF-SA) design for small ICell SRAM in 20nm Fin-shaped field effect transistor (FinFET) technology. The CFSA incorporates isolated DL current sensing approach which provides the higher Current Ratio Amplification (CRA) factor. The CF-SA significantly outperforms with 66.89% and 31.47% lower sensing delay than CCSA [13] and HSA [8] respectively under similar ICell and bit-line and data-line capacitance. Our results show that even at the worst corner the CF-SA demonstrates 2.15x and 3.02x higher differential current and 2.23x and 1.7x higher data-line differential voltage with 66.6% and 34.32% higher mean (μ) than those of the best prior arts. Furthermore, failure probability of the proposed design against process parameter variations is rigorously analyzed through Monte Carlo simulations. Bhupendra Singh Reniwal, Vikas Vijayvargiya, Pooran Singh, Santosh Kumar Vishvakarma, Devesh Dwivedi |
ACM Great Lakes Symposium on VLSI | 4 |