EDBT 2026 Demo / reviewers in the wild / expert
Gopal Raut
dblp:247/3715
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-1046-9457ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive-precision SIMD architecture for high-throughput and resource-efficient DNN acceleration
Vasundhara Trivedi, Harman Singh Bagga, Gopal Raut, Santosh Kumar Vishvakarma |
Integr. | 3 |
| 2025 | LPRE: Logarithmic Posit-enabled Reconfigurable edge-AI EngineabstractEdge-AI applications face huge challenges in resource-constrained environments, particularly in enhancing computational efficiency within bandwidth limitations. This work proposes the Logarithmic-Posit-enabled Reconfigurable edgeAI Engine (LPRE) that enhances hardware efficiency without compromising accuracy. The proposed architecture utilizes time-multiplexed dynamically configurable single-layer hardware to balance resource reuse and bandwidth for multi-layer perceptron and CNN models. Evaluations on LeNet-5 using MNIST demonstrate that LPRE achieves up to 4× throughput enhancement at 8-bit precision with negligible accuracy loss (compared to FP32 baseline), while requiring up to 80% and 50% fewer resources than fixed-point arithmetic and state-of-the-art works, respectively. The design is viable for various edge-AI applications, such as real-time number plate recognition, offering scalable, energy-efficient IoT solutions. Omkar Kokane, Mukul Lokhande, Gopal Raut, Adam Teman, Santosh Kumar Vishvakarma |
ISCAS | 3 |
| 2025 | Flex-PE: Flexible and SIMD Multiprecision Processing Element for AI WorkloadsabstractThe rapid evolution of artificial intelligence (AI) models, from deep neural networks (DNNs) to transformers/large-language models (LLMs), demands flexible hardware solutions to meet diverse execution needs across edge and cloud platforms. Existing accelerators lack unified support for multiprecision arithmetic and runtime-configurable activation functions (AFs). This work proposes Flex-PE, a single instruction, multiple data (SIMD)-enabled multiprecision processing element that efficiently integrates multiply-and-accumulate operations with configurable AFs using unified hardware, including Sigmoid, Tanh, ReLU, and SoftMax. The proposed design achieves throughput improvements of up to$16\times $FxP4,$8\times $FxP8,$4\times $FxP16, and$1\times $FxP32, with maximum hardware efficiency for both iterative and pipelined architectures. An area-efficient iterative Flex-PE-based SIMD systolic array reduces DMA reads by up to$62\times $and$371\times $for input feature maps and weight filters in VGG-16, achieving 8.42 GOPS/W energy efficiency with minimal accuracy loss (<2%). Flex-PE scales from 4-bit edge inference to FxP8/16/32, supporting edge and cloud high-performance computing (HPC) while providing high-performance adaptable AI hardware with optimal precision, throughput, and energy efficiency. Mukul Lokhande, Gopal Raut, Santosh Kumar Vishvakarma |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | An Empirical Approach to Enhance Performance for Scalable CORDIC-Based Deep Neural NetworksabstractPractical implementation of deep neural networks (DNNs) demands significant hardware resources, necessitating high computational power and memory bandwidth. While existing field-programmable gate array (FPGA)–based DNN accelerators are primarily optimized for fast single-task performance, cost, energy efficiency, and overall throughput are crucial considerations for their practical use in various applications. This article proposes a performance-centric pipeline Coordinate Rotation Digital Computer (CORDIC)–based MAC unit and implements a scalable CORDIC-based DNN architecture that is area- and power-efficient and has high throughput. The CORDIC-based neuron engine uses bit-rounding to maintain input-output precision and minimal hardware resource overhead. The results demonstrate the versatility of the proposed pipelined MAC, which operates at 460 MHz and allows for higher network throughput. A software-based implementation platform evaluates the proposed MAC operation’s accuracy for more extensive neural networks and complex datasets. The DNN accelerator with parameterized and modular layer-multiplexed architecture is designed. Empirical evaluation through Pareto analysis is used to improve the efficiency of DNN implementations by fixing the arithmetic precision and optimal pipeline stages. The proposed architecture utilizes layer-multiplexing, a technique that effectively reuses a single DNN layer to enhance efficiency while maintaining modularity and adaptability for integrating various network configurations. The proposed CORDIC MAC-based DNN architecture is scalable for any bit-precision network size, and the DNN accelerator is prototyped using the Xilinx Virtex-7 VC707 FPGA board, operating at 66 MHz. The proposed design does not use any Xilinx macros, making it easily adaptable for ASIC implementation. Compared with state-of-the-art designs, the proposed design reduces resource use by 45% and power consumption by 4× without sacrificing performance. The accelerator is validated using the MNIST dataset, achieving 95.06% accuracy, only 0.35% less than other cutting-edge implementations. Gopal Raut, Saurabh Karkun, Santosh Kumar Vishvakarma |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Loading Effect Free MOS-only Voltage Reference Ladder for ADC in RRAM-crossbar ArrayabstractIn the analog domain, with the increase in ReRAM m × n crossbar array, the Loading Effect (LE) seems to grow at the input of the comparator stage in analog to digital converter (ADC). The reference voltage generating ladder nodes for ADC are susceptible to design parameters due to small input voltages. We used the PMOS transistor to design this ladder circuitry. Further, sleep mode is applied using the power-gating (PG) technique to lower power dissipation. In this article, a Pareto study has been performed to evaluate robust and stable circuitry with minimum LE in the reference voltage ladder for ADC. An NMOS-based Current mirror is also designed and used with the proposed reference voltage ladder to achieve better stability in terms of power supply and reference voltage variations. Further, we analyzed the Process, Voltage, and Temperature (PVT) variation impact on the proposed circuitry. Finally, the power consumption of the proposed ladder at the 180nm technology node, is 0.7uW. Also, the circuit supports the power-gating technique in sleep mode, saving 43% of total power. Circuit's Monte-Carlo simulation for node voltage variation shows minimum mean and σ deviation. The circuit supports the power-gating technique in sleep mode, saving 43% of total power. Varun Bhatnagar, Gopal Raut, Santosh Kumar Vishvakarma |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Data multiplexed and hardware reused architecture for deep neural network acceleratorabstractDespite many decades of research on high-performance Deep Neural Network (DNN) accelerators, their massive computational demand still requires resource-efficient, optimized and parallel architecture for computational acceleration. Contemporary hardware implementations of DNNs face the burden of excess area requirement due to resource-intensive elements such as multipliers and non-linear Activation Functions (AFs). This paper proposes DNN with reused hardware-costly AF by multiplexing data using shift-register. The on-chip quantized log2 based memory addressing with an optimized technique is used to access input features, weights, and biases. This way the external memory bandwidth requirement is reduced and dynamically adjusted for DNNs. Further, high-throughput and resource-efficient memory elements for sigmoid activation function are extracted using the Taylor series and its order expansion have been tuned for better test accuracy. The performance is validated and compared with previous works for the MNIST dataset. Besides, the digital design of AF is synthesized at 45 nm technology node and physical parameters are compared with previous works. The proposed hardware reused architecture is verified for neural network 16:16:10:4 using 8-bit dynamic fixed-point arithmetic and implemented on Xilinx Zynq xc7z010clg400 SoC using 100 MHz clock. The implemented architecture uses 25% less hardware resources and consumes 12% less power without performance loss, compared to other state-of-the-art implementations, as lower hardware resources and power consumption are especially important for increasingly important edge computing solutions. Gopal Raut, Anton Biasizzo, Narendra Singh Dhakad, Gregor Papa, Santosh Kumar Vishvakarma |
Neurocomputing | 1 |