VLDB 2026 Research / reviewers in the wild / expert
Mark A. Indovina
dblp:357/6893
· DBLP profile ↗
10ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-6645-4596ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analysis of Area and Power Overhead for a Synthesizable 4-Phase Bundled-Data Asynchronous RISC-V ProcessorabstractGlobal clock routing in modern semiconductor nodes introduces significant design challenges, such as managing clock skew and possibly substantial dynamic power consumption — especially when pipeline stages are idle. To circumvent these issues, this paper presents the design and detailed analysis of a synthesizable 32-bit RISC-V processor that eliminates the global system clock by utilizing a 4-phase bundled-data asynchronous handshaking protocol. The proposed architecture features a 4-stage in-order execution pipeline synthesized using a 28 nm technology library. The handshaking control logic utilizes a typical Muller C-element implemented using standard library cells, specifically an XNOR gate and a D-latch, to ensure full synthesizability within a standard design flow. Area results post place and route reveal that this dedicated asynchronous control circuitry introduces a negligible area overhead of only 0.1% to the total processor design. Timing analysis post place and route indicate a functional operational frequency of 250 MHz, limited by physical routing parasitics in the multiplier logic. The total power consumption of the asynchronous implementation is 6.13 mW, representing a dramatic reduction compared to the 28.14 mW consumed by the synchronous baseline. This validates a significant power-efficiency advantage, as the asynchronous architecture eliminates the dynamic power overhead required for global clock distribution in the 28 nm node. Furthermore, the handshaking protocol enables innate, fine-grained power-gating, allowing idle pipeline stages to drop to near-leakage levels, an efficiency unavailable in conventionally clocked architectures. Sean Jacobs, Mark A. Indovina, Yashaswini Suresha |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | The State of Simulation Frameworks for Evaluating Emerging LLM Accelerators
Stefan Maczynski, Amlan Ganguly, Mark A. Indovina |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | BrIM: A Branching In-Memory Accelerator
Stefan Maczynski, Amlan Ganguly, Mark A. Indovina, Purab Ranjan Sutradhar, Sai Manoj Pudukotai Dinakarrao, Sathwika Bavikadi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | ReApprox-PIM: Reconfigurable Approximate Lookup-Table (LUT)-Based Processing-in-Memory (PIM) Machine Learning AcceleratorabstractConvolutional neural networks (CNNs) have achieved significant success in various applications. Numerous hardware accelerators are introduced to accelerate CNN execution with improved energy efficiency compared to traditional software implementations. Despite the achieved success, deploying traditional hardware accelerators for bulky CNNs on current and emerging smart devices is impeded by limited resources, including memory, power, area, and computational capabilities. Recent works introduced processing-in-memory (PIM), a non-Von-Neumann architecture, which is a promising approach to tackle the problem of data movement between logic and memory blocks. However, as observed from the literature, the existing PIM architectures cannot congregate all the computational operations due to limited programmability and flexibility. Furthermore, the capabilities of the PIM are challenged by the limited available on-chip memory. To enable faster computations and address the limited on-chip memory constraints, this work introduces a novel reconfigurable approximate computing-based PIM, termed ReApprox-PIM. The proposed ReApprox-PIM is capable of addressing the two challenges mentioned above in the following manner: (i) it utilizes a programmable look-up-table (LUT)-based processing architecture that can support different approximate computing techniques via programmability, and (ii) followed by resource-efficient, fast CNN computing via the implementation of highly-optimized approximate computing techniques. This results in improved computing footprint, operational parallelism, and reduced computational latency and power consumption compared to prior PIMs relying on exact computations for CNN inference acceleration at a minimal sacrifice of accuracy. We have evaluated the proposed ReApprox-PIM on various CNN architectures, for inference applications including standard LeNet, AlexNet, ResNet-18, -34, and -50. Our experimental results show that the ReApprox-PIM achieves a speedup of 1.63× with 1.66 × lower area for the processing components compared to the existing PIM architectures. Furthermore, the proposed ReApprox-PIM achieves 2.5× higher energy efficiency and 1.3× higher throughput compared to the state-of-the-art LUT-based PIM architectures. Sathwika Bavikadi, Purab Ranjan Sutradhar, Mark A. Indovina, Amlan Ganguly, Sai Manoj Pudukotai Dinakarrao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | FlutPIM: : A Look-up Table-based Processing in Memory Architecture with Floating-point Computation Support for Deep Learning ApplicationsabstractProcessing-in-Memory (PIM) has shown great potential for a wide range of data-driven applications, especially Deep Learning and AI. However, it is a challenge to facilitate the computational sophistication of a standard processor (i.e. CPU or GPU) within the limited scope of a memory chip without contributing significant circuit overheads. To address the challenge, we propose a programmable LUT-based area-efficient PIM architecture capable of performing various low-precision floating point (FP) computations using a novel LUT-oriented operand-decomposition technique. We incorporate such compact computational units within the memory banks in a large count to achieve impressive parallel processing capabilities, up to 4x higher than state-of-the-art FP-capable PIM. Additionally, we adopt a highly-optimized low-precision FP format that maximizes computational performance at a minimal compromise of computational precision, especially for Deep Learning Applications. The overall result is a 17% higher throughput and an impressive 8-20x higher compute Bandwidth/bank compared to the state-of-the-art of in-memory acceleration. Purab Ranjan Sutradhar, Sathwika Bavikadi, Mark A. Indovina, Sai Manoj Pudukotai Dinakarrao, Amlan Ganguly |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | POLAR: Performance-aware On-device Learning Capable Programmable Processing-in-Memory Architecture for Low-Power ML ApplicationsabstractImproving the performance of real-time Traffic Sign Recognition (TSR) applications using Deep Learning (DL) algorithms such as Convolutional Neural Networks (CNN) on software platforms is challenging due to the sheer computational complexity of these algorithms. In this work, we adopt a hardware-software combined approach to address this issue. We introduce a data-centric Processing-in-Memory (PIM) architecture that leverages Look-up-Table (LUT)-based processing for minimal data movement and superior performance and efficiency. Despite the superior performance, the limited available memory in PIM makes it complex to deploy deep CNNs. We propose merging CNN layers in this work to meet the limited resource constraints. One specific challenge in the TSR is the continuous change in the deployed environment, which makes a CNN model train over static data, leading to performance degradation over time. To address these challenges, we introduce a lightweight, performance-aware Generative Adversarial Network (GAN)-based on-device learning on PIM architecture. This compact CNN on PIM architecture attains data-level parallelism and reduces pipelining delays and makes it easier for on-device training and inference. Evaluation is performed on multiple state-of-the-art DL networks such as LeNet, AlexNet, ResNet using the German Traffic Sign Recognition Benchmark (GTSRB) Dataset, and the Belgium Traffic Sign Dataset (BTSD). With the proposed learning technique, it is observed to achieve maximum accuracy of 92.8% and 89.27% on GTSRB, and BTSD datasets. Also, it is observed the proposed mechanism maintains an average accuracy to be above 85% despite changes in the environment on all the CNNs deployed on the PIM accelerator. Sathwika Bavikadi, Purab Ranjan Sutradhar, Mark A. Indovina, Amlan Ganguly, Sai Manoj Pudukotai Dinakarrao |
DSD | 3 |
| 2022 | Look-up-Table Based Processing-in-Memory Architecture With Programmable Precision-Scaling for Deep Learning ApplicationsabstractProcessing in memory (PIM) architecture, with its ability to perform ultra-low-latency parallel processing, is regarded as a more suitable alternative to von Neumann computing architectures for implementing data-intensive applications such as Deep Neural Networks (DNN) and Convolutional Neural Networks (CNN). In this article, we present a Look-up Table (LUT) based PIM architecture aimed at CNN/DNN acceleration that replaces logic-based processing with pre-calculated results stored inside the LUTs in order to perform complex computations on the DRAM memory platform. Our LUT-based DRAM-PIM architecture offers superior performance at a significantly higher energy-efficiency compared to the more conventional bit-wise parallel PIM architectures, while at the same time avoids fabrication challenges associated with the in-memory implementation of logic circuits. Alongside, the processing elements can be programmed and re-programmed to perform virtually any operation, including operations of Convolutional, Fully Connected, Pooling, and Activating Layers of CNN/DNN. Furthermore, it is capable of operating on several combinations of bit-widths of the operand data and thereby offers a wider range of flexibility across performance, precision, and efficiency. Transmission Gate (TG) realization of the circuitry ensures minimal footprint from the PIM architecture. Our simulations demonstrate that the proposed architecture can perform AlexNet inference at a nearly 13× faster rate and 125× more efficiency compared to state-of-the-art GPU and also provides 1.35× higher throughput at 2.5× higher energy-efficiency than another recent DRAM-implemented LUT-based PIM architecture in its baseline operation mode. Moreover, it offers 12× higher frame-rate at 9× more efficiency per frame for the lowest operand precision setting, with respect to its own baseline operation mode. Purab Ranjan Sutradhar, Sathwika Bavikadi, Mark Connolly, Savankumar Prajapati, Mark A. Indovina, Sai Manoj Pudukotai Dinakarrao, Amlan Ganguly |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Flexible Instruction Set Architecture for Programmable Look-up Table based Processing-in-MemoryabstractProcessing in Memory (PIM) is a recent novel computing paradigm that is still in its nascent stage of development. Therefore, there has been an observable lack of standardized and modular Instruction Set Architectures (ISA) for the PIM devices. In this work, we present the design of an ISA which is primarily aimed at a recent programmable Look-up Table (LUT) based PIM architecture. Our ISA performs the three major tasks of i) controlling the flow of data between the memory and the PIM units, ii) reprogramming the LUTs to perform various operations required for a particular application, and iii) executing sequential steps of operation within the PIM device. A microcoded architecture of the Controller/Sequencer unit ensures minimum circuit overhead as well as offers programmability to support any custom operation. We provide a case study of CNN inferences, large matrix multiplications, and bitwise computations on the PIM architecture equipped with our ISA and present performance evaluations based on this setup. We also compare the performances with several other PIM architectures. Mark Connolly, Purab Ranjan Sutradhar, Mark A. Indovina, Amlan Ganguly |
ICCD | 3 |
| 2018 | A 0.24pJ/bit, 16Gbps OOK Transmitter Circuit in 45-nm CMOS for Inter and Intra-Chip Wireless InterconnectsabstractResearch in recent years has demonstrated that intra and inter-chip wireless interconnects are capable of establishing energy-efficient data communications within as well as between multiple chips. This paper presents a circuit level design of an energy-efficient millimeter wave (mm-wave) on-off keying (OOK) transmitter suitable for such wireless interconnects in 45-nm CMOS process. The transmitter consists of an NMOS cross-coupled VCO, an OOK modulator and a power amplifier. The transmitter is able to achieve maximum modulation data rate of 16Gb/s at 60GHz with the output power of -3dBm consuming a total power of 3.9mW, which translates to a bit-energy efficiency of 0.24pJ/bit. Tanmay Shinde, Suryanarayanan Subramaniam, Padmanabh Deshmukh, M. Meraj Ahmed, Mark A. Indovina, Amlan Ganguly |
ACM Great Lakes Symposium on VLSI | 5 |
| 2018 | Testing WiNoC-Enabled Multicore Chips with BIST for Wireless InterconnectsabstractThe following topics are dealt with: network-on-chip; multiprocessing systems; network routing; system-on-chip; integrated circuit interconnections; microprocessor chips; telecommunication traffic; wireless channels; cache storage; and telecommunication network routing. Abhishek Vashist, Amlan Ganguly, Mark A. Indovina |
NOCS | 3 |