EDBT 2026 Demo / reviewers in the wild / expert
Vishal Sharma 0004
dblp:20/6234-4
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-2618-0655ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlexDCIM: A 400 MHz 249.1 TOPS/W 64 Kb Flexible Digital Compute-in-Memory SRAM Macro for CNN AccelerationabstractThis work proposes a 64Kb fully reconfigurable SRAM compute-in-memory (CIM) macro for convolutional neural network (CNN) acceleration using a 65nm node. It supports operation up to 400 MHz. The fully digital operation of the proposed macro effectively removes the analog CIM design issues related to process variations, noise susceptibility, and data-conversion overhead. Hence, it offers no accuracy loss, high energy efficiency, and large area saving for computation. To support the digital computation, a new area-efficient Digital Processing Unit (DPU) is proposed which is equivalent to 8.75T per bit storage. Moreover, the proposed macro features full precision reconfigurability (1b to 8b) for both input and weight, and fully flexible input activation ranging from 1 to 64 parallel inputs. It makes the proposed macro feasible for different neural network topologies. Removing sense amplifiers (SAs) for the memory mode of the proposed design suggests additional area and power savings. The proposed CIM macro achieves an energy efficiency of 249.1TOPS/W and a throughput of 819.2 GOPS. Vishal Sharma 0004, Xin Zhang 0025, Narendra Singh Dhakad, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | A 64 Kb Reconfigurable Full-Precision Digital ReRAM-Based Compute-In-Memory for Artificial Intelligence ApplicationsabstractThis work presents a fully-digital 64 Kb non-volatile ReRAM based compute-in-memory (CIM) macro for the modern artificial intelligence (AI) edge devices, using 65 nm technology. This digital CIM architecture effectively removes the analog-design issues, related to process variations, noise susceptibility, and data-conversion overhead. Hence, it offers no accuracy loss and high energy-efficiency for the computation. To incorporate the digital computation, a novel NAND logic based 3.25T1R bitcell is proposed. The digital behaviour of this cell makes it superior to the conventional 1T1R based analog bitcell. Also, with the inherent non-volatility of ReRAM, the proposed cell can be a good substitute for SRAM-based CIM architectures with$4.62\times $,$1.96\times $,$3.96\times $, and$5.12\times $lower area than the XNOR-based 12T, Twin-8T, 8T, and 6T SRAM cell respectively. Moreover, the proposed CIM architecture allows full reconfigurabiliy from 1 to 16b precision for both input and weight. It also allows activating any number of parallel inputs, ranging from 1 to 128. According to simulation results, the proposed macro successfully operates up to 166.6 MHz for 1/8/15b input/weight/output precision and achieves 27.28 TOPS/W without any accuracy loss. Removing sense amplifiers for the ReRAM mode of the proposed work claims additional area and power savings. Vishal Sharma 0004, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | AND8T SRAM Macro with Improved Linearity for Multi-Bit In-Memory ComputingabstractIn this work, we propose a multi-bit precision (4b input, 4b weight and 4b output) in-memory computing (IMC) architecture, based on the voltage scaling and charge sharing scheme, for the artificial intelligence (AI) edge devices. To achieve the efficient computation, a new AND logic based 8T SRAM cell (AND8T) has been used which employs the charge-domain based computation. For such computation, AND8T incorporates an overlaying metal-oxide-metal capacitor (MOM cap) with no bit-cell area overhead. The proposed cell mitigates the linearity issue of multiply and accumulate (MAC) operation for the IMC unit which is highly desirable for the reliable operation of complex neural networks (CNN). Moreover, our high precision AND8T based IMC architecture allows 128 parallel MAC operations avoiding the need of serial multi-bits input implementation through multiple cycles. The proposed design has been successfully verified by the monte carlo simulation results while working at 50MHz clock frequency and 1V supply using standard 65nm node. Vishal Sharma 0004, Ju Eon Kim, Yuzong Chen 0001, Tony Tae-Hyoung Kim |
ISCAS | 1 |