Soumil Jain

dblp:274/3512 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-9549-9724ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Adiabatic Energy-Recycling Charge-Redistribution Array for Ultra-Low Power Massively Parallel Cross-Correlation
Shashank Bansal, Pål Gunnar Hogganvik, Soumil Jain, Gopabandu Hota, Bouchaib Cherif, Johannes Leugering, Gert Cauwenberghs
ISCAS3
2025 Clo-HDnn: Continual On-Device Learning Accelerator with Hyperdimensional Computing via Progressive Search
abstract
Clo-HDnn is an on-device learning (ODL) accelerator designed for emerging continual learning (CL) tasks. Clo-HDnn integrates hyperdimensional computing (HDC) along with low-cost Kronecker HD Encoder and weight clustering feature extraction (WCFE) to optimize accuracy and efficiency. Clo-HDnn adopts gradient-free CL to efficiently update and store the learned knowledge in the form of class hypervectors. Its dual-mode operation enables bypassing costly feature ex- traction for simpler datasets, while progressive search reduces complexity by up to $61 \%$ by encoding and comparing only partial query hypervectors. Achieving 4.66 TFLOPS/W (FE) and 3.78 TOPS/W (classifier), Clo-HDnn delivers $7.77 \times$ and $4.85 \times$ higher energy efficiency compared to SOTA ODL accelerators.
Chang Eun Song, Keming Fan, Soumil Jain, Gopabandhu Hota, Haichao Yang, Leo Liu, Meng-Fan Chang, Carlos H. Diaz, Gert Cauwenberghs, Tajana Rosing, Mingu Kang
HCS4
2024 Bio-plausible Learning-on-Chip with Selector-less Memristive Crossbars
abstract
One of the practical realizations of large-scale neuromorphic systems requires an area-efficient memristive crossbar array as a key building block supporting high-density synaptic connectivity. Conventional memristor-based AI accelerators rely on selector transistors to reduce sneak path-induced cross-talks, although other means can be equally effective. Removing the selector element on each memristor cross-point significantly improves array density (down to 4F2) and lowers power consumption. We present an integrated reconfigurable neuromorphic platform interfacing a selector-less 16x16 RRAM memristor crossbar array with peripheral row and column instrumentation for robust learning and inference with applications to AI on the edge. Bio-plausible local Hebbian-like incremental outer-product learning rules are mapped onto direct implementation across the memristive crossbar array, updated in a sequence of partial outer-product combinations presented at the periphery of the array. Our system provides a user-configurable platform to accommodate a broad spectrum of emerging non-volatile memory device technologies for synaptic crossbar arrays with embedded adaptive functionality for general AI and cognitive neuromorphic computing.
Jeong-Hoon Kim, Soumil Jain, Gopabandhu Hota, Jaeseoung Park, Duygu Kuzum, Gert Cauwenberghs
ISCAS2
2023 A Versatile and Efficient Neuromorphic Platform for Compute-in-Memory with Selector-less Memristive Crossbars
abstract
Memristive crossbar arrays have become essential building blocks in the realization of large-scale neuromorphic systems with high-density synaptic connectivity. Traditionally, memristor-based accelerators are equipped with selector elements to reduce cross-talk through sneak paths along unselected lines. However, due to the large drive strength required for selector elements, it comes at the cost of synaptic crossbar density. Selector- less alternatives require careful design of crossbar peripheral circuits to mitigate or eliminate sneak path-induced cross-talk. We propose a hybrid integrated platform that interfaces a selector- less memristor crossbar array with peripheral row and column instrumentation for array-parallel programming and readout for AI learning and inference applications. The proposed switched-capacitor voltage-sensing instrumentation avoids the need for current-sensing schemes with voltage-clamped sense lines that are typically used to mitigate the sneak path issues in selector- less crossbars but are substantially less energy-efficient than voltage-sensing. Our board-level platform is implemented using commercial off-the-shelf (COTS) data converters and switched capacitors, and is controlled by a Xilinx Spartan-6 FPGA. The system offers programmable sense times to characterize memristors over a wide range of resistances and the capability to switch between a transient-domain measurement and steady-state measurement to offer the desired trade-off between accuracy and energy efficiency during inference parallel readout. We implement a differential weight-encoding scheme to improve the accuracy of matrix-vector multiplication. The system also supports an array-level programming scheme for parallel write access as well as online learning-in-memory for neuromorphic applications through outer-product incremental decomposition of the weight matrix. Thus, our system offers a generic, user-configurable, and versatile platform to support wide dynamic range measurements of synaptic crossbar arrays and cognitive neuromorphic computing with emerging non-volatile memory devices.
Soumil Jain, Gopabandhu Hota, Sangheon Oh, Jiajia Wu 0008, Preston Fowler, Duygu Kuzum, Gert Cauwenberghs
ISCAS1