Shanmukha Mangadahalli Siddaramu

dblp:390/5032 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0008-9226-894XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SHOUT-Trainer: Closed-loop Trainer for Silent Data Corruption Hunting and Observation Using Transformers
Seyedeh Maryam Ghasemi, Shanmukha Mangadahalli Siddaramu, Mehdi Baradaran Tahoori
IOLTS2
2026 LOFT: Latent-Fault Optimization Training for Yield Boost in Resistive Crossbar AI Accelerators
Shanmukha Mangadahalli Siddaramu, Mehdi Baradaran Tahoori
IOLTS1
2026 SHOUT - Silent Data Corruption Hunting and Observation Using Transformers
Seyedeh Maryam Ghasemi, Shanmukha Mangadahalli Siddaramu, Tara Gheshlaghi, Sani R. Nassif, Mehdi Baradaran Tahoori
VTS2
2026 Temporal Reference Scouting Logic for PVT Reliable Logic Computation-in-Memory
Shanmukha Mangadahalli Siddaramu, Ali Nezhadi, Mahta Mayahinia, Sule Ozev, Mehdi Baradaran Tahoori
VTS1
2025 Testing of Passive Memristive Crossbars in AI Hardware Accelerators
abstract
Memristor-based computation-in-memory (CiM) architectures address the growing computational demands of AI accelerators by enabling analog matrix-vector multiplication (MVM) directly within the memory array. Passive (selectorless) memristive crossbars are particularly attractive for such architectures due to their high density and compatibility with back-end-of-line fabrication. Here, AI model parameters are stored as memristor conductances, and MVM is performed in situ by activating multiple rows simultaneously and sensing the resulting column currents. As a result, column currents become the primary observable, making column-level fault detection more relevant for AI workloads than conventional March-based cell-level tests, which are specially time - and energy-intensive for passive crossbars due to specialized biasing methods. To address this, we propose a current-based testing methodology tailored to passive crossbars that detects and diagnoses faulty columns while accounting for non-idealities such as sneak-path currents, line resistance, and process variations. Faulty columns are detected by programming all cells to a uniform state and measuring column-current deviations, followed by targeted test patterns to estimate per-column fault density. On a $64 \times 64$ array, the proposed approach achieves 100% faulty column detection, 98% overall fault coverage, and a $2.3 \times$ speed-up in test time compared to conventional March tests. Hence, providing a fast and efficient solution for manufacturing screening of passive crossbars with sufficient diagnostic resolution to support systemlevel fault tolerance for AI applications.
Shanmukha Mangadahalli Siddaramu, Mahta Mayahinia, Surendra Hemaram, Sule Ozev, Mehdi Baradaran Tahoori
ATS1
2024 Hardware and Software Co-Design for Optimized Decoding Schemes and Application Mapping in NVM Compute-in-Memory Architectures
abstract
The computation-in nonvolatile memory (NVM-CiM) approach addresses the growing computational demands and the memory-wall problem faced by traditional processor-centric architectures. Computation-in-memory (CiM) capitalizes on the parallel nature of memory arrays enabling effective computation through multirow memristor reading and sensing. In this context, the conventional design of memory decoders needs to be accordingly modified for efficient multirow activation and parallel data processing. This article presents the design and optimization of address decoders for NVM-CiM system architectures, employing a cross-layer co-optimization approach that integrates circuit and architecture design with application requirements. Our methodology starts at the circuit level, examining various decoder designs, including cascaded, hierarchical, latched, and hybrid models. An in-depth application-level characterization follows, utilizing an extended NVM-CiM-capable gem5 simulator to assess the impact of these decoders on the mapping of CiM-friendly applications and the resulting system performance, particularly in facilitating rapid and efficient activation of multirow memory configurations. This holistic analysis allows us to identify the bottlenecks and requirements from the application side and adjust the design of the decoder accordingly. Our analysis reveals that Hybrid Decoders significantly decrease latency and power consumption compared to other decoder designs within NVM-CiM systems. This highlights the crucial role of the decoder’s row selection flexibility, reducing additional system-level data movement even at the expense of its performance, can substantially improve the overall efficiency of NVM-CiM systems.
Shanmukha Mangadahalli Siddaramu, Ali Nezhadi, Mahta Mayahinia, Seyedeh Maryam Ghasemi, Mehdi Baradaran Tahoori
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1