VLDB 2026 Research / reviewers in the wild / expert
Simranjeet Singh
dblp:278/2675
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8297-1470ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | veriSiM: Formal Verification of SPICE Netlists for MAGIC-Based Logic-in-MemoryabstractAdvancements in emerging technologies have recently increased the traction of non-von Neumann design styles. One of the most popular design styles in this domain involves using memristors to perform logic operations in memory, known as Logic-in-Memory (LiM). Memristor Aided Logic (MAGIC) is one of such LiM based design style that is widely used given its benefits in latency and energy. Several prior works have focused on the generation of logic operations, also called microoperations, for LiM based on the MAGIC design style. Recently, the generation of SPICE netlists for MAGIC design style has been achieved by the MemSPICE tool. While this represents a significant step forward, verifying the correctness of the generated netlists still depends on SPICE-level simulations. These simulations become particularly impractical for medium-to-large designs presenting a bottleneck in the validation process. To address this limitation, in this paper, we introduce veriSiM, an automated formal verification methodology for MAGIC-based LiM. More concretely, it ensures the correctness of the generated LiM SPICE netlists against the golden reference Verilog design. Our methodology involves generating clauses from the SPICE netlists and verifying them against clauses generated from the golden reference Verilog design, using the high-performance Z3 solver to perform the equivalence checking. The clause generation process from the SPICE netlists needs to be based on several conditions, which have been identified and discussed in detail. We have used several benchmarks from ISCAS’85, ISCAS’89, and ITC’99 to demonstrate the efficacy of the veri Chandan Kumar Jha 0001, Simranjeet Singh, Khushboo Qayyum, Ankit Bende, Muhammad Hassan 0002, Vikas Rana, Farhad Merchant, Rolf Drechsler |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Dependable Neuromorphic Computing-in-Memory Architectures
Farhad Merchant, Ankit Bende, Markus Fritscher, Shahar Kvatinsky, Simranjeet Singh, Vikas Rana, Regina Dittmann, Keerthi Dorai Swamy Reddy, Christian Wenger, Fouwad Jamil Mir, Mottaqiallah Taouil, Manil Dev Gomony, Said Hamdioui, Henk Corporaal |
ETS | 5 |
| 2024 | MemSPICE: Automated Simulation and Energy Estimation Framework for MAGIC-Based Logic-in-MemoryabstractExisting logic-in-memory (LiM) research is limited to generating mappings and micro-operations. In this paper, we present MemSPICE, a novel framework that addresses this gap by automatically generating both the netlist and testbench needed to evaluate the LiM on a memristive crossbar. MemSPICE goes beyond conventional approaches by providing energy estimation scripts to calculate the precise energy consumption of the testbench at the SPICE level. We propose an automated framework that utilizes the mapping obtained from the SIMPLER tool to perform accurate energy estimation through SPICE simulations. To the best of our knowledge, no existing framework is capable of generating a SPICE netlist from a hardware description language. By offering a comprehensive solution for SPICE-based netlist generation, testbench creation, and accurate energy estimation, MemSPICE empowers researchers and engineers working on memristor-based LiM to enhance their understanding and optimization of energy usage in these systems. Finally, we tested the circuits from the ISCAS’85 benchmark on MemSPICE and conducted a detailed energy analysis. Simranjeet Singh, Chandan Kumar Jha 0001, Ankit Bende, Vikas Rana, Sachin B. Patkar, Rolf Drechsler, Farhad Merchant |
ASPDAC | 1 |
| 2024 | Error Detection and Correction Codes for Safe In-Memory ComputationsabstractIn-Memory Computing (IMC) introduces a new paradigm of computation that offers high efficiency in terms of latency and power consumption for AI accelerators. However, the non-idealities and defects of emerging technologies used in advanced IMC can severely degrade the accuracy of inferred Neural Networks (NN) and lead to malfunctions in safety-critical applications. In this paper, we investigate an architectural-level mitigation technique based on the coordinated action of multiple checksum codes, to detect and correct errors at run-time. This implementation demonstrates higher efficiency in recovering accuracy across different AI algorithms and technologies compared to more traditional methods such as Triple Modular Redundancy (TMR). The results show that several configurations of our implementation recover more than 91% of the original accuracy with less than half of the area required by TMR and less than 40% of latency overhead. Luca Parrini, Taha Soliman, Benjamin Hettwer, Jan Micha Borrmann, Simranjeet Singh, Ankit Bende, Vikas Rana, Farhad Merchant, Norbert Wehn |
ETS | 5 |
| 2024 | In-Memory Mirroring: Cloning Without ReadingabstractIn-memory computing (IMC) has gained signifi- cant attention recently as it attempts to reduce the impact of memory bottlenecks. Numerous schemes for digital IMC are presented in the literature, focusing on logic operations. Often, an application's description has data dependencies that must be resolved. Contemporary IMC architectures perform read followed by write operations for this purpose, which results in performance and energy penalties. To solve this fundamental problem, this paper presents in-memory mirroring (IMM). IMM eliminates the need for read and write-back steps, thus avoiding energy and performance penalties. Instead, we perform data movement within memory, involving row-wise and column-wise data transfers. Additionally, the IMM scheme enables parallel cloning of entire row (word) with a complexity of O(1). Moreover, we analyzed the energy consumption of the proposed technique on an RRAM crossbar with an experimentally validated JART VCM v1b model. The IMM increases energy efficiency and shows 2x performance improvement compared to conventional data movement methods. Simranjeet Singh, Ankit Bende, Chandan Kumar Jha 0001, Vikas Rana, Rolf Drechsler, Sachin B. Patkar, Farhad Merchant |
VLSI-SoC | 1 |
| 2024 | veriSIMPLER: An Automated Formal Verification Methodology for SIMPLER MAGIC Design Style Based In-Memory ComputingabstractIn-Memory Computing (IMC) using memristors has gained significant interest in recent years as it addresses the issue of memory bottleneck in the von Neumann architectures. One of the most popular design styles that have been developed to perform memristor-based IMC is Memristor-Aided loGIC (MAGIC). MAGIC design style based NOR and NOT operations can be used to perform IMC on memristor crossbars. The state-of-the-art SIMPLER MAGIC tool is used to generate a mapping of any arbitrary Boolean function to MAGIC operations that are suitable for high-throughput applications. The correctness of the mapping is examined by tedious manual inspections and functional simulations which may not account for all the edge cases which is not desirable. In this work, we alleviate this issue to the best of our knowledge for the first time by proposing veriSIMPLER. veriSIMPLER is an automated formal verification methodology to ensure the functional correctness of the mapping obtained using the SIMPLER MAGIC tool. The veriSIMPLER methodology generates Boolean Satisfiability Formulas (SAT) of the mapping obtained using the SIMPLER MAGIC tool and the golden reference Verilog designs. These SAT formulas are verified against each other using the Z3 solver. The veriSIMPLER methodology identified a critical bug in the mapping obtained from the SIMPLER MAGIC tool when buffers are connected between input and output. We also propose a methodology to patch this bug to generate the correct mapping, which in turn extends the capability of the SIMPLER MAGIC tool to handle buffers on top of providing a formal verification methodology. We have used a variety of benchmark circuits from the widely used ISCAS’85, ISCAS’89, ITC’99, and IWLS’93 to show the efficacy of the veriSIMPLER methodology. We aim to make the formally verified mapping obtained using the veriSIMPLER methodology open-source to promote further research in this direction. Chandan Kumar Jha 0001, Khushboo Qayyum, Kemal Çaglar Coskun, Simranjeet Singh, Muhammad Hassan 0002, Rainer Leupers, Farhad Merchant, Rolf Drechsler |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Hardware Security Primitives Using Passive RRAM Crossbar Array: Novel TRNG and PUF DesignsabstractWith rapid advancements in electronic gadgets, the security and privacy aspects of these devices are significant. For the design of secure systems, physical unclonable function (PUF) and true random number generator (TRNG) are critical hardware security primitives for security applications. This paper proposes novel implementations of PUF and TRNGs on the RRAM crossbar structure. Firstly, two techniques to implement the TRNG in the RRAM crossbar are presented based on write-back and 50% switching probability pulse. The randomness of the proposed TRNGs is evaluated using the NIST test suite. Next, an architecture to implement the PUF in the RRAM crossbar is presented. The initial entropy source for the PUF is used from TRNGs, and challenge-response pairs (CRPs) are collected. The proposed PUF exploits the device variations and sneak-path current to produce unique CRPs. We demonstrate, through extensive experiments, reliability of 100%, uniqueness of 47.78%, uniformity of 49.79%, and bit-aliasing of 48.57% without any post-processing techniques. Finally, the design is compared with the literature to evaluate its implementation efficiency, which is clearly found to be superior to the state-of-the-art. Simranjeet Singh, Furqan Zahoor, Gokulnath Rajendran, Sachin B. Patkar, Anupam Chattopadhyay, Farhad Merchant |
ASP-DAC | 1 |
| 2023 | IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin MachinesabstractIn-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively. Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISLPED | 2 |
| 2023 | PR-PUF: A Reconfigurable Strong RRAM PUFabstractPhysical Unclonable Functions (PUFs) offer the natural advantage of built-in key generation, thus eliminating the costly process of embedding unique key after manufacturing millions of integrated circuits. When PUFs are deployed for the application, all the Challenge-Response Pairs (CRP) are collected and stored in a trusted server, and the responses are compared with the one from the device during the run-time - forming the crux of various security protocols. Two issues are commonly faced during PUF designs. First, to enhance the applicability of PUF, larger set of CRP is desirable, which is referred to as a strong PUF. Second, due to the emergence of machine learning-based PUF modelling attacks, it is now imperative to have a PUF demonstrating resistance against such attacks. In this paper, we propose a novel Parity Resistive RAM PUF (PR-PUF) implemented using RRAM crossbar architecture. PR-PUF supports low-overhead reconfiguration, where both the original and reconfigured CRP space enhances the CRP size, with average uniqueness between reconfiguration of 49.98%. The construction also demonstrates excellent robustness against various modeling attacks. We present detailed design analysis and circuit-level simulation studies. Gokulnath Rajendran, Furqan Zahoor, Simranjeet Singh, Farhad Merchant, Vikas Rana, Anupam Chattopadhyay |
VLSI-SoC | 3 |
| 2022 | LCPP: Low Computational Processing Pipeline for Delivery Robots
Soofiyan Atar, Simranjeet Singh, Srijan Agrawal, Ravikumar Chaurasia, Shreyas Sule, Sravya Gadamsetty, Aditya Panwar, Amit Kumar 0027, Kavi Arya |
ICAART (3) | 2 |
| 2022 | P2Ag: Perception Pipeline in Agriculture for Robotic Harvesting of Tomatoes
Soofiyan Atar, Simranjeet Singh, Jaison Jose, Kavi Arya |
ICAART (3) | 2 |
| 2022 | PA-PUF: A Novel Priority Arbiter PUFabstractThis paper proposes a 3-input arbiter-based novel physically unclonable function (PUF) design. Firstly, a 3-input priority arbiter is designed using a simple arbiter, two multiplexers (2:1), and an XOR logic gate. The priority arbiter has an equal probability of 0’s and 1’s at the output, which results in excellent uniformity (49.45%) while retrieving the PUF response. Secondly, a new PUF design based on priority arbiter PUF (PA-PUF) is presented. The PA-PUF design is evaluated for uniqueness, non-linearity, and uniformity against the standard tests. The proposed PA-PUF design is configurable in challenge-response pairs through an arbitrary number of feed-forward priority arbiters introduced to the design. We demonstrate, through extensive experiments, reliability of 100% after performing the error correction techniques and uniqueness of 49.63%. Finally, the design is compared with the literature to evaluate its implementation efficiency, where it is clearly found to be superior compared to the state-of-the-art. Simranjeet Singh, Srinivasu Bodapati 0001, Sachin B. Patkar, Rainer Leupers, Anupam Chattopadhyay, Farhad Merchant |
VLSI-SoC | 1 |
| 2020 | Real time System Implementation for Stereo 3D Mapping and Visual OdometryabstractDepth Estimation and localisation is a critical and computation heavy task. Due to its computation complexity, it is hard to achieve the high rate performance on the normal CPU. We here have as purpose compact, low power architecture for real time stereo depth estimation and stereo visual odometry that can be easily used on the UAV (unmanned aerial vehicle) and other autonomous navigation vehicles. A novel implementation of Stereo odometry based on careful feature selection and tracking [2] on GPU is described. It accelerates core computations like feature tracking, RANSAC based non linear solver using the GPU. 3D Mapping using stereo disparity estimation based on More Global Matching (MGM) a variant of SGM is implemented on FPGA. A Pipeline Architecture is introduced to increase throughput of the 3D map by leveraging multiple ARM cores. The programs are tested on Jetson Nano (an embedded GPU) and Ultra96 (ARM-FPGA Soc) which have less form factor and consume less power. Update rate of 20 is achieved for 6 degrees of freedom pose on Jetson Nano (60 on i7 core with Nvidia 1050ti) and an update rate of 16 fps is achieved for 3D point cloud on Ultra96 making the system desirable for above mentioned applications. Yashwant Temburu, Mandar Datar 0001, Simranjeet Singh, Vaibhav Malviya, Sachin B. Patkar |
IPAS | 3 |