EDBT 2026 Demo / reviewers in the wild / expert
Sachin B. Patkar
dblp:70/178
· DBLP profile ↗
19ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 since 2021Theory of computation · 8 · 6 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MemSPICE: Automated Simulation and Energy Estimation Framework for MAGIC-Based Logic-in-MemoryabstractExisting logic-in-memory (LiM) research is limited to generating mappings and micro-operations. In this paper, we present MemSPICE, a novel framework that addresses this gap by automatically generating both the netlist and testbench needed to evaluate the LiM on a memristive crossbar. MemSPICE goes beyond conventional approaches by providing energy estimation scripts to calculate the precise energy consumption of the testbench at the SPICE level. We propose an automated framework that utilizes the mapping obtained from the SIMPLER tool to perform accurate energy estimation through SPICE simulations. To the best of our knowledge, no existing framework is capable of generating a SPICE netlist from a hardware description language. By offering a comprehensive solution for SPICE-based netlist generation, testbench creation, and accurate energy estimation, MemSPICE empowers researchers and engineers working on memristor-based LiM to enhance their understanding and optimization of energy usage in these systems. Finally, we tested the circuits from the ISCAS’85 benchmark on MemSPICE and conducted a detailed energy analysis. Simranjeet Singh, Chandan Kumar Jha 0001, Ankit Bende, Vikas Rana, Sachin B. Patkar, Rolf Drechsler, Farhad Merchant |
ASPDAC | 5 |
| 2024 | In-Memory Mirroring: Cloning Without ReadingabstractIn-memory computing (IMC) has gained signifi- cant attention recently as it attempts to reduce the impact of memory bottlenecks. Numerous schemes for digital IMC are presented in the literature, focusing on logic operations. Often, an application's description has data dependencies that must be resolved. Contemporary IMC architectures perform read followed by write operations for this purpose, which results in performance and energy penalties. To solve this fundamental problem, this paper presents in-memory mirroring (IMM). IMM eliminates the need for read and write-back steps, thus avoiding energy and performance penalties. Instead, we perform data movement within memory, involving row-wise and column-wise data transfers. Additionally, the IMM scheme enables parallel cloning of entire row (word) with a complexity of O(1). Moreover, we analyzed the energy consumption of the proposed technique on an RRAM crossbar with an experimentally validated JART VCM v1b model. The IMM increases energy efficiency and shows 2x performance improvement compared to conventional data movement methods. Simranjeet Singh, Ankit Bende, Chandan Kumar Jha 0001, Vikas Rana, Rolf Drechsler, Sachin B. Patkar, Farhad Merchant |
VLSI-SoC | 6 |
| 2023 | Hardware Security Primitives Using Passive RRAM Crossbar Array: Novel TRNG and PUF DesignsabstractWith rapid advancements in electronic gadgets, the security and privacy aspects of these devices are significant. For the design of secure systems, physical unclonable function (PUF) and true random number generator (TRNG) are critical hardware security primitives for security applications. This paper proposes novel implementations of PUF and TRNGs on the RRAM crossbar structure. Firstly, two techniques to implement the TRNG in the RRAM crossbar are presented based on write-back and 50% switching probability pulse. The randomness of the proposed TRNGs is evaluated using the NIST test suite. Next, an architecture to implement the PUF in the RRAM crossbar is presented. The initial entropy source for the PUF is used from TRNGs, and challenge-response pairs (CRPs) are collected. The proposed PUF exploits the device variations and sneak-path current to produce unique CRPs. We demonstrate, through extensive experiments, reliability of 100%, uniqueness of 47.78%, uniformity of 49.79%, and bit-aliasing of 48.57% without any post-processing techniques. Finally, the design is compared with the literature to evaluate its implementation efficiency, which is clearly found to be superior to the state-of-the-art. Simranjeet Singh, Furqan Zahoor, Gokulnath Rajendran, Sachin B. Patkar, Anupam Chattopadhyay, Farhad Merchant |
ASP-DAC | 4 |
| 2023 | IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin MachinesabstractIn-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively. Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISLPED | 7 |
| 2023 | CLARINET: A quire-enabled RISC-V-based framework for posit arithmetic empiricism
Niraj N. Sharma, Riya Jain, Mohana Madhumita Pokkuluri, Sachin B. Patkar, Rainer Leupers, Rishiyur S. Nikhil, Farhad Merchant |
J. Syst. Archit. | 4 |
| 2022 | PA-PUF: A Novel Priority Arbiter PUFabstractThis paper proposes a 3-input arbiter-based novel physically unclonable function (PUF) design. Firstly, a 3-input priority arbiter is designed using a simple arbiter, two multiplexers (2:1), and an XOR logic gate. The priority arbiter has an equal probability of 0’s and 1’s at the output, which results in excellent uniformity (49.45%) while retrieving the PUF response. Secondly, a new PUF design based on priority arbiter PUF (PA-PUF) is presented. The PA-PUF design is evaluated for uniqueness, non-linearity, and uniformity against the standard tests. The proposed PA-PUF design is configurable in challenge-response pairs through an arbitrary number of feed-forward priority arbiters introduced to the design. We demonstrate, through extensive experiments, reliability of 100% after performing the error correction techniques and uniqueness of 49.63%. Finally, the design is compared with the literature to evaluate its implementation efficiency, where it is clearly found to be superior compared to the state-of-the-art. Simranjeet Singh, Srinivasu Bodapati 0001, Sachin B. Patkar, Rainer Leupers, Anupam Chattopadhyay, Farhad Merchant |
VLSI-SoC | 3 |
| 2020 | Real time System Implementation for Stereo 3D Mapping and Visual OdometryabstractDepth Estimation and localisation is a critical and computation heavy task. Due to its computation complexity, it is hard to achieve the high rate performance on the normal CPU. We here have as purpose compact, low power architecture for real time stereo depth estimation and stereo visual odometry that can be easily used on the UAV (unmanned aerial vehicle) and other autonomous navigation vehicles. A novel implementation of Stereo odometry based on careful feature selection and tracking [2] on GPU is described. It accelerates core computations like feature tracking, RANSAC based non linear solver using the GPU. 3D Mapping using stereo disparity estimation based on More Global Matching (MGM) a variant of SGM is implemented on FPGA. A Pipeline Architecture is introduced to increase throughput of the 3D map by leveraging multiple ARM cores. The programs are tested on Jetson Nano (an embedded GPU) and Ultra96 (ARM-FPGA Soc) which have less form factor and consume less power. Update rate of 20 is achieved for 6 degrees of freedom pose on Jetson Nano (60 on i7 core with Nvidia 1050ti) and an update rate of 16 fps is achieved for 3D point cloud on Ultra96 making the system desirable for above mentioned applications. Yashwant Temburu, Mandar Datar 0001, Simranjeet Singh, Vaibhav Malviya, Sachin B. Patkar |
IPAS | 5 |
| 2014 | Storage-allocation to sequential structures in High-Level Synthesis-assisted prototypingabstractAlgorithm-to-hardware High-Level Synthesis (HLS) tools are becoming increasingly practical, particularly the domain specific approaches to HLS. Storage allocation is an important step in HLS where variables are mapped to onchip storage structures (OSS). HLS flows almost exclusively do this allocation to random access OSS whereas custom designers often pick from a repertoire of intuitive algorithm-appropriate OSS. In this work, revisiting a sparsely addressed problem of storage allocation to sequential access OSS, we report tractable algorithms for storage allocation to three kinds of sequential access style memories-Queue, Queue-Read Sequential-Write memory (QRSWM), and Sequential-Read Sequential-Write memory (SRSWM)-suitable for domains such as signal processing and matrix computations. A basic C to Verilog HLS flow was developed integrating these new allocations options to evaluate their impact on the overall design metrics of interest such as area and power. On application cases such as matrix multiplication and 2D/3D wavelet filtering, a comparison vis-a-vis RAM shows significant improvements in power consumption when targeting TSMC 0.18μ technology. Vinay B. Y. Kumar, Shovan Maity, Sachin B. Patkar |
ICCD | 3 |
| 2013 | Solution of PDEs-electrically coupled systems with electrical analogy
Yogesh Dilip Save, H. Narayanan, Sachin B. Patkar |
Integr. | 3 |
| 2009 | Acceleration of conjugate gradient method for circuit simulation using CUDAabstractThe Conjugate Gradient method is a popular iterative method to solve a system of linear equations and is used in a variety of applications. The DC Analyser is a circuit simulator built at IIT Bombay to solve large circuits containing resistances, voltage and current sources and which employs the conjugate gradient method. Current generation of graphics cards offer extremely high raw processing power and memory bandwidths compared to conventional CPUs. We have accelerated the conjugate gradient part of the DC Analyser using an Nvidia GTX 280 GPU and the new CUDA technology and successfully obtained a speedup of over 10× for the CG method and more than 4× for the entire application for very large circuits when compared to a single-threaded CPU implementation. Anirudh Maringanti, Viraj Athavale, Sachin B. Patkar |
HiPC | 3 |
| 2009 | A Pipelined Simulation Approach for Logic Emulation using Multi-FPGA PlatformsabstractEmulation of a large system on a multi-FPGA platform not only involves partitioning the system into multiple modules subject to given capacity and resource constraints, but also involves achieving higher throughput, lower cost of emulation and less communication overhead. Many good scheduling algorithms have been reported, however due to the lack of pipelining they fail to achieve high system throughput. An intelligent hardware scheduling approach is essential for obtaining high system throughput with possibly lower overheads. In this paper, we propose a scalable, high performance, low cost approach for simulation of multi-FPGA systems. We convert the unbalanced partitioned system into a balanced pipeline and maximize the throughput of the system. Our experiments on reference designs have shown a speed-up of up to 8.57times with a 10% hardware overhead over the conventional simulation approaches. Dinesh Baviskar, Sachin B. Patkar |
ISCAS | 2 |
| 2003 | The realization of finite state machines by decomposition and the principal lattice of partitions of a submodular function
Madhav P. Desai, H. Narayanan, Sachin B. Patkar |
Discret. Appl. Math. | 3 |
| 2003 | Improving graph partitions using submodular functions
Sachin B. Patkar, H. Narayanan |
Discret. Appl. Math. | 1 |
| 2001 | A note on optimal covering augmentation for graphic polymatroids
Sachin B. Patkar, H. Narayanan |
Inf. Process. Lett. | 1 |
| 2000 | Fast On-Line/Off-Line Algorithms for Optimal Reinforcement of a Network and Its Connections with Principal Partition
Sachin B. Patkar, H. Narayanan |
FSTTCS | 1 |
| 1994 | Approximation Algorithms for Min-k-overlap Problems Using the Principal Lattice of Partitions Approach
H. Narayanan, Subir K. Roy, Sachin B. Patkar |
MFCS | 3 |
| 1992 | Fast Sequential and Randomised Parallel Algorithms for Rigidity and approximate Min k-cut
Sachin B. Patkar, H. Narayanan |
FSTTCS | 1 |
| 1992 | Principal Lattice of Partition of submodular functions on Graphs: Fast algorithms for Principal Partition and Generic Rigidity
Sachin B. Patkar, H. Narayanan |
ISAAC | 1 |
| 1991 | A Fast Algorithm for the Principle Partition of a Graph
Sachin B. Patkar, H. Narayanan |
FSTTCS | 1 |