Akhilesh Jaiswal 0001

dblp:176/5075 · also Akhilesh R. Jaiswal · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-9911-2624ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 GEM3D-CIM: General Purpose Matrix Computation Using 3-D-Integrated SRAM-eDRAM Hybrid Compute-In-Memory-on-Memory Architecture
abstract
With the rapid growth of deep neural networks (DNNs), compute-in-memory (CIM) has emerged as a promising energy-efficient paradigm for accelerating multiply-and-accumulate (MAC) operations. Yet, current CIM architectures are largely limited to dot-product computations and struggle to efficiently support general-purpose matrix operations, such as transpose, element-wise addition, and multiplication. This work presents a 3D-integrated, memory-on-memory SRAM-eDRAM hybrid CIM architecture, implemented in GlobalFoundries 22 nm FDSOI technology, capable of performing general matrix operations directly within the memory crossbar with 4-bit precision. By leveraging a specialized transpose-based architecture, in-memory arithmetic operations, peripheral-aware design, and 3D SRAM–eDRAM integration, the proposed architecture balances latency, energy efficiency, and compute density for general purpose matrix operations while remaining compatible with the conventional CIM dot product architectures. Overall, this memory-on-memory CIM framework generalizes CIM beyond dot products, enabling versatile matrix processing and paving the way for broader applications in AI acceleration and general-purpose high performance computing.
Subhradip Chakraborty, Akhilesh Jaiswal 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 X-pSRAM: A Photonic SRAM with Embedded XOR Logic for Ultra-Fast in-Memory Computing
abstract
Traditional von Neumann architectures suffer from fundamental bottlenecks due to continuous data movement between memory and processing units, a challenge that worsens with technology scaling as electrical interconnect delays become more significant. These limitations impede the performance and energy efficiency required for modern data-intensive applications. In contrast, photonic in-memory computing presents a promising alternative by harnessing the advantages of light, enabling ultra-fast data propagation without length-dependent impedance, thereby significantly reducing computational latency and energy consumption. This work proposes a novel differential photonic static random access memory (pSRAM) bitcell that facilitates electro-optic data storage while enabling ultra-fast in-memory Boolean XOR computation. By employing cross-coupled microring resonators and differential photodiodes, the XOR-augmented pSRAM (X-pSRAM) bitcell achieves at least 10 GHz read, write, and compute operations entirely in the optical domain. Additionally, wavelength-division multiplexing (WDM) enables n-bit XOR computation in a single-shot operation, supporting massively parallel processing and enhanced computational efficiency. Validated on GlobalFoundries' 45SPCLO node, the XpSRAM consumed 13.2 fJ energy per bit for XOR computation, representing a significant advancement toward next-generation optical computing with applications in cryptography, hyperdimensional computing, and neural networks.
Md. Abdullah-Al Kaiser, Sugeet Sunder, Ajey P. Jacob, Akhilesh Jaiswal 0001
ASAP4
2025 A Mixed-Signal Photonic SRAM-based High-Speed Energy-Efficient Photonic Tensor Core with Novel Electro-Optic ADC
abstract
The rapid surge in data generated by Internet of Things (IoT), artificial intelligence (AI), and machine learning (ML) applications demands ultra-fast, scalable, and energy-efficient hardware, as traditional von Neumann architectures face significant latency and power challenges due to data transfer bottlenecks between memory and processing units. Furthermore, conventional electrical memory technologies are increasingly constrained by rising bitline and wordline capacitance, as well as the resistance of compact and long interconnects, as technology scales. In contrast, photonics-based in-memory computing systems offer substantial speed and energy improvements over traditional transistor-based systems, owing to their ultra-fast operating frequencies, low crosstalk, and high data bandwidth. Hence, we present a novel differential photonic SRAM (pSRAM) bitcell-augmented scalable mixed-signal multi-bit photonic tensor core, enabling high-speed, energy-efficient matrix multiplication operations using fabrication-friendly integrated photonic components. Additionally, we propose a novel 1-hot encoding electro-optic analog-to-digital converter (eoADC) architecture to convert the multiplication outputs into digital bitstreams, supporting processing in the electrical domain. Our designed photonic tensor core, utilizing GlobalFoundries’ monolithic 45SPCLO technology node, achieves computation speeds of 4.10 tera-operations per second (TOPS) and a power efficiency of 3.02 TOPS/W.
Md. Abdullah-Al Kaiser, Sugeet Sunder, Ajey P. Jacob, Akhilesh Jaiswal 0001
DAC4
2025 Reconfigurable Retina-Inspired Looming Detection
abstract
Recent advances in retinal neuroscience inspired the development of hardware and software systems that leverage evolutionarily derived retinal computations for real-world computer vision applications. In this work, we propose a novel, reconfigurable CMOS circuit designed specifically for Looming Detection (LD), a key retinal computation associated with detecting rapidly approaching objects and potential threats. We analyze the circuit's performance using real-world data from controlled laboratory environments and hardware-aware algorithmic simulations, demonstrating its accuracy in identifying potential collisions and avoiding false-positives from mundane object movements. Furthermore, we evaluate the CMOS hardware characteristics using GlobalFoundries' 22nm FD-SOI technology, exhibiting a 0.36 pJ energy consumption per looming spike for a unit kernel size of 5 × 5 pixels. This work contributes a foundational approach to adaptive hardware-software co-design by integrating advances in retinal neuroscience with modern CMOS technology to address real-time LD with in-sensor computations.
Jason Sinaga, Shay Snyder, Md. Abdullah-Al Kaiser, Dan Jinoy, Gregory W. Schwartz, Maryam Parsa, Akhilesh Jaiswal 0001
FCCM7
2024 Toward High-Accuracy, Programmable Extreme-Edge Intelligence for Neuromorphic Vision Sensors utilizing Magnetic Domain Wall Motion-based MTJ
abstract
The desire to empower resource-limited edge devices with computer vision (CV) must overcome the high energy consumption of collecting and processing vast sensory data. To address the challenge, this work proposes an energy-efficient non-von-Neumann in-pixel processing solution for neuromorphic vision sensors employing emerging (X) magnetic domain wall magnetic tunnel junction (MDWMTJ) for the first time, in conjunction with CMOS-based neuromorphic pixels. Our hybrid CMOS+X approach performs in-situ massively parallel asynchronous analog convolution, exhibiting low power consumption and high accuracy across various CV applications by leveraging the non-volatility and programmability of the MDWMTJ. Moreover, our developed device-circuit-algorithm co-design framework captures device constraints (low tunnel-magnetoresistance, low dynamic range) and circuit constraints (non-linearity, process variation, area consideration) based on monte-carlo simulations and device parameters utilizing GF22nm FD-SOI technology. Our experimental results suggest we can achieve an average of 45.3% reduction in backend-processor energy, maintaining similar front-end energy compared to the state-of-the-art and high accuracy of 79.17% and 95.99% on the DVS-CIFAR10 and IBM DVS128-Gesture datasets, respectively.
Md. Abdullah-Al Kaiser, Gourav Datta, Peter A. Beerel, Akhilesh Jaiswal 0001
DAC4
2024 A 9 Transistor SRAM Featuring Array-level XOR Parallelism with Secure Data Toggling Operation
abstract
Security, speed, and energy efficiency are critical for computing applications in general and for edge applications in particular. Digital In-Memory Computing (IMC) in SRAM cells has widely been studied to accelerate inference tasks to maximize both throughput and energy efficiency for intelligent computing at the edge. XOR operations have been of particular interest due to their wide applicability in numerous applications that include binary neural networks and encryption. Based on monte-carlo simulation on commercial Globalfoundries 22nm node, we are proposing a novel 9T SRAM cell enabling multiple rows of data (entire array) to be XORed in a massively parallel single cycle fashion, compared to XOR of two rows in prior IMC works. The new cell also supports array-level data-toggling, using the least number of transistors compared to previous works, within the SRAM cell efficiently to circumvent imprinting attacks and erase the SRAM value in case of a remanence attack. Our proposed SRAM thus combines security with in-memory computing while achieving significant area efficiency.
Zihan Yin, Annewsha Datta, Shwetha Vijayakumar, Ajey P. Jacob, Akhilesh Jaiswal 0001
ACM Great Lakes Symposium on VLSI5
2024 Recent Advances in Scalable Energy-Efficient and Trustworthy Spiking Neural Networks: from Algorithms to Technology
abstract
Neuromorphic computing and, in particular, spiking neural networks (SNNs) have become an attractive alternative to deep neural networks for a broad range of signal processing applications, processing static and/or temporal inputs from different sensory modalities, including audio and vision sensors. In this paper, we start with a description of recent advances in algorithmic and optimization innovations to efficiently train and scale low-latency, and energy-efficient spiking neural networks (SNNs) for complex machine learning applications. We then discuss the recent efforts in algorithm-architecture co-design that explores the inherent trade-offs between achieving high energy-efficiency and low latency while still providing high accuracy and trustworthiness. We then describe the underlying hardware that has been developed to leverage such algorithmic innovations in an efficient way. In particular, we describe a hybrid method to integrate significant portions of the model’s computation within both memory components as well as the sensor itself. Finally, we discuss the potential path forward for research in building deployable SNN systems identifying key challenges in the algorithm-hardware-application co-design space with an emphasis on trustworthiness.
Souvik Kundu 0002, Rui-Jie Zhu 0003, Akhilesh Jaiswal 0001, Peter A. Beerel
ICASSP3
2024 Hardware-Algorithm Co-Design Enabling Processing-In-Pixel-In-Memory (P2M) for Neuromorphic Vision Sensors
abstract
The high volume of data transmission between the edge sensor and the cloud processor leads to energy and throughput bottlenecks for resource-constrained edge devices focused on computer vision. Hence, researchers are investigating different approaches (e.g., near-sensor processing, in-sensor processing, in-pixel processing) by executing computations closer to the sensor to reduce the transmission bandwidth. Specifically, in-pixel processing for neuromorphic vision sensors (e.g., dynamic vision sensors (DVS)) involves incorporating asynchronous multiply-accumulate (MAC) operations within the pixel array, resulting in improved energy efficiency. In a CMOS implementation, low overhead energy-efficient analog MAC accumulates charges on a passive capacitor; however, the capacitor’s limited charge retention time affects the algorithmic integration time choices, impacting the algorithmic accuracy, bandwidth, energy, and training efficiency. Consequently, this results in a design trade-off on the hardware aspect- creating a need for a low-leakage compute unit while maintaining the area and energy benefits. In this work, we present a holistic analysis of the hardware-algorithm codesign trade-off based on the limited integration time posed by the hardware and techniques to improve the leakage performance of the in-pixel analog MAC operations.
Md. Abdullah-Al Kaiser, Akhilesh Jaiswal 0001
ICASSP2
2023 Technology-Circuit-Algorithm Tri-Design for Processing-in-Pixel-in-Memory (P2M)
abstract
The massive amounts of data generated by camera sensors motivate data processing inside pixel arrays, i.e., at the extreme-edge. Several critical developments have fueled recent interest in the processing-in-pixel-in-memory paradigm for a wide range of visual machine intelligence tasks, including (1) advances in 3D integration technology to enable complex processing inside each pixel in a 3D integrated manner while maintaining pixel density, (2) analog processing circuit techniques for massively parallel low-energy in-pixel computations, and (3) algorithmic techniques to mitigate non-idealities associated with analog processing through hardware-aware training schemes. This article presents a comprehensive technology-circuit-algorithm landscape that connects technology capabilities, circuit design strategies, and algorithmic optimizations to power, performance, area, bandwidth reduction, and application-level accuracy metrics. We present our results using a comprehensive co-design framework incorporating hardware and algorithmic optimizations for various complex real-life visual intelligence tasks mapped onto our P2M paradigm.
Md. Abdullah-Al Kaiser, Gourav Datta, Sreetama Sarkar, Souvik Kundu 0002, Zihan Yin, Manas Garg, Ajey P. Jacob, Peter A. Beerel, Akhilesh Jaiswal 0001
ACM Great Lakes Symposium on VLSI9
2023 A Context-Switching/Dual-Context ROM Augmented RAM using Standard 8T SRAM
abstract
The landscape of emerging applications has been continually widening, encompassing various data-intensive applications like artificial intelligence, machine learning, secure encryption, Internet-of-Things, etc. A sustainable approach toward creating dedicated hardware platforms that can cater to multiple applications often requires the underlying hardware to context-switch or support more than one context simultaneously. This paper presents a context-switching and dual-context memory based on the standard 8T SRAM bit-cell. Specifically, we exploit the availability of multi-VT transistors by selectively choosing the read-port transistors of the 8T SRAM cell to be either high-VT or low-VT. The 8T SRAM cell is thus augmented to store ROM data (represented as the VT of the transistors constituting the read-port) while simultaneously storing RAM data. Further, we propose specific sensing methodologies such that the memory array can support RAM-only or ROM-only mode (context-switching (CS) mode) or RAM and ROM mode simultaneously (dual-context (DC) mode). Extensive Monte-Carlo simulations have verified the robustness of our proposed ROM-augmented CS/DC memory on the Globalfoundries 22nm-FDX technology node.
Md. Abdullah-Al Kaiser, Edwin Tieu, Ajey P. Jacob, Akhilesh Jaiswal 0001
ACM Great Lakes Symposium on VLSI4
2023 In-Sensor & Neuromorphic Computing Are all You Need for Energy Efficient Computer Vision
abstract
Due to the high activation sparsity and use of accumulates (AC) instead of expensive multiply-and-accumulates (MAC), neuromorphic spiking neural networks (SNNs) have emerged as a promising low-power alternative to traditional DNNs for several computer vision (CV) applications. However, most existing SNNs require multiple time steps for acceptable inference accuracy, hindering real-time deployment and increasing spiking activity and, consequently, energy consumption. Recent works proposed direct encoding that directly feeds the analog pixel values in the first layer of the SNN in order to significantly reduce the number of time steps. Although the overhead for the first layer MACs with direct encoding is negligible for deep SNNs and the CV processing is efficient using SNNs, the data transfer between the image sensors and the downstream processing costs significant bandwidth and may dominate the total energy. To mitigate this concern, we propose an in-sensor computing hardware-software co-design framework for SNNs targeting image recognition tasks. Our approach reduces the bandwidth between sensing and processing by 12−96× and the resulting total energy by 2.32× compared to traditional CV processing, with a 3.8% reduction in accuracy on ImageNet.
Gourav Datta, Zeyu Liu 0003, Md. Abdullah-Al Kaiser, Souvik Kundu 0002, Joe Mathai, Zihan Yin, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel
ICASSP8
2023 Enabling ISPless Low-Power Computer Vision
abstract
Current computer vision (CV) systems use an image signal processing (ISP) unit to convert the high resolution raw images captured by image sensors to visually pleasing RGB images. Typically, CV models are trained on these RGB images and have yielded state-of-the-art (SOTA) performance on a wide range of complex vision tasks, such as object detection. In addition, in order to deploy these models on resource-constrained low-power devices, recent works have proposed in-sensor and in-pixel computing approaches that try to partly/fully bypass the ISP and yield significant bandwidth reduction between the image sensor and the CV processing unit by downsampling the activation maps in the initial convolutional neural network (CNN) layers. However, direct inference on the raw images degrades the test accuracy due to the difference in covariance of the raw images captured by the image sensors compared to the ISP-processed images used for training. Moreover, it is difficult to train deep CV models on raw images, because most (if not all) large-scale open-source datasets consist of RGB images. To mitigate this concern, we propose to invert the ISP pipeline, which can convert the RGB images of any dataset to its raw counterparts, and enable model training on raw images. We release the raw version of the COCO dataset, a large-scale benchmark for generic high-level vision tasks. For ISP-less CV systems, training on these raw images result in a ∼7.1% increase in test accuracy on the visual wake works (VWW) dataset compared to relying on training with traditional ISP-processed RGB datasets. To further improve the accuracy of ISP-less CV models and to increase the energy and bandwidth benefits obtained by in-sensor/in-pixel computing, we propose an energy-efficient form of analog in-pixel demosaicing that may be coupled with in-pixel CNN computations. When evaluated on raw images captured by real sensors from the PASCALRAW dataset, our approach results in a 8.1% increase in mAP. Lastly, we demonstrate a further 20.5% increase in mAP by using a novel application of few-shot learning with thirty shots each for the novel PASCALRAW dataset, constituting 3 classes. Codes are available at https://github.com/godatta/ISP-less-CV.
Gourav Datta, Zeyu Liu 0003, Zihan Yin, Linyu Sun, Akhilesh Jaiswal 0001, Peter A. Beerel
WACV5
2022 P2M-DeTrack: Processing-in-Pixel-in-Memory for Energy-efficient and Real-Time Multi-Object Detection and Tracking
abstract
Today’s high resolution, high frame rate cameras in autonomous vehicles generate a large volume of data that needs to be transferred and processed by a downstream processor or machine learning (ML) accelerator to enable intelligent computing tasks, such as multi-object detection and tracking. The massive amount of data transfer incurs significant energy, latency, and bandwidth bottlenecks, which hinders real-time processing. To mitigate this problem, we propose an algorithm-hardware co-design framework called Processing-in-Pixel-in-Memory-based object Detection and Tracking (P2M-DeTrack). P2M-DeTrack is based on a custom faster R-CNN-based model that is distributed partly inside the pixel array (front-end) and partly in a separate FPGA/ASIC (back-end). The proposed front-end in-pixel processing down-samples the input feature maps significantly with judiciously optimized strided convolution and pooling. Compared to a conventional baseline design that transfers frames of RGB pixels to the back-end, the resulting P2M-DeTrack designs reduce the data bandwidth between sensor and back-end by up to 24×. The designs also reduce the sensor and total energy (obtained from in-house circuit simulations at Globalfoundries 22nm technology node) per frame by 5.7× and 1.14×, respectively. Lastly, they reduce the sensing and total frame latency by an estimated 1.7× and 3×, respectively. We evaluate our approach on the multi-object object detection (tracking) task of the large-scale BDD100K dataset and observe only a 0.5% reduction in the mean average precision (0.8% reduction in the identification F1 score) compared to the state-of-the-art.
Gourav Datta, Souvik Kundu 0002, Zihan Yin, Joe Mathai, Zeyu Liu 0003, Mulin Tian, Shunlin Lu, Ravi Teja Lakkireddy, Andrew G. Schmidt, Wael Abd-Almageed, Ajey P. Jacob, Akhilesh Jaiswal 0001, Peter A. Beerel
VLSI-SoC13
2019 Digital and Analog-Mixed-Signal In-Memory Processing in CMOS SRAM
abstract
No abstract available.
Akhilesh Jaiswal 0001, Amogh Agrawal, Indranil Chakraborty, Mustafa Fayez Ali, Kaushik Roy 0001
ACM Great Lakes Symposium on VLSI1
2019 On Robustness of Spin-Orbit-Torque Based Stochastic Sigmoid Neurons for Spiking Neural Networks
abstract
Nano-scale neuro-mimetic devices have recently gained wide research interest in the quest to enable brain-like energy-efficiency with cognitive computing abilities. Traditionally, neuromorphic devices have exploited deterministic nano-scale devices for emulating the intrinsic neuronal and synaptic behavior. However, of particular interest are stochastic neuromorphic devices owing to - 1) availability of nano-scale devices that are inherently stochastic based on intrinsic device physics 2) various neuroscience experiments have demonstrated that cortical neurons are stochastic in nature. In this paper, we focus on spin orbit torque based Magnetic Tunnel Junction (SOT-MTJ) that exhibit stochastic sigmoid behavior with respect to the switching process. We first discuss the modeling framework that was used to study the effect of dimensional variations in SOT-MTJs and the resulting changes in the stochastic sigmoid behavior. Our model is based on the well-known stochastic-Landau-Lifshitz-Gilbert-Slonczewski equation under mono-domain approximation. Subsequently, we abstract the sigmoid characteristic of the device into a behavioral model and study the effect of variations in sigmoid characteristics on a deep binary network. Our results show that the variations in the sigmoidal neuron behavior results in a minimal loss in accuracy (for CIFAR 10 dataset). Additionally, the degradation in accuracy monotonically increases with increase in induced variations. This highlights the robustness of stochastic neural networks based on SOT-MTJs in presence of dimensional variations.
Akhilesh Jaiswal 0001, Amogh Agrawal, Indranil Chakraborty, Deboleena Roy, Kaushik Roy 0001
IJCNN1
2019 8T SRAM Cell as a Multibit Dot-Product Engine for Beyond Von Neumann Computing
abstract
Large-scale digital computing almost exclusively relies on the von Neumann architecture, which comprises separate units for storage and computations. The energy-expensive transfer of data from the memory units to the computing cores results in the well-known von Neumann bottleneck. Various approaches aimed toward bypassing the von Neumann bottleneck are being extensively explored in the literature. These include in-memory computing based on CMOS and beyond CMOS technologies, wherein by making modifications to the memory array, vector computations can be carried out as close to the memory units as possible. Interestingly, in-memory techniques based on CMOS technology are of special importance due to the ubiquitous presence of field-effect transistors and the resultant ease of large-scale manufacturing and commercialization. On the other hand, perhaps the most important computation required for applications such as machine learning, etc., comprises the dot-product operation. Emerging nonvolatile memristive technologies have been shown to be very efficient in computing analog dot products in an in situ fashion. The memristive analog computation of the dot product results in much faster operation as opposed to digital vector in-memory bitwise Boolean computations. However, challenges with respect to large-scale manufacturing coupled with the limited endurance of memristors have hindered rapid commercialization of memristive-based computing solutions. In this paper, we show that the standard 8 transistor (8T) digital SRAM array can be configured as an analoglike in-memory multibit dot-product engine (DPE). By applying appropriate analog voltages to the read ports of the 8T SRAM array and sensing the output current, an approximate analog-digital DPE can be implemented. We present two different configurations for enabling multibit dot-product computations in the 8T SRAM cell array, without modifying the standard bit-cell structure. We also demonstrate the robustness of the present proposal in presence of nonidealities such as the effect of line resistances and transistor threshold voltage variations. Since our proposal preserves the standard 8T-SRAM array structure, it can be used as a storage element with standard read-write instructions and also as an on-demand analoglike dot-product accelerator.
Akhilesh Jaiswal 0001, Indranil Chakraborty, Amogh Agrawal, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Designing Energy-Efficient Intermittently Powered Systems Using Spin-Hall-Effect-Based Nonvolatile SRAM
abstract
Intermittently powered systems represent a new class of batteryless devices that operate solely on energy harvested from their environment. Due to the unreliable nature of ambient energy sources, these devices experience frequent intervals of power loss, leading to sudden reboots. Tolerating such power supply disruptions require the ability to rapidly checkpoint/save system state when power loss is imminent and restore it at the start of the next power cycle to continue computations in a seamless manner. A typical microcontroller used in these systems consists of a fast nonvolatile SRAM and a nonvolatile Flash storage. Prior work has shown how emerging nonvolatile memory technologies such as STT-MRAM can improve the energy efficiency of these systems, either by using STT-MRAM as a drop-in replacement for Flash (henceforth referred to as the SRAM+STT-MRAM memory configuration) or using STT-MRAM as unified memory (henceforth referred to as the unified STT-MRAM memory configuration). However, both these configurations have significant drawbacks. Using the SRAM+STT-MRAM configuration leads to high checkpointing overhead due to the inefficient write operations of STT-MRAM whereas using the unified STT-MRAM configuration is inefficient due to executing every program instruction directly from STT-MRAM. This paper proposes a novel Spin Hall Effect-based nonvolatile-SRAM (SNVRAM) bit-cell that combines the nonvolatility of spin devices with the speed and energy efficiency of conventional 6T SRAM cells. We explore the use of the proposed SNVRAM to replace the SRAM in a transiently powered system to mitigate the drawbacks of the aforementioned memory configurations. Simulation results using a set of evaluation benchmarks demonstrate that the SNVRAM+STT-MRAM configuration leads to significant memory energy benefits of $2.6\times $ and $2.8\times $ on average, compared to the SRAM+STT-MRAM and unified STT-MRAM memory configurations, respectively.
Arnab Raha, Akhilesh Jaiswal 0001, Syed Shakib Sarwar, Hrishikesh Jayakumar, Vijay Raghunathan, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Stochastic Switching of SHE-MTJ as a Natural Annealer for Efficient Combinatorial Optimization
abstract
Efficient computing models for combinatorial optimization problems, like the Ising spin model, have been researched intensively as an alternative to the von-Neumann based general-purpose computing. The bottleneck mainly stems from the fact that the computing complexity of such optimization problems increases exponentially with the size of the problem. Although efficient heuristic algorithms have been designed for such combinatorial problems, yet hardware implementations of such top-down approaches suffer from complex control requirements and frequent memory accesses. Interestingly, unique device characteristics of the recent emerging devices, such as stochastic spintronic devices, can potentially pave the way for efficient hardware implementation of such combinatorial optimization problems. In this work, we leverage stochastic switching of nano-magnets in presence of thermal noise to implement an efficient combinatorial optimization solver and demonstrate its feasibility by solving realistic NP-complete problems.
Yong Shim, Akhilesh Jaiswal 0001, Kaushik Roy 0001
ICCD2
2017 Image segmentation with stochastic magnetic tunnel junctions and spiking neurons
abstract
Image segmentation is a crucial pre-processing stage used in many object identification problems. The purpose of image segmentation is to simplify the representation of an image such that it can be more conveniently analyzed in the later stages of a problem. This is generally achieved through partitioning a complicated image into specific groups based on color, intensity or texture of the pixels of that image. Locally Excitatory Globally Inhibitory Oscillator Network or LEGION is one such segmentation algorithm, where synchronization and desynchronization between coupled oscillators are used to segment an image. To extract maximum benefits from the fast parallel processing nature of LEGION, one must resort to a hardware implementation of this architecture. Unfortunately, the present structure of LEGION with relaxation oscillators as nodes, is not ideal for scalable and energy efficient hardware realization of the network. In this work we propose two different networks for image segmentation, one with leaky integrate and fire neurons and the other with stochastic Magnetic Tunneling Junctions (MTJs), both inspired by the operating principles of LEGION. The structure of the proposed networks allows them to be translated into energy efficient and scalable hardware platforms. We demonstrate that the proposed networks can effectively and efficiently segment binary and gray-scale images with multiple objects.
Chamika M. Liyanagedera, Parami Wijesinghe, Akhilesh Jaiswal 0001, Kaushik Roy 0001
IJCNN3
2016 Significance driven hybrid 8T-6T SRAM for energy-efficient synaptic storage in artificial neural networks
Gopalakrishnan Srinivasan, Parami Wijesinghe, Syed Shakib Sarwar, Akhilesh Jaiswal 0001, Kaushik Roy 0001
DATE4