VLDB 2026 Research / reviewers in the wild / expert
Shaahin Angizi
dblp:149/0425
· DBLP profile ↗
81ranked-venue papers
18as first author
52since 2021 · last 2026
0000-0003-2289-6381ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 18 first-author · 49 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INSPIRE: In-Sensor Compressed Weight Retrieval for Enhancing ViT Efficiency at EdgeabstractDeploying Vision Transformer (ViT) models on edge devices poses significant challenges due to the high bandwidth, energy demands, and latency associated with transmitting large weight parameter sets to the sensing unit, along with limited on-chip memory resources, which are often insufficient for storing these parameters. To address these constraints, we present a software-hardware co-design framework that incorporates a novel in-sensor Compressed Weight Retrieval mechanism within an intelligent vision sensor. This framework offers two key contributions. First, we propose an innovative hardware-friendly weight compression algorithm that substantially reduces bandwidth and power consumption by optimizing on-chip memory usage for storing weight parameters. Second, we leverage the exceptional efficiency of Silicon Photonic (SiPh) devices and design a novel in-sensor accelerator called INSPIRE for the first time to perform in-sensor retrieval of the compressed weights and parallel fine-grained convolution operations next to the pixel array, enabling low-power adaptable ViT inference on resource-constrained edge platforms. Our extensive simulation results show that INSPIRE can remarkably reduce the memory footprint of ViT results with favorable accuracy. Besides, INSPIRE significantly reduces the bandwidth and power requirements associated with storing weight parameters in on-chip memory. INSPIRE achieves up to 245.4 Kilo FPS/W and reduces the data transfer energy by a factor of ∼11× on average compared with 4-bit quantized ViTs. Deniz Najafi, Mohaiminul Al Nahian, Navid Khoshavi, Abdullah Al Arafat, Mamshad Nayeem Rizve, Mahdi Nikdast, Adnan Siraj Rakin, Shaahin Angizi |
DATE | 9 |
| 2026 | Late Breaking Results: Thermally Assisted RowPress-RowHammer Synergy for Cross-Row Bit FlipsabstractIn this work, we shed light on a previously uncharacterized thermal-assisted disturbance vulnerability in modern DDR4 DRAM by exploiting the temperature sensitivity and dense cell layout of scaled memory devices, a phenomenon we call HeatHammer. HeatHammer operates by first RowPressing the near aggressor, keeping it activated long enough to raise its local temperature, delay refresh operations, and erode its electrical isolation, and then RowHammering the far aggressor to induce bit flips in the victim row. This thermally weakened state significantly amplifies disturbance propagation, closely resembling the Half-Double effect [1]. Elevated temperatures accelerate charge leakage, shrink retention time, and reduce sense margins, thereby enabling disturbance effects to traverse through the compromised near aggressor via capacitive coupling and charge sharing. We evaluate this vulnerability on DRAM chips from two leading DRAM manufacturers and demonstrate that HeatHammer can substantially degrade the effectiveness of existing mitigation mechanisms such as Target Row Refresh (TRR). Filip Roth Trønnes-Christensen, Ranyang Zhou, Gamana Aragonda, Abeer Matar A. Almalky, Mohaiminul Al Nahian, Adnan Siraj Rakin, Shaahin Angizi |
DATE | 7 |
| 2026 | SENTRY: Spiking Event Reasoning for Selective Deep Inference in Event-Driven Edge VisionabstractEvent-driven cameras are well suited for always-on edge vision, but forwarding all events creates unnecessary overhead. Events outside user-defined ROIs are semantically irrelevant, while many ROI events arise from background motion, lighting changes, or sensor noise rather than meaningful activity. We propose SENTRY, a near-sensor pipeline that addresses both issues through hierarchical spatial reasoning. A coarse spike-rate gate discards frames with low ROI activity, while a spatially-aware confirmation stage compares each region’s activity against the global background rate to suppress diffuse noise. Evaluated on a real indoor surveillance sequence, SENTRY achieves +64% precision and − 57% false positive rate relative to a global spike-rate baseline, targeting resource-constrained, always-on platforms where energy efficiency and rapid response must be achieved simultaneously. Shayan Gerami, Sepehr Tabrizchi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | Photonics-Enabled Edge Processing: A Vision for Near-Sensor Optical IntelligenceabstractEdge intelligence is rapidly shifting computation from centralized cloud infrastructure toward the point of data generation. This shift is especially important for visual sensing systems, where continuous streams of high-dimensional pixel data must be converted, stored, transmitted, and processed under strict energy and latency constraints. While processing-in-sensor and processing-near-sensor architectures have reduced data movement, they remain limited by analog-to-digital conversion, memory access, electronic bandwidth, and the difficulty of supporting increasingly complex models near the sensor. This invited paper argues that integrated photonics can provide a new substrate for edge processing by enabling high-bandwidth, low-latency, and naturally parallel analog computation close to the sensing interface. We review the basic principles of photonic computing and discuss how it can be used to realize near-sensor multiply-and-accumulate operations. We then use recent work from our group as representative case studies, including optical in-sensor acceleration, optical near-sensor acceleration with compressive acquisition, near-sensor neuro-symbolic photonic computing, and in-sensor compressed weight retrieval for vision transformers. These examples motivate a broader research agenda in which photonics is not only a fast accelerator for neural operations, but also a system-level enabler for data-centric, energy-aware, and real-time edge intelligence. Deniz Najafi, Shaahin Angizi, Mahdi Nikdast |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Shallow Enough? A Cross-Architecture Study of Ultra-Low-Depth Neural Networks for Edge InferenceabstractEdge deployment imposes strict latency, memory, and energy constraints that scale directly with network depth, yet the question of which architectural paradigm offers the best accuracy-efficiency tradeoff at ultra-low depth (four to six layers) remains open. We present the first systematic comparison of convolutional, pure transformer, and hybrid CNN-transformer models under fixed shallow depth budgets. We develop an analytical framework characterizing the representational cost of local convolutional versus global attention-based feature mixing as a function of depth, and validate it with experiments on ImageNet-1K across architectures spanning MobileNetV2, DeiT, Swin, MobileViT, and EfficientFormer. Beyond FLOPs and accuracy, we report real-world latency and energy measurements on an NVIDIA Jetson Nano and examine deployment feasibility on Cortex-M class microcontrollers. Our experiments reveal how accuracy, latency, and memory scale across all three architecture families as depth decreases, providing practitioners with direct guidance on which model class to choose for a given depth and hardware budget. Chengwei Zhou, Haotian Yu, Shoma Yukawa, Deniz Najafi, Shaahin Angizi, Gourav Datta |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | A 40 μW 8-bit Accelerator Wake-Up Circuit for Always-on Smart IoT Sensor Monitoring
Andrew Ding, Sungheon Jeong 0001, Hamza Errahmouni Barkam, Shaahin Angizi, Nader Bagherzadeh, Mohsen Imani |
ISCAS | 5 |
| 2026 | Closing the Loop in LLM-Based Hardware Generation: An Autonomous Agentic Workflow for Robust TPU Design
Deepak Vungarala, Kartik Pandit, Gamana Aragonda, Jeremy McLynch, Adeola Adeoye-Davids, Bryan Galecio, NhatHai Phan, Abdallah Khreishah, Ramtin Zand, Arnob Ghosh, Shaahin Angizi |
VTS | 11 |
| 2025 | DeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the EdgeabstractVision Transformers (ViTs) excel in tackling complex vision tasks, yet their substantial size poses significant challenges for applications on resource-constrained edge devices. The increased size of these models leads to higher overhead (e.g., energy, latency) when transmitting model weights between the edge device and the server. Hence, ViTs are not ideal for edge devices where the entire model may not fit on the device. Current model compression techniques often achieve high compression ratios at the expense of performance degradation, particularly for ViTs. To overcome the limitations of existing works, we rethink model compression strategy for ViTs from first principle approach and develop an orthogonal strategy called DeepCompress-ViT. The objective of the DeepCompress-ViT is to encode the model weights to a highly compressed encoded representation using a novel training method, denoted as Unified Compression Training (UCT). Proposed UCT is accompanied by a decoding mechanism during inference, which helps to gain any loss of accuracy due to high compression ratio. We further optimize this decoding step by reordering the decoding operation using associative property of matrix multiplication, ensuring that the compressed weights can be decoded during inference without incurring any computational overhead. Our extensive experiments across multiple ViT models on modern edge devices show that DeepCompress-ViT can successfully compress ViTs at high compression ratios (> 14×). DeepCompress-ViT enables the entire model to be stored on edge device, resulting in unprecedented reductions in energy consumption (> 1470×) and latency (> 68×) for edge ViT inference. Our code is available at https://github.com/ML-Security-Research-LAB/DeepCompress-ViT. Abdullah Al Arafat, Deniz Najafi, Akhlak Mahmood, Mamshad Nayeem Rizve, Mohaiminul Al Nahian, Ranyang Zhou, Shaahin Angizi, Adnan Siraj Rakin |
CVPR | 8 |
| 2025 | ResISC: Residue Number System-Based Integrated Sensing and Computing for Efficient Edge AIabstractThis paper presents ResISC, an RNS-based integrated sensing and computing architecture enabling efficient edge AI. ResISC platform features (i) an in-sensor residue encoder converting images directly to RNS in the analog domain, (ii) an energy-efficient RNS-based processing-near-sensor CNN accelerator utilizing SOT-MRAM, and (iii) an innovative mixed-radix unit for efficient activation operations. By employing selective channel deactivation, ResISC reduces computation overhead by up to $89 \%$, while achieving a $3.4 \times$ improvement in power efficiency and up to a $71 \times$ reduction in execution time compared to processing-in-MRAM platforms. Experiments on various datasets demonstrate that ResISC achieves competitive accuracy levels (up to $94.63 \%$ on CIFAR-10) with minimal degradation, making it an ideal solution for power-constrained, real-time edge applications. Sepehr Tabrizchi, Samin Sohrabi, Mohamadreza Mohammadi, Ramtin Zand, Shaahin Angizi, Arman Roohi |
DAC | 5 |
| 2025 | Exploiting Boosting in Hyperdimensional Computing for Enhanced Reliability in HealthcareabstractHyperdimensional computing (HDC) enables efficient data encoding and processing in high-dimensional spaces, benefiting machine learning and data analysis. However, under-utilization of these spaces can lead to overfitting and reduced model reliability, especially in data-limited systems-a critical issue in sectors like healthcare that demand robustness and consistent performance. We introduce BoostHD, an approach that applies boosting algorithms to partition the hyperdimensional space into subspaces, creating an ensemble of weak learners. By integrating boosting with HDC, BoostHD enhances performance and reliability beyond existing HDC methods. Our analysis highlights the importance of efficient utilization of hyperdimensional spaces for improved model performance. Experiments on healthcare datasets show that BoostHD outperforms state-of-the-art methods. On the WESAD dataset, it achieved an accuracy of 98.37% ± 0.32%, surpassing Random Forest, XGBoost, and On-lineHD. BoostHD also demonstrated superior inference efficiency and stability, maintaining high accuracy under data imbalance and noise. In person-specific evaluations, it achieved an average accuracy of 96.19%, outperforming other models. By addressing the limitations of both boosting and HDC, BoostHD expands the applicability of HDC in critical domains where reliability and precision are paramount. Sungheon Jeong 0001, Hamza Errahmouni Barkam, Sanggeon Yun, Yeseong Kim, Shaahin Angizi, Mohsen Imani |
DATE | 5 |
| 2025 | Compromising the Intelligence of Modern DNNs: On the Effectiveness of Targeted RowPressabstractRecent advancements in side-channel attacks have revealed the vulnerability of modern Deep Neural Networks (DNNs) to malicious adversarial weight attacks. The well-studied RowHammer attack has effectively compromised DNN performance by inducing precise and deterministic bit-flips in the main memory (e.g., DRAM). Similarly, RowPress has emerged as another effective strategy for flipping targeted bits in DRAM. However, the impact of RowPress on deep learning applications has yet to be explored in the existing literature, leaving a fundamental research question unanswered: How does RowPress compare to RowHammer in leveraging bit-flip attacks to compromise DNN performance? This paper is the first to address this question and evaluate the impact of RowPress on DNN applications. We conduct a comparative analysis utilizing a novel DRAM-profile-aware attack designed to capture the distinct bit-flip patterns caused by RowHammer and RowPress. Eleven widely-used DNN architectures trained on different benchmark datasets deployed on a Samsung DRAM chip conclusively demonstrate that they suffer from a drastically more rapid performance degradation under the RowPress attack compared to RowHammer. The difference in the underlying attack mechanism of RowHammer and RowPress also renders existing RowHammer mitigation mechanisms ineffective under RowPress. As a result, RowPress introduces a new vulnerability paradigm for DNN compute platforms and unveils the urgent need for corresponding protective measures. Ranyang Zhou, Jacqueline Tiffany Liu, Shaahin Angizi, Adnan Siraj Rakin |
DATE | 4 |
| 2025 | iSEW: in-Sensor Embedded Watermarking for Secure ImagingabstractThis paper proposes an analog-domain watermarking approach implemented directly in CMOS sensors. By embedding the watermark at the readout stage before ADC, we achieve a robust, tamper-resistant mechanism with minimal impact on image quality. Modifications to the column amplifier architecture and a secure pattern generator enable effective watermark embedding. Experimental results on a 64×64 pixel array showcase high watermark detection rates (> 85%) under various attacks. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
FCCM | 2 |
| 2025 | LLM-IMC: Automating Analog In-Memory Computing Architecture Generation with Large Language ModelsabstractResistive crossbars enabling analog In-Memory Computing (IMC) have garnered significant attention from academia and industry as a promising architecture for Deep Neural Network (DNN) acceleration, thanks to their high memory access bandwidth and in-situ computing capabilities. However, the knowledge-intensive hardware design process and the lack of high-quality circuit netlists have constrained design space exploration and optimization of analog IMC to behavioral system-level tools. In this one-page abstract, we introduce LLM-IMC, a novel fine-tune-free Large Language Model (LLM) framework, supported by a Python-based tool, designed for analog IMC SPICE code generation. LLM-IMC systematically addresses these limitations by automating the creation of diverse IMC simulation scripts, enabling efficient design space exploration through LLM-driven performance, and outlining an integration roadmap for hardware-oriented neuromorphic crossbar design flows. Deepak Vungarala, Md Hasibul Amin, Pietro Mercati, Arman Roohi, Ramtin Zand, Shaahin Angizi |
FCCM | 6 |
| 2025 | How Vulnerable are Large Language Models (LLMs) against Adversarial Bit-Flip Attacks?
Abeer Matar A. Almalky, Ranyang Zhou, Shaahin Angizi, Adnan Siraj Rakin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Maximizing Sub-Array Resource Utilization in Digital Processing-in-Memory: A Versatile Hardware-Aware Approach
Gamana Aragonda, Deniz Najafi, Deepak Vungarala, Sepehr Tabrizchi, Arman Roohi, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | PixelPrune: Optimizing AIoT Vision Systems via In-Sensor Segmentation and Adaptive Data Transfer
Mohammadreza Mohammadi, Mehrdad Morsali, Sepehr Tabrizchi, Brendan Reidy, Arman Roohi, Shaahin Angizi, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | Event-Driven Spatiotemporal Processing-In-Sensor with Phase Change Memory-based Optical Acceleration
Mehrdad Morsali, Deniz Najafi, Amin Shafiee, Sepehr Tabrizchi, Pietro Mercati, Mohsen Imani, Arman Roohi, Navid Khoshavi, Mahdi Nikdast, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 10 |
| 2025 | SenGuard: A Novel Processing In-Sensor Method for Privacy-Enhanced Smart Imaging
Neeraj Solanki, Sepehr Tabrizchi, Ali Shafiee Sarvestani, Shaahin Angizi, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Magnetic In/Near-Sensor Architectures: From Raw Sensing to Smart Processing
Sepehr Tabrizchi, Ali Shafiee Sarvestani, Md Hasibul Amin, Deniz Najafi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | From Prompt to Accelerator: A Perspective on LLM-Based Analog In-Memory Accelerator Design Automation
Deepak Vungarala, Md Hasibul Amin, Arman Roohi, Arnob Ghosh, Ramtin Zand, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon PhotonicsabstractVision Transformers (ViTs) have emerged as a powerful architecture for computer vision tasks due to their ability to model long-range dependencies and global contextual relationships. However, their substantial compute and memory demands hinder efficient deployment in scenarios with strict energy and bandwidth limitations. In this work, we propose Opto-ViT, the first near-sensor, region-aware ViT accelerator leveraging silicon photonics (SiPh) for real-time and energy-efficient vision processing. Opto-ViT features a hybrid electronic-photonic architecture, where the optical core handles compute-intensive matrix multiplications using Vertical-Cavity Surface-Emitting Lasers (VCSELs) and Microring Resonators (MRs), while nonlinear functions and normalization are executed electronically. To reduce redundant computation and patch processing, we introduce a lightweight Mask Generation Network (MGNet) that identifies regions of interest in the current frame and prunes irrelevant patches before ViT encoding. We further co-optimize the ViT backbone using quantization-aware training and matrix decomposition tailored for photonic constraints. Experiments across device fabrication, circuit and architecture co-design, to classification, detection, and video tasks demonstrate that Opto-ViT achieves 100.4 KFPS/W with up to 84% energy savings with less than 1.6% accuracy loss, while enabling scalable and efficient ViT deployment at the edge. Mehrdad Morsali, Chengwei Zhou, Deniz Najafi, Sreetama Sarkar, Pietro Mercati, Navid Khoshavi, Peter A. Beerel, Mahdi Nikdast, Gourav Datta, Shaahin Angizi |
ICCAD | 10 |
| 2025 | SA-DS: A Dataset for Large Language Model-Driven AI Accelerator Design GenerationabstractIn the ever-evolving landscape of Deep Neural Networks (DNN) hardware acceleration, unlocking the true potential of systolic array accelerators has long been hindered by the daunting challenges of expertise and time investment. Large Language Models (LLMs) offer a promising solution for automating code generation, which is key to unlocking unprecedented efficiency and performance in various domains, including hardware descriptive code. The generative power of LLMs can enable the effective utilization of preexisting designs and dedicated hardware generators. However, the successful application of LLMs to hardware accelerator design is contingent upon the availability of specialized datasets tailored for this purpose. To bridge this gap, we introduce the Systolic Array-based Accelerator DataSet (SA-DS). SA-DS comprises a diverse collection of spatial array designs following the standardized Berkeley’s Gemmini accelerator generator template, enabling design reuse, adaptation, and customization. SA-DS is intended to spark LLM-centered research on DNN hardware accelerator architecture. We envision that SA-DS provides a framework that will shape the course of DNN hardware acceleration research for generations to come. SA-DS is open-sourced under the permissive MIT license at https://github.com/ACADLab/SA-DS. Deepak Vungarala, Mahmoud Nazzal, Mehrdad Morsali, Chao Zhang 0014, Arnob Ghosh, Abdallah Khreishah, Shaahin Angizi |
ISCAS | 7 |
| 2025 | Assessing the Potential of Escalating RowHammer Attack Distance to Bypass-Counter-Based DefensesabstractThis brief studies the impact of escalating DRAM RowHammer (RH) attack distance to potentially bypass well-developed counter-based defenses leveraging a multisided fault injection mechanism. By conducting systematic experimentation on 128 commercial DDR4 products, our results challenge recent research findings, showing that cells positioned at a greater physical distance from the target rows do not significantly affect performance across chips sourced from leading DRAM manufacturers. This implies such RH models are unable to reliably bypass the latest counter-based defense mechanisms. We conduct an extensive attack design space exploration and compare the performance efficiency between this mechanism and the well-known double-sided attack. Ranyang Zhou, Jacqueline Tiffany Liu, Nakul Kochar, Adnan Siraj Rakin, Shaahin Angizi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Deep-TROJ: An Inference Stage Trojan Insertion Algorithm Through Efficient Weight Replacement AttackabstractTo insert Trojan into a Deep Neural Network (DNN), the existing attack assumes the attacker can access the victim's training facilities. However, a realistic threat model was recently developed by leveraging memory fault to inject Trojans at the inference stage. In this work, we develop a novel Trojan attack by adopting a unique memory fault injection technique that can inject bit-flip into the page table of the main memory. In the main memory, each weight block consists of a group of weights located at a specific address of a DRAM row. A bit-flip in the page frame number replaces a target weight block of a DNN model with another replacement weight block. To develop a successful Trojan attack leveraging this unique fault model, the attacker must solve three key challenges: i) how to identify a minimum set of target weight blocks to be modified? ii) how to identify the corresponding optimal replacement weight block? iii) how to optimize the trigger to maximize the attacker's objective given a target and replacement weight block set? We address them by proposing a novel Deep- Troj attack algorithm that can identify a minimum set of vulnerable target and corresponding replacement weight blocks while optimizing the trigger at the same time. We evaluate the performance of our proposed Deep-TROJ on CIFAR-IO, CIFAR-IOO, and ImageNet dataset for fifteen different DNN architectures, including vision transformers. Proposed Deep- Troj is the most successful one to date that does not require access to training facilities while successfully bypassing the existing defenses. Our code is available at https://github.com/ML-Security-Research-LABIDeep-TROJ. Ranyang Zhou, Shaahin Angizi, Adnan Siraj Rakin |
CVPR | 3 |
| 2024 | Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image ProcessingabstractThis paper proposes a high-performance and energy-efficient optical near-sensor accelerator for vision applications, called Lightator. Harnessing the promising efficiency offered by photonic devices, Lightator features innovative compressive acquisition of input frames and fine-grained convolution operations for low-power and versatile image processing at the edge for the first time. This will substantially diminish the energy consumption and latency of conversion, transmission, and processing within the established cloud-centric architecture as well as recently designed edge accelerators. Our device-to-architecture simulation results show that with favorable accuracy, Lightator achieves 84.4 Kilo FPS/W and reduces power consumption by a factor of ~24× and 73× on average compared with existing photonic accelerators and GPU baseline. Mehrdad Morsali, Brendan Reidy, Deniz Najafi, Sepehr Tabrizchi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Ramtin Zand, Shaahin Angizi |
DAC | 9 |
| 2024 | HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROIabstractWith the rise of tiny IoT devices powered by machine learning (ML), many researchers have directed their focus toward compressing models to fit on tiny edge devices. Recent works have achieved remarkable success in compressing ML models for object detection and image classification on microcontrollers with small memory, e.g., 512kB SRAM. However, there remain many challenges prohibiting the deployment of ML systems that require high-resolution images. Due to fundamental limits in memory capacity for tiny IoT devices, it may be physically impossible to store large images without external hardware. To this end, we propose a high-resolution image scaling system for edge ML, called HiRISE, which is equipped with selective region-of-interest (ROI) capability leveraging analog in-sensor image scaling. Our methodology not only significantly reduces the peak memory requirements, but also achieves up to 17.7× reduction in data transfer and energy consumption. Brendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi, Arman Roohi, Ramtin Zand |
DAC | 4 |
| 2024 | DNN-Defender: A Victim-Focused In-DRAM Defense Mechanism for Taming Adversarial Weight Attack on DNNsabstractWith deep learning deployed in many security-sensitive areas, machine learning security is becoming progressively important. Recent studies demonstrate attackers can exploit system-level techniques exploiting the RowHammer vulnerability of DRAM to deterministically and precisely flip bits in Deep Neural Networks (DNN) model weights to affect inference accuracy. The existing defense mechanisms are software-based, such as weight reconstruction requiring expensive training overhead or performance degradation. On the other hand, generic hardware-based victim-/aggressor-focused mechanisms impose expensive hardware overheads and preserve the spatial connection between victim and aggressor rows. In this paper, we present the first DRAM-based victim-focused defense mechanism tailored for quantized DNNs, named DNN-Defender that leverages the potential of in-DRAM swapping to withstand the targeted bit-flip attacks with a priority protection mechanism. Our results indicate that DNN-Defender can deliver a high level of protection downgrading the performance of targeted RowHammer attacks to a random attack level. In addition, the proposed defense has no accuracy drop on CIFAR-10 and ImageNet datasets without requiring any software training or incurring hardware overhead. Ranyang Zhou, Adnan Siraj Rakin, Shaahin Angizi |
DAC | 4 |
| 2024 | OISA: Architecting an Optical In-Sensor Accelerator for Efficient Visual ComputingabstractTargeting vision applications at the edge, in this work, we systematically explore and propose a high-performance and energy-efficient Optical In-Sensor Accelerator architecture called OISA for the first time. Taking advantage of the promising efficiency of photonic devices, the OISA intrinsically implements a coarse-grained convolution operation on the input frames in an innovative minimum-conversion fashion in low-bit-width neural networks. Such a design remarkably reduces the power consumption of data conversion, transmission, and processing in the conventional cloud-centric architecture as well as recently-presented edge accelerators. Our device-to-architecture simulation results on various image data-sets demonstrate acceptable accuracy while OISA achieves 6.68 TOp/s/W efficiency. OISA reduces power consumption by a factor of 7.9 and 18.4 on average compared with existing electronic in-/near-sensor and ASIC accelerators. Mehrdad Morsali, Sepehr Tabrizchi, Deniz Najafi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Shaahin Angizi |
DATE | 7 |
| 2024 | DIAC: Design Exploration of Intermittent-Aware Computing Realizing Batteryless SystemsabstractBattery-powered IoT devices face challenges like cost, maintenance, and environmental sustainability, prompting the emergence of batteryless energy-harvesting systems that harness ambient sources. However, their intermittent behavior can disrupt program execution and cause data loss, leading to unpredictable outcomes. Despite exhaustive studies employing conventional checkpoint methods and intricate programming paradigms to address these pitfalls, this paper proposes an innovative systematic methodology, namely DIAC. The DIAC synthesis procedure enhances the performance and efficiency of intermittent computing systems, with a focus on maximizing forward progress and minimizing the energy overhead imposed by distinct memory arrays for backup. Then, a finite-state machine is delineated, encapsulating the core operations of an IoT node, sense, compute, transmit, and sleep states. First, we validate the robustness and functionalities of a DIAC-based design in the presence of power disruptions. DIAC is then applied to a wide range of benchmarks, including ISCAS-89, MCNS, and ITC-99. The simulation results substantiate the power-delay-product (PDP) benefits. For example, results for complex MCNC benchmarks indicate a PDP improvement of 61%, 56%, and 38% on average compared to three alternative techniques, evaluated at 45 nm. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
DATE | 2 |
| 2024 | DRAM-Locker: A General-Purpose DRAM Protection Mechanism Against Adversarial DNN Weight AttacksabstractIn this work, we propose DRAM-Locker as a robust general-purpose defense mechanism that can protect DRAM against various adversarial Deep Neural Network (DNN) weight attacks affecting data or page tables. DRAM-Locker harnesses the capabilities of in-DRAM swapping combined with a lock-table to prevent attackers from singling out specific DRAM rows to safeguard DNN's weight parameters. Our results indicate that DRAM-Locker can deliver a high level of protection downgrading the performance of targeted weight attacks to a random attack level. Furthermore, the proposed defense mechanism demonstrates no reduction in accuracy when applied to CIFAR-I0 and CIFAR-100. Importantly, DRAM-Locker does not necessitate any software retraining or result in extra hardware burden. Ranyang Zhou, Arman Roohi, Adnan Siraj Rakin, Shaahin Angizi |
DATE | 5 |
| 2024 | Hybrid Magneto-electric FET-CMOS Integrated Memory Design for Instant-on ComputingabstractThe surge in the number of normally-off power-constraint Internet of Things (IoT) devices in recent years has amplified the demand for high-performance and energy-efficient in-memory computing architectures built on top of various non-volatile memories. Magneto-Electric Field Effect Transistors (MEFETs) have presented compelling design features suitable for logic and memory integration as an emerging post-CMOS FET. These include high-speed switching, minimal power usage, and non-volatility. This work introduces a new in-memory computing architecture designed for edge applications, leveraging emerging MEFETs. The proposed architecture enables the execution of both Boolean logic operations and Binary Content Addressable Memory (BCAM) operations within a single cycle. Furthermore, the energy consumption during the write operation of the proposed cell is optimized by introducing a new write circuitry. The outcomes of our device-to-architecture evaluation reveal approximately 43.5% and 96.9% reduction in read and write energy consumption, respectively, compared to the counterpart non-volatile memories. At the application level, the proposed architecture is applied to implement Binary Neural Networks (BNNs) based on AlexNet and VGG16. Our results showcase a decrease of approximately 54% in the overall energy consumption when implementing these networks using the proposed design compared to non-volatile in-memory computing designs. Deniz Najafi, Sepehr Tabrizchi, Ranyang Zhou, Mohammadreza Amel Solouki, Andrew Marshall, Arman Roohi, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 7 |
| 2024 | RACSen: Residue Arithmetic and Chaotic Processing in Sensors to Enhance CMOS Imager SecurityabstractThe widespread adoption of vision sensors raises significant security and privacy concerns. In this paper, we present RACSen as a novel architecture that can increase the security and efficiency of conventional image sensors. RACSen leverages the intricate mathematical properties of the residue number system (RNS) with analog scrambling techniques to create a sophisticated dual-layered encryption mechanism. Incorporating RNS within analog-to-digital converters further strengthens security by mitigating replay attacks and preserving data transmission integrity and confidentiality. Our results demonstrate exceptional encryption, with a perfect pixel change rate of 99.90 and high intensity change of 45.77. This offers robust image data protection with minimal overhead of 11.11%. Sepehr Tabrizchi, Nedasadat Taheri, Shaahin Angizi, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | ChaoSen: Security Enhancement of Image Sensor through in-Sensor Chaotic ComputingabstractWireless Sensor Networks (WSN) are integral to diverse applications, ranging from environmental monitoring to urban smart infrastructure. In the realm of WSNs, security remains a critical challenge owing to the complex nature of the sensor environment. As a result, WSN security has become a research focus in recent years. In this paper, we introduce ChaoSen, a novel image sensor system incorporating analog chaotic circuits within the sensor, thereby enhancing the overall system security. The system utilizes a scrambler module, which intricately intertwines with the chaotic encryption process, to reduce the predictability of pixel values and enhance the security of the system. Comparative evaluations demonstrate that the system achieves an NPCR value of 99.5562% and a UACI of 35.81900%, indicating high sensitivity to input changes and significant alteration in pixel intensity. Our approach also demonstrates its resilience against common cyber attacks, balancing enhanced security with resource efficiency. Nedasadat Taheri, Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ICCD | 3 |
| 2024 | PiPSim: A Behavior-Level Modeling Tool for CNN Processing-in-Pixel AcceleratorsabstractConvolutional neural networks (CNNs) have been gaining popularity in recent years, and researchers have designed specialized architectures to speed up the inference process. However, despite the promising potential of processing near-/in- sensor architectures actively explored in the visual Internet of Things, there is still a need to develop a behavior-level simulator to model performance and facilitate early design exploration. This article proposes a stand-alone simulation platform for processing-in-pixel (PiP) systems, namely, PiPSim. It offers a flexible interface and a wide range of design options for customizing the efficiency and accuracy of PiP-based accelerators using a hierarchical structure. Its organization spans from the device level, e.g., memory technology, upward to the circuit level, e.g., compute-add on architecture, and then to the algorithm level, e.g., DNN workloads. PiPSim realizes instruction-accurate evaluation of circuit-level performance metrics as well as learning accuracy at run-time. Compared to SPICE simulation, PiPSim achieves over 25$000\times $speed-up with less than a 2.5% error rate on average. Furthermore, PiPSim can optimize the design and estimate the tradeoff relationships among different performance metrics. Arman Roohi, Sepehr Tabrizchi, Mehrdad Morsali, David Z. Pan, Shaahin Angizi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | P-PIM: A Parallel Processing-in-DRAM Framework Enabling Row Hammer ProtectionabstractIn this work, we propose a Parallel Processing-In-DRAM architecture named P-PIM leveraging the high density of DRAM to enable fast and flexible computation. P-PIM enables bulk bit-wise in-DRAM logic between operands in the same bit-line by elevating the analog operation of the memory sub-array based on a novel dual-row activation mechanism. With this, P-PIM can opportunistically perform a complete and inexpensive in-DRAM RowHammer (RH) self-tracking and mitigation technique to protect the memory unit against such a challenging security vulnerability. Our results show that P-PIM achieves ~72% higher energy efficiency than the fastest charge-sharing-based designs. As for the RH protection, with a worst-case slowdown of ~0.8%, P-PIM archives up to 71% energy-saving over the SRAM/CAM-based frameworks and about 90% saving over DRAM-based frameworks. Ranyang Zhou, Sepehr Tabrizchi, Mehrdad Morsali, Arman Roohi, Shaahin Angizi |
DATE | 5 |
| 2023 | Accelerating Low Bit-width Neural Networks at the Edge, PIM or FPGA: A Comparative StudyabstractDeep Neural Network (DNN) acceleration with digital Processing-in-Memory (PIM) platforms at the edge is an actively-explored domain with great potential to not only address memory-wall bottlenecks but to offer orders of performance improvement in comparison to the von-Neumann architecture. On the other side, FPGA-based edge computing has been followed as a potential solution to accelerate compute-intensive workloads. In this work, adopting low-bit-width neural networks, we perform a solid and comparative inference performance analysis of a recent processing-in-SRAM tape-out with a low-resource FPGA board and a high-performance GPU to provide a guideline for the research community. We explore and highlight the key architectural constraints of these edge candidates that impact their overall performance. Our experimental data demonstrate that the processing-in-SRAM can obtain up to ~160x speed-up and up to 228x higher efficiency (img/s/W) compared to the under-test FPGA on the CIFAR-10 dataset. Nakul Kochar, Lucas Ekiert, Deniz Najafi, Deliang Fan, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | IMA-GNN: In-Memory Acceleration of Centralized and Decentralized Graph Neural Networks at the EdgeabstractIn this paper, we propose IMA-GNN as an In-Memory Accelerator for centralized and decentralized Graph Neural Network inference, explore its potential in both settings and provide a guideline for the community targeting flexible and efficient edge computation. Leveraging IMA-GNN, we first model the computation and communication latencies of edge devices. We then present practical case studies on GNN-based taxi demand and supply prediction and also adopt four large graph datasets to quantitatively compare and analyze centralized and decentralized settings. Our cross-layer simulation results demonstrate that on average, IMA-GNN in the centralized setting can obtain ~790x communication speed-up compared to the decentralized GNN setting. However, the decentralized setting performs computation ~1400x faster while reducing the power consumption per device. This further underlines the need for a hybrid semi-decentralized GNN approach. Mehrdad Morsali, Mahmoud Nazzal, Abdallah Khreishah, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | SenTer: A Reconfigurable Processing-in-Sensor Architecture Enabling Efficient Ternary MLPabstractRecently, Intelligent IoT (IIoT), including various sensors, has gained significant attention due to its capability of sensing, deciding, and acting by leveraging artificial neural networks (ANN). Nevertheless, to achieve acceptable accuracy and high performance in visual systems, a power-delay-efficient architecture is required. In this paper, we propose an ultra-low-power processing in-sensor architecture, namely SenTer, realizing low-precision ternary multi-layer perceptron networks, which can operate in detection and classification modes. Moreover, SenTer supports two activation functions based on user needs and the desired accuracy-energy trade-off. SenTer is capable of performing all the required computations for the MLP's first layer in the analog domain and then submitting its results to a co-processor. Therefore, SenTer significantly reduces the overhead of analog buffers, data conversion, and transmission power consumption by using only one ADC. Additionally, our simulation results demonstrate acceptable accuracy on various datasets compared to the full precision models. Sepehr Tabrizchi, Rebati Raman Gaire, Shaahin Angizi, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | NeSe: Near-Sensor Event-Driven Scheme for Low Power Energy Harvesting SensorsabstractDigital technologies have made it possible to deploy visual sensor nodes capable of detecting motion events in the coverage area cost-effectively. However, background subtraction, as a widely used approach, remains an intractable task due to its inability to achieve competitive accuracy and reduced computation cost simultaneously. In this paper, an effective background subtraction approach, namely NeSe, for tiny energy-harvested sensors is proposed leveraging non-volatile memory (NVM). Using the developed software/hardware method, the accuracy and efficiency of event detection can be adjusted at runtime by changing the precision depending on the application's needs. Due to the near-sensor implementation of background subtraction and NVM usage, the proposed design reduces the data movement overhead while ensuring intermittent resiliency. The background is stored for a specific time interval within NVMs and compared with the next frame. If the power is cut, the background remains unchanged and is updated after the interval passes. Once the moving object is detected, the device switches to the high-powered sensor mode to capture the image. Sepehr Tabrizchi, Mehrdad Morsali, Shaahin Angizi, Arman Roohi |
ISCAS | 3 |
| 2023 | Ocellus: Highly Parallel Convolution-in-Pixel Scheme Realizing Power-Delay-Efficient Edge IntelligenceabstractWith the advent of Edge Intelligence (EI) devices, always-on intelligent and self-powered visual perception systems are receiving considerable attention. These emerging systems require continuous sensing and instant processing; however, the high energy data conversion/transmission of raw data and the limited available energy and computation resources make designing energy-efficient and low bandwidth CMOS vision sensors vital but challenging. This paper proposes a low-power integrated sensing and computing engine, namely Ocellus, which considerably decreases power costs of data movement/conversion and enables data/compute -intensive neural network tasks. Ocellus offers several unique features, including a highly parallel analog convolution-in-pixel scheme and reconfigurable filtering modes with filter pruning capability. These features realize low-precision ternary weight neural networks to mitigate the overhead of analog-to-digital converters and analog buffers. Moreover, the proposed structure supports a zero-skipping scheme to further reduce power consumption. Our circuit-to-application cosimulation results demonstrate comparable, even better, accuracy to the full-precision baseline on object classification tasks, while it achieves a frame rate of 1000 and efficiency of ~1.45 TOp/s/W. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ISLPED | 2 |
| 2022 | Work-in-Progress: A Processing-in-Pixel Accelerator based on Multi-level HfOx ReRAMabstractThis work paves the way to realize a processing-in-pixel accelerator based on a multi-level HfOxReRAM as a flexible, energy-efficient, and high-performance solution for real-time and smart image processing at edge devices. The proposed design intrinsically implements and supports a coarse-grained convolution operation in low-bit-width neural networks leveraging a novel compute-pixel with non-volatile weight storage at the sensor side. Our evaluations show that such a design can remarkably reduce the power consumption of data conversion and transmission to an off-chip processor maintaining accuracy compared with the recent in-sensor computing designs. Minhaz Abedin, Arman Roohi, Nathaniel C. Cady, Shaahin Angizi |
CASES | 4 |
| 2022 | ReD-LUT: Reconfigurable In-DRAM LUTs Enabling Massive Parallel ComputationabstractIn this paper, we propose a reconfigurable processing-in-DRAM architecture named ReD-LUT leveraging the high density of commodity main memory to enable a flexible, general-purpose, and massively parallel computation. ReD-LUT supports lookup table (LUT) queries to efficiently execute complex arithmetic operations (e.g., multiplication, division, etc.) via only memory read operation. In addition, ReD-LUT enables bulk bit-wise in-memory logic by elevating the analog operation of the DRAM sub-array to implement Boolean functions between operands stored in the same bit-line beyond the scope of prior DRAM-based proposals. We explore the efficacy of ReD-LUT in two computationally-intensive applications, i.e., low-precision deep learning acceleration, and the Advanced Encryption Standard (AES) computation. Our circuit-to-architecture simulation results show that for a quantized deep learning workload, ReD-LUT reduces the energy consumption per image by a factor of 21.4× compared with the GPU and achieves ~37.8× speedup and 2.1× energy-efficiency over the best in-DRAM bit-wise accelerators. As for AES data-encryption, it reduces energy consumption by a factor of ~2.2× compared to an ASIC implementation. Ranyang Zhou, Arman Roohi, Durga Misra, Shaahin Angizi |
ICCAD | 4 |
| 2022 | TizBin: A Low-Power Image Sensor with Event and Object Detection Using Efficient Processing-in-Pixel SchemesabstractIn the Artificial Intelligence of Things (AIoT) era, always-on intelligent and self-powered visual perception systems have gained considerable attention and are widely used. Thus, this paper proposes TizBin, a low-power processing in-sensor scheme with event and object detection capabilities to eliminate power costs of data conversion and transmission and enable data-intensive neural network tasks. Once the moving object is detected, TizBin architecture switches to the high-power object detection mode to capture the image. TizBin offers several unique features, such as analog convolutions enabling low-precision ternary weight neural networks (TWNN) to mitigate the overhead of analog buffer and analog-to-digital converters. Moreover, TizBin exploits non-volatile magnetic RAMs to store NN’s weights, remarkably reducing static power consumption. Our circuit-to-application co-simulation results for TWNNs demonstrate minor accuracy degradation on various image datasets, while TizBin achieves a frame rate of 1000 and efficiency of ∼1.83 TOp/s/W. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ICCD | 2 |
| 2022 | semiMul: Floating-Point Free Implementations for Efficient and Accurate Neural Network TrainingabstractMultiply–accumulate operation (MAC) is a fundamental component of machine learning tasks, where multiplication (either integer or float multiplication) compared to addition is costly in terms of hardware implementation or power consumption. In this paper, we approximate floating-point multiplication by converting it to integer addition while preserving the test accuracy of shallow and deep neural networks. We mathematically show and prove that our proposed method can be utilized with any floating-point format (e.g., FP8, FP16, FP32, etc.). It is also highly compatible with conventional hardware architectures and can be employed in CPU, GPU, or ASIC accelerators for neural network tasks with minimum hardware cost. Moreover, the proposed method can be utilized in embedded processors without a floating-point unit to perform neural network tasks. We evaluated our method on various datasets such as MNIST, FashionMNIST, SVHN, Cifar-10, and Cifar-100, with both FP16 and FP32 arithmetics. The proposed method preserves the test accuracy and, in some cases, overcomes the overfitting problem and improves the test accuracy. Ali Nezhadi, Shaahin Angizi, Arman Roohi |
ICMLA | 2 |
| 2022 | SCiMA: A Generic Single-Cycle Compute-in-Memory Acceleration Scheme for Matrix ComputationsabstractThis work proposes a new generic Single-cycle Compute-in-Memory (CiM) Accelerator for matrix computation named SCiMA. SCiMA is developed on top of the existing commodity Spin-Orbit Torque Magnetic Random-Access Memory chip. Every sub-array’s peripherals are transformed to realize a full set of single-cycle 2- and 3-input in-memory bulk bitwise functions specifically designed to accelerate a wide variety of graph and matrix multiplication tasks. We explore SCiMA’s efficiency by selecting a complex matrix processing operation, i.e., calculating determinant as an essential and under-explored application in the CiM domain. The cross-layer device-to-architecture simulation framework shows the presented platform can reduce energy consumption by 70.43% compared with the most recent CiM designs implemented with the same memory technology. SCiMA also achieves up to 2.5x speedup compared with current CiM platforms. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ISCAS | 2 |
| 2022 | FlexiDRAM: A Flexible in-DRAM Framework to Enable Parallel General-Purpose ComputationabstractIn this paper, we propose a Flexible processing-in-DRAM framework named FlexiDRAM that supports the efficient implementation of complex bulk bitwise operations. This framework is developed on top of a new reconfigurable in-DRAM accelerator that leverages the analog operation of DRAM sub-arrays and elevates it to implement XOR2-MAJ3 operations between operands stored in the same bit-line. FlexiDRAM first generates an efficient XOR-MAJ representation of the desired logic and then appropriately allocates DRAM rows to the operands to execute any in-DRAM computation. We develop ISA and software support required to compute in-DRAM operation. FlexiDRAM transforms current memory architecture to a massively parallel computational unit and can be leveraged to significantly reduce the latency and energy consumption of complex workloads. Our extensive circuit-to-architecture simulation results show that averaged across two well-known deep learning workloads, FlexiDRAM achieves ∼ 15 × energy-saving and 13 × speedup over the GPU outperforming recent processing-in-DRAM platforms. Ranyang Zhou, Arman Roohi, Durga Misra, Shaahin Angizi |
ISLPED | 4 |
| 2022 | MeF-RAM: A New Non-Volatile Cache Memory Based on Magneto-Electric FETabstractMagneto-Electric FET ( MEFET ) is a recently developed post-CMOS FET, which offers intriguing characteristics for high-speed and low-power design in both logic and memory applications. In this article, we present MeF-RAM , a non-volatile cache memory design based on 2-Transistor-1-MEFET ( 2T1M ) memory bit-cell with separate read and write paths. We show that with proper co-design across MEFET device, memory cell circuit, and array architecture, MeF-RAM is a promising candidate for fast non-volatile memory ( NVM ). To evaluate its cache performance in the memory system, we, for the first time, build a device-to-architecture cross-layer evaluation framework to quantitatively analyze and benchmark the MeF-RAM design with other memory technologies, including both volatile memory (i.e., SRAM, eDRAM) and other popular non-volatile emerging memory (i.e., ReRAM, STT-MRAM, and SOT-MRAM). The experiment results for the PARSEC benchmark suite indicate that, as an L2 cache memory, MeF-RAM reduces Energy Area Latency ( EAT ) product on average by ~98% and ~70% compared with typical 6T-SRAM and 2T1R SOT-MRAM counterparts, respectively. Shaahin Angizi, Navid Khoshavi, Andrew Marshall, Peter Dowben, Deliang Fan |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2021 | Max-PIM: Fast and Efficient Max/Min Searching in DRAMabstractRecently, in-DRAM computing is becoming one promising technique to address the notorious ‘memory-wall’ issue for big data processing. In this work, for the first time, we propose a novel ‘Min/Max-in-memory’ algorithm based on iterative XNOR bit-wise comparison, which supports parallel inmemory searching for minimum and maximum of bulk data stored in DRAM as unsigned & signed integers, fixed-point and floating numbers. We then develop a new processing-in-DRAM architecture, called Max-PIM, that supports complete bit-wise Boolean logic and beyond. Differentiating from prior works, Max-PIM is optimized with one-cycle fast XNOR logicin-DRAM operation and in-memory data transpose, which are heavily used and keys to accelerate the proposed Min/Max-in-memory algorithm efficiently. Extensive experiments of utilizing Max-PIM in big data sorting and graph processing applications show that it could speed up~50X and~1000X than GPU and CPU, while only consuming 10% and 1% energy, respectively. Moreover, comparing with recent representative In-DRAM computing platforms, i.e., Ambit [1], DRISA [2], our design could speed up~3X - 10X. Fan Zhang 0069, Shaahin Angizi, Deliang Fan |
DAC | 2 |
| 2021 | PIM-Quantifier: A Processing-in-Memory Platform for mRNA QuantificationabstractProcessing-in-memory (PIM) architecture has been considered as a promising solution for the “memory-wall” issue in many data-intensive applications, especially in bioinformatics. Recent works of developing PIM for genome alignment and assembling have achieved tremendous improvement, while another important genome analysis - mRNA quantification has not been explored. Efficient and accurate mRNA quantification is a crucial step for molecular signature identification, disease outcome prediction and drug development. In this paper, for the first time, we propose a SOT-MRAM based PIM platform, named PIM-Quantifier, for efficient mRNA quantification. A PIM-friendly alignment-free quantification algorithm is first proposed. Then, we present the optimized PIM architecture/circuit designs and mapping method to efficiently accelerate mRNA quantification. Extensive experiments show that PIM-Quantifier significantly improves mRNA quantification performance than CPU and recent other PIM platforms in efficiency defined as throughput/power. Fan Zhang 0069, Shaahin Angizi, Naima Ahmed Fahmi, Wei Zhang 0076, Deliang Fan |
DAC | 2 |
| 2021 | Processing-in-Memory Acceleration of MAC-based Applications Using Residue Number System: A Comparative StudyabstractProcessing-in-memory (PIM) has raised as a viable solution for the memory wall crisis and has attracted great interest in accelerating computationally intensive AI applications ranging from filtering to complex neural networks. In this paper, we try to take advantage of both PIM and the residue number system (RNS) as an alternative for the conventional binary number representation to accelerate multiplication-and-accumulations (MACs), primary operations of target applications. The PIM architecture utilizes the maximum internal bandwidth of memory chips to realize a local and parallel computation to eliminates the off-chip data transfer. Moreover, RNS limits inter-digit carry propagation by performing arithmetic operations on small residues independently and in parallel. Thus, we develop a PIM-RNS, entitled PRIMS, and analyze the potential of intertwining PIM architecture with the inherent parallelism of the RNS arithmetic to delineate the opportunities and challenges. To this end, we build a comprehensive device-to-architecture evaluation framework to quantitatively study this problem considering the impact of PIM technology for a well-known three-moduli set as a case study. Shaahin Angizi, Arman Roohi, MohammadReza Taheri, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2021 | RNSiM: Efficient Deep Neural Network Accelerator Using Residue Number SystemsabstractIn this paper, we propose an efficient convolutional neural network (CNN) accelerator design, entitled RNSiM, based on the Residue Number System (RNS) as an alternative for the conventional binary number representation. Instead of traditional arithmetic implementation that suffers from the inevitable lengthy carry propagation chain, the novelty of RNSiM lies in that all the data, including stored weights and communication/computation, are performed in the RNS domain. Due to the inherent parallelism of the RNS arithmetic, power and latency are significantly reduced. Moreover, an enhanced integrated intermodulo operation core is developed to decrease the overhead imposed by non-modular operations. Further improvement in systems' performance efficiency is achieved by developing efficient Processing-in-Memory (PIM) designs using various volatile CMOS and non-volatile Post-CMOS technologies to accelerate RNS-based multiplication-and-accumulations (MACs). The RN-SiM accelerator's performance on different datasets, including MNIST, SVHN, and CIFAR-10, is evaluated. With almost the same accuracy to the baseline CNN, the RNSiM accelerator can significantly increase both energy-efficiency and speedup compared with the state-of-the-art FPGA, GPU, and PIM designs. RNSiM and other RNS-PIMs, based on our method, reduce the energy consumption by orders of$28-77\times$and$331-897\times$compared with the FPGA and the GPU platforms, respectively. Arman Roohi, MohammadReza Taheri, Shaahin Angizi, Deliang Fan |
ICCAD | 3 |
| 2021 | Non-Volatile Approximate Arithmetic Circuits Using Scalable Hybrid Spin-CMOS Majority GatesabstractIn the nanoscale era, leakage/static power dissipation has become an inevitable and important issue for CMOS devices. To alleviate this issue, we propose to use spintronic devices with near-zero leakage power and non-volatility as key components in arithmetic circuits for error-resilient applications. To this end, spintronic threshold devices are first utilized to construct highly-scalable majority gates (MGs) based on spin-CMOS technology. These MGs are then used in the design of compressors for constructing multipliers and accumulators. For an MG-based compressor, the truth table of a conventional compressor is transformed to ensure that the outputs depend only on the number of input “1”s. To synthesize and optimize the MG-based circuits, a heuristic majority-inverter graph (HMIG) is further proposed for the design of an accurate and two approximate non-volatile 4-2 compressors (denoted as MG-EC, MG-AC1 and MG-AC2). Due to the high scalability of the MGs, approximate compressors with a larger number of inputs can be devised using the same method. Compared to previous designs, the proposed 4-2 compressors show shorter critical path delays and lower energy consumption; MG-AC1 and MG-AC2 also achieve a higher accuracy than state-of-the-art approximate designs. For achieving a similar image quality in image compression, the multiplier implementations using MG-AC1 and MG-AC2 result in more significant reductions in delay and energy than those using other approximate designs. Honglan Jiang, Shaahin Angizi, Deliang Fan, Jie Han 0001, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | A Flexible Processing-in-Memory Accelerator for Dynamic Channel-Adaptive Deep Neural NetworksabstractWith the success of deep neural networks (DNN), many recent works have been focusing on developing hardware accelerator for power and resource-limited embedded system via model compression techniques, such as quantization, pruning, low-rank approximation, etc. However, almost all existing DNN structure is fixed after deployment, which lacks runtime adaptive DNN structure to adapt to its dynamic hardware resource, power budget, throughput requirement, as well as dynamic workload. Correspondingly, there is no runtime adaptive hardware platform to support dynamic DNN structure. To address this problem, we first propose a dynamic channel-adaptive deep neural network (CA-DNN) which can adjust the involved convolution channel (i.e. model size, computing load) at run-time (i.e. at inference stage without retraining) to dynamically trade off between power, speed, computing load and accuracy. Further, we utilize knowledge distillation method to optimize the model and quantize the model to 8-bits and 16-bits, respectively, for hardware friendly mapping. We test the proposed model on CIFAR-10 and ImageNet dataset by using ResNet. Comparing with the same model size of individual model, our CA-DNN achieves better accuracy. Moreover, as far as we know, we are the first to propose a Processing-in-Memory accelerator for such adaptive neural networks structure based on Spin Orbit Torque Magnetic Random Access Memory(SOT-MRAM) computational adaptive sub-arrays. Then, we comprehensively analyze the trade-off of the model with different channel-width between the accuracy and the hardware parameters, eg., energy, memory, and area overhead. Li Yang 0009, Shaahin Angizi, Deliang Fan |
ASP-DAC | 2 |
| 2020 | PIM-Assembler: A Processing-in-Memory Platform for Genome AssemblyabstractIn this paper, for the first time, we propose a high-throughput and energy-efficient Processing-in-DRAM-accelerated genome assembler called PIM-Assembler based on an optimized and hardware-friendly genome assembly algorithm. PIM-Assembler can assemble large-scale DNA sequence dataset from all-pair overlaps. We first develop PIM-Assembler platform that harnesses DRAM as computational memory and transforms it to a fundamental processing unit for genome assembly. PIM-Assembler can perform efficient X(N)OR-based operations inside DRAM incurring low cost on top of commodity DRAM designs (~5% of chip area). PIM-Assembler is then optimized through a correlated data partitioning and mapping methodology that allows local storage and processing of DNA short reads to fully exploit the genome assembly algorithm-level's parallelism. The simulation results show that PIM-Assembler achieves on average 8.4× and 2.3 wise× higher throughput for performing bulk bit-XNOR-based comparison operations compared with CPU and recent processing-in-DRAM platforms, respectively. As for comparison/addition-extensive genome assembly application, it reduces the execution time and power by ~5× and ~ 7.5× compared to GPU. Shaahin Angizi, Naima Ahmed Fahmi, Wei Zhang 0076, Deliang Fan |
DAC | 1 |
| 2020 | PIM-Aligner: A Processing-in-MRAM Platform for Biological Sequence AlignmentabstractIn this paper, we propose a high-throughput and energy-efficient Processing-in-Memory accelerator (PIM-Aligner) to execute DNA short read alignment based on an optimized and hardware-friendly alignment algorithm. We first reconstruct the existing sequence alignment algorithm based on BWT and FM-index such that it can be fully implemented in PIM platforms. It supports exact alignment and also handles mismatches to reduce excessive backtracking. We then develop PIM-Aligner platform that transforms SOT-MRAM array to a potential computational memory to accelerate the reconstructed alignment-in-memory algorithm incurring a low cost on top of original SOT-MRAM chips (less than 10% of chip area). Accordingly, we present a local data partitioning, mapping, and pipeline technique to maximize the parallelism in multiple computational sub-array while doing the alignment task. The simulation results show that PIM-Aligner outperforms recent platforms based on dynamic programming with ~ 3.1× higher throughput per Watt. Besides, PIM-Aligner improves the short read alignment throughput per Watt per mm2by ~ 9× and 1.9× compared to FM-index-based ASIC and processing-in-ReRAM designs, respectively. Shaahin Angizi, Jiao Sun, Wei Zhang 0076, Deliang Fan |
DATE | 1 |
| 2020 | Exploring DNA Alignment-in-Memory Leveraging Emerging SOT-MRAMabstractIn this work, we review two alternative Processing-in-Memory (PIM) accelerators based on Spin-Orbit-Torque Magnetic Random Access Memory (SOT-MRAM) to execute DNA short read alignment based on an optimized and hardware-friendly alignment algorithm. We first discuss the reconstruction of the existing sequence alignment algorithm based on BWT and FM-index such that it can be fully implemented leveraging PIM functions. We then transform SOT-MRAM array to a potential computational memory by presenting two different reconfigurable sense amplifiers to accelerate the reconstructed alignment-in-memory algorithm. The cross-layer simulation results show that such PIM platforms are able to achieve a nearly ten-fold and two-fold increases in throughput/power/area measure compared with recent ASIC and processing-in-ReRAM designs, respectively. Shaahin Angizi, Wei Zhang 0076, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | Modeling and Benchmarking Computing-in-Memory for Design Space ExplorationabstractThe bottleneck between the limited memory bandwidth and high speed processing demands is the main cause of problems associated with high volume of data transfers in data-intensive applications. As a possible remedy to these issues, computing-in-memory (CiM) enables a subset of logic and arithmetic operations to be performed where the data resides, i.e., inside the memory. Various CiM designs have been proposed to date, based on different technologies. Given the variety of options available, picking the right design option for a system/application can be a complex task. When choosing a CiM design, it is important to establish evaluation conditions that are as uniform as possible to make a fair choice between available design options. In this paper, we describe a methodology for an uniform benchmarking of CiM designs. Our approach evaluates devices/circuits, arrays and the overall impact of CiM to a system with a framework based on Eva-CiM. As a case study, we analyze the array-level performance of 7 recent CiM designs implemented with SRAM, DRAM, FeFET-RAM, STT-MRAM, SOT-MRAM, and RRAM. After we identify that the FeFET-RAM-based design shows promising energy and delay savings at the array level, we carry out a system level evaluation showing that FeFET-RAM-based CiM outperforms a CMOS SRAM CiM baseline by an average of 60% across a set of 17 benchmarks (with respect to energy savings). Regarding speedups, both technologies offer virtually the same benefit of about 1.5X when compared to a situation where processing does not happen in memory. Dayane Reis, Shaahin Angizi, Xunzhao Yin, Deliang Fan, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Sparse BD-Net: A Multiplication-less DNN with Sparse Binarized Depth-wise Separable ConvolutionabstractIn this work, we propose a multiplication-less binarized depthwise-separable convolution neural network, called BD-Net. BD-Net is designed to use binarized depthwise separable convolution block as the drop-in replacement of conventional spatial-convolution in deep convolution neural network (DNN). In BD-Net, the computation-expensive convolution operations (i.e., Multiplication and Accumulation) are converted into energy-efficient Addition/Subtraction operations. For further compressing the model size while maintaining the dominant computation in addition/subtraction, we propose a brand-new sparse binarization method with a hardware-oriented structured sparsity pattern. To successfully train such sparse BD-Net, we propose and leverage two techniques: (1) a modified group-lasso regularization whose group size is identical to the capacity of basic computing core in accelerator and (2) a weight penalty clipping technique to solve the disharmony issue between weight binarization and lasso regularization. The experiment results show that the proposed sparse BD-Net can achieve comparable or even better inference accuracy, in comparison to the full precision CNN baseline. Beyond that, a BD-Net customized process-in-memory accelerator is designed using SOT-MRAM, which owns characteristics of high channel expansion flexibility and computation parallelism. Through the detailed analysis from both software and hardware perspectives, we provide an intuitive design guidance for software/hardware co-design of DNN acceleration on mobile embedded systems. Note that this journal submission is the extended version of our previous published paper in ISVLSI 2018 [24]. Zhezhi He, Li Yang 0009, Shaahin Angizi, Adnan Siraj Rakin, Deliang Fan |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | ApGAN: Approximate GAN for Robust Low Energy Learning From Imprecise ComponentsabstractA Generative Adversarial Network (GAN) is an adversarial learning approach which empowers conventional deep learning methods by alleviating the demands of massive labeled datasets. However, GAN training can be computationally-intensive limiting its feasibility in resource-limited edge devices. In this paper, we propose an approximate GAN (ApGAN) for accelerating GANs from both algorithm and hardware implementation perspectives. First, inspired by the binary pattern feature extraction method along with binarized representation entropy, the existing Deep Convolutional GAN (DCGAN) algorithm is modified by binarizing the weights for a specific portion of layers within both the generator and discriminator models. Further reduction in storage and computation resources is achieved by leveraging a novel hardware-configurable in-memory addition scheme, which can operate in the accurate and approximate modes. Finally, a memristor-based processing-in-memory accelerator for ApGAN is developed. The performance of the ApGAN accelerator on different data-sets such as Fashion-MNIST, CIFAR-10, STL-10, and celeb-A is evaluated and compared with recent GAN accelerator designs. With almost the same Inception Score (IS) to the baseline GAN, the ApGAN accelerator can increase the energy-efficiency by ~28.6× achieving 35-fold speedup compared with a baseline GPU platform. Additionally, it shows 2.5× and 5.8× higher energy-efficiency and speedup over CMOS-ASIC accelerator subject to an 11 percent reduction in IS. Arman Roohi, Shadi Sheikhfaal, Shaahin Angizi, Deliang Fan, Ronald F. DeMara |
IEEE Trans. Computers | 3 |
| 2020 | MRIMA: An MRAM-Based In-Memory AcceleratorabstractIn this paper, we propose MRIMA, as a novel magnetic RAM (MRAM)-based in-memory accelerator for nonvolatile, flexible, and efficient in-memory computing. MRIMA transforms current spin transfer torque magnetic random access memory (STT-MRAM) arrays to massively parallel computational units capable of working as both nonvolatile memory and in-memory logic. Instead of integrating complex logic units in cost-sensitive memory, MRIMA exploits hardware-friendly bit-line computing methods to implement complete Boolean logic functions between operands within a memory array in a single clock cycle, overcoming the multicycle logic issue in contemporary processing-in-memory (PIM) platforms. We present practical case studies to demonstrate MRIMA's acceleration for binary-weight and low bit-width convolutional neural networks (CNNs) as well as data encryption. Our device-to-architecture co-simulation results on CNN acceleration demonstrate that MRIMA can obtain 1.7× better energy-efficiency and 11.2× speed-up compared to ASICs, and 1.8× better energy-efficiency and 2.4× speed-up over the best DRAM-based PIM solutions. As an advanced encryption standard (AES) in-memory encryption engine, MRIMA shows ~77% and 21% lower energy consumption compared to CMOS-ASIC and recent domain-wall-based design, respectively. Shaahin Angizi, Zhezhi He, Amro Awad, Deliang Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | ParaPIM: a parallel processing-in-memory accelerator for binary-weight deep neural networksabstractRecent algorithmic progression has brought competitive classification accuracy despite constraining neural networks to binary weights (+1/-1). These findings show remarkable optimization opportunities to eliminate the need for computationally-intensive multiplications, reducing memory access and storage. In this paper, we present ParaPIM architecture, which transforms current Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) sub-arrays to massively parallel computational units capable of running inferences for Binary-Weight Deep Neural Networks (BWNNs). ParaPIM's in-situ computing architecture can be leveraged to greatly reduce energy consumption dealing with convolutional layers, accelerate BWNNs inference, eliminate unnecessary off-chip accesses and provide ultra-high internal bandwidth. The device-to-architecture co-simulation results indicate ~4x higher energy efficiency and 7.3x speedup over recent processing-in-DRAM acceleration, or roughly 5x higher energy-efficiency and 20.5x speedup over recent ASIC approaches, while maintaining inference accuracy comparable to baseline designs. Shaahin Angizi, Zhezhi He, Deliang Fan |
ASP-DAC | 1 |
| 2019 | AlignS: A Processing-In-Memory Accelerator for DNA Short Read Alignment Leveraging SOT-MRAMabstractClassified as a complex big data analytics problem, DNA short read alignment serves as a major sequential bottleneck to massive amounts of data generated by next-generation sequencing platforms. With Von-Neumann computing architectures struggling to address such computationally-expensive and memory-intensive task today, Processing-in-Memory (PIM) platforms are gaining growing interests. In this paper, an energy-efficient and parallel PIM accelerator (AlignS) is proposed to execute DNA short read alignment based on an optimized and hardware-friendly alignment algorithm. We first develop AlignS platform that harnesses SOT-MRAM as computational memory and transforms it to a fundamental processing unit for short read alignment. Accordingly, we present a novel, customized, highly parallel read alignment algorithm that only seeks the proposed simple and parallel in-memory operations (i.e. comparisons and additions). AlignS is then optimized through a new correlated data partitioning and mapping methodology that allows local storage and processing of DNA sequence to fully exploit the algorithm-level's parallelism, and to accelerate both exact and inexact matches. The device-to-architecture co-simulation results show that AlignS improves the short read alignment throughput per Watt per mm2 by ~12× compared to the ASIC accelerator. Compared to recent FM-index-based ReRAM platform, AlignS achieves 1.6× higher throughput per Watt. Shaahin Angizi, Jiao Sun, Wei Zhang 0076, Deliang Fan |
DAC | 1 |
| 2019 | GraphS: A Graph Processing Accelerator Leveraging SOT-MRAMabstractIn this work, we present GraphS architecture, which transforms current Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) to massively parallel computational units capable of accelerating graph processing applications. GraphS can be leveraged to greatly reduce energy consumption dealing with underlying adjacency matrix computations, eliminating unnecessary off-chip accesses and providing ultra-high internal bandwidth. The device-to-architecture co-simulation for three social network data-sets indicate roughly 3.6× higher energy-efficiency and 5.3× speed-up over recent ReRAM crossbar. It achieves ~4× higher energy-efficiency and 5.1× speed-up over recent processing-in-DRAM acceleration methods. Shaahin Angizi, Jiao Sun, Wei Zhang 0076, Deliang Fan |
DATE | 1 |
| 2019 | GraphiDe: A Graph Processing Accelerator leveraging In-DRAM-ComputingabstractIn this paper, we propose GraphiDe, a novel DRAM-based processing-in-memory (PIM) accelerator for graph processing. It transforms current DRAM architecture to massively parallel computational units exploiting the high internal bandwidth of the modern memory chips to accelerate various graph processing applications. GraphiDe can be leveraged to greatly reduce energy consumption and latency dealing with underlying adjacency matrix computations by eliminating unnecessary off-chip accesses. The extensive circuit-architecture simulations over three social network data-sets indicate that GraphiDe achieves on average 3.1x energy-efficiency improvement and 4.2x speed-up over the recent DRAM based PIM platform. It achieves ~59x higher energy-efficiency and 83x speed-up over GPU-based acceleration methods. Shaahin Angizi, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | ReDRAM: A Reconfigurable Processing-in-DRAM Platform for Accelerating Bulk Bit-Wise OperationsabstractIn this paper, we propose ReDRAM, as a reconfigurable DRAM-based processing-in-memory (PIM) accelerator, which transforms current DRAM architecture to massively parallel computational units exploiting the high internal bandwidth of modern memory chips. ReDRAM uses the analog operation of DRAM sub-arrays and elevates it to implement a full set of 1- and 2-input bulk bit-wise operations (NOT, (N)AND, (N)OR, and even X(N)OR) between operands stored in the same bit-line, based on a new dual-row activation mechanism with a modest change to peripheral circuits such sense amplifiers. ReDRAM can be leveraged to greatly reduce energy consumption and latency of complex in-DRAM logic computations relying on state-of-the-art mechanisms based on triple-row activation, dual-contact cells, row initialization, NOR style, etc. The extensive circuit-architecture simulations show that ReDRAM achieves on average 54× and 7.1× higher throughput for performing bulk bit-wise operations compared with CPU and GPU, respectively. Besides, ReDRAM outperforms recent processing-in-DRAM platforms with up to 3.7× better performance. Shaahin Angizi, Deliang Fan |
ICCAD | 1 |
| 2018 | IMCE: Energy-efficient bit-wise in-memory convolution engine for deep neural networkabstractIn this paper, we pave a novel way towards the concept of bit-wise In-Memory Convolution Engine (IMCE) that could implement the dominant convolution computation of Deep Convolutional Neural Networks (CNN) within memory. IMCE employs parallel computational memory sub-array as a fundamental unit based on our proposed Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) design. Then, we propose an accelerator system architecture based on IMCE to efficiently process low bit-width CNNs. This architecture can be leveraged to greatly reduce energy consumption dealing with convolutional layers and also accelerate CNN inference. The device to architecture co-simulation results show that the proposed system architecture can process low bit-width AlexNet on ImageNet data-set favorably with 785.25μJ/img, which consumes ~3× less energy than that of recent RRAM based counterpart. Besides, the chip area is ~4× smaller. Shaahin Angizi, Zhezhi He, Farhana Parveen, Deliang Fan |
ASP-DAC | 1 |
| 2018 | HielM: Highly flexible in-memory computing using STT MRAMabstractIn this paper we propose a Highly Flexible InMemory (HieIM) computing platform using STT MRAM, which can be leveraged to implement Boolean logic functions without sacrificing memory functionality. It could pre-process data within memory to further reduce power hungry long distance communication between memory and processing units as in Von-Neumann computing system. HieIM can implement all the Boolean logic functions (AND/NAND, OR/NOR, XOR/XNOR) between any two cells in the same memory array, thus overcoming the `operand locality' problem in contemporary in-memory computing platform designs. To investigate the performance of HieIM, we test in-memory bulk bit-wise Boolean logic operations using different vector datasets, which shows ~ 8x energy saving and ~ 5x speedup compared to recent DRAM based in-memory computing platform. We further implement an in-memory data encryption engine design based on HieIM as another case study. With AES algorithm, it shows 51.5% and 68.9% lower energy consumption compared to CMOS-ASIC and CMOL based implementations, respectively. Farhana Parveen, Zhezhi He, Shaahin Angizi, Deliang Fan |
ASP-DAC | 3 |
| 2018 | PIMA-logic: a novel processing-in-memory architecture for highly flexible and energy-efficient logic computationabstractIn this paper, we propose PIMA-Logic, as a novel Processing-in-Memory Architecture for highly flexible and efficient Logic computation. Insteadof integrating complex logic units in cost-sensitive memory, PIMA-Logic exploits a hardware-friendly approach to implement Boolean logic functions between operands either located in the same row or the same column within entire memory arrays. Furthermore, it can efficiently process more complex logic functions between multiple operands to further reduce the latency and power-hungry data movement. The proposed architecture is developed based on Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) array and it can simultaneously work as a non-volatile memory and a reconfigurable in-memory logic. The device-to-architecture co-simulation results show that PIMA-Logic can achieve up to 56% and 31.6% improvements with respect to overall energy and delay on combinational logic benchmarks compared to recent Pinatubo architecture. We further implement an in-memory data encryption engine based on PIMA-Logic as a case study. With AES application, it shows 77.2% and 21% lower energy consumption compared to CMOS-ASIC and recent RIMPA implementation, respectively. Shaahin Angizi, Zhezhi He, Deliang Fan |
DAC | 1 |
| 2018 | CMP-PIM: an energy-efficient comparator-based processing-in-memory neural network acceleratorabstractIn this paper, an energy-efficient and high-speed comparator-based processing-in-memory accelerator (CMP-PIM) is proposed to efficiently execute a novel hardware-oriented comparator-based deep neural network called CMPNET. Inspired by local binary pattern feature extraction method combined with depthwise separable convolution, we first modify the existing Convolutional Neural Network (CNN) algorithm by replacing the computationally-intensive multiplications in convolution layers with more efficient and less complex comparison and addition. Then, we propose a CMP-PIM that employs parallel computational memory sub-array as a fundamental processing unit based on SOT-MRAM. We compare CMP-PIM accelerator performance on different data-sets with recent CNN accelerator designs. With the close inference accuracy on SVHN data-set, CMP-PIM can get ∼ 94× and 3× better energy efficiency compared to CNN and Local Binary CNN (LBCNN), respectively. Besides, it achieves 4.3× speed-up compared to CNN-baseline with identical network configuration. Shaahin Angizi, Zhezhi He, Adnan Siraj Rakin, Deliang Fan |
DAC | 1 |
| 2018 | Leveraging Spintronic Devices for Efficient Approximate Logic and Stochastic Neural NetworksabstractITRS has identified nano-magnet based spintronic devices as promising post-CMOS technologies for information processing and data storage due to their ultra-low switching energy, non-volatility, superior endurance, excellent retention time, high integration density and compatibility with CMOS technology. As for data storage, spintronic memory has been widely accepted as a universal high performance next-generation non-volatile memory candidate. As for information processing, spintronic computing remains complementary in its features to CMOS technology. In this paper, we present two innovative spintronic computing primitives, i.e. spintronic approximate logic and spintronic stochastic neural network, which both leverage the intrinsic spintronic device physics to achieve much more compact and efficient designs than CMOS counterparts. In spintronic approximate logic, we employ the intrinsic current-mode thresholding operation to implement an accuracy-configurable adder and further demonstrate its application in approximate DSP applications. In spintronic stochastic neural networks, we leverage the stochastic properties of domain wall devices and magnetic tunnel junction to implement a low-power and robust artificial neural network design. Shaahin Angizi, Zhezhi He, Yu Bai 0004, Jie Han 0001, Mingjie Lin, Ronald F. DeMara, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | DIMA: a depthwise CNN in-memory acceleratorabstractIn this work, we first propose a deep depthwise Convolutional Neural Network (CNN) structure, called Add-Net, which uses bi-narized depthwise separable convolution to replace conventional spatial-convolution. In Add-Net, the computationally expensive convolution operations (i.e. Multiplication and Accumulation) are converted into hardware-friendly Addition operations. We meticulously investigate and analyze the Add-Net's performance (i.e. accuracy, parameter size and computational cost) in object recognition application compared to traditional baseline CNN using the most popular large scale ImageNet dataset. Accordingly, we propose a Depthwise CNN In-Memory Accelerator (DIMA) based on SOT-MRAM computational sub-arrays to efficiently accelerate Add-Net within non-volatile MRAM. Our device-to-architecture co-simulation results show that, with almost the same inference accuracy to the baseline CNN on different data-sets, DIMA can obtain ∼1.4× better energy-efficiency and 15.7× speedup compared to ASICs, and, ∼1.6× better energy-efficiency and 5.6× speedup over the best processing-in-DRAM accelerators. Shaahin Angizi, Zhezhi He, Deliang Fan |
ICCAD | 1 |
| 2018 | PIM-TGAN: A Processing-in-Memory Accelerator for Ternary Generative Adversarial NetworksabstractGenerative Adversarial Network (GAN) has emerged as one of the most promising semi-supervised learning methods where two neural nets train themselves in a competitive environment. In this paper, as far as we know, we are the first to present a statistically trained Ternarized Generative Adversarial Network (TGAN) with fully ternarized weights (i.e. -1,0,+1) to massively reduce the need for computation and storage resources in the conventional GAN structures. In the proposed TGAN, the computationally expensive convolution operations (i.e. Multiplication and Accumulation) in both generator and discriminator’s forward path are converted into hardwarefriendly Addition/Subtraction operations. Accordingly, we propose a Processing-in-Memory accelerator for TGAN called (PIM-TGAN) based on Spin-Orbit Torque Magnetic Random Access Memory (SOT-MRAM) computational sub-arrays to efficiently accelerate the training process of GAN within non-volatile memory. In addition, we propose a parallelism technique to further enhance the training efficiency of TGAN. Our device-to-architecture co-simulation results show that, with almost the same inception score to the baseline GAN with floating point number weights on different data-sets, the proposed PIM-TGAN can obtain ~25.6× better energy-efficiency and 22× speedup compared to GPU platform averagely, and, 9.2× better energy-efficiency and 5.4× speedup over the best processing-in-ReRAM accelerators. Adnan Siraj Rakin, Shaahin Angizi, Zhezhi He, Deliang Fan |
ICCD | 2 |
| 2018 | IMFlexCom: Energy Efficient In-Memory Flexible Computing Using Dual-Mode SOT-MRAMabstractIn this article, we propose an In-Memory Flexible Computing platform (IMFlexCom) using a novel Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) array architecture, which could work in dual mode: memory mode and computing mode. Such intrinsic in-memory logic (AND/OR/XOR) could be used to process data within memory to greatly reduce power-hungry and long distance massive data communication in conventional Von Neumann computing systems. A comprehensive reliability analysis is performed, which confirms ∼90mV and ∼10mV (worst-case) sense margin for memory and in-memory logic operation in variations on resistance-area product and tunnel magnetoresistance. We further show that sense margin for in-memory logic computation can be significantly increased by increasing the oxide thickness. Furthermore, we employ bulk bitwise vector operation and data encryption engine as case studies to investigate the performance of our proposed design. IMFlexCom shows ∼35× energy saving and ∼18× speedup for bulk bitwise in-memory vector AND/OR operation compared to DRAM-based in-memory logic. Again, IMFlexCom can achieve 77.27% and 85.4% lower energy consumption compared to CMOS-ASIC- and CMOL-based Advanced Encryption Standard (AES) implementations, respectively. It offers almost similar energy consumption as recent DW-AES implementation with 66.7% less area overhead. Farhana Parveen, Shaahin Angizi, Deliang Fan |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | Design and Evaluation of a Spintronic In-Memory Processing Platform for Nonvolatile Data EncryptionabstractIn this paper, we propose an energy-efficient reconfigurable platform for in-memory processing based on novel four-terminal spin Hall effect-driven domain wall motion devices that could be employed as both nonvolatile memory cell and in-memory logic unit. The proposed designs lead to unity of memory and logic. The device to system level simulation results show that, with 28% area increase in memory structure, the proposed in-memory processing platform achieves a write energy ~15.6 fJ/bit with 79% reduction compared to that of SOT-MRAM counterpart while keeping the identical 1 ns writing speed. In addition, the proposed in-memory logic scheme improves the operating energy by 61.3%, as compared with the recent nonvolatile in-memory logic designs. An extensive reliability analysis is also performed over the proposed circuits. We employ advanced encryption standard (AES) algorithm as a case study to elucidate the efficiency of the proposed platform at application level. Simulation results exhibit that the proposed platform can show up to 75.7% and 30.4% lower energy consumption compared to CMOS-ASIC and recent pipelined domain wall (DW) AES implementations, respectively. In addition, the AES energy-delay product can show 15.1% and 6.1% improvements compared to the DW-AES and CMOS-ASIC implementations, respectively. Shaahin Angizi, Zhezhi He, Nader Bagherzadeh, Deliang Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Energy Efficient In-Memory Computing Platform Based on 4-Terminal Spin Hall Effect-Driven Domain Wall Motion DevicesabstractIn this paper, we propose an energy efficient in-memory computing platform based on novel 4-terminal spin Hall effect-driven domain wall motion devices that could be employed as both non-volatile memory cell and in-memory logic unit. The proposed designs lead to unity of memory and logic. The device to architecture level simulation results show that, with 45% area increase, the proposed in-memory computing platform achieves the write energy 15.6 ~ fJ/bit which is more than one order lower than that of standard 1-transistor 1-magnetic tunnel junction counterpart while keeping the identical 1ns writing speed. In addition, the proposed in-memory logic scheme improves the operating energy by 61.3% as compared with the conventional nonvolatile in-memory logic designs. Shaahin Angizi, Zhezhi He, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 1 |
| 2017 | Leveraging Dual-Mode Magnetic Crossbar for Ultra-low Energy In-memory Data EncryptionabstractThe logic-in-memory architecture is highly promising for high-throughput data-driven applications. This paper presents a novel dual-mode magnetic crossbar architecture consisting of perpendicularly cross-coupled magnetic racetrack nanowires, which could morph between non-volatile multi-bit racetrack memory mode and in-memory data encryption mode. The proposed magnetic crossbar is able to automatically perform parallel in-memory bit-wise XOR computations of the data stored in the racetrack memories with the help of magnetic coupling physics without complex peripheral circuits, which could be leveraged to design energy efficient in-memory data encryption engine. We employ Advanced Encryption Standard (AES) algorithm to elucidate the efficiency of the proposed design. The device-to-architecture level simulation results show that the proposed architecture can achieve 70% and 17.5% lower energy consumption compared to CMOS-ASIC and recent domain wall (DW) AES implementations, respectively. In addition, the AES encryption speed increases by 29.7% compared to the DW-AES implementation. Zhezhi He, Shaahin Angizi, Farhana Parveen, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Energy Efficient In-Memory Binary Deep Neural Network Accelerator with Dual-Mode SOT-MRAMabstractIn this paper, we explore potentials of leveraging spin-based in-memory computing platform as an accelerator for Binary Convolutional Neural Networks (BCNN). Such platform can implement the dominant convolution computation based on presented Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) array. The proposed array architecture could simultaneously work as non-volatile memory and a reconfigurable in-memory logic (AND, OR) without add-on logic circuits to memory chip as in conventional logic-in-memory designs. The computed logic output could be also simply read out like a normal MRAM bit-cell using the shared memory peripheral circuits. We employ such intrinsic in-memory computing architecture to efficiently process data within memory to greatly reduce power hungry and omit long distance data communication concerning state-of-the-art BCNN hardware. Deliang Fan, Shaahin Angizi |
ICCD | 2 |
| 2017 | Exploring STT-MRAM Based In-Memory Computing Paradigm with Application of Image Edge ExtractionabstractIn this paper, we propose a novel Spin-Transfer Torque Magnetic Random-Access Memory (STT-MRAM) array design that could simultaneously work as non-volatile memory and implement a reconfigure in-memory logic operation without add-on logic circuits to the memory chip. The computed output could be simply read out like a typical MRAM bit-cell through the modified peripheral circuit. Such intrinsic in-memory computation can be used to process data locally and transfers the "cooked" data to the primary processing unit (i.e. CPU or GPU) for complex computation with high precision requirement. It greatly reduces power-hungry and long distance data communication, and further leads to extreme parallelism within memory. In this work, we further propose an in-memory edge extraction algorithm as a case study to demonstrate the efficiency of in-memory preprocessing methodology. The simulation results show that our edge extraction method reduces data communication as much as 8x for grayscale image, thus greatly reducing system energy consumption. Meanwhile, the F-measure result shows only ∼10% degradation compared to conventional edge detection operators, such as Prewitt, Sobel and Roberts. Zhezhi He, Shaahin Angizi, Deliang Fan |
ICCD | 2 |
| 2017 | Hybrid polymorphic logic gate using 6 terminal magnetic domain wall motion deviceabstractPolymorphic gates are capable of adapting to multiple functionalities depending on the application and need. In this paper, we propose a hybrid spin-CMOS polymorphic logic gate based on a novel 6 terminal composite magnetic domain wall motion device structure. As far as we know, we are the first to present a single polymorphic gate that is able to perform a full set of 2-input Boolean logic functions (i.e. AND/NAND, OR/NOR, NOT, XOR/XNOR) by configuring the applied keys. The SPICE device-circuit co-simulation indicates that a full adder design using our proposed polymorphic logic gate shows 45.74% power reduction compared with traditional CMOS full adder design. Moreover, it can be a promising hardware security primitive by implementing logic locking and polymorphic transformation to protect Integrated Circuit (IC) against counterfeiting and reverse engineering. To summarize, our proposed design simultaneously provides non-volatility, low power consumption, compactness and polymorphism to logic circuits, which opens a new paradigm for future power efficient and secured computing. Farhana Parveen, Shaahin Angizi, Zhezhi He, Deliang Fan |
ISCAS | 2 |
| 2017 | Low power in-memory computing based on dual-mode SOT-MRAMabstractIn this paper, we propose a novel Spin Orbit Torque Magnetic Random Access Memory (SOT-MRAM) array design that could simultaneously work as non-volatile memory and implement a reconfigurable in-memory logic (AND, OR) without add-on logic circuits to memory chip as in traditional logic-in-memory designs. The computed logic output could be simply read out like a normal MRAM bit-cell using the shared memory peripheral circuits. Such intrinsic in-memory logic could be used to process data within memory to greatly reduce power-hungry and long distance data communication in conventional Von-Neumann computing systems. We further employ in-memory data encryption using Advanced Encryption Standard (AES) algorithm as a case study to demonstrate the efficiency of the proposed design. The device to architecture co-simulation results show that the proposed design can achieve 70.15% and 80.87% lower energy consumption compared to CMOS-ASIC and CMOL-AES implementations, respectively. It offers almost similar energy consumption as recent DW-AES implementation, but with 60.65% less area overhead. Farhana Parveen, Shaahin Angizi, Zhezhi He, Deliang Fan |
ISLPED | 2 |
| 2015 | Restoring and non-restoring array divider designs in Quantum-dot Cellular Automata
Samira Sayedsalehi, Mostafa Rahimi Azghadi, Shaahin Angizi, Keivan Navi |
Inf. Sci. | 3 |