EDBT 2026 Demo / reviewers in the wild / expert
Ramtin Zand
dblp:166/3057
· DBLP profile ↗
39ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PiCASO: A Real-time Pi-Integrated Conversational AI System with Optimized LLMsabstractReal-time conversational AI holds immense promise for applications such as social robotics and human–computer interaction; however, its deployment on edge devices is constrained by the latency and energy demands of large language models (LLMs) and speech recognition systems. This work addresses these challenges by presenting PiCASo, a three-part modular system consisting of a Whisper-based Speech-to-Text (STT) component, a quantized LLM-powered Question-Answering (QA) module, and a lightweight Text-to-Speech (TTS) module. To enable efficient deployment on resource-constrained platforms such as the Raspberry Pi 5, we apply model compression techniques including post-training quantization at 8-bit, 4-bit, and 2-bit precision, along with ternary 1.58-bit models trained using quantization-aware training. We introduce a multi-objective optimization framework that balances accuracy (semantic similarity), measured by the NUBIA score, and throughput, measured in tokens per second (TPS). Experimental evaluations on noisy and noiseless variants of the SQuAD dataset show that Phi-3 models achieve the strongest semantic accuracy, and Llama-8B with 1.58-bit precision delivers robust performance while supporting real-time inference with throughput exceeding 3 TPS on resource-constrained edge devices. Our results demonstrate that strategic compression combined with edge-aware optimization enables scalable, low-latency, and resilient conversational AI suitable for real-world, noise-prone environments. Mahsa Ardakani, Jinendra Malekar, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | SENTRY: Spiking Event Reasoning for Selective Deep Inference in Event-Driven Edge VisionabstractEvent-driven cameras are well suited for always-on edge vision, but forwarding all events creates unnecessary overhead. Events outside user-defined ROIs are semantically irrelevant, while many ROI events arise from background motion, lighting changes, or sensor noise rather than meaningful activity. We propose SENTRY, a near-sensor pipeline that addresses both issues through hierarchical spatial reasoning. A coarse spike-rate gate discards frames with low ROI activity, while a spatially-aware confirmation stage compares each region’s activity against the global background rate to suppress diffuse noise. Evaluated on a real indoor surveillance sequence, SENTRY achieves +64% precision and − 57% false positive rate relative to a global spike-rate baseline, targeting resource-constrained, always-on platforms where energy efficiency and rapid response must be achieved simultaneously. Shayan Gerami, Sepehr Tabrizchi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | ReCQ: Residual Compensation for Quantized Convolutional Neural NetworksabstractPost-training quantization (PTQ) is widely used to reduce the computational and memory requirements of deep neural networks for deployment on resource-constrained platforms. However, aggressive low-bit quantization often introduces significant accuracy degradation, particularly in deep convolutional architectures. To address this challenge, we propose ReCQ, a lightweight residual compensation framework designed to recover quantization-induced errors in convolutional neural networks (CNNs). ReCQ augments quantized models with small depthwise-separable compensation modules that operate in parallel with existing convolutional blocks, learning residual corrections while preserving the original network structure. We evaluate ReCQ across a diverse set of CNN architectures, including VGG, ResNet, MobileNetV2, and ConvNeXt, representing different design paradigms such as plain convolutional networks, residual networks, depthwise-separable architectures, and modern convolutional backbones. Experimental results show that ReCQ consistently improves the accuracy of quantized models compared to standard PTQ and the recent QwT method while introducing only minimal model size overhead. In many configurations, ReCQ achieves higher accuracy with comparable or even smaller parameter overhead, demonstrating its effectiveness and general applicability across modern CNN architectures. Mohammadreza Mohammadi, Matthew Grenier, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | SpikeViT: A Memory-Efficient Mobile Spiking Vision TransformabstractSpiking Transformers constitute an emerging class of neural architectures that seek to unify the representational power of Transformer-based models with the computational efficiency of spiking neural networks (SNNs). By leveraging discrete spike-based communication and event-driven processing, Spiking Transformers enable temporally sparse computation while maintaining the global context modeling and scalability inherent to self-attention mechanisms. This integration facilitates energy-efficient sequence modeling and opens new avenues for deploying large-scale attention-based models on neuromorphic hardware. However, existing Spiking Transformer architectures often incur substantial memory overhead, limiting their suitability for deployment in resource-constrained environments such as edge devices. To address this limitation, we propose SpikeViT, an efficient Spiking Transformer architecture designed to minimize memory consumption while preserving representational capacity. The architecture adopts a parallel design, combining a convolutional SNN with a lightweight transformer, connected via bidirectional cross-modal bridges that enable efficient tokenization and integration of spike-based features. Experimental results on the CIFAR10-DVS dataset show that SpikeViT achieves competitive accuracy while reducing memory footprint by up to 50% compared to state-of-the-art models, making it well-suited for deployment in energy- and memory-constrained neuromorphic systems. James Seekings, Hasti Zanganeh, Brendan Reidy, Jason Kamran Eshraghian, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | Closing the Loop in LLM-Based Hardware Generation: An Autonomous Agentic Workflow for Robust TPU Design
Deepak Vungarala, Kartik Pandit, Gamana Aragonda, Jeremy McLynch, Adeola Adeoye-Davids, Bryan Galecio, NhatHai Phan, Abdallah Khreishah, Ramtin Zand, Arnob Ghosh, Shaahin Angizi |
VTS | 9 |
| 2025 | ResISC: Residue Number System-Based Integrated Sensing and Computing for Efficient Edge AIabstractThis paper presents ResISC, an RNS-based integrated sensing and computing architecture enabling efficient edge AI. ResISC platform features (i) an in-sensor residue encoder converting images directly to RNS in the analog domain, (ii) an energy-efficient RNS-based processing-near-sensor CNN accelerator utilizing SOT-MRAM, and (iii) an innovative mixed-radix unit for efficient activation operations. By employing selective channel deactivation, ResISC reduces computation overhead by up to $89 \%$, while achieving a $3.4 \times$ improvement in power efficiency and up to a $71 \times$ reduction in execution time compared to processing-in-MRAM platforms. Experiments on various datasets demonstrate that ResISC achieves competitive accuracy levels (up to $94.63 \%$ on CIFAR-10) with minimal degradation, making it an ideal solution for power-constrained, real-time edge applications. Sepehr Tabrizchi, Samin Sohrabi, Mohamadreza Mohammadi, Ramtin Zand, Shaahin Angizi, Arman Roohi |
DAC | 4 |
| 2025 | LLM-IMC: Automating Analog In-Memory Computing Architecture Generation with Large Language ModelsabstractResistive crossbars enabling analog In-Memory Computing (IMC) have garnered significant attention from academia and industry as a promising architecture for Deep Neural Network (DNN) acceleration, thanks to their high memory access bandwidth and in-situ computing capabilities. However, the knowledge-intensive hardware design process and the lack of high-quality circuit netlists have constrained design space exploration and optimization of analog IMC to behavioral system-level tools. In this one-page abstract, we introduce LLM-IMC, a novel fine-tune-free Large Language Model (LLM) framework, supported by a Python-based tool, designed for analog IMC SPICE code generation. LLM-IMC systematically addresses these limitations by automating the creation of diverse IMC simulation scripts, enabling efficient design space exploration through LLM-driven performance, and outlining an integration roadmap for hardware-oriented neuromorphic crossbar design flows. Deepak Vungarala, Md Hasibul Amin, Pietro Mercati, Arman Roohi, Ramtin Zand, Shaahin Angizi |
FCCM | 5 |
| 2025 | CrossNAS: A Cross-Layer Neural Architecture Search Framework for PIM SystemsabstractIn this paper, we propose the CrossNAS framework, an automated approach for exploring a vast, multidimensional search space that spans various design abstraction layers-circuits, architecture, and systems-to optimize the deployment of machine learning workloads on analog processing-in-memory (PIM) systems. CrossNAS leverages the single-path one-shot weight-sharing strategy combined with the evolutionary search for the first time in the context of PIM system mapping and optimization. CrossNAS sets a new benchmark for PIM neural architecture search (NAS), outperforming previous methods in both accuracy and energy efficiency while maintaining comparable or shorter search times. Md Hasibul Amin, Mohammadreza Mohammadi, Jason D. Bakos, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | PixelPrune: Optimizing AIoT Vision Systems via In-Sensor Segmentation and Adaptive Data Transfer
Mohammadreza Mohammadi, Mehrdad Morsali, Sepehr Tabrizchi, Brendan Reidy, Arman Roohi, Shaahin Angizi, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | A Decomposition-Based Memristive Crossbar Solver and FPGA-Accelerated Hardware Implementation
Suyash Vardhan Singh, Anzhelika Kolinko, Md Hasibul Amin, Ramtin Zand, Jason D. Bakos |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Magnetic In/Near-Sensor Architectures: From Raw Sensing to Smart Processing
Sepehr Tabrizchi, Ali Shafiee Sarvestani, Md Hasibul Amin, Deniz Najafi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | From Prompt to Accelerator: A Perspective on LLM-Based Analog In-Memory Accelerator Design Automation
Deepak Vungarala, Md Hasibul Amin, Arman Roohi, Arnob Ghosh, Ramtin Zand, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Rep Smarter, Not Harder: AI Hypertrophy Coaching with Wearable Sensors and Edge Neural NetworksabstractOptimizing resistance training for hypertrophy requires balancing proximity to muscular failure, often quantified by Repetitions in Reserve (RiR), with fatigue management. However, subjective RiR assessment is unreliable, leading to suboptimal training stimuli or excessive fatigue. This paper introduces a novel system for real-time feedback on near-failure states (RiR $\le$ 2) during resistance exercise using only a single wrist-mounted Inertial Measurement Unit (IMU). We propose a two-stage pipeline suitable for edge deployment: first, a ResNet-based model segments repetitions from the 6-axis IMU data in real-time. Second, features derived from this segmentation, alongside direct convolutional features and historical context captured by an LSTM, are used by a classification model to identify exercise windows corresponding to near-failure states. Using a newly collected dataset from 13 diverse participants performing preacher curls to failure (631 total reps), our segmentation model achieved an F1 score of 0.83, and the near-failure classifier achieved an F1 score of 0.82 under simulated real-time evaluation conditions (1.6 Hz inference rate). Deployment on a Raspberry Pi 5 yielded an average inference latency of 112 ms, and on an iPhone 16 yielded 23.5 ms, confirming the feasibility for edge computation. This work demonstrates a practical approach for objective, real-time training intensity feedback using minimal hardware, paving the way for accessible AI-driven hypertrophy coaching tools that help users manage intensity and fatigue effectively. Grant King, Musa Azeem, Savannah Noblitt, Ramtin Zand, Homayoun Valafar |
ICMLA | 4 |
| 2025 | NSF-MAP: Neurosymbolic Multimodal Fusion for Robust and Interpretable Anomaly Prediction in Assembly PipelinesabstractIn modern assembly pipelines, identifying anomalies is crucial in ensuring product quality and operational efficiency. Conventional single-modality methods fail to capture the intricate relationships required for precise anomaly prediction in complex predictive environments with abundant data and multiple modalities. This paper proposes a neurosymbolic AI and fusion-based approach for multimodal anomaly prediction in assembly pipelines. We introduce a time series and image-based fusion model that leverages decision-level fusion techniques. Our research builds upon three primary novel approaches in multimodal learning: time series and image-based decision-level fusion modeling, transfer learning for fusion, and knowledge-infused learning. We evaluate the novel method using our derived and publicly available multimodal dataset and conduct comprehensive ablation studies to assess the impact of our preprocessing techniques and fusion model compared to traditional baselines. The results demonstrate that a neurosymbolic AI-based fusion approach that uses transfer learning can effectively harness the complementary strengths of time series and image data, offering a robust and interpretable approach for anomaly prediction in assembly pipelines with enhanced performance. \noindent The datasets, codes to reproduce the results, supplementary materials, and demo are available at https://github.com/ChathurangiShyalika/NSF-MAP. Chathurangi Shyalika, Renjith Prasad, Fadi El Kalach, Revathy Venkataramanan, Ramtin Zand, Ramy F. Harik, Amit P. Sheth |
IJCAI | 5 |
| 2024 | Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image ProcessingabstractThis paper proposes a high-performance and energy-efficient optical near-sensor accelerator for vision applications, called Lightator. Harnessing the promising efficiency offered by photonic devices, Lightator features innovative compressive acquisition of input frames and fine-grained convolution operations for low-power and versatile image processing at the edge for the first time. This will substantially diminish the energy consumption and latency of conversion, transmission, and processing within the established cloud-centric architecture as well as recently designed edge accelerators. Our device-to-architecture simulation results show that with favorable accuracy, Lightator achieves 84.4 Kilo FPS/W and reduces power consumption by a factor of ~24× and 73× on average compared with existing photonic accelerators and GPU baseline. Mehrdad Morsali, Brendan Reidy, Deniz Najafi, Sepehr Tabrizchi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Ramtin Zand, Shaahin Angizi |
DAC | 8 |
| 2024 | HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROIabstractWith the rise of tiny IoT devices powered by machine learning (ML), many researchers have directed their focus toward compressing models to fit on tiny edge devices. Recent works have achieved remarkable success in compressing ML models for object detection and image classification on microcontrollers with small memory, e.g., 512kB SRAM. However, there remain many challenges prohibiting the deployment of ML systems that require high-resolution images. Due to fundamental limits in memory capacity for tiny IoT devices, it may be physically impossible to store large images without external hardware. To this end, we propose a high-resolution image scaling system for edge ML, called HiRISE, which is equipped with selective region-of-interest (ROI) capability leveraging analog in-sensor image scaling. Our methodology not only significantly reduces the peak memory requirements, but also achieves up to 17.7× reduction in data transfer and energy consumption. Brendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi, Arman Roohi, Ramtin Zand |
DAC | 6 |
| 2024 | Edge-Centric Real-Time Segmentation for Autonomous Underwater Cave ExplorationabstractThis paper addresses the challenge of deploying machine learning (ML)-based segmentation models on edge platforms to facilitate real-time scene segmentation for Autonomous Underwater Vehicles (AUVs) in underwater cave exploration and mapping scenarios. We focus on three ML models-U-Net, CaveSeg, and YOLOv8n-deployed on four edge platforms: Raspberry Pi-4, Intel Neural Compute Stick 2 (NCS2), Google Edge TPU, and NVIDIA Jetson Nano. Experimental results reveal that mobile models with modern architectures, such as YOLOv8n, and specialized models for semantic segmentation, like U-Net, offer higher accuracy with lower latency. YOLOv8n emerged as the most accurate model, achieving a 72.5 Intersection Over Union (IoU) score. Meanwhile, the U-Net model deployed on the Coral Dev board delivered the highest speed at 79.24 FPS and the lowest energy consumption at 6.23 mJ. The detailed quantitative analyses and comparative results presented in this paper offer critical insights for deploying cave segmentation systems on underwater robots, ensuring safe and reliable AUV navigation during cave exploration and mapping missions. Mohammadreza Mohammadi, Adnan Abdullah, Aishneet Juneja, Ioannis M. Rekleitis, Md Jahidul Islam, Ramtin Zand |
ICMLA | 6 |
| 2024 | AssemAI: Interpretable Image-Based Anomaly Detection for Manufacturing PipelinesabstractAnomaly detection in manufacturing pipelines remains a critical challenge, intensified by the complexity and variability of industrial environments. This paper introduces AssemAI, an interpretable image-based anomaly detection system tailored for smart manufacturing pipelines. Utilizing a curated image dataset from an industry-focused rocket assembly pipeline, we address the challenge of imbalanced image data and demonstrate the importance of image-based methods in anomaly detection. Our primary contributions include deriving an image dataset, fine-tuning an object detection model YOLO-FF, and implementing a custom anomaly detection model for assembly pipelines. The proposed approach leverages domain knowledge in data preparation, model development and reasoning. We implement several anomaly detection models on the derived image dataset, including a Convolutional Neural Network, Vision Transformer (ViT), and pretrained versions of these models. Additionally, we incorporate explainability techniques at both user and model levels, utilizing ontology for user-level explanations and SCORE-CAM for indepth feature and model analysis. Finally, the best-performing anomaly detection model and YOLO-FF are deployed in a real-time setting. Our results include ablation studies on the baselines and a comprehensive evaluation of the proposed system. This work highlights the broader impact of advanced image-based anomaly detection in enhancing the reliability and efficiency of smart manufacturing processes. The image dataset, codes to reproduce the results and additional experiments are available at https:/github.com/renjithk4/AssemAI. Renjith Prasad, Chathurangi Shyalika, Fadi El Kalach, Revathy Venkataramanan, Ramtin Zand, Ramy F. Harik, Amit P. Sheth |
ICMLA | 5 |
| 2023 | IMAC-Sim: : A Circuit-level Simulator For In-Memory Analog Computing ArchitecturesabstractWith the increased attention to memristive-based in-memory analog computing (IMAC) architectures as an alternative for energy-hungry computer systems for machine learning applications, a tool that enables exploring their device- and circuit-level design space can significantly boost the research and development in this area. Thus, in this paper, we develop IMAC-Sim, a circuit-level simulator for the design space exploration of IMAC architectures. IMAC-Sim is a Python-based simulation framework, which creates the SPICE netlist of the IMAC circuit based on various device- and circuit-level hyperparameters selected by the user, and automatically evaluates the accuracy, power consumption, and latency of the developed circuit using a user-specified dataset. Moreover, IMAC-Sim simulates the interconnect parasitic resistance and capacitance in the IMAC architectures and is also equipped with horizontal and vertical partitioning techniques to surmount these reliability challenges. IMAC-Sim is a flexible tool that supports a broad range of device- and circuit-level hyperparameters. In this paper, we perform controlled experiments to exhibit some of the important capabilities of the IMAC-Sim, while the entirety of its features is available for researchers via an open-source tool at https://github.com/iCAS-Lab/IMAC-Sim. Md Hasibul Amin, Mohammed E. Elbtity, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Heterogeneous Integration of In-Memory Analog Computing Architectures with Tensor Processing UnitsabstractTensor processing units (TPUs), specialized hardware accelerators for machine learning tasks, have shown significant performance improvements when executing convolutional layers in convolutional neural networks (CNNs). However, they struggle to maintain the same efficiency in fully connected (FC) layers, leading to suboptimal hardware utilization. In-memory analog computing (IMAC) architectures, on the other hand, have demonstrated notable speedup in executing FC layers. This paper introduces a novel, heterogeneous, mixed-signal, and mixed-precision architecture that integrates an IMAC unit with an edge TPU to enhance mobile CNN performance. To leverage the strengths of TPUs for convolutional layers and IMAC circuits for dense layers, we propose a unified learning algorithm that incorporates mixed-precision training techniques to mitigate potential accuracy drops when deploying models on the TPU-IMAC architecture. The simulations demonstrate that the TPU-IMAC configuration achieves up to 2.59× performance improvements, and 88% memory reductions compared to conventional TPU architectures for various CNN models while maintaining comparable accuracy. The TPU-IMAC architecture shows potential for various applications where energy efficiency and high performance are essential, such as edge computing and real-time processing in mobile devices. The unified training algorithm and the integration of IMAC and TPU architectures contribute to the potential impact of this research on the broader machine learning landscape. Mohammed E. Elbtity, Brendan Reidy, Md Hasibul Amin, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | Facial Expression Recognition at the Edge: CPU vs GPU vs VPU vs TPUabstractFacial Expression Recognition (FER) plays an important role in human-computer interactions and is used in a wide range of applications. Convolutional Neural Networks (CNN) have shown promise in their ability to classify human facial expressions, however, large CNNs are not well-suited to be implemented on resource-and energy-constrained IoT devices. In this work, we present a hierarchical framework for developing and optimizing hardware-aware CNNs tuned for deployment at the edge. We perform a comprehensive analysis across various edge AI accelerators including NVIDIA Jetson Nano, Intel Neural Compute Stick, and Coral TPU. Using the proposed strategy, we achieved a peak accuracy of 99.49% when testing on the CK+ facial expression recognition dataset. Additionally, we achieved a minimum inference latency of 0.39 milliseconds and a minimum power consumption of 0.52 Watts. Mohammadreza Mohammadi, Heath Smith, Lareb Khan, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | Caveline Detection at the Edge for Autonomous Underwater Cave Exploration and MappingabstractThis paper explores the problem of deploying machine learning (ML)-based object detection and segmentation models on edge platforms to enable realtime caveline detection for Autonomous Underwater Vehicles (AUVs) used for under-water cave exploration and mapping. We specifically investigate three ML models, i.e., U-Net, Vision Transformer (ViT), and YOLOv8, deployed on three edge platforms: Raspberry Pi-4, Intel Neural Compute Stick 2 (NCS2), and NVIDIA Jetson Nano. The experimental results unveil clear tradeoffs between model accuracy, processing speed, and energy consumption. The most accurate model has shown to be U-Net with an 85.53 F1-score and 85.38 Intersection Over Union (IoU) value. Meanwhile, the highest inference speed and lowest energy consumption are achieved by the YOLOv8 model deployed on Jetson Nano operating in the high-power and low-power modes, respectively. The comprehensive quantitative analyses and comparative results provided in the paper highlight important nuances that can guide the deployment of caveline detection systems on underwater robots for ensuring safe and reliable AUV navigation during underwater cave exploration and mapping missions. Mohammadreza Mohammadi, Sheng-En Huang, Titon Barua, Ioannis M. Rekleitis, Md Jahidul Islam, Ramtin Zand |
ICMLA | 6 |
| 2023 | Realtime Facial Expression Recognition: Neuromorphic Hardware vs. Edge AI AcceleratorsabstractThe paper focuses on real-time facial expression recognition (FER) systems as an important component in various real-world applications such as social robotics. We investigate two hardware options for the deployment of FER machine learning (ML) models at the edge: neuromorphic hardware versus edge AI accelerators. Our study includes exhaustive experiments providing comparative analyses between the Intel Loihi neuromorphic processor and four distinct edge platforms: Raspberry Pi-4, Intel Neural Compute Stick (NSC), Jetson Nano, and Coral TPU. The results obtained show that Loihi can achieve approximately two orders of magnitude reduction in power dissipation and one order of magnitude energy savings compared to Coral TPU which happens to be the least power-intensive and energy-consuming edge AI accelerator. These reductions in power and energy are achieved while the neuromorphic solution maintains a comparable level of accuracy with the edge accelerators, all within the real-time latency requirements. Heath Smith, James Seekings, Mohammadreza Mohammadi, Ramtin Zand |
ICMLA | 4 |
| 2023 | Work in Progress: Real-time Transformer Inference on Edge AI AcceleratorsabstractTransformer models have become a dominant architecture in the world of machine learning. From natural language processing to more recent computer vision applications, Transformers have shown remarkable results and established a new state-of-the-art in many domains. However, this increase in performance has come at the cost of ever-increasing model sizes requiring more resources to deploy. Machine learning (ML) models are used in many real-world systems, such as robotics, mobile devices, and internet of things (IoT) devices, that require fast inference with low energy consumption. For batterypowered devices, lower energy consumption directly translates into longer battery life. To address these issues, several edge AI accelerators have been developed. Among these, the Coral Edge TPU has shown promising results for image classification while maintaining very low energy consumption. Many of these devices, including the Coral TPU, were originally designed to accelerate convolutional neural networks, making deployment of Transformers challenging. Here, we propose a methodology to deploy Transformers on Edge TPU. We provide extensive latency, power, and energy comparisons among the leading edge devices and show that our methodology allows for real-time inference of Transformers while maintaining the lowest power and energy consumption of other edge devices on the market. Brendan Reidy, Mohammadreza Mohammadi, Mohammed E. Elbtity, Heath Smith, Ramtin Zand |
RTAS | 5 |
| 2022 | MRAM-based Analog Sigmoid Function for In-memory ComputingabstractWe propose an analog implementation of the transcendental activation function leveraging two spin-orbit torque magnetoresistive random-access memory (SOT-MRAM) devices and a CMOS inverter. The proposed analog neuron circuit consumes 1.8-27x less power, and occupies 2.5-4931x smaller area, compared to the state-of-the-art analog and digital implementations. Moreover, the developed neuron can be readily integrated with memristive crossbars without requiring any intermediate signal conversion units. The architecture-level analyses show that a fully-analog in-memory computing (IMC) circuit that use our SOT-MRAM neuron along with an SOT-MRAM based crossbar can achieve more than 1.1x, 12x, and 13.3x reduction in power, latency, and energy, respectively, compared to a mixed-signal implementation with analog memristive crossbars and digital neurons. Finally, through cross-layer analyses, we provide a guide on how varying the device-level parameters in our neuron can affect the accuracy of multilayer perceptron (MLP) for MNIST classification. Md Hasibul Amin, Mohammed E. Elbtity, Mohammadreza Mohammadi, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | Interconnect Parasitics and Partitioning in Fully-Analog In-Memory Computing ArchitecturesabstractFully-analog in-memory computing (IMC) architectures that implement both matrix-vector multiplication and nonlinear vector operations within the same memory array have shown promising performance benefits over conventional IMC systems due to the removal of energy-hungry signal conversion units. However, maintaining the computation in the analog domain for the entire deep neural network (DNN) comes with potential sensitivity to interconnect parasitics. Thus, in this paper, we investigate the effect of wire parasitic resistance and capacitance on the accuracy of DNN models deployed on fully-analog IMC architectures. Moreover, we propose a partitioning mechanism to alleviate the impact of the parasitic while keeping the computation in the analog domain through dividing large results for a $400 \times 120 \times 84 \times 10$ DNN model deployed on a results for a $400 \times 120 \times 84 \times 10$ DNN model deployed on a fully-analog IMC circuit show that a 94.84 % accuracy could be achieved for MNIST classification application with 16,8, and 8 horizontal partitions, as well as 8,8, and 1 vertical partitions for first, second, and third layers of the DNN, respectively, which is comparable to the $\sim 97$ % accuracy realized by digital implementation on CPU. It is shown that accuracy benefits are extra circuitry required for handling partitioning. Md Hasibul Amin, Mohammed E. Elbtity, Ramtin Zand |
ISCAS | 3 |
| 2022 | APTPU: Approximate Computing Based Tensor Processing UnitabstractWe propose an approximate tensor processing unit (APTPU), which includes two main components: (1) approximate processing elements (APEs) consisting of a low-precision multiplier and an approximate adder, and (2) pre-approximate units (PAUs) which are shared among the APEs in the APTPU’s systolic array, functioning as the steering logic to pre-process the operands and feed them to the APEs. We conduct extensive experiments to evaluate the performance of the APTPU across various configurations and various workloads. The results show that the APTPU’s systolic array achieves up to$5.2\times \textit {TOPS}/mm^{2}$and$4.4\times \textit {TOPS}/W$improvements compared to that of a conventional systolic array design. The comparison between the proposed APTPU and in-house TPU designs shows that we can achieve approximately$2.5\times $and$1.2\times $area and power reduction, respectively, while realizing comparable accuracy. Finally, a comparison with the state-of-the-art approximate systolic arrays shows that the APTPU can realize up to$1.58\times $,$2\times $, and$1.78\times $, reduction in delay, power, and area, respectively, while using similar design specifications and synthesis constraints. Mohammed E. Elbtity, Peyton Chandarana, Brendan Reidy, Jason Kamran Eshraghian, Ramtin Zand |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | TSV Extrusion Morphology Classification Using Deep Convolutional Neural NetworksabstractIn this paper, we utilize deep convolutional neural networks (CNNs) to classify the morphology of through-silicon via (TSV) extrusion in three dimensional (3D) integrated circuits (ICs). TSV extrusion is a crucial reliability concern which can deform and crack interconnect layers in 3D ICs and cause device failures. Herein, the white light interferometry (WLI) technique is used to obtain the surface profile of the extruded TSVs. We have developed a program that uses raw data obtained from WLI to create a TSV extrusion morphology dataset, including TSV images with 54 × 54 pixels that are labeled and categorized into three morphology classes. Four CNN architectures with different network complexities are implemented and trained for TSV extrusion morphology classification application. Data augmentation and dropout approaches are utilized to realize a balance between overfitting and underfitting in the CNN models. Results obtained show that the CNN model with optimized complexity, dropout, and data augmentation can achieve a classification accuracy comparable to that of a human expert. Brendan Reidy, Golareh Jalilvand, Tengfei Jiang, Ramtin Zand |
ICMLA | 4 |
| 2019 | HSC-FPGAabstractThe HSC-FPGA offers an intriguing feasible architecture for the next generation of configurable fabrics, which allows embracing the advantages of both CMOS and beyond-CMOS technologies without requiring significant modification to the routing structure, programming paradigms, and synthesis tool-chain of the commercial FPGAs. In the HSC-FPGA, the intrinsic characteristics of magnetic random access memory (MRAM)-look-up table (LUT) circuits are used to implement sequential logic, while combinational logic circuits are implemented by static random access memory (SRAM)-LUTs. Fabric-level simulation results for the developed HSC-FPGA show that it can achieve at least 18%, 70%, and 15% reduction in terms of area, standby power, and read power consumption, respectively, for various ISCAS-89 and ITC-99 benchmark circuits compared to conventional SRAM-based FPGAs. The power consumption values can be further decreased by the power-gating allowed by the non-volatility feature of MRAM-LUTs. Moreover, the benefits of increased heterogeneity for reconfigurable computing is extended along realizing probabilistic computing paradigms within a fabric, which is enabled by probabilistic spin logic devices. The cooperating strengths of technology-heterogeneity and heterogeneity in computing paradigm in the proposed HSC-FPGA are leveraged to develop energy-efficient and reliability-aware training and evaluation circuits for deep belief networks with memristive crossbar arrays and p-bit based probabilistic neurons. Ramtin Zand, Ronald F. DeMara |
FPGA | 1 |
| 2019 | Clockless Spin-based Look-Up Tables with Wide Read MarginabstractIn this paper, we develop a 6-input fracturable non-volatile Clockless LUT (C-LUT) using spin Hall effect (SHE)-based Magnetic Tunnel Junctions (MTJs) and provide a detailed comparison between the SHE-MTJ-based C-LUT and Spin Transfer Torque (STT)-MTJ-based C-LUT. The proposed C-LUT offers an attractive alternative for implementing combinational logic as well as sequential logic versus previous spin-based LUT designs in the literature. Foremost, C-LUT eliminates the sense amplifier typically employed by using a differential polarity dual MTJ design, as opposed to a static reference resistance MTJ. This realizes a much wider read margin and the Monte Carlo simulation of the proposed fracturable C-LUT indicates no read and write errors in the presence of a variety of process variations scenarios involving MOS transistors as well as MTJs. Additionally, simulation results indicate that the proposed C-LUT reduces the standby power dissipation by 5.4-fold compared to the SRAM-based LUT. Furthermore, the proposed SHE-MTJ-based C-LUT reduces the area by 1.3-fold and 2-fold compared to the SRAM-based LUT and the STT-MTJ-based C-LUT, respectively. Soheil Salehi, Ramtin Zand, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | AQuRate: MRAM-based Stochastic Oscillator for Adaptive Quantization Rate Sampling of Sparse SignalsabstractRecently, the promising aspects of compressive sensing have inspired new circuit-level approaches for their efficient realization within the literature. However, most of these recent advances involving novel sampling techniques have been proposed without considering hardware and signal constraints. Additionally, traditional hardware designs for generating non-uniform sampling clock incur large area overhead and power dissipation. Herein, we propose a novel non-uniform clock generator called Adaptive Quantization Rate (AQR) generator using Magnetic Random Access Memory (MRAM)-based stochastic oscillator devices. Our proposed AQR generator provides ~25-fold reduction in area, on average, while offering ~6-fold reduced power dissipation, on average, compared to the state-of-the-art non-uniform clock generators. Soheil Salehi, Ramtin Zand, Alireza Zaeemzadeh, Nazanin Rahnavard, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | Self-Organized Sub-bank SHE-MRAM-based LLC: An energy-efficient and variation-immune read and write architecture
Soheil Salehi, Navid Khoshavi, Ramtin Zand, Ronald F. DeMara |
Integr. | 3 |
| 2019 | Composable Probabilistic Inference Networks Using MRAM-based Stochastic NeuronsabstractMagnetoresistive random access memory (MRAM) technologies with thermally unstable nanomagnets are leveraged to develop an intrinsic stochastic neuron as a building block for restricted Boltzmann machines (RBMs) to form deep belief networks (DBNs). The embedded MRAM-based neuron is modeled using precise physics equations. The simulation results exhibit the desired sigmoidal relation between the input voltages and probability of the output state. A probabilistic inference network simulator (PIN-Sim) is developed to realize a circuit-level model of an RBM utilizing resistive crossbar arrays along with differential amplifiers to implement the positive and negative weight values. The PIN-Sim is composed of five main blocks to train a DBN, evaluate its accuracy, and measure its power consumption. The MNIST dataset is leveraged to investigate the energy and accuracy tradeoffs of seven distinct network topologies in SPICE using the 14nm HP-FinFET technology library with the nominal voltage of 0.8V, in which an MRAM-based neuron is used as the activation function. The software and hardware level simulations indicate that a 784× 200× 10 topology can achieve less than 5% error rates with ∼400pJ energy consumption. The error rates can be reduced to 2.5% by using a 784× 500× 500× 500× 10 DBN at the cost of ∼10× higher energy consumption and significant area overhead. Finally, the effects of specific hardware-level parameters on power dissipation and accuracy tradeoffs are identified via the developed PIN-Sim framework. Ramtin Zand, Kerem Yunus Çamsari, Supriyo Datta, Ronald F. DeMara |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2018 | Logic-Encrypted Synthesis for Energy-Harvesting-Powered Spintronic-Embedded Datapath DesignabstractThe objectives of advancing secure, intermittency-tolerant, and energy-aware logic datapaths are addressed herein by developing a spin-based design methodology and its corresponding synthesis steps. The approach selectively-inserts Non-Volatile (NV) Polymorphic Gates (PGs) to realize datapaths which are suitable for intrinsic operation in Energy-Harvesting-Powered (EHP) devices. Spin Hall Effect (SHE)-based Magnetic Tunnel (MTJs) are utilized to design NV-PGs, which are combined within a Flip-Flop (FF) circuit to develop a PG-FF realizing Boolean logic functions with inherent state-holding capability. The reconfigurability of PGs is leveraged for logic-encryption to enhance the security of the developed intermittency-resilient circuits, which are applied to ISCAS-89, MCNS, and ITC-99 benchmarks. The results obtained indicate that the PG-FF based design can achieve up to 7.1% and 13.6% improvements in terms of area and Power Delay Product (PDP), respectively, compared to NV-FF based methodologies that replace the CMOS-based FFs with NV-FFs. Further PDP improvements are achieved by using low-energy barrier SHE-MTJ devices within the PG-FF circuit. SHE-MTJs with 30kT energy exhibit 40.5% reduction in PDP at the cost of lower retention times in the range of minutes, which is still sufficient to achieve forward progress in EHP devices having more than hundreds of power-on and power-off cycles per minute. Arman Roohi, Ramtin Zand, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Low-Energy Deep Belief Networks Using Intrinsic Sigmoidal Spintronic-based Probabilistic NeuronsabstractA low-energy hardware implementation of deep belief network (DBN) architecture is developed using near-zero energy barrier probabilistic spin logic devices (p-bits), which are modeled to realize an intrinsic sigmoidal activation function. A CMOS/spin based weighted array structure is designed to implement a restricted Boltzmann machine (RBM). Device-level simulations based on precise physics relations are used to validate the sigmoidal relation between the output probability of a p-bit and its input currents. Characteristics of the resistive networks and p-bits are modeled in SPICE to perform a circuit-level simulation investigating the performance, area, and power consumption tradeoffs of the weighted array. In the application-level simulation, a DBN is implemented in MATLAB for digit recognition using the extracted device and circuit behavioral models. The MNIST data set is used to assess the accuracy of the DBN using 5,000 training images for five distinct network topologies. The results indicate that a baseline error rate of 36.8% for a 784x10 DBN trained by 100 samples can be reduced to only 3.7% using a 784x800x800x10 DBN trained by 5,000 input samples. Finally, Power dissipation and accuracy tradeoffs for probabilistic computing mechanisms using resistive devices are identified. Ramtin Zand, Kerem Yunus Çamsari, Steven D. Pyle, Ibrahim Ahmed 0002, Chris H. Kim, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | Survivability Modeling and Resource Planning for Self-Repairing Reconfigurable Device FabricsabstractA resilient system design problem is formulated as the quantification of uncommitted reconfigurable resources required for a system of components to survive its lifetime within mission availability specifications. We show that this survivability metric can be calculated according to the residual functionality obtained from pools of dynamically configurable elements constituting the amorphous resource pool (ARP). The ARP is depleted based on the failure rate to replenish the functionality lost in a reconfigurable fabric due to the occurrence of permanent faults during the mission lifetime. While genetic algorithms are selected for the reparation method, any probabilistic or deterministic active repair strategy is covered without loss of generality. Parameters of this model are correlated with reliability specifications of Xilinx Virtex-4 field programmable gate array devices, which are then utilized for MCNC benchmark circuits along with a realistic space mission. Calculation of the spare fabric resources which must be budgeted for a mission, maximum mission lifetime, and repair policy parameters are realized using the proposed probabilistic survivability model for soft computing-based repair strategies. Rashad S. Oreifej, Rawad N. Al-Haddad, Ramtin Zand, Rizwan A. Ashraf, Ronald F. DeMara |
IEEE Trans. Cybern. | 3 |
| 2017 | Voltage-Based Concatenatable Full Adder Using Spin Hall Effect SwitchingabstractMagnetic tunnel junction (MTJ)-based devices have been studied extensively as a promising candidate to implement hybrid energy-efficient computing circuits due to their nonvolatility, high integration density, and CMOS compatibility. In this paper, MTJs are leveraged to develop a novel full adder (FA) based on 3- and 5-input majority gates. Spin Hall effect (SHE) is utilized for changing the MTJ states resulting in low-energy switching behavior. SHE-MTJ devices are modeled in Verilog-A using precise physical equations. SPICE circuit simulator is used to validate the functionality of 1-bit SHE-based FA. The simulation results show 76% and 32% improvement over previous voltage-mode MTJ-based FA in terms of energy consumption and device count, respectively. The concatanatability of our proposed 1-bit SHE-FA is investigated through developing a 4-bit SHE-FA. Finally, delay and power consumption of an n-bit SHE-based adder has been formulated to provide a basis for developing an energy efficient SHE-based n-bit arithmetic logic unit. Arman Roohi, Ramtin Zand, Deliang Fan, Ronald F. DeMara |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Energy-Efficient and Process-Variation-Resilient Write Circuit Schemes for Spin Hall Effect MRAM DeviceabstractIn this paper, various energy-efficient write schemes are proposed for switching operation of spin hall effect (SHE)-based magnetic tunnel junctions (MTJs). A transmission gate (TG)-based write scheme is proposed, which provides a symmetric and energy-efficient switching behavior. We have modeled an SHE-MTJ using precise physics equations, and then leveraged the model in SPICE circuit simulator to verify the functionality of our designs. Simulation results show the TG-based write scheme advantages in terms of device count and switching energy. In particular, it can operate at 12% higher clock frequency while realizing at least 13% reduction in energy consumption compared to the most energy-efficient write circuits. We have analyzed the performance of the implemented write circuits in presence of process variation (PV) in the transistors' threshold voltage and SHE-MTJ dimensions. Results show that the proposed TG-based design is the second most PV-resilient write circuit scheme for SHE-MTJs among the implemented designs. Finally, we have proposed the 1TG-1T-1R SHE-based magnetic random access memory (MRAM) bit cell based on the TG-based write circuit. Comparisons with several of the most energy-efficient and variation-resilient SHE-MRAM cells indicate that 1TG-1T-1R delivers reduced energy consumption with 43.9% and 10.7% energy-delay product improvement, while incurring low area overhead. Ramtin Zand, Arman Roohi, Ronald F. DeMara |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Reactive rejuvenation of CMOS logic paths using self-activating voltage domainsabstractAlthough the trend of technology scaling is sought to realize higher performance computer systems, it also results in Integrated Circuits (ICs) suffering from increasing Process, Voltage, and Temperature (PVT) variations and adverse aging effects. In most cases, these reliability threats manifest themselves as timing errors on critical speed-paths of the circuit, if a large design guardband is not reserved. In this work, we propose the Reactive Rejuvenation (RR) architectural approach consisting of detection and recovery phases to mitigate circuit from BTI-induced aging. The BTI impact on the critical and near critical paths performance is continuously examined through a lightweight logic circuit which asserts an error signal in the case of any timing violation in those paths. By utilizing timing violation occurrence in the system, the timing-sensitive portion of the circuit is recovered from BTI through switching computations to redundant aging-critical voltage domain. The proposed technique achieves aging mitigation and reduced energy consumption as compared to a baseline circuit. Thus, significant voltage guardbands to meet the desired timing specification are avoided. Rizwan A. Ashraf, Ahmad Alzahrani 0001, Navid Khoshavi, Ramtin Zand, Soheil Salehi, Arman Roohi, Mingjie Lin, Ronald F. DeMara |
ISCAS | 4 |