VLDB 2026 Research / reviewers in the wild / expert
Arman Roohi
dblp:161/8102
· DBLP profile ↗
50ranked-venue papers
6as first author
40since 2021 · last 2026
0000-0002-0900-8768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 6 first-author · 36 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: ADC-FIST: ADC-Free In/Near-Sensor Stochastic Object Tracking
Mehran Shoushtari Moghadam, Sepehr Tabrizchi, Ali Shafiee Sarvestani, Sercan Aygün, Arman Roohi, M. Hassan Najafi |
DATE | 5 |
| 2026 | SENTRY: Spiking Event Reasoning for Selective Deep Inference in Event-Driven Edge VisionabstractEvent-driven cameras are well suited for always-on edge vision, but forwarding all events creates unnecessary overhead. Events outside user-defined ROIs are semantically irrelevant, while many ROI events arise from background motion, lighting changes, or sensor noise rather than meaningful activity. We propose SENTRY, a near-sensor pipeline that addresses both issues through hierarchical spatial reasoning. A coarse spike-rate gate discards frames with low ROI activity, while a spatially-aware confirmation stage compares each region’s activity against the global background rate to suppress diffuse noise. Evaluated on a real indoor surveillance sequence, SENTRY achieves +64% precision and − 57% false positive rate relative to a global spike-rate baseline, targeting resource-constrained, always-on platforms where energy efficiency and rapid response must be achieved simultaneously. Shayan Gerami, Sepehr Tabrizchi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | A2CD-Gaze: Decoder-Free Compressed-Domain Eye Parsing via In-Sensor Pre-ADC Analog Convolution
Sepehr Tabrizchi, Arman Roohi |
ISLPED | 3 |
| 2026 | GAN-enhanced machine learning and metabolic modeling identify reprogramming in pancreatic cancerabstractPancreatic ductal adenocarcinoma is one of the deadliest forms of cancer, presenting significant clinical challenges due to poor prognosis and limited treatment options. Understanding the metabolic reprogramming that drives this disease is crucial for identifying new therapeutic targets and improving patient outcomes. We developed a novel computational framework integrating genome-scale metabolic modeling with machine learning to identify metabolic signatures and therapeutic vulnerabilities in pancreatic cancer. To address the inherent class imbalance in cancer datasets, we generated synthetic healthy samples using a Wasserstein Generative Adversarial Network with Gradient Penalty, implementing a rigorous three-step biological filtration process to ensure their validity. This approach enabled the creation of a balanced dataset for robust comparison of healthy versus cancerous metabolic states. Our machine learning classifier achieved 94.83% accuracy in distinguishing between these states, demonstrating the effectiveness of our integrated approach. Systems-level analysis revealed three key dysregulated pathways: heparan sulfate degradation, O-glycan metabolism, and heme degradation. We identified impaired lysosomal degradation of heparan sulfate proteoglycans as a potential contributor to disease pathogenesis, providing a mechanistic explanation for the previously observed association between lysosomal storage disorders and pancreatic cancer. Additionally, nervonic acid transport emerged as the most discriminative reaction between healthy and cancerous states, with gene-level analysis highlighting fatty acid binding proteins, fatty acid transporters, and acyl-CoA synthetases as key molecular drivers of metabolic reprogramming. Our multi-level approach connected genetic drivers to functional metabolic consequences, revealing coordinated upregulation of fatty acid transport and activation processes. These findings enhance our understanding of pancreatic cancer metabolism and present potential therapeutic targets, demonstrating the value of integrated computational approaches in cancer research. Tahereh Razmpour, Masoud Tabibian, Arman Roohi, Rajib Saha |
PLoS Comput. Biol. | 3 |
| 2026 | BISen: A Robust Framework for Efficient CNN Inference on Battery-Free Intelligent Sensory NodesabstractWe present BISen, a framework for efficient and reliable convolutional neural network (CNN) inference on battery-free, energy-harvesting IoT sensor nodes. Battery-powered deployments suffer from limited lifetimes, high replacement costs, and environmental impacts, problems that will intensify as IoT scales to billions of devices. Energy-harvesting nodes remove batteries but face intermittent power, resulting in frequent failures that corrupt the intermediate CNN state, require costly checkpointing and rollback, and amplify non-volatile memory (NVM) traffic under tight on-chip memory constraints, leaving little harvested energy for useful sensing and inference. BISen introduces a reactive intermittent execution model for CNN workloads on off-the-shelf ultra-low-power microcontrollers. An energy-aware state machine with a safe-stop mechanism halts execution before brownout, while selective checkpointing preserves only the minimal CNN state needed for forward progress. This enables seamless resumption across power cycles while sharply reducing NVM reads/writes and memory-access overheads. Across two commercial MCU+radio platforms, three real harvested power traces, and nine CNNs, BISen cuts NVM operations by up to 86.4%, reduces standby/load/store operations by up to 94.1%, 94.5%, and 90.7%, and improves sensing throughput by about 1.3−1.4× compared to a state-of-the-art reactive baseline under the same energy budget, enabling long-lived, battery-free, carbon-aware IoT deployments. Sepehr Tabrizchi, Shayan Gerami, Justin Feng, Nader Sehatbakhsh, David Z. Pan, Arman Roohi |
IEEE Trans. Computers | 6 |
| 2025 | Late Breaking Results: Automated Topology Generation for Power Amplifier Designs through BiLSTM-based DNN and Multi-objective OptimizationsabstractThis work presents an automated, intelligent methodology for optimizing power amplifier (PA) design by predicting the most suitable circuit topology-specifically, the input and output matching networks-for a given high electron mobility transistor (HEMT). A classification-based bidirectional long short-term memory (BiLSTM) deep neural network (DNN) is trained to determine the optimal PA topology, while multi-objective Pareto front-based optimization techniques refine the network’s hyperparameters, including the number of hidden layers and neurons. The proposed approach is adaptable to various HEMT models and is validated through the design and optimization of high-performance PAs using lumped elements and transmission lines, operating within the $1-2 \mathrm{GHz}$ frequency range. The method is demonstrated using the Cree CGH40010 GaN HEMT on a Rogers RO4350B substrate, achieving a power output of approximately 40 dBm, a power-added efficiency (PAE) of at least 50%, and a power gain exceeding 10dB. Lida Kouhalvandi, Sercan Aygün, M. Hassan Najafi, Arman Roohi |
DAC | 4 |
| 2025 | ResISC: Residue Number System-Based Integrated Sensing and Computing for Efficient Edge AIabstractThis paper presents ResISC, an RNS-based integrated sensing and computing architecture enabling efficient edge AI. ResISC platform features (i) an in-sensor residue encoder converting images directly to RNS in the analog domain, (ii) an energy-efficient RNS-based processing-near-sensor CNN accelerator utilizing SOT-MRAM, and (iii) an innovative mixed-radix unit for efficient activation operations. By employing selective channel deactivation, ResISC reduces computation overhead by up to $89 \%$, while achieving a $3.4 \times$ improvement in power efficiency and up to a $71 \times$ reduction in execution time compared to processing-in-MRAM platforms. Experiments on various datasets demonstrate that ResISC achieves competitive accuracy levels (up to $94.63 \%$ on CIFAR-10) with minimal degradation, making it an ideal solution for power-constrained, real-time edge applications. Sepehr Tabrizchi, Samin Sohrabi, Mohamadreza Mohammadi, Ramtin Zand, Shaahin Angizi, Arman Roohi |
DAC | 6 |
| 2025 | iSEW: in-Sensor Embedded Watermarking for Secure ImagingabstractThis paper proposes an analog-domain watermarking approach implemented directly in CMOS sensors. By embedding the watermark at the readout stage before ADC, we achieve a robust, tamper-resistant mechanism with minimal impact on image quality. Modifications to the column amplifier architecture and a secure pattern generator enable effective watermark embedding. Experimental results on a 64×64 pixel array showcase high watermark detection rates (> 85%) under various attacks. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
FCCM | 3 |
| 2025 | LLM-IMC: Automating Analog In-Memory Computing Architecture Generation with Large Language ModelsabstractResistive crossbars enabling analog In-Memory Computing (IMC) have garnered significant attention from academia and industry as a promising architecture for Deep Neural Network (DNN) acceleration, thanks to their high memory access bandwidth and in-situ computing capabilities. However, the knowledge-intensive hardware design process and the lack of high-quality circuit netlists have constrained design space exploration and optimization of analog IMC to behavioral system-level tools. In this one-page abstract, we introduce LLM-IMC, a novel fine-tune-free Large Language Model (LLM) framework, supported by a Python-based tool, designed for analog IMC SPICE code generation. LLM-IMC systematically addresses these limitations by automating the creation of diverse IMC simulation scripts, enabling efficient design space exploration through LLM-driven performance, and outlining an integration roadmap for hardware-oriented neuromorphic crossbar design flows. Deepak Vungarala, Md Hasibul Amin, Pietro Mercati, Arman Roohi, Ramtin Zand, Shaahin Angizi |
FCCM | 4 |
| 2025 | Maximizing Sub-Array Resource Utilization in Digital Processing-in-Memory: A Versatile Hardware-Aware Approach
Gamana Aragonda, Deniz Najafi, Deepak Vungarala, Sepehr Tabrizchi, Arman Roohi, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | PixelPrune: Optimizing AIoT Vision Systems via In-Sensor Segmentation and Adaptive Data Transfer
Mohammadreza Mohammadi, Mehrdad Morsali, Sepehr Tabrizchi, Brendan Reidy, Arman Roohi, Shaahin Angizi, Ramtin Zand |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Event-Driven Spatiotemporal Processing-In-Sensor with Phase Change Memory-based Optical Acceleration
Mehrdad Morsali, Deniz Najafi, Amin Shafiee, Sepehr Tabrizchi, Pietro Mercati, Mohsen Imani, Arman Roohi, Navid Khoshavi, Mahdi Nikdast, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | SenGuard: A Novel Processing In-Sensor Method for Privacy-Enhanced Smart Imaging
Neeraj Solanki, Sepehr Tabrizchi, Ali Shafiee Sarvestani, Shaahin Angizi, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Magnetic In/Near-Sensor Architectures: From Raw Sensing to Smart Processing
Sepehr Tabrizchi, Ali Shafiee Sarvestani, Md Hasibul Amin, Deniz Najafi, Shaahin Angizi, Ramtin Zand, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | From Prompt to Accelerator: A Perspective on LLM-Based Analog In-Memory Accelerator Design Automation
Deepak Vungarala, Md Hasibul Amin, Arman Roohi, Arnob Ghosh, Ramtin Zand, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Always-On Sensing in Energy-Harvested Systems via Stochastic Intermittent ComputingabstractThis paper introduces Stochastic Intermittent Computing (STIC), a framework that integrates intermittent computing (ImC) and stochastic computing (SC) to enable always-on sensing in energy-harvested systems. STIC dynamically adjusts computational precision based on available energy, eliminating the need for non-volatile memory checkpointing traditionally used in ImC systems. By adapting precision in real-time, STIC ensures continuous operation even under severe power fluctuations, significantly improving energy efficiency and system resilience. Evaluation results demonstrate that STIC achieves substantial reductions in area, power, and energy consumption owing to the simplicity of SC and its tolerance to aggressive voltage scaling. Evaluations across multiple neural networks and charging traces confirm that STIC enables robust, low-power edge intelligence for resource-constrained environments. Sepehr Tabrizchi, Mehran Moghadam, Ali Shafiee Sarvestani, Sercan Aygün, M. Hassan Najafi, Arman Roohi |
ISLPED | 6 |
| 2025 | Poster Abstract: RL-SEP: RL -Based S mart E xit Point Selection for Enhancing Energy Harvested System LongevityabstractRL-SEP is a reinforcement learning scheduler that optimizes neural network execution in energy-harvesting devices. By dynamically selecting quantization levels and early exit points, it improves active operation time by up to 11% over the reactive method while achieving 136% better accuracy-to-energy ratio and maintaining higher energy reserves. Testing on ResNet-18 and DenseNet-121 shows robust performance across various harvesting sources. Ali Shafiee Sarvestani, Sepehr Tabrizchi, Nader Sehatbakhsh, Arman Roohi |
SenSys | 4 |
| 2024 | Lightator: An Optical Near-Sensor Accelerator with Compressive Acquisition Enabling Versatile Image ProcessingabstractThis paper proposes a high-performance and energy-efficient optical near-sensor accelerator for vision applications, called Lightator. Harnessing the promising efficiency offered by photonic devices, Lightator features innovative compressive acquisition of input frames and fine-grained convolution operations for low-power and versatile image processing at the edge for the first time. This will substantially diminish the energy consumption and latency of conversion, transmission, and processing within the established cloud-centric architecture as well as recently designed edge accelerators. Our device-to-architecture simulation results show that with favorable accuracy, Lightator achieves 84.4 Kilo FPS/W and reduces power consumption by a factor of ~24× and 73× on average compared with existing photonic accelerators and GPU baseline. Mehrdad Morsali, Brendan Reidy, Deniz Najafi, Sepehr Tabrizchi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Ramtin Zand, Shaahin Angizi |
DAC | 7 |
| 2024 | HiRISE: High-Resolution Image Scaling for Edge ML via In-Sensor Compression and Selective ROIabstractWith the rise of tiny IoT devices powered by machine learning (ML), many researchers have directed their focus toward compressing models to fit on tiny edge devices. Recent works have achieved remarkable success in compressing ML models for object detection and image classification on microcontrollers with small memory, e.g., 512kB SRAM. However, there remain many challenges prohibiting the deployment of ML systems that require high-resolution images. Due to fundamental limits in memory capacity for tiny IoT devices, it may be physically impossible to store large images without external hardware. To this end, we propose a high-resolution image scaling system for edge ML, called HiRISE, which is equipped with selective region-of-interest (ROI) capability leveraging analog in-sensor image scaling. Our methodology not only significantly reduces the peak memory requirements, but also achieves up to 17.7× reduction in data transfer and energy consumption. Brendan Reidy, Sepehr Tabrizchi, Mohammadreza Mohammadi, Shaahin Angizi, Arman Roohi, Ramtin Zand |
DAC | 5 |
| 2024 | OISA: Architecting an Optical In-Sensor Accelerator for Efficient Visual ComputingabstractTargeting vision applications at the edge, in this work, we systematically explore and propose a high-performance and energy-efficient Optical In-Sensor Accelerator architecture called OISA for the first time. Taking advantage of the promising efficiency of photonic devices, the OISA intrinsically implements a coarse-grained convolution operation on the input frames in an innovative minimum-conversion fashion in low-bit-width neural networks. Such a design remarkably reduces the power consumption of data conversion, transmission, and processing in the conventional cloud-centric architecture as well as recently-presented edge accelerators. Our device-to-architecture simulation results on various image data-sets demonstrate acceptable accuracy while OISA achieves 6.68 TOp/s/W efficiency. OISA reduces power consumption by a factor of 7.9 and 18.4 on average compared with existing electronic in-/near-sensor and ASIC accelerators. Mehrdad Morsali, Sepehr Tabrizchi, Deniz Najafi, Mohsen Imani, Mahdi Nikdast, Arman Roohi, Shaahin Angizi |
DATE | 6 |
| 2024 | DIAC: Design Exploration of Intermittent-Aware Computing Realizing Batteryless SystemsabstractBattery-powered IoT devices face challenges like cost, maintenance, and environmental sustainability, prompting the emergence of batteryless energy-harvesting systems that harness ambient sources. However, their intermittent behavior can disrupt program execution and cause data loss, leading to unpredictable outcomes. Despite exhaustive studies employing conventional checkpoint methods and intricate programming paradigms to address these pitfalls, this paper proposes an innovative systematic methodology, namely DIAC. The DIAC synthesis procedure enhances the performance and efficiency of intermittent computing systems, with a focus on maximizing forward progress and minimizing the energy overhead imposed by distinct memory arrays for backup. Then, a finite-state machine is delineated, encapsulating the core operations of an IoT node, sense, compute, transmit, and sleep states. First, we validate the robustness and functionalities of a DIAC-based design in the presence of power disruptions. DIAC is then applied to a wide range of benchmarks, including ISCAS-89, MCNS, and ITC-99. The simulation results substantiate the power-delay-product (PDP) benefits. For example, results for complex MCNC benchmarks indicate a PDP improvement of 61%, 56%, and 38% on average compared to three alternative techniques, evaluated at 45 nm. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
DATE | 3 |
| 2024 | DRAM-Locker: A General-Purpose DRAM Protection Mechanism Against Adversarial DNN Weight AttacksabstractIn this work, we propose DRAM-Locker as a robust general-purpose defense mechanism that can protect DRAM against various adversarial Deep Neural Network (DNN) weight attacks affecting data or page tables. DRAM-Locker harnesses the capabilities of in-DRAM swapping combined with a lock-table to prevent attackers from singling out specific DRAM rows to safeguard DNN's weight parameters. Our results indicate that DRAM-Locker can deliver a high level of protection downgrading the performance of targeted weight attacks to a random attack level. Furthermore, the proposed defense mechanism demonstrates no reduction in accuracy when applied to CIFAR-I0 and CIFAR-100. Importantly, DRAM-Locker does not necessitate any software retraining or result in extra hardware burden. Ranyang Zhou, Arman Roohi, Adnan Siraj Rakin, Shaahin Angizi |
DATE | 3 |
| 2024 | Hybrid Magneto-electric FET-CMOS Integrated Memory Design for Instant-on ComputingabstractThe surge in the number of normally-off power-constraint Internet of Things (IoT) devices in recent years has amplified the demand for high-performance and energy-efficient in-memory computing architectures built on top of various non-volatile memories. Magneto-Electric Field Effect Transistors (MEFETs) have presented compelling design features suitable for logic and memory integration as an emerging post-CMOS FET. These include high-speed switching, minimal power usage, and non-volatility. This work introduces a new in-memory computing architecture designed for edge applications, leveraging emerging MEFETs. The proposed architecture enables the execution of both Boolean logic operations and Binary Content Addressable Memory (BCAM) operations within a single cycle. Furthermore, the energy consumption during the write operation of the proposed cell is optimized by introducing a new write circuitry. The outcomes of our device-to-architecture evaluation reveal approximately 43.5% and 96.9% reduction in read and write energy consumption, respectively, compared to the counterpart non-volatile memories. At the application level, the proposed architecture is applied to implement Binary Neural Networks (BNNs) based on AlexNet and VGG16. Our results showcase a decrease of approximately 54% in the overall energy consumption when implementing these networks using the proposed design compared to non-volatile in-memory computing designs. Deniz Najafi, Sepehr Tabrizchi, Ranyang Zhou, Mohammadreza Amel Solouki, Andrew Marshall, Arman Roohi, Shaahin Angizi |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | RACSen: Residue Arithmetic and Chaotic Processing in Sensors to Enhance CMOS Imager SecurityabstractThe widespread adoption of vision sensors raises significant security and privacy concerns. In this paper, we present RACSen as a novel architecture that can increase the security and efficiency of conventional image sensors. RACSen leverages the intricate mathematical properties of the residue number system (RNS) with analog scrambling techniques to create a sophisticated dual-layered encryption mechanism. Incorporating RNS within analog-to-digital converters further strengthens security by mitigating replay attacks and preserving data transmission integrity and confidentiality. Our results demonstrate exceptional encryption, with a perfect pixel change rate of 99.90 and high intensity change of 45.77. This offers robust image data protection with minimal overhead of 11.11%. Sepehr Tabrizchi, Nedasadat Taheri, Shaahin Angizi, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2024 | ChaoSen: Security Enhancement of Image Sensor through in-Sensor Chaotic ComputingabstractWireless Sensor Networks (WSN) are integral to diverse applications, ranging from environmental monitoring to urban smart infrastructure. In the realm of WSNs, security remains a critical challenge owing to the complex nature of the sensor environment. As a result, WSN security has become a research focus in recent years. In this paper, we introduce ChaoSen, a novel image sensor system incorporating analog chaotic circuits within the sensor, thereby enhancing the overall system security. The system utilizes a scrambler module, which intricately intertwines with the chaotic encryption process, to reduce the predictability of pixel values and enhance the security of the system. Comparative evaluations demonstrate that the system achieves an NPCR value of 99.5562% and a UACI of 35.81900%, indicating high sensitivity to input changes and significant alteration in pixel intensity. Our approach also demonstrates its resilience against common cyber attacks, balancing enhanced security with resource efficiency. Nedasadat Taheri, Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ICCD | 4 |
| 2024 | PiPSim: A Behavior-Level Modeling Tool for CNN Processing-in-Pixel AcceleratorsabstractConvolutional neural networks (CNNs) have been gaining popularity in recent years, and researchers have designed specialized architectures to speed up the inference process. However, despite the promising potential of processing near-/in- sensor architectures actively explored in the visual Internet of Things, there is still a need to develop a behavior-level simulator to model performance and facilitate early design exploration. This article proposes a stand-alone simulation platform for processing-in-pixel (PiP) systems, namely, PiPSim. It offers a flexible interface and a wide range of design options for customizing the efficiency and accuracy of PiP-based accelerators using a hierarchical structure. Its organization spans from the device level, e.g., memory technology, upward to the circuit level, e.g., compute-add on architecture, and then to the algorithm level, e.g., DNN workloads. PiPSim realizes instruction-accurate evaluation of circuit-level performance metrics as well as learning accuracy at run-time. Compared to SPICE simulation, PiPSim achieves over 25$000\times $speed-up with less than a 2.5% error rate on average. Furthermore, PiPSim can optimize the design and estimate the tradeoff relationships among different performance metrics. Arman Roohi, Sepehr Tabrizchi, Mehrdad Morsali, David Z. Pan, Shaahin Angizi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | P-PIM: A Parallel Processing-in-DRAM Framework Enabling Row Hammer ProtectionabstractIn this work, we propose a Parallel Processing-In-DRAM architecture named P-PIM leveraging the high density of DRAM to enable fast and flexible computation. P-PIM enables bulk bit-wise in-DRAM logic between operands in the same bit-line by elevating the analog operation of the memory sub-array based on a novel dual-row activation mechanism. With this, P-PIM can opportunistically perform a complete and inexpensive in-DRAM RowHammer (RH) self-tracking and mitigation technique to protect the memory unit against such a challenging security vulnerability. Our results show that P-PIM achieves ~72% higher energy efficiency than the fastest charge-sharing-based designs. As for the RH protection, with a worst-case slowdown of ~0.8%, P-PIM archives up to 71% energy-saving over the SRAM/CAM-based frameworks and about 90% saving over DRAM-based frameworks. Ranyang Zhou, Sepehr Tabrizchi, Mehrdad Morsali, Arman Roohi, Shaahin Angizi |
DATE | 4 |
| 2023 | SenTer: A Reconfigurable Processing-in-Sensor Architecture Enabling Efficient Ternary MLPabstractRecently, Intelligent IoT (IIoT), including various sensors, has gained significant attention due to its capability of sensing, deciding, and acting by leveraging artificial neural networks (ANN). Nevertheless, to achieve acceptable accuracy and high performance in visual systems, a power-delay-efficient architecture is required. In this paper, we propose an ultra-low-power processing in-sensor architecture, namely SenTer, realizing low-precision ternary multi-layer perceptron networks, which can operate in detection and classification modes. Moreover, SenTer supports two activation functions based on user needs and the desired accuracy-energy trade-off. SenTer is capable of performing all the required computations for the MLP's first layer in the analog domain and then submitting its results to a co-processor. Therefore, SenTer significantly reduces the overhead of analog buffers, data conversion, and transmission power consumption by using only one ADC. Additionally, our simulation results demonstrate acceptable accuracy on various datasets compared to the full precision models. Sepehr Tabrizchi, Rebati Raman Gaire, Shaahin Angizi, Arman Roohi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | EnCoDe: Enhancing Compressed Deep Learning Models Through Feature - - - Distillation and Informative Sample SelectionabstractThis paper presents Encode, a novel technique that merges active learning, model compression, and knowledge distillation to optimize deep learning models. The method tackles issues such as generalization loss, resource intensity, and data redundancy that usually impede compressed models' performance. It actively integrates valuable samples for labeling, thus enhancing the student model's performance while economizing on labeled data and computational resources. Encode's utility is empirically validated using SVHN and CIFAR-10 datasets, demonstrating improved model compactness, enhanced generalization, reduced computational complexity, and lessened labeling efforts. In our evaluations, applied to compressed versions of VGGll and AlexNet models, Encode consistently outperforms baselines even when trained with 60% of the total training samples. Thus, it establishes an effective framework for enhancing the accuracy and generalization capabilities of compressed models, which is especially beneficial in situations with limited resources and scarce labeled data. Rebati Raman Gaire, Sepehr Tabrizchi, Arman Roohi |
ICMLA | 3 |
| 2023 | NeSe: Near-Sensor Event-Driven Scheme for Low Power Energy Harvesting SensorsabstractDigital technologies have made it possible to deploy visual sensor nodes capable of detecting motion events in the coverage area cost-effectively. However, background subtraction, as a widely used approach, remains an intractable task due to its inability to achieve competitive accuracy and reduced computation cost simultaneously. In this paper, an effective background subtraction approach, namely NeSe, for tiny energy-harvested sensors is proposed leveraging non-volatile memory (NVM). Using the developed software/hardware method, the accuracy and efficiency of event detection can be adjusted at runtime by changing the precision depending on the application's needs. Due to the near-sensor implementation of background subtraction and NVM usage, the proposed design reduces the data movement overhead while ensuring intermittent resiliency. The background is stored for a specific time interval within NVMs and compared with the next frame. If the power is cut, the background remains unchanged and is updated after the interval passes. Once the moving object is detected, the device switches to the high-powered sensor mode to capture the image. Sepehr Tabrizchi, Mehrdad Morsali, Shaahin Angizi, Arman Roohi |
ISCAS | 4 |
| 2023 | Ocellus: Highly Parallel Convolution-in-Pixel Scheme Realizing Power-Delay-Efficient Edge IntelligenceabstractWith the advent of Edge Intelligence (EI) devices, always-on intelligent and self-powered visual perception systems are receiving considerable attention. These emerging systems require continuous sensing and instant processing; however, the high energy data conversion/transmission of raw data and the limited available energy and computation resources make designing energy-efficient and low bandwidth CMOS vision sensors vital but challenging. This paper proposes a low-power integrated sensing and computing engine, namely Ocellus, which considerably decreases power costs of data movement/conversion and enables data/compute -intensive neural network tasks. Ocellus offers several unique features, including a highly parallel analog convolution-in-pixel scheme and reconfigurable filtering modes with filter pruning capability. These features realize low-precision ternary weight neural networks to mitigate the overhead of analog-to-digital converters and analog buffers. Moreover, the proposed structure supports a zero-skipping scheme to further reduce power consumption. Our circuit-to-application cosimulation results demonstrate comparable, even better, accuracy to the full-precision baseline on object classification tasks, while it achieves a frame rate of 1000 and efficiency of ~1.45 TOp/s/W. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ISLPED | 3 |
| 2022 | Work-in-Progress: A Processing-in-Pixel Accelerator based on Multi-level HfOx ReRAMabstractThis work paves the way to realize a processing-in-pixel accelerator based on a multi-level HfOxReRAM as a flexible, energy-efficient, and high-performance solution for real-time and smart image processing at edge devices. The proposed design intrinsically implements and supports a coarse-grained convolution operation in low-bit-width neural networks leveraging a novel compute-pixel with non-volatile weight storage at the sensor side. Our evaluations show that such a design can remarkably reduce the power consumption of data conversion and transmission to an off-chip processor maintaining accuracy compared with the recent in-sensor computing designs. Minhaz Abedin, Arman Roohi, Nathaniel C. Cady, Shaahin Angizi |
CASES | 2 |
| 2022 | ReD-LUT: Reconfigurable In-DRAM LUTs Enabling Massive Parallel ComputationabstractIn this paper, we propose a reconfigurable processing-in-DRAM architecture named ReD-LUT leveraging the high density of commodity main memory to enable a flexible, general-purpose, and massively parallel computation. ReD-LUT supports lookup table (LUT) queries to efficiently execute complex arithmetic operations (e.g., multiplication, division, etc.) via only memory read operation. In addition, ReD-LUT enables bulk bit-wise in-memory logic by elevating the analog operation of the DRAM sub-array to implement Boolean functions between operands stored in the same bit-line beyond the scope of prior DRAM-based proposals. We explore the efficacy of ReD-LUT in two computationally-intensive applications, i.e., low-precision deep learning acceleration, and the Advanced Encryption Standard (AES) computation. Our circuit-to-architecture simulation results show that for a quantized deep learning workload, ReD-LUT reduces the energy consumption per image by a factor of 21.4× compared with the GPU and achieves ~37.8× speedup and 2.1× energy-efficiency over the best in-DRAM bit-wise accelerators. As for AES data-encryption, it reduces energy consumption by a factor of ~2.2× compared to an ASIC implementation. Ranyang Zhou, Arman Roohi, Durga Misra, Shaahin Angizi |
ICCAD | 2 |
| 2022 | TizBin: A Low-Power Image Sensor with Event and Object Detection Using Efficient Processing-in-Pixel SchemesabstractIn the Artificial Intelligence of Things (AIoT) era, always-on intelligent and self-powered visual perception systems have gained considerable attention and are widely used. Thus, this paper proposes TizBin, a low-power processing in-sensor scheme with event and object detection capabilities to eliminate power costs of data conversion and transmission and enable data-intensive neural network tasks. Once the moving object is detected, TizBin architecture switches to the high-power object detection mode to capture the image. TizBin offers several unique features, such as analog convolutions enabling low-precision ternary weight neural networks (TWNN) to mitigate the overhead of analog buffer and analog-to-digital converters. Moreover, TizBin exploits non-volatile magnetic RAMs to store NN’s weights, remarkably reducing static power consumption. Our circuit-to-application co-simulation results for TWNNs demonstrate minor accuracy degradation on various image datasets, while TizBin achieves a frame rate of 1000 and efficiency of ∼1.83 TOp/s/W. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ICCD | 3 |
| 2022 | semiMul: Floating-Point Free Implementations for Efficient and Accurate Neural Network TrainingabstractMultiply–accumulate operation (MAC) is a fundamental component of machine learning tasks, where multiplication (either integer or float multiplication) compared to addition is costly in terms of hardware implementation or power consumption. In this paper, we approximate floating-point multiplication by converting it to integer addition while preserving the test accuracy of shallow and deep neural networks. We mathematically show and prove that our proposed method can be utilized with any floating-point format (e.g., FP8, FP16, FP32, etc.). It is also highly compatible with conventional hardware architectures and can be employed in CPU, GPU, or ASIC accelerators for neural network tasks with minimum hardware cost. Moreover, the proposed method can be utilized in embedded processors without a floating-point unit to perform neural network tasks. We evaluated our method on various datasets such as MNIST, FashionMNIST, SVHN, Cifar-10, and Cifar-100, with both FP16 and FP32 arithmetics. The proposed method preserves the test accuracy and, in some cases, overcomes the overfitting problem and improves the test accuracy. Ali Nezhadi, Shaahin Angizi, Arman Roohi |
ICMLA | 3 |
| 2022 | SCiMA: A Generic Single-Cycle Compute-in-Memory Acceleration Scheme for Matrix ComputationsabstractThis work proposes a new generic Single-cycle Compute-in-Memory (CiM) Accelerator for matrix computation named SCiMA. SCiMA is developed on top of the existing commodity Spin-Orbit Torque Magnetic Random-Access Memory chip. Every sub-array’s peripherals are transformed to realize a full set of single-cycle 2- and 3-input in-memory bulk bitwise functions specifically designed to accelerate a wide variety of graph and matrix multiplication tasks. We explore SCiMA’s efficiency by selecting a complex matrix processing operation, i.e., calculating determinant as an essential and under-explored application in the CiM domain. The cross-layer device-to-architecture simulation framework shows the presented platform can reduce energy consumption by 70.43% compared with the most recent CiM designs implemented with the same memory technology. SCiMA also achieves up to 2.5x speedup compared with current CiM platforms. Sepehr Tabrizchi, Shaahin Angizi, Arman Roohi |
ISCAS | 3 |
| 2022 | FlexiDRAM: A Flexible in-DRAM Framework to Enable Parallel General-Purpose ComputationabstractIn this paper, we propose a Flexible processing-in-DRAM framework named FlexiDRAM that supports the efficient implementation of complex bulk bitwise operations. This framework is developed on top of a new reconfigurable in-DRAM accelerator that leverages the analog operation of DRAM sub-arrays and elevates it to implement XOR2-MAJ3 operations between operands stored in the same bit-line. FlexiDRAM first generates an efficient XOR-MAJ representation of the desired logic and then appropriately allocates DRAM rows to the operands to execute any in-DRAM computation. We develop ISA and software support required to compute in-DRAM operation. FlexiDRAM transforms current memory architecture to a massively parallel computational unit and can be leveraged to significantly reduce the latency and energy consumption of complex workloads. Our extensive circuit-to-architecture simulation results show that averaged across two well-known deep learning workloads, FlexiDRAM achieves ∼ 15 × energy-saving and 13 × speedup over the GPU outperforming recent processing-in-DRAM platforms. Ranyang Zhou, Arman Roohi, Durga Misra, Shaahin Angizi |
ISLPED | 2 |
| 2021 | Entropy-Based Modeling for Estimating Adversarial Bit-flip Attack Impact on Binarized Neural NetworkabstractOver past years, the high demand to efficiently process deep learning (DL) models has driven the market of the chip design companies. However, the new Deep Chip architectures, a common term to refer to DL hardware accelerator, have slightly paid attention to the security requirements in quantized neural networks (QNNs), while the black/white -box adversarial attacks can jeopardize the integrity of the inference accelerator. Therefore in this paper, a comprehensive study of the resiliency of QNN topologies to black-box attacks is examined. Herein, different attack scenarios are performed on an FPGA-processor co-design, and the collected results are extensively analyzed to give an estimation of the impact's degree of different types of attacks on the QNN topology. To be specific, we evaluated the sensitivity of the QNN accelerator to a range number of bit-flip attacks (BFAs) that might occur in the operational lifetime of the device. The BFAs are injected at uniformly distributed times either across the entire QNN or per individual layer during the image classification. The acquired results are utilized to build the entropy-based model that can be leveraged to construct resilient QNN architectures to bit-flip attacks. Navid Khoshavi, Saman Sargolzaei, Yu Bi, Arman Roohi |
ASP-DAC | 4 |
| 2021 | Processing-in-Memory Acceleration of MAC-based Applications Using Residue Number System: A Comparative StudyabstractProcessing-in-memory (PIM) has raised as a viable solution for the memory wall crisis and has attracted great interest in accelerating computationally intensive AI applications ranging from filtering to complex neural networks. In this paper, we try to take advantage of both PIM and the residue number system (RNS) as an alternative for the conventional binary number representation to accelerate multiplication-and-accumulations (MACs), primary operations of target applications. The PIM architecture utilizes the maximum internal bandwidth of memory chips to realize a local and parallel computation to eliminates the off-chip data transfer. Moreover, RNS limits inter-digit carry propagation by performing arithmetic operations on small residues independently and in parallel. Thus, we develop a PIM-RNS, entitled PRIMS, and analyze the potential of intertwining PIM architecture with the inherent parallelism of the RNS arithmetic to delineate the opportunities and challenges. To this end, we build a comprehensive device-to-architecture evaluation framework to quantitatively study this problem considering the impact of PIM technology for a well-known three-moduli set as a case study. Shaahin Angizi, Arman Roohi, MohammadReza Taheri, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | RNSiM: Efficient Deep Neural Network Accelerator Using Residue Number SystemsabstractIn this paper, we propose an efficient convolutional neural network (CNN) accelerator design, entitled RNSiM, based on the Residue Number System (RNS) as an alternative for the conventional binary number representation. Instead of traditional arithmetic implementation that suffers from the inevitable lengthy carry propagation chain, the novelty of RNSiM lies in that all the data, including stored weights and communication/computation, are performed in the RNS domain. Due to the inherent parallelism of the RNS arithmetic, power and latency are significantly reduced. Moreover, an enhanced integrated intermodulo operation core is developed to decrease the overhead imposed by non-modular operations. Further improvement in systems' performance efficiency is achieved by developing efficient Processing-in-Memory (PIM) designs using various volatile CMOS and non-volatile Post-CMOS technologies to accelerate RNS-based multiplication-and-accumulations (MACs). The RN-SiM accelerator's performance on different datasets, including MNIST, SVHN, and CIFAR-10, is evaluated. With almost the same accuracy to the baseline CNN, the RNSiM accelerator can significantly increase both energy-efficiency and speedup compared with the state-of-the-art FPGA, GPU, and PIM designs. RNSiM and other RNS-PIMs, based on our method, reduce the energy consumption by orders of$28-77\times$and$331-897\times$compared with the FPGA and the GPU platforms, respectively. Arman Roohi, MohammadReza Taheri, Shaahin Angizi, Deliang Fan |
ICCAD | 1 |
| 2020 | SHIELDeNN: Online Accelerated Framework for Fault-Tolerant Deep Neural Network ArchitecturesabstractWe propose SHIELDeNN, an end-to-end inference accelerator frame-work that synergizes the mitigation approach and computational resources to realize a low-overhead error-resilient Neural Network (NN) overlay. We develop a rigorous fault assessment paradigm to delineate a ground-truth fault-skeleton map for revealing the most vulnerable parameters in NN. The error-susceptible parameters and resource constraints are given to a function to find superior design. The error-resiliency magnitude offered by SHIELDeNN can be adjusted based on the given boundaries. SHIELDeNN methodology improves the error-resiliency magnitude of cnvW1A1 by 17.19% and 96.15% for 100 MBUs that target weight and activation layers, respectively. Navid Khoshavi, Arman Roohi, Connor Broyles, Saman Sargolzaei, Yu Bi, David Z. Pan |
DAC | 2 |
| 2020 | Fiji-FIN: A Fault Injection Framework on Quantized Neural Network Inference AcceleratorabstractIn recent years, the big data booming has boosted the development of highly accurate prediction models driven from machine learning (ML) and deep learning (DL) algorithms. These models can be orchestrated on the customized hardware in the safety-critical missions to accelerate the inference process in ML/DL -powered IoT. However, the radiation-induced transient faults and black/white -box attacks can potentially impact the individual parameters in ML/DL models which may result in generating noisy data/labels or compromising the pre-trained model. In this paper, we propose Fiji-FIN1, a suitable framework for evaluating the resiliency of IoT devices during the ML/DL model execution with respect to the major security challenges such as bit perturbation attacks and soft errors. Fiji-FIN is capable of injecting both single bit/event flip/upset and multi-bit flip/upset faults on the architectural ML/DL accelerator embedded in ML/DL -powered IoT. Fiji-FIN is significantly more accurate compared to the existing software-level fault injections paradigms on ML/DL -driven IoT devices. Navid Khoshavi, Connor Broyles, Yu Bi, Arman Roohi |
ICMLA | 4 |
| 2020 | ApGAN: Approximate GAN for Robust Low Energy Learning From Imprecise ComponentsabstractA Generative Adversarial Network (GAN) is an adversarial learning approach which empowers conventional deep learning methods by alleviating the demands of massive labeled datasets. However, GAN training can be computationally-intensive limiting its feasibility in resource-limited edge devices. In this paper, we propose an approximate GAN (ApGAN) for accelerating GANs from both algorithm and hardware implementation perspectives. First, inspired by the binary pattern feature extraction method along with binarized representation entropy, the existing Deep Convolutional GAN (DCGAN) algorithm is modified by binarizing the weights for a specific portion of layers within both the generator and discriminator models. Further reduction in storage and computation resources is achieved by leveraging a novel hardware-configurable in-memory addition scheme, which can operate in the accurate and approximate modes. Finally, a memristor-based processing-in-memory accelerator for ApGAN is developed. The performance of the ApGAN accelerator on different data-sets such as Fashion-MNIST, CIFAR-10, STL-10, and celeb-A is evaluated and compared with recent GAN accelerator designs. With almost the same Inception Score (IS) to the baseline GAN, the ApGAN accelerator can increase the energy-efficiency by ~28.6× achieving 35-fold speedup compared with a baseline GPU platform. Additionally, it shows 2.5× and 5.8× higher energy-efficiency and speedup over CMOS-ASIC accelerator subject to an 11 percent reduction in IS. Arman Roohi, Shadi Sheikhfaal, Shaahin Angizi, Deliang Fan, Ronald F. DeMara |
IEEE Trans. Computers | 1 |
| 2018 | Logic-Encrypted Synthesis for Energy-Harvesting-Powered Spintronic-Embedded Datapath DesignabstractThe objectives of advancing secure, intermittency-tolerant, and energy-aware logic datapaths are addressed herein by developing a spin-based design methodology and its corresponding synthesis steps. The approach selectively-inserts Non-Volatile (NV) Polymorphic Gates (PGs) to realize datapaths which are suitable for intrinsic operation in Energy-Harvesting-Powered (EHP) devices. Spin Hall Effect (SHE)-based Magnetic Tunnel (MTJs) are utilized to design NV-PGs, which are combined within a Flip-Flop (FF) circuit to develop a PG-FF realizing Boolean logic functions with inherent state-holding capability. The reconfigurability of PGs is leveraged for logic-encryption to enhance the security of the developed intermittency-resilient circuits, which are applied to ISCAS-89, MCNS, and ITC-99 benchmarks. The results obtained indicate that the PG-FF based design can achieve up to 7.1% and 13.6% improvements in terms of area and Power Delay Product (PDP), respectively, compared to NV-FF based methodologies that replace the CMOS-based FFs with NV-FFs. Further PDP improvements are achieved by using low-energy barrier SHE-MTJ devices within the PG-FF circuit. SHE-MTJs with 30kT energy exhibit 40.5% reduction in PDP at the cost of lower retention times in the range of minutes, which is still sufficient to achieve forward progress in EHP devices having more than hundreds of power-on and power-off cycles per minute. Arman Roohi, Ramtin Zand, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | NV-Clustering: Normally-Off Computing Using Non-Volatile DatapathsabstractWith technology downscaling, static power dissipation presents a crucial challenge to multicore, many-core, and System-on-Chip (SoC) architectures due to the increased role of leakage currents in overall energy consumption and the need to support power-gating schemes. Herein, a non-Volatile (NV) flip-flop design approach, referred to as NV Clustering, is developed to realize middleware-transparent intermittent computing. First, a Logic-Embedded Flip-Flop (LE-FF) is developed to realize rudimentary Boolean logic functions along with an inherent state-holding capability within a compact footprint. Second, the NV-Clustering synthesis procedure and corresponding tool module are utilized to instantiate the LE-FF library cells within conventional Register Transfer Language (RTL) specifications. This selectively clusters together logic and NV state-holding functionality, based on energy and area minimization criteria. NV-Clustering is applied to a wide range of benchmarks including ISCAS-89, MCNS, and ITC-99 computational circuits using a LE-FF based on the Spin Hall Effect (SHE)-assisted Spin Transfer Torque (STT) Magnetic Tunnel Junction (MTJ). Simulation results validate functionality and power dissipation, area, and delay benefits. For instance, results for ISCAS-89 benchmarks indicate 15 percent area reduction on average, up to 22 percent reduction in energy consumption, and up to 14 percent reduction in delay as compared to alternative NV-FF based designs, as evaluated via SPICE simulation at the 45-nm technology node. Arman Roohi, Ronald F. DeMara |
IEEE Trans. Computers | 1 |
| 2017 | Voltage-Based Concatenatable Full Adder Using Spin Hall Effect SwitchingabstractMagnetic tunnel junction (MTJ)-based devices have been studied extensively as a promising candidate to implement hybrid energy-efficient computing circuits due to their nonvolatility, high integration density, and CMOS compatibility. In this paper, MTJs are leveraged to develop a novel full adder (FA) based on 3- and 5-input majority gates. Spin Hall effect (SHE) is utilized for changing the MTJ states resulting in low-energy switching behavior. SHE-MTJ devices are modeled in Verilog-A using precise physical equations. SPICE circuit simulator is used to validate the functionality of 1-bit SHE-based FA. The simulation results show 76% and 32% improvement over previous voltage-mode MTJ-based FA in terms of energy consumption and device count, respectively. The concatanatability of our proposed 1-bit SHE-FA is investigated through developing a 4-bit SHE-FA. Finally, delay and power consumption of an n-bit SHE-based adder has been formulated to provide a basis for developing an energy efficient SHE-based n-bit arithmetic logic unit. Arman Roohi, Ramtin Zand, Deliang Fan, Ronald F. DeMara |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Energy-Efficient and Process-Variation-Resilient Write Circuit Schemes for Spin Hall Effect MRAM DeviceabstractIn this paper, various energy-efficient write schemes are proposed for switching operation of spin hall effect (SHE)-based magnetic tunnel junctions (MTJs). A transmission gate (TG)-based write scheme is proposed, which provides a symmetric and energy-efficient switching behavior. We have modeled an SHE-MTJ using precise physics equations, and then leveraged the model in SPICE circuit simulator to verify the functionality of our designs. Simulation results show the TG-based write scheme advantages in terms of device count and switching energy. In particular, it can operate at 12% higher clock frequency while realizing at least 13% reduction in energy consumption compared to the most energy-efficient write circuits. We have analyzed the performance of the implemented write circuits in presence of process variation (PV) in the transistors' threshold voltage and SHE-MTJ dimensions. Results show that the proposed TG-based design is the second most PV-resilient write circuit scheme for SHE-MTJs among the implemented designs. Finally, we have proposed the 1TG-1T-1R SHE-based magnetic random access memory (MRAM) bit cell based on the TG-based write circuit. Comparisons with several of the most energy-efficient and variation-resilient SHE-MRAM cells indicate that 1TG-1T-1R delivers reduced energy consumption with 43.9% and 10.7% energy-delay product improvement, while incurring low area overhead. Ramtin Zand, Arman Roohi, Ronald F. DeMara |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Loss-Aware Switch Design and Non-Blocking Detection Algorithm for Intra-Chip Scale Photonic Interconnection NetworksabstractAs the number of on-chip processor cores increases, power-efficient solutions are sought for data communication between cores. TheHelix-hnon-blocking photonic switch is developed to improve physical-layer and network performance parameters for a wide range of silicon nano-photonic multicore interconnection topologies. Traffic benchmarks and practical case studies using a cycle-accurate simulation environment indicate significantly reduced insertion loss providing improved bandwidth density and scalability to manycore plurality. Improvements in system performance parameters are quantified for network bandwidth, transmission efficiency, and latency in popular photonic internconnection topologies, in comparison to previous switch designs. For instance, utilizing the Helix-h switch in a mesh topology, the bandwidth is increased by 112 percent compared to the previously highest performing switch design. Execution time and energy efficiency are improved by up to 92 and 99 percent, respectively, for representative multicore applications. Finally, the technique is generalized to a novel graph-theoretic method for articulating blocking conditions in photonic switches. Hesam Shabani, Arman Roohi, Akram Reza, Midia Reshadi, Nader Bagherzadeh, Ronald F. DeMara |
IEEE Trans. Computers | 2 |
| 2015 | Reactive rejuvenation of CMOS logic paths using self-activating voltage domainsabstractAlthough the trend of technology scaling is sought to realize higher performance computer systems, it also results in Integrated Circuits (ICs) suffering from increasing Process, Voltage, and Temperature (PVT) variations and adverse aging effects. In most cases, these reliability threats manifest themselves as timing errors on critical speed-paths of the circuit, if a large design guardband is not reserved. In this work, we propose the Reactive Rejuvenation (RR) architectural approach consisting of detection and recovery phases to mitigate circuit from BTI-induced aging. The BTI impact on the critical and near critical paths performance is continuously examined through a lightweight logic circuit which asserts an error signal in the case of any timing violation in those paths. By utilizing timing violation occurrence in the system, the timing-sensitive portion of the circuit is recovered from BTI through switching computations to redundant aging-critical voltage domain. The proposed technique achieves aging mitigation and reduced energy consumption as compared to a baseline circuit. Thus, significant voltage guardbands to meet the desired timing specification are avoided. Rizwan A. Ashraf, Ahmad Alzahrani 0001, Navid Khoshavi, Ramtin Zand, Soheil Salehi, Arman Roohi, Mingjie Lin, Ronald F. DeMara |
ISCAS | 6 |
| 2015 | Modeling an Improved Modified Type in Metallic Quantum-Dot Fixed Cell for Nano Structure ImplementationabstractQuantum-dot cellular automata (QCA) is a transistor-less computation approach which encodes binary information via configuration of charges among quantum dots. The fundamental QCA logic primitives are the majority gate and the inverter gate which can be employed to design various QCA circuits. In this study by applying some fixed predefined level of polarization, a detailed modeling of a modified type of fixed metal-dots QCA cell will be explored. An efficient architecture controlled by predefined polarization of fixed cells that position next to the input cells is presented for implementing a desired nano structure. The efficiency of the proposed approach is verified by implementing of several important examples of Boolean function. Samira Sayedsalehi, Arman Roohi |
PDP | 2 |