VLDB 2026 Research / reviewers in the wild / expert
Farshad Firouzi
dblp:15/9350
· DBLP profile ↗
63ranked-venue papers
25as first author
37since 2021 · last 2026
0000-0002-8359-4304ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 18 first-author · 20 since 2021Computer networks · 9 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Focus Session: Do Agentic LLMs Change the Paradigm of Hardware Test Generation?abstractTechnology scaling and increasing System-on-Chip (SoC) complexity exacerbate reliability challenges arising from both structural defects and runtime-dependent failures, including Silent Data Corruptions (SDCs) that evade traditional error detection mechanisms. Structural testing remains essential for detecting modeled faults such as stuck-at faults; however, it is inherently limited in capturing failures that arise under dynamic operating conditions. In contrast, functional testing can expose workload-dependent failures, albeit at the cost of high testing overhead and largely unguided workload generation. This paper presents an agentic testing framework that integrates Large Language Models (LLMs) with Reinforcement Learning (RL) and Tree-structured Parzen Estimators (TPE) to guide functional workload generation and Automatic Test Pattern Generation (ATPG) settings under user-defined constraints. The proposed approach leverages feedback-driven optimization to steer test generation toward failure-prone behaviors while reducing reliance on manual expertise. Experimental evaluation on a RISC-V processor core demonstrates that the method outperforms manually generated workloads for functional testing, while experiments on six benchmark circuits show test quality comparable to expert-generated ATPG scripts for structural testing, with improved efficiency and scalability. Farshad Firouzi, Agastya Seth, Peter Domanski, Bahareh J. Farahani, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty |
DATE | 1 |
| 2026 | The Fracture Bits in Large Language Models
Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 2 |
| 2026 | MoD-CiM: A Mixture-of-Defenses Framework Against Power-Hammering Attacks in Multi-Tenant Compute-in-Memory
Ashish Reddy Bommana, Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 3 |
| 2026 | Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection
Sujan Ghimire, Parsa Mirfasihi, Md Muhtasim Alam Chowdhury, Veeramani Pugazhenthi, Harish Kumar Dharavath, Farshad Firouzi, Rozhin Yasaei, Pratik Satam, Soheil Salehi |
VTS | 6 |
| 2026 | Late Breaking Results - A Systematic Vulnerability Analysis of MRAM-Based Compute-in-Memory against Side-Channel Attacks
Hossein Pourmehrani, Yashas Krishnamohan, Sumukh Prashant Bhanushali, Saurabh Dhiman, Rajendra Bishnoi, Arindam Sanyal, Farshad Firouzi, Naghmeh Karimi |
VTS | 7 |
| 2026 | HAT-FI: Hardware-Aware Training for Fault-Tolerant LLM Inference on RRAM-Based Compute-in-Memory
Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 3 |
| 2026 | Defending Against Model Inversion Attacks for Biomedical Images via Learnable Data PerturbationabstractThe increasing need for sharing healthcare data and collaborating on clinical research has raised privacy concerns. Health information leakage due to malicious attacks can lead to serious problems such as misdiagnoses and patient identification issues. Privacy-preserving machine learning (PPML) and privacy-enhancing technologies, particularly federated learning (FL), have emerged in recent years as innovative solutions to balance privacy protection with data utility; however, they also suffer from inherent privacy vulnerabilities. Model inversion attacks constitute major threats to data sharing in federated learning. Researchers have proposed many defenses against model inversion attacks. However, current defense methods for healthcare data lack generalizability, i.e., existing solutions may not be applicable to data from a broader range of populations. In addition, most existing defense methods are tested using non-healthcare data, which raises concerns about their applicability to real-world healthcare systems. In this study, we present a defense against model inversion attacks in federated learning. We achieve this using latent data perturbation and minimax optimization, utilizing both general and medical image datasets. We compare our method against two baselines and observe a reduction of at least 4% in the attacker’s accuracy when classifying reconstructed images, while maintaining model utility within a 1.5% drop of client classification accuracy of the undefended model. These results demonstrate improved privacy protection with minimal utility loss and suggest the potential for a generalizable defense in healthcare settings. Shiyi Jiang, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Internet Things J. | 2 |
| 2026 | Reflection on the convergence and interplay of edge, fog, and cloud in the AI-driven Internet of Things (IoT)
Farshad Firouzi, Bahareh J. Farahani, Alexander Marinsek |
Inf. Syst. | 1 |
| 2026 | COMET-3D: Compute-in-Memory-Based Transformer Accelerator With Optimized Pipeline and 3D Heterogeneous IntegrationabstractTransformers have become the backbone of large language decoder and encoder models, but their compute- and memory-intensive nature makes them inefficient on traditional von Neumann architectures. Compute-in-memory (CIM) architectures offer a promising path forward by enablingin situmatrix operations and reducing memory access overhead. However, existing CIM-based accelerators suffer from: (1) unrealistic assumption of single-cycle activation of entire crossbar; (2) inefficient pipelines that are not optimized for operational unit (OU)-based execution; and (3) high analog-to-digital converter (ADC) cost. To address these limitations, we propose a latency-optimized pipeline tailored specifically for OU-based CIM execution and introduce COMET-3D—a 3D heterogeneous architecture that integrates SRAM and ReRAM-based CIM arrays with a logic die containing digital MAC units and softmax modules. The proposed architecture and dataflow maximize hardware utilization and enable efficient acceleration of the multi-head self-attention (MHSA) layer. Experimental results across BERT, GPT2, and LLAMA demonstrate that COMET-3D outperforms baseline architectures with similar compute resources by up to 34× in energy-delay product (EDP) for LLAMA, with gains of 7.4× for GPT2 and 4.3× for BERT-Large. Ashish Reddy Bommana, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | PHANTOM: Power Hammering Attack and Countermeasure on Multi-Tenant ReRAM Compute-in-Memory AcceleratorsabstractThe increasing demand for efficient and low-power deep neural network (DNN) inference has advanced the adoption of ReRAM-based compute-in-memory (CiM) accelerators, which perform computations directly within memory to reduce energy consumption and enhance throughput. However, such architectures are vulnerable to security threats, especially in a multi-tenant environment where multiple users share the same physical resources. This paper introduces a new attack model for multi-tenant ReRAM-based CiM, power hammering, that exploits the temperature sensitivity of ReRAM cells, inducing local temperature increases that lead to conductance drift and ultimately result in erroneous inference outcomes. This serves as a denial-of-service (DoS) attack, where malicious co-tenants degrade inferencing accuracy and system reliability for legitimate users in a shared environment, ultimately undermining trust and causing potential losses to the service provider. Additionally, we propose a novel strategy to counter this security vulnerability. In this technique, we focus on selectively protecting important weights with error compensation hardware. These important weights are treated as faults, and their computation is offloaded to compensation hardware. Simulation results confirm the effectiveness of the proposed method in ensuring accurate classification results even under adversarial conditions, thereby enabling secure multi-tenant inference on ReRAM-based CiM accelerators. Ashish Reddy Bommana, Rajendra Bishnoi, Naghmeh Karimi, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | DiabLLM: An LLM-Based Framework for Blood Glucose Prediction in Type 1 DiabetesabstractAccurate Blood Glucose (BG) prediction is essential for enabling glycemic control in individuals with Type 1 Diabetes Mellitus (T1DM), particularly within Smart and Connected Health (SCH) systems that integrate Continuous Glucose Monitoring (CGM) and automated insulin delivery. The adaptability of Large Language Models (LLMs) provides a promising foundation for unified, fine-tunable forecasting models. We introduce DiabLLM, a framework based on two recent LLM-based architectures: Time-LLM, which incorporates a lightweight projection layer and alignment techniques to transform time-series data into embeddings interpretable by pre-trained LLMs, and Chronos, which employs time-series-aware tokenization and quantization to convert continuous inputs into discrete sequences for forecasting. Both models process 30-minute sequences of six historical BG values and predict 30- and 45-minute horizons. Experimental results on the OhioT1DM and D1NAMO datasets demonstrate that DiabLLM outperforms state-of-the-art baselines, including a Deep Reinforcement Learning model and an ensemble of LSTM, GRU, and WaveNet, achieving up to 27% improvement in RMSE and 37% in MAE. To enhance robustness to noisy and missing input data, a denoising autoencoder was employed for input reconstruction, yielding improved predictive performance. In addition, knowledge distillation was shown to significantly compress the model, making it a practical candidate for efficient deployment on resource-constrained edge devices without compromising accuracy. Amirhossein Mahmoudi, Ghazal Farahani, Peter Domanski, Bahareh J. Farahani, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | TIDE-S: Telemetry Informed Delay Testing With Optimized Sensor PlacementabstractSilent data corruption (SDC) refers to undetected errors that yield incorrect results without triggering system alerts or error logs. Existing test methodologies are inadequate for capturing dynamic voltage fluctuations that occur under realistic workload conditions, thereby limiting their effectiveness for detecting SDCs. We present TIDE-S, a telemetry-informed delay testing (TIDE) framework that integrates presilicon sensor placement strategies to improve telemetry accuracy. By evaluating different sensor allocation schemes—uniform,$K$-means, and energy-aware clustering—TIDE-S improves the spatial granularity of voltage observation, enabling more accurate correlation between voltage fluctuations and path delay behavior. The combined telemetry- and sensor-aware methodology significantly improves the detection of timing-sensitive SDCs. The proposed framework incurs minimal infrastructure overhead by leveraging standard pad-based voltage observation, making it practical for real system-on-chip (SoC) designs. We demonstrate the effectiveness of TIDE-S across multiple RISC-V-based SoCs and a diverse set of real-world benchmarks, showing quantifiable improvements in voltage estimation, slack prediction, and test coverage. Deepesh Sahoo, Eduardo Ortega, Peter Domanski, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | AGILE: A Multi-Task Contrastive Learning Framework with Adversarial Gradient Iterative Learning for Bio-Signal Anonymization
Tamonash Bhattacharyya, Farshad Firouzi, Amir-Mohammad Rahmani, Sanaz R. Mousavi, Krishnendu Chakrabarty |
BSN | 2 |
| 2025 | DEAR: Dependable 3D Architecture for Robust DNN TrainingabstractReRAM-based compute-in-memory (CiM) architectures present an attractive design choice for accelerating deep neural network (DNN) training. However, these architectures are susceptible to stuck-at faults (SAFs) in ReRAM cells, which arise from manufacturing defects and cell wearout over time, particularly due to the continuous weight updates during DNN training. These faults significantly degrade accuracy and compromise dependability. To address this issue, we propose DEAR: dependable 3D architecture for robust DNN training. DEAR introduces a novel online compensation method that employs a digital compensation unit to correct SAF-induced errors dynamically during both forward and backward propagation. This approach mitigates errors induced by SAFs during both the forward and backward phases of DNN training. Additionally, DEAR leverages an HBM-based 3D memory structure to store fault-related error information efficiently. Experimental results show that DEAR limits inferencing accuracy loss to under 2% even when up to 10% of cells are faulty with uniformly distributed faults, and under 2% for up to 5% faulty cells in clustered distributions. This high fault tolerance is achieved with an area overhead of 11.5% and energy overhead of less than 6% for VGG networks and less than 12% for ResNet networks. Ashish Reddy Bommana, Farshad Firouzi, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 2 |
| 2025 | Prompt, Fab, Flex: Agentic LLMs for Flexible Electronics DesignabstractFlexible Electronics (FE) have emerged as a promising platform for extreme edge applications that demand attributes tailored to the application domain, such as ultra-low cost, low power consumption, mechanical flexibility, biocompatibility, and environmental sustainability. While advances in printed and flexible device technologies have demonstrated the feasibility of sensing, computing, and communication on deformable substrates, the design and implementation of FE-based systems remain limited by traditional Electronic Design Automation (EDA) workflows, which are complex, time-intensive, and largely inaccessible to non-experts. In parallel, recent progress in Large Language Models (LLMs) has enabled automation across multiple stages of integrated circuit design; however, existing approaches exclusively target conventional silicon technologies and are not designed to address the unique constraints of FE. This work introduces the first LLM-driven framework for end-to-end hardware design automation in flexible electronics. The proposed methodology supports Register-Transfer Level (RTL) generation, logic synthesis, and cross-layer Power–Performance–Area (PPA) Design Space Exploration (DSE) for bespoke Machine Learning (ML) classifiers. Experimental results demonstrate the feasibility and effectiveness of the approach in generating resource-efficient hardware designs optimized for FE, thereby lowering barriers to adoption and accelerating the development of personalized, application-specific FEs. Farshad Firouzi, Bahareh J. Farahani, Polykarpos Vergos, Deepesh Sahoo, Nathaniel Bleier, Krishnendu Chakrabarty |
ICCAD | 1 |
| 2025 | Taming Sparse Giants: Deploying Mixture-of-Experts on 3D Heterogeneous Compute-in-Memory SystemsabstractThe deployment of large Mixture-of-Experts (MoE) models on 3D heterogeneous integrated (3D-HI) Compute-in-Memory (CiM) architectures presents unique challenges, requiring joint optimization of area, energy, latency, and perplexity (PPL). We first introduce a detailed 3D-stacked CiM architecture model, incorporating both SRAM and ReRAM tiers with thermal and device-level considerations. Building on this foundation, we present OPTIMEX, a multi-objective optimization framework that efficiently maps MoE expert projections onto heterogeneous tiers. Our evaluation demonstrates substantial benefits: up to$\text{6 0. 9 \%}$area and$\text{5 4. 7 \%}$energy reduction versus an all-SRAM baseline, while lowering PPL by as much as 98.4% compared to all-ReRAM configurations. Furthermore, OPTIMEX outperforms common heuristics, delivering improvements of up to 73.1% in area, 67.0% in energy, 96.1% in PPL, and 74.9% in latency. Together, these contributions highlight a path toward scalable, energy-efficient, and reliable MoE deployment on advanced CiM platforms. Ashish Reddy Bommana, Farshad Firouzi, Krishnendu Chakrabarty |
ICCD | 3 |
| 2025 | MALLS: Multi-Agent LLMs for Synthetic Hardware Vulnerability Generation and DetectionabstractLLMs have demonstrated promising capabilities in generating RTL code from high-level functional descriptions of hardware modules. However, their effectiveness is constrained by the lack of high-quality, diverse datasets particularly for applications in IP design, verification, and security analysis. To address this limitation, we introduce MALLS, a multi-agent framework in which specialized LLM agents namely, a generator and a discriminator collaborate in an adversarial yet cooperative setting to improve the quality and correctness of RTL designs and curate a high quality synthetic hardware vulnerability dataset. In this architecture, the generator agent is responsible for producing RTL implementations from initial seed examples through in-context learning, while the discriminator agent assesses the generator's output for functional correctness and the presence of security vulnerabilities. This interaction creates a dynamic feedback loop, enabling both agents to iteratively improve through each other's responses leading to self-supervised learning. A hard bank of examples is maintained in a database which include instances that were difficult to generate or detect by either agents. By generating paired positive (correct) and negative (buggy) examples, the system learns to distinguish subtle design flaws and generalize across diverse RTL patterns while generating high quality synthetic examples of hardware vulnerabilities. Experimental results show that using the hard bank of examples produced by the adversarial multi-LLM setup improves both vulnerability generation and detection performance. Jonti Talukdar, Agastya Seth, Sanmitra Banerjee, Farshad Firouzi, Krishnendu Chakrabarty |
ICCD | 4 |
| 2025 | Exploration of Low-Power Flexible Stress Monitoring Classifiers for Conformal WearablesabstractConventional stress monitoring relies on episodic, symptom-focused interventions, missing the need for continuous, accessible, and cost-efficient solutions. State-of-the-art approaches use rigid, silicon-based wearables, which, though capable of multitasking, are not optimized for lightweight, flexible wear, limiting their practicality for continuous monitoring. In contrast, flexible electronics (FE) offer flexibility and low manufacturing costs, enabling real-time stress monitoring circuits. However, implementing complex circuits like machine learning (ML) classifiers in FE is challenging due to integration and power constraints. Previous research has explored flexible biosensors and ADCs, but classifier design for stress detection remains underexplored. This work presents the first comprehensive design space exploration of low-power, flexible stress classifiers. We cover various ML classifiers, feature selection, and neural simplification algorithms, with over 1200 flexible classifiers. To optimize hardware efficiency, fully customized circuits with low-precision arithmetic are designed in each case. Our exploration provides insights into designing real-time stress classifiers that offer higher accuracy than current methods, while being low-cost, conformable, and ensuring low power and compact size. Florentia Afentaki, Sri Sai Rakesh Nakkilla, Konstantinos Balaskas, Paula L. Duarte, Shiyi Jiang, Georgios Zervakis 0001, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ISLPED | 7 |
| 2025 | LLM-Aided In-Field Workload Generation for Detecting Silent Data Corruptions at ScaleabstractComputational integrity is crucial in large-scale data centers where Silent Data Corruptions (SDCs) pose a growing reliability challenge. SDCs can lead to incorrect computation results not captured by traditional error detection mechanisms, making their detection and mitigation essential. However, existing post-manufacturing and in-field testing methods, such as opportunistic and ripple testing, face significant scalability challenges due to high computational costs and test times. We propose an LLM-aided approach for generating targeted test cases to detect SDCs. As a case study, we focus on the functional blocks of a RISC-V CV32E40P processor core. Our method generates targeted test cases that maximize voltage droops in given hardware modules, such as functional units, increasing the likelihood of triggering SDCs in-field. Additionally, our approach is architecture-aware and layout-aware, enhancing fault activation and enabling automated optimization of generated test cases. Experimental evaluations demonstrate that the proposed method significantly improves SDC detection efficiency by reducing the number of required test cases while preserving high test coverage. Compared to commonly used test cases, the proposed approach increases average voltage droops by up to 38%. By integrating LLM-aided test case generation, the proposed approach achieves voltage droops of up to 9% relative to the supply voltage, improving the effectiveness of in-field SDC detection and mitigation strategies. Peter Domanski, Deepesh Sahoo, Eduardo Ortega, Farshad Firouzi, Krishnendu Chakrabarty |
ITC | 4 |
| 2025 | TIDE: Telemetry-Informed Delay Testing for Silent Data Corruption *abstractSilent Data Corruption (SDC) is caused by undetected errors that yield incorrect results without triggering system alerts or error logs. Existing test methodologies are inadequate for capturing dynamic voltage fluctuations that occur under realistic workload conditions, thereby limiting their effectiveness for detecting SDCs. To address these limitations, we introduce Telemetry-Informed Delay Testing (TIDE), a novel methodology that enhances SDC detection by leveraging telemetry sensors to monitor voltage fluctuations and their impact on timing integrity. By incorporating dynamic, workload-aware test generation, the proposed framework overcomes key limitations of traditional approaches and facilitates early detection of SDCs. The effectiveness of TIDE is demonstrated through case studies conducted on two RISC-V-based SoCs and multiple workloads. Deepesh Sahoo, Eduardo Ortega, Peter Domanski, Farshad Firouzi, Krishnendu Chakrabarty |
ITC | 4 |
| 2025 | Silent Data Corruption: Advancing Detection, Diagnosis, and Mitigation StrategiesabstractSilent Data Corruptions (SDCs) pose a critical challenge to computer system reliability, arising from vulnerabilities across different layers of the computing stack. This paper addresses this challenge through three complementary contributions that systematically target SDCs from hardware manufacturing to application-level resilience. First, we analyze timing failures caused by random process variations in advanced technology nodes, revealing that extreme slow paths at lower voltages are dominated by single weak transistors—insights crucial for manufacturing and in-field testing. Second, we introduce an LLM-driven framework that generates targeted functional test programs to induce SDCs, demonstrating its effectiveness in stressing hardware, uncovering latent vulnerabilities, and increasing energy consumption in a given device under test (DUT), making it a valuable tool for in-field testing. Third, as machine learning continues to drive advancements across critical domains such as healthcare, finance, and autonomous systems, ensuring its reliability is paramount. However, the susceptibility of these applications to SDCs threatens their reliability and robustness. To address this, we propose Fidelity-Q, a novel fault injection methodology to evaluate the impact of SDCs on Quantized Neural Networks (QNNs), showing that lower-bit quantization increases error susceptibility. Collectively, these contributions provide a comprehensive approach to identifying, analyzing, and mitigating SDCs across the computing stack, from hardware testing to machine learning applications. Peter Domanski, Mukarram Ali Faridi, Gabriel Kaunang, Wilson Pradeep, Adit D. Singh, Alfian Amrizal, Yanjing Li, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 8 |
| 2025 | ChipMnd: LLMs for Agile Chip DesignabstractThe increasing complexity of semiconductor design, along with stringent performance, power, and time-to-market requirements, has outpaced the capabilities of traditional Electronic Design Automation (EDA) methodologies. Conventional design workflows rely on manual intervention for critical tasks such as hardware description, synthesis optimization, and verification, leading to inefficiencies and scalability limitations. Large Language Models (LLMs) present a transformative approach by automating key stages of the design pipeline, enabling intelligent synthesis tuning, test generation, and security analysis. This paper introduces ChipMind, an LLM-driven framework comprising specialized agents and modules for digital and analog chip design. ChipMind integrates AI-driven methodologies to enhance design efficiency, accelerate prototyping, and optimize key design trade-offs, thereby addressing fundamental challenges in modern semiconductor development. Farshad Firouzi, David Z. Pan, Jiaqi Gu 0002, Bahareh J. Farahani, Jayeeta Chaudhuri, Ziang Yin, Pingchuan Ma 0012, Peter Domanski, Krishnendu Chakrabarty |
VTS | 1 |
| 2025 | SPICED+: Syntactical Bug Pattern Identification and Correction of Trojans in A/MS Circuits Using LLM-Enhanced DetectionabstractAnalog and mixed-signal (A/MS) integrated circuits (ICs) are crucial in modern electronics, playing key roles in signal processing, amplification, sensing, and power management. Many IC companies outsource manufacturing to third-party foundries, creating security risks such as syntactical bugs and stealthy analog Trojans. Traditional Trojan detection methods, including embedding circuit watermarks and hardware-based monitoring, impose significant area and power overheads while failing to effectively identify and localize the Trojans. To overcome these shortcomings, we present SPICED+, a software-based framework designed for syntactical bug pattern identification and the correction of Trojans in A/MS circuits, leveraging large language model (LLM)-enhanced detection. It uses LLM-aided techniques to detect, localize, and iteratively correct analog Trojans in SPICE netlists, without requiring explicit model training, and thus incurs zero area overhead. The framework leverages chain-of-thought reasoning and few-shot learning to guide the LLMs in understanding and applying anomaly detection rules, enabling accurate identification and correction of Trojan-impacted nodes. With the proposed method, we achieve an average Trojan coverage of 93.3%, average Trojan correction rate of 91.2%, and an average false-positive rate of 1.4%. Jayeeta Chaudhuri, Dhruv Thapar, Arjun Chaudhuri, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Neural Architecture Search for Blood Glucose Prediction in Type-1 DiabeticsabstractFor subjects affected with type-1 diabetes mellitus, accurately predicting future blood glucose values helps regulate insulin delivery. This paper introduces a dual Q-network-based neural architecture search approach to develop and train per-sonalized BG prediction models for individuals affected with type-1 diabetes mellitus. Utilizing historical blood glucose data collected via body sensor networks, the proposed model forecasts future blood glucose levels. When evaluated on the OhioTlDM dataset, the proposed approach shows significant improvements over the state-of-the-art, achieving a 46.78% reduction in root mean square error and a 56.05% reduction in mean absolute error while predicting blood glucose values 5 minutes into the future. Anthony Liardo, Aritra Ray, Farshad Firouzi, Kyle J. Lafata, Krishnendu Chakrabarty |
BSN | 3 |
| 2024 | LLM-AID: Leveraging Large Language Models for Rapid Domain-Specific Accelerator DevelopmentabstractThe challenges posed by the Dark Silicon era, combined with the escalating computational demands of emerging applications, such as Deep Learning (DL), have strained the capabilities of traditional CPUs and GPUs, necessitating the development of Domain-Specific Accelerators (DSAs). Despite offering substantial enhancements in Power, Performance, and Area (PPA), DSAs encounter significant challenges, including the rapid evolution of applications that necessitate the frequent development of new architectures. This, coupled with the expertise-intensive nature of the design process, often leads to reduced flexibility and extended development cycles, ultimately hindering the broader adoption and efficient deployment of DSAs. To address these challenges, this paper introduces LLM-AID, an agile framework that streamlines the DSA design flow by transforming high-level abstract specifications into Hardware Description Language (HDL) code and facilitating backend Computer-Aided Design (CAD) tool operations. By synergistically combining Large Language Models (LLMs), High-Level Synthesis (HLS) tools, design exploration techniques, and symbolic AI, LLM-AID dramatically accelerates design iterations, optimizes hardware performance, and significantly reduces time-to-market. This innovative approach democratizes DSA development, empowering designers to achieve unprecedented productivity while delivering high-quality DSA solutions. Farshad Firouzi, Sri Sai Rakesh Nakkilla, Chenghao Fu, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty |
ICCAD | 1 |
| 2024 | KG-Infused LLM for Virtual Health Assistant: Accelerated Inference and Enhanced PerformanceabstractVirtual Health Assistants (VHAs) represent a significant advancement in patient care, leveraging artificial intelligence to provide continuous support and interaction. Despite the integration of advanced Large Language Models (LLMs), which have greatly improved conversational capabilities, VHAs still struggle to fully replicate the nuanced expertise of human medical professionals. This limitation is primarily due to their reliance on broad training data, which can result in responses that lack the necessary reliability and contextual relevance required in health-care, where precision is paramount. To address this challenge, recent research has focused on integrating healthcare databases, such as the Unified Medical Language System (UMLS), a comprehensive biomedical knowledge source, to enhance the reasoning capabilities and reliability of LLMs. However, previous efforts have been impeded by high inference times and suboptimal performance on various evaluation metrics due to inefficient retrieval of pertinent information from extensive databases. In this study, we propose a novel methodology that involves constructing a Knowledge Graph (KG) within the Neo4j graph database using UMLS data to facilitate faster retrieval. This approach is augmented by employing named entity recognition techniques to accurately identify relevant entities and by applying learning-to-rank and semantic matching algorithms to effectively rank the retrieved information. We validated our approach using various LLMs, including GPT-3.5 Turbo, GPT-4, LLaMA-7b, and LLaMA-13b, across the BioASQ, MedicationQA, and ExpertQA datasets. Our experiments demonstrated a 23% improvement in ROUGE-L scores, an 8-10% improvement in BERTScores, and a 20-30% improvement in BLEU scores, achieving up to an 80% reduction in inference time. Siva Kumar Katta, Aritra Ray, Farshad Firouzi, Krishnendu Chakrabarty |
ICMLA | 3 |
| 2024 | Preserving Accuracy While Stealing Watermarked Deep Neural NetworksabstractThe deployment of Deep Neural Networks (DNNs) as cloud services has accelerated significantly over the years. Training an application-specific DNN for cloud deployment requires substantial computational resources and costs associated with hyper-parameter tuning and model selection. To preserve Intellectual Property (IP) rights, model owners embed watermarks into publicly deployed DNNs. These trigger inputs and labels are uniquely selected and embedded into the watermarked DNN by the model owner, remaining undisclosed during deployment. If a watermarked DNN (target classifier) is stolen via white-box access and re-deployed by an adversary (pirated classifier) without securing the IP rights from the model owner, the model owner can identify their IP by sending trigger inputs to retrieve trigger labels. Typically, adversaries tamper with the model weights of the target classifier prior to deployment, which in turn reduces the utility of the well-trained DNN. The authors proposes re-deploying the target classifier without altering the model weights to preserve model utility, and using a small sample of non-identical in-distribution inputs (used for training the target classifier) to train a Siamese neural network to evade detection, at inference stage. Experimental evaluations on standard benchmark datasets- MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100- using ResNet architectures with varying triggers demonstrate that the proposed method achieves zero false positive rate (fraction of clean testing input incorrectly labelled as trigger inputs) and false negative rate (fraction of trigger inputs incorrectly labelled as clean in-distribution inputs) in nearly all cases, proving its efficacy. Aritra Ray, Farshad Firouzi, Kyle J. Lafata, Krishnendu Chakrabarty |
ICMLA | 2 |
| 2024 | SEC-CiM: Selective Error Compensation for ReRAM-based Compute-in-Memory*abstractReRAM-based Compute-in-Memory (CiM) architectures offer an attractive design choice for accelerating Convolutional Neural Network (CNN) inferencing in edge computing environments. However, these architectures are susceptible to stuck-at-faults (SAFs) in ReRAM cells stemming from manufacturing defects and cell wearout over time, significantly degrading CNN inferencing accuracy. To address this challenge, we propose a technique called Selective Error Compensation for CiM (SEC-CiM). This technique strategically mitigates errors by leveraging the insight that compensating for errors in a limited number of selected columns in a crossbar is sufficient to maintain CNN inferencing accuracy. With this strategy, SEC-CiM achieves significantly lower overhead compared to previous work. Notably, it effectively addresses errors resulting from stuck-at intermediate levels, a critical aspect that was previously overlooked. We develop a theoretical framework to determine the minimum number of columns requiring error compensation. Simulation results demonstrate that SEC-CiM limits the drop in inferencing accuracy to 2% for the ResNet18 and VGG16 models, even when up to 30% of the ReRAM cells in the crossbar are faulty. Similarly, for the Densenet121 CNN, comparable accuracy results are obtained when up to 15% of the ReRAM cells are faulty. We achieve this high level of fault tolerance with moderate area and power consumption overhead of 12.2% and 10.2%, respectively. Ashish Reddy Bommana, Farshad Firouzi, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ITC | 2 |
| 2024 | Low-Overhead Clustered Federated Learning for Personalized Stress MonitoringabstractStress, recognized widely as a substantial health concern, adversely affects individuals by undermining both their physical and mental well being. Prior studies on stress monitoring and management utilize a centralized cloud-based approach that combines data from each client for modeling. However, such a centralized approach raises data privacy concerns. To preserve privacy, decentralized federated learning (FL) has been proposed as a potential alternative framework. Nevertheless, existing FL algorithms have to deal with data heterogeneity; data skewness in each participant can significantly degrade the overall model performance. To tackle this challenge, we present a personalized, low-overhead clustered FL algorithm for stress-level recognition. The proposed algorithm outperforms two state-of-the-art baseline algorithms by providing over 7% and 12% increase in accuracy, respectively. The proposed algorithm also obtains a reduction of 37.5% and 9.6% in the training runtime compared to the two baseline algorithms. We also present a novel cold-start algorithm for new clients who join the trained system. Our results suggest that this cold-start algorithm is robust in terms of individual classification accuracy and total training time. Shiyi Jiang, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Internet Things J. | 2 |
| 2023 | Guest Editorial Special Issue on Empowering the Future Generation Systems: Opportunities by the Convergence of Cloud, Edge, AI, and IoTabstractThe future generation of the Internet of Things (IoT) systems is characterized by the fusion of technologies—from edge–fog–cloud computing to artificial intelligence (AI) and blockchain—closing the gap between the physical and digital worlds [A1]. Although these technologies have been developed separately over time, the synergy among them has taken a giant leap. We are witnessing a fast-paced convergence of these technologies resulting in a fundamental paradigm shift unlocking vast benefits and opportunities across vertical markets. However, there are still several barriers, such as a lack of consensus toward any reference models or best practices, hindering the full fusion of these technologies [A1]. To tackle these challenges and facilitate this promising transformation, this special issue was organized to provide a holistic multidisciplinary reference for solutions, architectures, protocols, services, and applications addressing all aspects of the future generation of IoT systems via the fusion of edge, cloud, AI, and blockchain, while considering the corresponding challenges. Thanks to the enormous support from the Editor-in-Chief, Prof. Honggang Wang, and the dedicated work of many reviewers, after a rigorous review process, 27 excellent articles out of 125 submissions were accepted for inclusion in this special issue of the IEEE Internet of Things Journal. We introduce these papers and highlight their key contributions below. Farshad Firouzi, Mahmoud Daneshmand, Jaeseung Song, Kunal Mankodiya |
IEEE Internet Things J. | 1 |
| 2023 | Fusion of IoT, AI, Edge-Fog-Cloud, and Blockchain: Challenges, Solutions, and a Case Study in Healthcare and MedicineabstractThe digital transformation is characterized by the convergence of technologies—from the Internet of Things (IoT) to edge–fog–cloud computing, artificial intelligence (AI), and Blockchain—in multiple dimensions, blurring the lines between the physical and digital worlds. Although these innovations have evolved independently over time, they are increasingly becoming more intertwined, driving the development of new business models. With more adaptation, embracement, and development, we are witnessing a steady convergence and fusion of these technologies resulting in an unprecedented paradigm shift that is expected to disrupt and reshape the next-generation systems in vertical domains in a way that the capabilities of the technologies are aligned in the best possible way to complement each other. Despite the fact that the convergence of the four technologies can potentially tackle the main shortcomings of the existing systems, its adoption is still in its infancy phase, suffering from several issues, such as the absence of consensus toward any reference models or best practices. This article provides a comprehensive insight into the fusions of these paradigms by discussing a blend of topics addressing all the importation aspects from design to deployment. We will begin this article by providing an in-depth discussion on the main requirements, state-of-the-art reference architectures, applications, and challenges. Following this, we will present a reference architecture and a case study on privacy-preserving stress monitoring and management to better elaborate on the corresponding details and considerations. Farshad Firouzi, Shiyi Jiang, Krishnendu Chakrabarty, Bahareh J. Farahani, Mahmoud Daneshmand, Jaeseung Song, Kunal Mankodiya |
IEEE Internet Things J. | 1 |
| 2022 | AI-Driven Data Monetization: The Other Face of Data in IoT-Based Smart and Connected HealthabstractAs the trajectory of the Internet of Things (IoT) moving at a rapid pace and with the rapid worldwide development and public embracement of wearable sensors, these days, most companies and organizations are awash in massive amounts of data. Determining how to profit from data deluge can give companies an edge in the market because data have the potential to add tremendous value to many aspects of a business. The market has already seen a level of monetization across vertical domains in the form of layering connected devices with a variety of Software-as-a-Service (SaaS) choices, such as subscription plans or smart device insights. Out of this arena is evolving a “machine economy” in which the ability to correctly monetize data rather than simply hoard it, will provide a significant advantage in a competitive digital environment. The recent advent of the technological advances in the fields of big data, analytics, and artificial intelligence (AI) has opened new avenues of competition, where data are utilized strategically and treated as a continuously changing asset able to unleash new revenue opportunities for monetization. Such growth has made room for an onslaught of new tools, architectures, business models, platforms, and marketplaces that enable organizations to successfully monetize data. In fact, emerging business models are striving to alter the power balance between users and companies that harvest information. Start-ups and organizations are offering to sell user data to data analytics companies and other businesses. Monetizing data goes beyond just selling data. It is also possible to include steps that add value to data. Generally, organizations can monetize data by: 1) utilizing it to make better business decisions or improve processes; 2) surrounding flagship services or products with data; or 3) selling information to current or new markets. This article will address all important aspects of IoT data monetization with more focus on the healthcare industry and discuss the corresponding challenges, such as data management, scalability, regulations, interoperability, security, and privacy. In addition, it presents a holistic reference architecture for the healthcare data economy with an in-depth case study on the detection and prediction of cardiac anomalies using multiparty computation (MPC) and privacy-preserving machine learning (PPML) techniques. Farshad Firouzi, Bahareh J. Farahani, Mojtaba Barzegari, Mahmoud Daneshmand |
IEEE Internet Things J. | 1 |
| 2022 | Guest Editorial Special Issue on AI-Driven IoT Data Monetization: A Transition From Value Islands to Value EcosystemsabstractAs The trajectory of the Internet of Things (IoT) is moving at a rapid pace, most companies and organizations are awash and drowning in massive amounts of data. Determining how to profit from data deluge and unlock its value can give companies an edge in the market because data have the potential to add tremendous value to many aspects of a business [A1]. The market has already seen a level of monetization across vertical domains e.g., in the form of layering connected devices with a variety of Insights-as-a- Service options. Out of this arena, the data economy concept has been evolving, characterized by correctly monetizing data rather than simply hoarding it, which will provide a significant advantage in a competitive digital environment [A1]. The recent advent of technological advances in the fields of Big Data, Analytics, and Artificial Intelligence (AI) has opened new avenues of competition, where IoT data is considered a living and evolving entity that can unlock enormous opportunities for monetization. Such growth brought forth a slew of new tools, architectures, business models, platforms, and marketplaces, enabling organizations to monetize data successfully. In this context, emerging business models also strive to alter the power balance between users and companies that harvest information by utilizing usage policy enforcement and privacypreserving machine learning techniques. Monetizing data goes beyond just selling data. It is also possible to include steps that add value to data. Generally, organizations can monetize data by 1) utilizing it to make better business decisions or improve processes; 2) surrounding flagship services or products with data; or 3) selling information to current or new markets [A1]. Farshad Firouzi, Bahareh J. Farahani, Mahmoud Daneshmand, Cesare Pautasso |
IEEE Internet Things J. | 1 |
| 2022 | A Resilient and Hierarchical IoT-Based Solution for Stress Monitoring in Everyday SettingsabstractThe conventional mental healthcare regime often follows a symptom-focused and episodic approach in a noncontinuous manner, wherein the individual discretely records their biomarker levels or vital signs for a short period prior to a subsequent doctor’s visit. Recognizing that each individual is unique and requires continuous stress monitoring and personally tailored treatment, we propose a holistic hybrid edge–cloud Wearable Internet of Things (WIoT)-based online stress monitoring solution to address the above needs. To eliminate the latency associated with cloud access, appropriate edge models—spiking neural network (SNN), Conditionally Parameterized Convolutions (CondConv), and support vector machine (SVM)—are trained, enabling low-energy real-time stress assessment near the subjects on the spot. This work leverages design-space exploration for the purpose of optimizing the performance and energy efficiency of machine learning inference at the edge. The cloud exploits a novel multimodal matching network model that outperforms six state-of-the-art stress recognition algorithms by 2%–7% in terms of accuracy. An offloading decision process is formulated to strike the right balance between accuracy, latency, and energy. By addressing the interplay of edge–cloud, the proposed hierarchical solution leads to a reduction of 77.89% in response time and 78.56% in energy consumption with only a 7.6% drop in accuracy compared to the Internet of Things (IoT)–Cloud scheme, and it achieves a 5.8% increase in accuracy on average compared to the IoT-Edge scheme. Shiyi Jiang, Farshad Firouzi, Krishnendu Chakrabarty, Eric B. Elbogen |
IEEE Internet Things J. | 2 |
| 2022 | The convergence and interplay of edge, fog, and cloud in the AI-driven Internet of Things (IoT)
Farshad Firouzi, Bahareh J. Farahani, Alexander Marinsek |
Inf. Syst. | 1 |
| 2021 | Harnessing the Power of Smart and Connected Health to Tackle COVID-19: IoT, AI, Robotics, and Blockchain for a Better WorldabstractAs COVID-19 hounds the world, the common cause of finding a swift solution to manage the pandemic has brought together researchers, institutions, governments, and society at large. The Internet of Things (IoT), artificial intelligence (AI)-including machine learning (ML) and Big Data analytics-as well as Robotics and Blockchain, are the four decisive areas of technological innovation that have been ingenuity harnessed to fight this pandemic and future ones. While these highly interrelated smart and connected health technologies cannot resolve the pandemic overnight and may not be the only answer to the crisis, they can provide greater insight into the disease and support frontline efforts to prevent and control the pandemic. This article provides a blend of discussions on the contribution of these digital technologies, propose several complementary and multidisciplinary techniques to combat COVID-19, offer opportunities for more holistic studies, and accelerate knowledge acquisition and scientific discoveries in pandemic research. First, four areas, where IoT can contribute are discussed, namely: 1) tracking and tracing; 2) remote patient monitoring (RPM) by wearable IoT (WIoT); 3) personal digital twins (PDTs); and 4) real-life use case: ICT/IoT solution in South Korea. Second, the role and novel applications of AI are explained, namely: 1) diagnosis and prognosis; 2) risk prediction; 3) vaccine and drug development; 4) research data set; 5) early warnings and alerts; 6) social control and fake news detection; and 7) communication and chatbot. Third, the main uses of robotics and drone technology are analyzed, including: 1) crowd surveillance; 2) public announcements; 3) screening and diagnosis; and 4) essential supply delivery. Finally, we discuss how distributed ledger technologies (DLTs), of which blockchain is a common example, can be combined with other technologies for tackling COVID-19. Farshad Firouzi, Bahareh J. Farahani, Mahmoud Daneshmand, Kathy Grise, Jaeseung Song, Roberto Saracco, Lucy Lu Wang, Kyle Lo, Plamen Angelov 0001, Eduardo A. Soares 0001, Po-Shen Loh, Zeynab Talebpour, Reza Moradi, Mohsen Goodarzi, Haleh Ashraf, Mohammad Talebpour, Alireza Talebpour, Luca Romeo, Rupam Das, Hadi Heidari, Dana K. Pasquale, James Moody, Chris Woods, Erich Huang, Payam M. Barnaghi, Majid Sarrafzadeh, Ron C. Li, Kristen L. Beck, Olexandr Isayev, NakMyoung Sung |
IEEE Internet Things J. | 1 |
| 2021 | The convergence of IoT and distributed ledger technologies (DLT): Opportunities, challenges, and solutions
Bahareh J. Farahani, Farshad Firouzi, Markus Lücking |
J. Netw. Comput. Appl. | 2 |
| 2018 | Towards fog-driven IoT eHealth: Promises and challenges of IoT in medicine and healthcare
Bahareh J. Farahani, Farshad Firouzi, Victor Chang 0001, Mustafa Badaroglu, Nicholas Constant, Kunal Mankodiya |
Future Gener. Comput. Syst. | 2 |
| 2018 | Internet-of-Things and big data for smarter healthcare: From device to architecture, applications and analytics
Farshad Firouzi, Amir-Mohammad Rahmani, Kunal Mankodiya, Mustafa Badaroglu, Geoff V. Merrett, Bahareh J. Farahani |
Future Gener. Comput. Syst. | 1 |
| 2018 | Keynote Paper: From EDA to IoT eHealth: Promises, Challenges, and SolutionsabstractThe interaction between technology and healthcare has a long history. However, recent years have witnessed the rapid growth and adoption of the Internet of Things (IoT) paradigm, the advent of miniature wearable biosensors, and research advances in big data techniques for effective manipulation of large, multiscale, multimodal, distributed, and heterogeneous data sets. These advances have generated new opportunities for personalized precision eHealth and mHealth services. IoT heralds a paradigm shift in the healthcare horizon by providing many advantages, including availability and accessibility, ability to personalize and tailor content, and cost-effective delivery. Although IoT eHealth has vastly expanded the possibilities to fulfill a number of existing healthcare needs, many challenges must still be addressed in order to develop consistent, suitable, safe, flexible and power-efficient systems that are suitable fit for medical needs. To enable this transformation, it is necessary for a large number of significant technological advancements in the hardware and software communities to come together. This keynote paper addresses all these important aspects of novel IoT technologies for smart healthcare-wearable sensors, body area sensors, advanced pervasive healthcare systems, and big data analytics. It identifies new perspectives and highlights compelling research issues and challenges, such as scalability, interoperability, device-network-human interfaces, and security, with various case studies. In addition, with the help of examples, we show how knowledge from CAD areas, such as large scale analysis and optimization techniques can be applied to the important problems of eHealth. Farshad Firouzi, Bahareh J. Farahani, Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Guest Editorial: Alternative Computing and Machine Learning for Internet of ThingsabstractThe impending Internet of Things (IoT) wave is promising to affect every aspect of our daily lives, ranging from smart things to smart buildings, smart cities, and smart environments. A lot of attention has been devoted to the tsunami of data produced by IoT, and the related means of extracting useful actionable information from it, spawning efforts in Big Data processing and machine learning. Yet, all of this does little to address the need for IoT to capture, interpret, and act on this wall of (noisy) information at the right time, at the right place, and in the right form. Conventional computing systems are a poor match to the needs of this emerging massively distributed real-time system. Hence, alternative computing techniques present an attractive alternative, trading off computational resolution for significant gains in quality-of-service energy efficiency and robustness. This observation is based on the conjecture that most applications related to IoT have an inherent error resilience and are evolutionary (that is, learning-based). Alternative computing strategies may be conceived at every level of the design hierarchy, starting from the device level with novel 3-D nonvolatile memory/logic combinations, or at the architectural level by shifting away from the traditional von Neumann architecture to different computing paradigms such as neuromorphic and/or stochastic computation all the way up to the algorithmic and data representation levels. Farshad Firouzi, Bahareh J. Farahani, Andrew B. Kahng, Jan M. Rabaey, Natasha Balac |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | On-Chip Droop-Induced Circuit Delay Prediction Based on Support-Vector MachinesabstractVoltage droop is a major reliability concern in nano-scale very large-scale integration designs. Undesirable voltage droop is often a result of excessive IR drop. On the other hand, Ldi/dt-induced droop occurs when logic gates in the circuit draw high-switching current from the on-chip power supply network, and this problem is exacerbated at high-clock frequencies and smaller technology nodes. A consequence of voltage droop is usually an increase in path delays and the occurrence of intermittent faults during circuit operation. The addition of conservative timing margins, also known as guardbands, is a common practice to tackle the problem of voltage droop. However, such static and pessimistic guardbands, which are calculated at design time based on worst-case conditions, lead to significant performance loss. Dynamic frequency scaling is an alternative approach that enables the dynamic adjustment of clock frequency based on the actual voltage droop seen during runtime. For dynamic voltage-frequency to be effective, accurate and real-time prediction of voltage droop is essential. We propose a support-vector machine (SVM)-based regression method to predict voltage droop due to pattern-dependent IR drop based on inputs to the chip at runtime. Moreover, we reduce the amount of data needed for accurate prediction by using correlation-based feature selection. Several benchmarks from ITC'99 and International Work on Logic and Synthesis'05 highlight the effectiveness of the proposed method in terms of delay-prediction accuracy. Since real-time droop prediction requires hardware implementation of the predictor, we present the hardware design and synthesis results to demonstrate that the hardware overhead for the SVM predictor is negligible for large circuits. Fangming Ye, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Stress-aware P/G TSV planning in 3D-ICsabstractPower/Ground (P/G) Through-Silicon-Vias (TSVs) in the Power Distribution Network (PDN) of Three-Dimensional-Integrated-Circuit (3D-IC) have a twofold impact on the delays of the surrounding gates. TSV fabrication causes thermal stress around TSVs, which results in significant carrier mobility variations in their vicinity. On the other hand, the insertion of P/G TSVs will change the voltage of each node in the power grid, which also impacts the delays of the connected gates. Thus, it is necessary to consider the combined effect on delay variation during the P/G TSV planning. In this work, we propose a methodology using Mixed-Integer-Bilinear-Programming (MIBLP) to optimize this delay variation by a refined P/G TSV allocation. Taking into account the impact of thermal stress as well as voltage drop on the circuit delay, we optimally plan the P/G TSVs to minimize the circuit delay for different keep-out zones (KOZs) and PDN pitches. Shengcheng Wang, Farshad Firouzi, Fabian Oboril, Mehdi Baradaran Tahoori |
ASP-DAC | 2 |
| 2015 | On-line prediction of NBTI-induced aging rates
Rafal Baranowski, Farshad Firouzi, Saman Kiamehr, Chang Liu 0010, Mehdi Baradaran Tahoori, Hans-Joachim Wunderlich |
DATE | 2 |
| 2015 | Re-using BIST for circuit aging monitoringabstractBias Temperature Instability (BTI)-induced transistor aging degrades path delay over time and may eventually induce circuit failure due to timing violations. Chip health monitoring is therefore necessary to track delay changes on a per-chip basis. We propose a method to accurately predict the fine-grained circuit-delay degradation with minimal area and performance overhead. It re-uses on-chip design-for-test (DfT) infrastructure to track the severity of run-time stress by periodiclly capturing system state and compacting it using a multiple input signature register (MISR). The captured stress information is fed to a software-based prediction model in realtime. The prediction model is trained offline using support vector regression. Aging prediction based on run-time stress monitoring can be used to proactively activate aging mitigation techniques. Experimental results for benchmark circuits highlight the accuracy of the proposed approach. Farshad Firouzi, Fangming Ye, Arunkumar Vijayan, Abhishek Koneru, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ETS | 1 |
| 2015 | Aging- and Variation-Aware Delay Monitoring Using Representative Critical Path SelectionabstractProcess together with runtime variations in temperature and voltage, as well as transistor aging, degrade path delay and may eventually induce circuit failure due to timing variations. Therefore, in-field tracking of path delays is essential, and to respond to this need, several delay sensor designs have been proposed in the literature. However, due to the significant overhead of these sensors and the large number of critical paths in today's IC, it is infeasible to monitor the delay of every critical path in silicon. We present an aging- and variationaware representative path selection technique based on machine learning that allows to measure the delay of a small set of paths and infer the delay of a larger pool of paths that are likely to fail due to delay variations. Simulation results for benchmark circuits highlight the accuracy of the proposed approach for predicting critical-path delay based on the selected representative paths. Farshad Firouzi, Fangming Ye, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2014 | Adaptive Mitigation of Parameter VariationsabstractIn the deep nanoscale regime, process and runtime variations have emerged as the major sources of uncertainty and unpredictability in circuit operation. Static mitigation approaches do not consider the dependence of variations on workload and chip usage, while adaptive techniques do not incorporate detailed circuit-level information. We propose a fine-grained adaptive technique in which machine learning is exploited to perform circuit clustering and obtain a representative for each cluster. By monitoring the representative in each cluster at runtime, performance variations in the entire cluster can be tracked such that appropriate fine-grained adaptation can be applied to each cluster. Experimental results for ISCAS'89, IWLS'05, and ITC'99 benchmarks as well as the LEON processor show that the proposed approach introduces negligible overhead significantly extends circuit lifetime, facilitates higher operating frequencies, and reduces the leakage power. Farshad Firouzi, Fangming Ye, Saman Kiamehr, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ATS | 1 |
| 2014 | Aging-aware standard cell library designabstractTransistor aging, mostly due to Bias Temperature Instability (BTI), is one of the major unreliability sources at nano-scale technology nodes. BTI causes the circuit delay to increase and eventually leads to a decrease in the circuit lifetime. Typically, standard cells in the library are optimized according to the design time delay, however, due to the asymmetric effect of BTI, the rise and fall delays might become significantly imbalanced over the lifetime. In this paper, the BTI effect is mitigated by balancing the rise and fall delays of the standard cells at the excepted lifetime. We find an optimal tradeoff between the increase in the size of the library and the lifetime improvement (timing margin reduction) by non-uniform extension of the library cells for various ranges of the input signal probabilities. The simulation results reveal that our technique can prolong the circuit lifetime by around 150% with a negligible area overhead. Saman Kiamehr, Farshad Firouzi, Mojtaba Ebrahimi, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2014 | P/G TSV planning for IR-drop reduction in 3D-ICsabstractIn recent years, interconnect issues emerged as major performance challenges for Two-Dimensional-Integrated-Circuits (2D-ICs). In this context, Three-Dimensional-ICs (3D-ICs), which consist of several active layers stacked above each other, offer a very attractive alternative to conventional 2D-ICs. However, 3D-ICs also face many challenges associated with the Power Distribution Network (PDN) design due to the increasing power density and larger supply current compared to 2D-ICs. As an important part of 3D-IC PDNs, Power/Ground (P/G) Through-Silicon-Vias (TSVs) should be well-managed. Excessive or ill-placed P/G TSVs impact the power integrity (e.g. IR-drop), and also consume a considerable amount of chip real estate. In this work, we propose a Mixed-Integer-Linear-Programming (MILP)-based technique to plan the P/G TSVs. The goal of our approach is to minimize the average IR-drop while satisfying the total area constraint of TSVs by optimizing the P/G TSV placement. Therefore, the locations, sizes and the total number of the P/G TSVs are co-optimized simultaneously. The experimental results show that the average IR-drop can be reduced by 11.8 % in average using the proposed method compared to a random placement technique with a much smaller runtime. Shengcheng Wang, Farshad Firouzi, Fabian Oboril, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2014 | On-chip voltage-droop prediction using support-vector machinesabstractVoltage droop is a major reliability concern in nano-scale VLSI designs. Undesirable voltage droop occurs when logic gates in the circuit draw high switching current from the on-chip power supply network, and this problem is exacerbated at high clock frequencies and smaller technology nodes. A consequence of voltage droop is an increase in path delays and the occurrence of intermittent faults during circuit operation. The addition of conservative timing margins, a.k.a. guardbands, is a common practice to tackle the problem of voltage droop. However, such static and pessimistic guardbands, which are calculated at design time based on worst-case conditions, lead to significant performance loss. Dynamic frequency scaling (DVF) is an alternative approach that enables the dynamic adjustment of clock frequency based on the actual voltage droop seen during runtime. For DVF to be effective, accurate and real-time prediction of voltage droop is essential. We propose a support-vector machine (SVM)-based regression method to predict voltage droop at runtime. Several benchmarks from ITC99 and IWLS'05 highlight the effectiveness of the proposed method in terms of delay-prediction accuracy. Since real-time droop prediction requires hardware implementation of the predictor, we present synthesis results to demonstrate that the hardware overhead for the SVM predictor is negligible for large circuits. Fangming Ye, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
VTS | 2 |
| 2013 | Statistical analysis of BTI in the presence of process-induced voltage and temperature variationsabstractIn nano-scale regime, there are various sources of uncertainty and unpredictability of VLSI designs such as transistor aging mainly due to Bias Temperature Instability (BTI) as well as Process-Voltage-Temperature (PVT) variations. BTI exponentially varies by temperature and the actual supply voltage seen by the transistors within the chip which are functions of leakage power. Leakage power is strongly impacted by PVT and BTI which in turn results in thermal-voltage variations. Hence, neglecting one or some of these aspects can lead to a considerable inaccuracy in the estimated BTI-induced delay degradation. However, a holistic approach to tackle all these issues and their interdependence is missing. In this paper, we develop an analytical model to predict the probability density function and covariance of temperatures and voltage droops of a die in the presence of the BTI and process variation. Based on this model, we propose a statistical method that characterizes the life-time of the circuit affected by BTI in the presence of process-induced temperature-voltage variations. We observe that for benchmark circuits, treating each aspect independently and ignoring their intrinsic interactions results in 16% over-design, translating to unnecessary yield and performance loss. Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori |
ASP-DAC | 1 |
| 2013 | Incorporating the impacts of workload-dependent runtime variations into timing analysisabstractIn the nanometer era, runtime variations due to workload dependent voltage and temperature variations as well as transistor aging introduce remarkable uncertainty and unpredictability to nanoscale VLSI designs. Consideration of short-term and long-term workload-dependent runtime variations at design time and the interdependence of various parameters remain as major challenges. Here, we propose a static timing analysis framework to accurately capture the combined effects of various workload-dependent runtime variations happening at different time scales, making the link between system-level runtime effects and circuit-level design. The proposed framework is fully integrated with existing commercial EDA toolset, making it scalable for very large designs. We observe that for benchmark circuits, treating each aspect independently and ignoring their intrinsic interactions is optimistic and results in considerable underestimation of timing margin. Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori, Sani R. Nassif |
DATE | 1 |
| 2013 | Instruction-set extension under process variation and aging effectsabstractWe propose a novel custom instruction (CI) selection technique for process variation and transistor aging aware instruction-set architecture synthesis. For aggressive clocking, we select CIs based on statistical static timing analysis (SSTA), which achieves efficient speedup during target lifetime while mitigating degradation of timing yield (i.e., probability of satisfying the timing). Furthermore, we consider process variation and aging on not only CIs but also basic instructions (BIs). Even if basic functional units (BFUs), e.g., ALU, get slower due to aging, only a few BIs with critical propagation delay may violate the timing, whereas the other BIs running on the same BFU can still satisfy the timing. We then introduce “customized BFUs”, which execute only such aging-critical BIs. The customized BFUs, used as spare BFUs of the aging-critical BIs, can extend lifetime of the system. Combining the two approaches enables speedup as well as lifetime extension with no or negligibly small area/power overhead. Experiments demonstrate that our work outperforms conventional worst-case work (by an average speedup of about 49%) and existing SSTA-based work (16x or more lifetime extension with comparable speedup). Yuko Hara-Azumi, Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2013 | A layout-aware x-filling approach for dynamic power supply noise reduction in at-speed scan testingabstractPower Supply Noise (PSN) has emerged as an important resilience issue in nano-scale CMOS technology. Due to simultaneous switching of various gates, the actual supply voltage seen by individual gates inside the circuit might be lower than the nominal supply voltage, leading to extra delays. Since in at-speed scan testing simultaneous switchings are higher than the functional mode, test invalidation due to excessive PSN can happen, which may impact yield loss. In this paper, we propose a Linear Programming-based X-filling approach to minimize PSN in at-speed scan test by assigning appropriate values to X-bits in partially specified test patterns. In this paper, spatial and transition time correlations due to circuit layout, power mesh, and netlist are taken into account to increase the accuracy of dynamic PSN estimation and for the first time the delay of the circuit is directly targeted to minimize the effect of PSN during the at-speed scan test. Saman Kiamehr, Farshad Firouzi, Mehdi Baradaran Tahoori |
ETS | 2 |
| 2013 | Representative critical-path selection for aging-induced delay monitoringabstractTransistor aging degrades path delay over time and may eventually induce circuit failure due to timing variations. Therefore, in-field tracking of path delays is essential and to respond to this need, several delay sensor designs have been proposed in the literature. However, due to the significant overhead of these designs and the large number of critical paths in today's IC, it is infeasible to monitor the delay of every critical path in silicon. We present an aging-aware representative path-selection method that allows us to measure the delay of a small set of paths and infer the delay of a larger pool of paths that are likely to fail due to transistor aging. Moreover, since aging is affected by process variations and runtime variations in temperature and voltage, we use machine learning and linear algebra to incorporate these variations during representative path selection. Simulation results for benchmark circuits highlight the accuracy of the proposed approach for predicting critical path delay based on the selected representative paths. Farshad Firouzi, Fangming Ye, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ITC | 1 |
| 2013 | Chip-level modeling and analysis of electrical masking of soft errorsabstractWith continuous downscaling of VLSI technologies, logic cells are becoming more susceptible to radiation-induced soft error. To accurately model this at chip-level, the impact of electrical masking should be accurately considered. Moreover, increasing complexity of VLSI chips at nanoscale results in voltage fluctuation across the chip which impacts the electrical masking. In this paper, we present a chip-level electrical masking analysis which accurately considers the impact of voltage fluctuation across the chip. Our analysis shows that neglecting voltage fluctuation in electrical masking can lead up to 152% inaccuracy in the overall soft error rate. We also present a technique based on backward pulse propagation to reduce the runtime of this analysis. Saman Kiamehr, Mojtaba Ebrahimi, Farshad Firouzi, Mehdi Baradaran Tahoori |
VTS | 3 |
| 2013 | Power-Aware Minimum NBTI Vector Selection Using a Linear Programming ApproachabstractTransistor aging is a major reliability concern for nanoscale CMOS technology that can significantly reduce the operation lifetime of very large-scale integration chips. Negative bias temperature instability (NBTI) is a major contributor to transistor aging that affects pMOS transistors. On the other hand, leakage power is becoming a dominant factor of the total power with successive technology scaling. Since the input combinations applied to a logic core have a significant impact on both NBTI and leakage power, input vector control can be used to optimize both phenomena during idle cycles. In this paper, we present an efficient input vector selection technique based on linear programming for cooptimizing the NBTI-induced delay degradation and leakage power consumption during standby mode. Since the NBTI-induced delay degradation and leakage power are not affected by the input vector in the same direction, we provide a pareto curve based on both phenomena. A suitable point from such a pareto curve is chosen based on circuit conditions and requirements during runtime. Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | NBTI mitigation by optimized NOP assignment and insertionabstractNegative Bias Temperature Instability (NBTI) is a major source of transistor aging in scaled CMOS, resulting in slower devices and shorter lifetime. NBTI is strongly dependent on the input vector. Moreover, a considerable fraction of execution time of an application is spent to execute NOP (No Operation) instructions. Based on these observations, we present a novel NOP assignment to minimize NBTI effect, i.e. maximum NBTI relaxation, on the processors. Our analysis shows that NBTI degradation is more impacted by the source operands rather than instruction opcodes. Given this, we obtain the instruction, along with the operands, with minimal NBTI degradation, to be used as NOP. We also proposed two methods, software-based and hardware-based, to replace the original NOP with this maximum aging reduction NOP. Experimental results based on SPEC2000 applications running on a MIPS processor show that this method can extend the lifetime by 37% in average while the overhead is negligible. Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori |
DATE | 1 |
| 2012 | Input and transistor reordering for NBTI and HCI reduction in complex CMOS gatesabstractAs CMOS feature size scales to the nanometer regime, transistor aging mostly due to Negative Bias Temperature Instability (NBTI) and Hot Carrier Injection (HCI), has emerged as a major reliability concern. Threshold voltage shift causes the circuit to fail, once the post-aging delay exceeds the timing constraint. In this paper, we investigate the stacking effect of transistors on aging and propose a novel input/transistor reordering approach to alleviate the effect of NBTI and HCI during the active mode operation of the circuit. According to the results, the circuit failing due to aging effect is postponed by increasing the operational lifetime for ISCAS benchmarks by 23.6%, in average, while it has a negligible effect on delay, area, and power compared to the original cell input ordering. Saman Kiamehr, Farshad Firouzi, Mehdi Baradaran Tahoori |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Dynamic Soft Error Hardening via Joint Body Biasing and Dynamic Voltage ScalingabstractShrinking feature sizes, reduced voltages, and higher transistor count of nano-scale silicon chips challenge designers in terms of performance, power consumption, and reliability. This paper investigates the effect of simultaneous use of dynamic voltage and frequency scaling (DVFS) and body biasing (BB) on power consumption, reliability, and performance. An analytical model of reliability as a function of body bias voltage, supply voltage, and frequency is proposed. We derive a three dimensional optimization problem by exploiting proposed reliability model in conjunction with power consumption and performance model. The resulting problem is solved using widely-used geometric optimization to identify optimal supply voltage and body bias voltage and then is validated using accurate simulation. Afterwards, it is demonstrated how joint energy-performance-reliability space optimization method can be used in an adaptive reliability-aware power management systems. Finally, we show that combined soft error aware BB and DVFS is capable of improving power consumption about 30% in comparison to reliability-aware DVFS only for the same level of reliability and performance constraints. Farshad Firouzi, Amir Yazdanbakhsh, Hamed Dorosti, Sied Mehdi Fakhraie |
DSD | 1 |
| 2011 | A linear programming approach for minimum NBTI vector selectionabstractTransistor aging is a serious reliability challenge for nanoscale CMOS technology which can significantly reduce the operation lifetime of VLSI chips. Negative Bias Temperature Instability (NBTI) is the major contributor to transistor aging which affect PMOS transistors. The input vectors applied to the logic core has a significant impact on the overall aging of the logic block. In this paper, we present an efficient input vector selection technique based on Linear Programming (LP) to be used for maximum relaxation during the standby phase. We consider an accurate delay model for post-aging critical paths. Our mixed-LP (binary-relaxed) formulation scales well for very large circuits and provides near-optimal solutions. Experimental results and comparison with Monte-Carlo simulations show the speedup (4-5 orders of magnitude) and further optimization (11%) of our approach. Using these input vectors for the standby phase, the aging effect can be postponed by 71% in average. Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Modeling and estimation of power supply noise using linear programmingabstractPower supply noise in nano-scale VLSI is one of the design concerns. Due to switching current of various logic gates, the actual supply voltage seen by different devices fluctuates, causing extra delays and ultimately intermittent faults during operation. Therefore, accurate estimation of worst case scenario, maximum noise and the vectors causing it, is extremely important for design, verification, and manufacturing test steps. In this paper we present a mixed-integer linear programming modeling of power supply noise in digital circuits to obtain fast and accurate solutions. Compared with accurate SPICE simulations of random vectors for a set of benchmark circuits, the proposed approach can achieve 13115× speedup while obtains 2.7% more optimization in average. Farshad Firouzi, Saman Kiamehr, Mehdi Baradaran Tahoori |
ICCAD | 1 |
| 2010 | Instruction reliability analysis for embedded processorsabstractAdvances in silicon technology and shrinking the feature size to nanometer scale make unreliability of nano devices the most important concern of fault-tolerant designs. Soft error analysis has been greatly aided by the concept of architectural vulnerability factor (AVF) and architecturally correct execution (ACE). In this work, we exploit the techniques of AVF analysis to introduce the instruction-level vulnerability metric for software reliability analysis. The proposed metric can be used to make judgments about the reliability of different programs on different processors with regard to architectural and compiler guidelines for improving the processor reliability. Ali Azarpeyvand, Mostafa E. Salehi, Farshad Firouzi, Amir Yazdanbakhsh, Sied Mehdi Fakhraie |
DDECS | 3 |