VLDB 2026 Research / reviewers in the wild / expert
Krishnendu Chakrabarty
dblp:c/KrishnenduChakrabarty
· DBLP profile ↗
802ranked-venue papers
42as first author
184since 2021 · last 2026
0000-0003-4475-6435ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 764 · 42 first-author · 169 since 2021Software engineering, systems software and programming languages · 82 · 20 since 2021Computer networks · 13 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Security and privacy · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise-Agnostic One-Shot Training and Retraining for Robust DNN Inferencing on Analog Compute-in-Memory SystemsabstractAnalog Compute-in-Memory (ACiM) architectures are a promising alternatives to traditional von Neumann-based systems for accelerating deep neural networks (DNNs), as they alleviate the memory bottleneck by performing in-situ matrixvector multiplications. However, the analog nature of computation in ACiM makes DNNs highly susceptible to noise and process variations. To mitigate the effects of analog noise, existing approaches rely on variation-aware or noise-aware training, retraining, or fine-tuning. These methods, however, are not scalable, as they require chip-specific retraining and typically involve separate training runs for different levels of noise tolerance. Moreover, they overlook the inherent fault tolerance of analog-to-digital converters (ADCs). To address these limitations, we propose a one-shot training and retraining strategy for robust DNN inferencing on ACiM platforms. Our method is guided by a detailed analysis of error propagation through ADCs, revealing that robustness can be enhanced by strategically reshaping the weight distribution to better align with ADC resilience characteristics. Simulation results and experimental results with fabricated chips show that the proposed method improves inferencing accuracy by $\mathbf{7 0 \%} \boldsymbol{-} \mathbf{9 0 \%}$ for ResNet-18 and DenseNet-121 under $\mathbf{7 0 \%}$ noise injection on CIFAR-10 and SVHN, and by $\mathbf{5 0 \%}$-80% for VGG-16 under $\mathbf{5 0 \%}$ noise. These gains are achieved with only a $5 \%$ energy overhead due to the modified weight distribution. Ashish Reddy Bommana, Ben Feinberg, T. Patrick Xiao, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty |
ASP-DAC | 6 |
| 2026 | WARP: Workload-Aware Reference Prediction for Reliable Multi-Bit FeFET Readout under Charge-Trapping DegradationabstractFerroelectric FET (FeFET)-based arrays are promising candidates for energy-efficient, high-density non-volatile memory in data-intensive applications. However, charge-trappinginduced degradation and process variations pose significant reliability challenges. These effects lead to reduced memory window and degraded read accuracy over time. We propose a workload-aware degradation modeling and readout framework for FeFET arrays. First, we select a small set of representative workloads to efficiently capture degradation trends across a large workload space. We apply a two-step method to reduce read error: (a) adjust intermediate state currents to widen the separation between states; (b) select optimum reference thresholds based on the shifted distributions. Next, we perform detailed tradeoff analysis involving degradation improvement, on-chip area, and the overhead of a memory-mapped CPU polling system for in-field workload tracking. This is the first work to propose adaptive reference prediction for FeFETs based on runtime workload characteristics. Our framework improves read reliability with minimum hardware overhead and enables scalable in-field monitoring for future FeFET-based systems. Dhruv Thapar, Ashish Reddy Bommana, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ASP-DAC | 5 |
| 2026 | Focus Session: Do Agentic LLMs Change the Paradigm of Hardware Test Generation?abstractTechnology scaling and increasing System-on-Chip (SoC) complexity exacerbate reliability challenges arising from both structural defects and runtime-dependent failures, including Silent Data Corruptions (SDCs) that evade traditional error detection mechanisms. Structural testing remains essential for detecting modeled faults such as stuck-at faults; however, it is inherently limited in capturing failures that arise under dynamic operating conditions. In contrast, functional testing can expose workload-dependent failures, albeit at the cost of high testing overhead and largely unguided workload generation. This paper presents an agentic testing framework that integrates Large Language Models (LLMs) with Reinforcement Learning (RL) and Tree-structured Parzen Estimators (TPE) to guide functional workload generation and Automatic Test Pattern Generation (ATPG) settings under user-defined constraints. The proposed approach leverages feedback-driven optimization to steer test generation toward failure-prone behaviors while reducing reliance on manual expertise. Experimental evaluation on a RISC-V processor core demonstrates that the method outperforms manually generated workloads for functional testing, while experiments on six benchmark circuits show test quality comparable to expert-generated ATPG scripts for structural testing, with improved efficiency and scalability. Farshad Firouzi, Agastya Seth, Peter Domanski, Bahareh J. Farahani, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty |
DATE | 8 |
| 2026 | Reliability, Test, and Security of Compute-In-Memories
Soyed Tuhin Ahmed, Krishnendu Chakrabarty, Jin-Fu Li 0001, Mottaqiallah Taouil, Fouwad Jamil Mir, Said Hamdioui, Mehdi Baradaran Tahoori, Martin Keim, Jongsin Yun |
ETS | 2 |
| 2026 | SOLDER: State-Space Oriented Learning with Denoising-Enhanced Representations for PCB Defect Analysis
Md Fahim Ul Islam, Krishnendu Chakrabarty |
ETS | 2 |
| 2026 | EXACT: Edge-eXplainable Autonomous Causal Telemetry for Silicon Lifecycle Management
Hsiao-Ping Ni, Eduardo Ortega, Krishnendu Chakrabarty |
ETS | 3 |
| 2026 | The Fracture Bits in Large Language Models
Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 3 |
| 2026 | VTS2026 Contest Publication: TTTC's E.J. McCluskey Best Doctoral Thesis Award
Luca Benini, Paolo Bernardi 0002, Alberto Bosio, Swarup Bhunia, Riccardo Cantoro, Degang Chen 0001, Krishnendu Chakrabarty, Jayeeta Chaudhuri, Bastien Deveautour, Gabriele Filipponi, Angelo Garofalo, Salvatore Pappalardo, Sudipta Paria, Michael Rogenmoser, Philippe Sauter, Michael Sekyere |
VTS | 7 |
| 2026 | DART: Dynamic Repair for Interconnect Fault Tolerance in Hybrid Bonding
Partho Bhoumik, Puneet Gupta 0001, Krishnendu Chakrabarty |
VTS | 4 |
| 2026 | MoD-CiM: A Mixture-of-Defenses Framework Against Power-Hammering Attacks in Multi-Tenant Compute-in-Memory
Ashish Reddy Bommana, Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 4 |
| 2026 | LEAD: Link Exploitability Analysis for Die-to-Die Interconnects in Heterogeneous Integration*
Arjun Hati, Eduardo Ortega, Jonti Talukdar, James F. Plusquellic, Krishnendu Chakrabarty |
VTS | 5 |
| 2026 | DefectVICL: Data-Efficient Wafer Defect Classification with Vision In Context Learning
Md Fahim Ul Islam, Soyed Tuhin Ahmed, John M. Carulli Jr., Krishnendu Chakrabarty |
VTS | 4 |
| 2026 | TRACK: Telemetry-based Representation Analysis via Centered Kernel Alignment for Silicon Lifecycle Management
Eduardo Ortega, Jonti Talukdar, Hsiao-Ping Ni, Krishnendu Chakrabarty |
VTS | 4 |
| 2026 | HAT-FI: Hardware-Aware Training for Fault-Tolerant LLM Inference on RRAM-Based Compute-in-Memory
Soyed Tuhin Ahmed, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 4 |
| 2026 | 2.5D/3D Chiplet-based Integration: New Dimensions in Design and Testing
Ganap A. Tewary, Partho Bhoumik, Pragnya S. Nalla, Yu Cao 0001, Krishnendu Chakrabarty, Jeff Zhang 0001 |
VTS | 5 |
| 2026 | NERT: Network- and Routing-aware Testing of Interconnects in Fanout Wafer-Level Packaging
Dhruv Thapar, Partho Bhoumik, Arjun Chaudhuri, Krishnendu Chakrabarty |
VTS | 4 |
| 2026 | WARP-TPG: Warpage-Aware Test Pattern Generation for Small-Delay Defects
Dhruv Thapar, Arjun Chaudhuri, Krishnendu Chakrabarty |
VTS | 3 |
| 2026 | Defending Against Model Inversion Attacks for Biomedical Images via Learnable Data PerturbationabstractThe increasing need for sharing healthcare data and collaborating on clinical research has raised privacy concerns. Health information leakage due to malicious attacks can lead to serious problems such as misdiagnoses and patient identification issues. Privacy-preserving machine learning (PPML) and privacy-enhancing technologies, particularly federated learning (FL), have emerged in recent years as innovative solutions to balance privacy protection with data utility; however, they also suffer from inherent privacy vulnerabilities. Model inversion attacks constitute major threats to data sharing in federated learning. Researchers have proposed many defenses against model inversion attacks. However, current defense methods for healthcare data lack generalizability, i.e., existing solutions may not be applicable to data from a broader range of populations. In addition, most existing defense methods are tested using non-healthcare data, which raises concerns about their applicability to real-world healthcare systems. In this study, we present a defense against model inversion attacks in federated learning. We achieve this using latent data perturbation and minimax optimization, utilizing both general and medical image datasets. We compare our method against two baselines and observe a reduction of at least 4% in the attacker’s accuracy when classifying reconstructed images, while maintaining model utility within a 1.5% drop of client classification accuracy of the undefended model. These results demonstrate improved privacy protection with minimal utility loss and suggest the potential for a generalizable defense in healthcare settings. Shiyi Jiang, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Internet Things J. | 3 |
| 2026 | COMET-3D: Compute-in-Memory-Based Transformer Accelerator With Optimized Pipeline and 3D Heterogeneous IntegrationabstractTransformers have become the backbone of large language decoder and encoder models, but their compute- and memory-intensive nature makes them inefficient on traditional von Neumann architectures. Compute-in-memory (CIM) architectures offer a promising path forward by enablingin situmatrix operations and reducing memory access overhead. However, existing CIM-based accelerators suffer from: (1) unrealistic assumption of single-cycle activation of entire crossbar; (2) inefficient pipelines that are not optimized for operational unit (OU)-based execution; and (3) high analog-to-digital converter (ADC) cost. To address these limitations, we propose a latency-optimized pipeline tailored specifically for OU-based CIM execution and introduce COMET-3D—a 3D heterogeneous architecture that integrates SRAM and ReRAM-based CIM arrays with a logic die containing digital MAC units and softmax modules. The proposed architecture and dataflow maximize hardware utilization and enable efficient acceleration of the multi-head self-attention (MHSA) layer. Experimental results across BERT, GPT2, and LLAMA demonstrate that COMET-3D outperforms baseline architectures with similar compute resources by up to 34× in energy-delay product (EDP) for LLAMA, with gains of 7.4× for GPT2 and 4.3× for BERT-Large. Ashish Reddy Bommana, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | CATCH: A Cost Analysis Tool for Co-Optimization of Chiplet-Based Heterogeneous SystemsabstractWith the increasing prevalence of chiplet systems in high-performance computing applications, the number of design options has increased dramatically. Instead of chips defaulting to a single die per package, now there are viable 3D stacking options to integrate multiple dies in a package either through vertical stacking or horizontal integration on a substrate along with a plethora of choices regarding configurations and processes. For chiplet-based designs, high-impact decisions such as those regarding the number of chiplets, the design partitions, the interconnect types, and other factors must be made early in the development process. In this work, we describe an open-source tool, CATCH, that can be used to guide these early design choices. We also present case studies showing some of the insights we can draw by using this tool. We look at case studies on optimal chip size, defect density, test cost, IO types, assembly processes, and substrates. Additionally, we include the cost breakdown for a specific case based on a prior work. Alexander Graening, Jonti Talukdar, Saptadeep Pal, Krishnendu Chakrabarty, Puneet Gupta 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Path-Driven Washing and Drying Co-Optimization in Continuous-Flow Lab-on-ChipsabstractRapid advances in microfluidics technologies have facilitated the emergence of highly integrated lab-on-a-chip (LoC) biochip systems. With such a coin-sized biochip, complicated bioassay procedures can be executed efficiently without any human intervention. To ensure the correctness of assay outcomes, however, cross-contamination among different fluid samples and reagents needs to be dealt with separately during assay execution. As a consequence, washing operations have to be introduced and a washing path network needs to be established on the chip to remove the residues left behind in flow channels/devices. Also, chip drying after washing operations is crucial for maintaining some properties (e.g., pH values) of the subsequent reagents, so that precision degradation caused by residual buffer fluids can be avoided for those concentration-sensitive assays. To realize optimized assay procedures, we consider both washing operations and chip drying for the first time and propose an integer linear programming (ILP)-based path-driven washing and drying cooptimization method called PathDriver-WD for continuous-flow LoC biochip systems. The proposed method includes the following four key techniques: 1) The necessity of contamination removals and channel drying is analyzed systemically to avoid unnecessary washing and drying operations, 2) washing and drying operations are integrated with the regular removal of excess fluids, so that extra channel occupation can be minimized, 3) practical computation models are adopted to evaluate the durations of different washing and drying operations, and 4) optimized washing/drying paths and time windows are computed and assigned so that the completion time of assays can be minimized. Simulation results on multiple benchmarks demonstrate that the proposed method leads to highly efficient washing and drying procedures as well as minimized assay completion time. Xing Huang 0001, Zhiwen Yu 0001, Bin Guo 0001, Hanbin Ma, Tsung-Yi Ho, Ulf Schlichtmann, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2026 | Dynamic Interposer Obfuscation Through Distributed Scramblers in Heterogeneously Integrated 2.5D ICsabstractRecent breakthroughs in heterogeneous integration (HI) using 2.5D and 3D ICs have been key to advances in the semiconductor industry. However, heterogeneous integration has also led to several sources of distrust due to the use of thirdparty IP, testing, and fabrication facilities in the design and manufacturing process. Recent work on 2.5D IC security has focused on attacks that can be mounted through rogue chiplets integrated in the design. Thus, existing solutions implement interchiplet communication protocols that prevent unauthorized data modification and interruption in a 2.5D system. However, none of the existing solutions offer inherent security against IP theft. We develop a comprehensive threat model for 2.5D systems indicating that such systems remain vulnerable to IP theft. We present a method that prevents IP theft by obfuscating the connectivity of chiplets on the interposer using reconfigurable interconnection networks. We further achieve dynamic obfuscation of chiplet interconnects while demonstrating resilience against removal, data-snooping, test-based data leakage attacks. We present an approach to implement reconfigurable chiplet interconnection networks through both centralized and distributed scrambler designs, evaluating the security and implementation benefits of both the architectures. We present a comprehensive methodology to design, implement, and integrate interconnect scramblers in a large 2.5D HI system. We also evaluate the power, performance, and area overhead for both distributed and centralized scramblers. Jonti Talukdar, Seungmin Woo, Sung Kyu Lim, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Modeling and Analysis of Defects and Variations in Multibit FeFET Devices and Crossbar ArchitecturesabstractFerroelectric field-effect transistors (FeFETs) have many promising applications, but the impact of manufacturing imperfections on these devices has yet to be studied comprehensively. We extend a previous FeFET compact model to combine the Preisach ferroelectric capacitor model with the BSIM-SOI MOSFET model. We calibrate this new compact model with data from a technology CAD model that is calibrated against a fabricated metal-ferroelectric-metal capacitor. We analyze polarization defects in the ferroelectric layer using this compact model. We address two classes of defects and map them to stuck-at-fault models, referred to as neutral faults, and stuck-at-plus and stuck-at-minus faults. We also present an analysis of the impact of device-level opens, shorts, coupling faults and extrinsic variations on the multi-level FeFET device in a 1T-1FeFET cell. Simulation results under a clustered fault distribution with a 2% fault rate for faults show a maximum accuracy degradation of 83.86% and 81.66% for ResNet18 and VGG16, respectively when inferencing is performed on faulty FeFET crossbar designs. Similarly, extrinsic variations and clustered polarization defects show a maximum accuracy degradation of 80% and 75% across both DNN models, respectively. We also evaluate the effectiveness of write-verify for mitigating the impact of extrinsic variations, polarization defects, and BEOL faults. Dhruv Thapar, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | PHANTOM: Power Hammering Attack and Countermeasure on Multi-Tenant ReRAM Compute-in-Memory AcceleratorsabstractThe increasing demand for efficient and low-power deep neural network (DNN) inference has advanced the adoption of ReRAM-based compute-in-memory (CiM) accelerators, which perform computations directly within memory to reduce energy consumption and enhance throughput. However, such architectures are vulnerable to security threats, especially in a multi-tenant environment where multiple users share the same physical resources. This paper introduces a new attack model for multi-tenant ReRAM-based CiM, power hammering, that exploits the temperature sensitivity of ReRAM cells, inducing local temperature increases that lead to conductance drift and ultimately result in erroneous inference outcomes. This serves as a denial-of-service (DoS) attack, where malicious co-tenants degrade inferencing accuracy and system reliability for legitimate users in a shared environment, ultimately undermining trust and causing potential losses to the service provider. Additionally, we propose a novel strategy to counter this security vulnerability. In this technique, we focus on selectively protecting important weights with error compensation hardware. These important weights are treated as faults, and their computation is offloaded to compensation hardware. Simulation results confirm the effectiveness of the proposed method in ensuring accurate classification results even under adversarial conditions, thereby enabling secure multi-tenant inference on ReRAM-based CiM accelerators. Ashish Reddy Bommana, Rajendra Bishnoi, Naghmeh Karimi, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | DiabLLM: An LLM-Based Framework for Blood Glucose Prediction in Type 1 DiabetesabstractAccurate Blood Glucose (BG) prediction is essential for enabling glycemic control in individuals with Type 1 Diabetes Mellitus (T1DM), particularly within Smart and Connected Health (SCH) systems that integrate Continuous Glucose Monitoring (CGM) and automated insulin delivery. The adaptability of Large Language Models (LLMs) provides a promising foundation for unified, fine-tunable forecasting models. We introduce DiabLLM, a framework based on two recent LLM-based architectures: Time-LLM, which incorporates a lightweight projection layer and alignment techniques to transform time-series data into embeddings interpretable by pre-trained LLMs, and Chronos, which employs time-series-aware tokenization and quantization to convert continuous inputs into discrete sequences for forecasting. Both models process 30-minute sequences of six historical BG values and predict 30- and 45-minute horizons. Experimental results on the OhioT1DM and D1NAMO datasets demonstrate that DiabLLM outperforms state-of-the-art baselines, including a Deep Reinforcement Learning model and an ensemble of LSTM, GRU, and WaveNet, achieving up to 27% improvement in RMSE and 37% in MAE. To enhance robustness to noisy and missing input data, a denoising autoencoder was employed for input reconstruction, yielding improved predictive performance. In addition, knowledge distillation was shown to significantly compress the model, making it a practical candidate for efficient deployment on resource-constrained edge devices without compromising accuracy. Amirhossein Mahmoudi, Ghazal Farahani, Peter Domanski, Bahareh J. Farahani, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Runtime Fault Localization in Deep Neural Network AcceleratorsabstractSystolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Although fault detection and repair techniques have been proposed to enhance the robustness of systolic arrays, fault localization remains an open problem. We propose a fault tolerance framework including run-time based fault detection and fault localization, both leveraging functional data to generate checksums on-the-fly. This approach enables error detection and localization during normal operation, avoiding the need for dedicated test patterns or additional downtime. Experimental evaluation shows that the proposed fault localization architecture incurs an area overhead less than 2% for a 256× 256 systolic array. In simulations, the proposed method achieves 100% fault detection and localization in a 256× 256 systolic array. Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2026 | Design Automation Techniques for Microfluidic Fully Programmable Valve Array Biochips: A Systematic SurveyabstractFlow-based microfluidic biochips have attracted much attention over the past two decades. By integrating diverse micro-components, e.g., mixers and filters, on a miniaturized planar substrate, complicated bioassays such as protein crystallization and drug screening can be executed automatically without requiring human invention, thus becoming a promising alternative to traditional cumbersome laboratory equipment. As manufacturing technology advances, it has become possible to implement hundreds of thousands of microvalves within a single chip. This breakthrough has given rise to fully programmable valve array (FPVA) biochips, representing a next-generation platform in flow-based microfluidics that offers enhanced reconfigurability and operational flexibility. Nevertheless, the exponential increase in valve density has introduced significant design complexity when implementing sophisticated assay protocols. As a result, the design automation of FPVAs has emerged as a critical research frontier, attracting considerable attention from both academia and industry. This review article systematically examines recent advances in FPVA design automation, involving computer-aided design methods for architectural synthesis, volume management, sample preparation, automated testing, fault localization, error recovery, and washing optimization. These techniques enable FPVA users to concentrate on assay protocol development while delegating implementation-specific design and optimization tasks to design automation tools. Furthermore, we analyze emerging security implications in FPVAs, particularly focusing on bioassay accuracy and reliability that ensure experimental reproducibility. Finally, potential trajectories for future research are discussed in detail to further promote the integration level and widespread application of FPVAs. Shuang Qi, Zhiwen Yu 0001, Bin Guo 0001, Sizhao Li, Hanbin Ma, Tsung-Yi Ho, Krishnendu Chakrabarty, Xing Huang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2026 | Defect-Aware Built-In Self-Test and Dynamic Repair for Fan-Out Wafer-Level PackagingabstractFan-out wafer-level packaging (FOWLP) addresses the demand for higher interconnect densities by offering reduced form factor, improved signal integrity, and enhanced performance. However, FOWLP faces several manufacturing challenges, such as coefficient of thermal expansion (CTE) mismatch, warpage, die shift, and postmolding protrusion, causing misalignment and bonding issues during redistribution layer (RDL) buildup. In order to address these challenges, we propose a comprehensive defect analysis and testing framework for FOWLP interconnects. We use Ansys Q3D to map defects to equivalent electrical circuit models and perform fault simulations to investigate the impacts of these defects on chiplet functionality. Additionally, we present a built-in self-test (BIST) architecture to detect stuck-at (SA) and bridging faults while accurately diagnosing the fault type and location. We also show how fault signatures can be propagated and decoded across multichiplet interfaces to initiate dynamic repair. To enhance postfabrication yield and reliability, we introduce a priority-based dynamic repair scheme aligned with the universal chiplet interconnect express (UCIe) specification that maximizes repairability without incurring additional hardware overhead in the I/O PHY. This framework offers a practical path toward robust and efficient testability in the next-generation advanced packaging. Partho Bhoumik, Chris Bailey 0001, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | TIDE-S: Telemetry Informed Delay Testing With Optimized Sensor PlacementabstractSilent data corruption (SDC) refers to undetected errors that yield incorrect results without triggering system alerts or error logs. Existing test methodologies are inadequate for capturing dynamic voltage fluctuations that occur under realistic workload conditions, thereby limiting their effectiveness for detecting SDCs. We present TIDE-S, a telemetry-informed delay testing (TIDE) framework that integrates presilicon sensor placement strategies to improve telemetry accuracy. By evaluating different sensor allocation schemes—uniform,$K$-means, and energy-aware clustering—TIDE-S improves the spatial granularity of voltage observation, enabling more accurate correlation between voltage fluctuations and path delay behavior. The combined telemetry- and sensor-aware methodology significantly improves the detection of timing-sensitive SDCs. The proposed framework incurs minimal infrastructure overhead by leveraging standard pad-based voltage observation, making it practical for real system-on-chip (SoC) designs. We demonstrate the effectiveness of TIDE-S across multiple RISC-V-based SoCs and a diverse set of real-world benchmarks, showing quantifiable improvements in voltage estimation, slack prediction, and test coverage. Deepesh Sahoo, Eduardo Ortega, Peter Domanski, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | The Unlikely Hero: Nonidealities in Analog Photonic Neural Networks as Built-in Adversarial DefendersabstractElectronic-photonic computing systems have emerged as a promising platform for accelerating deep neural network (DNN) workloads. Major efforts have been focused on countering hardware non-idealities and boosting efficiency with various hardware/algorithm co-design methods. However, the adversarial robustness of such photonic analog mixed-signal AI hardware remains unexplored. Though the hardware variations can be mitigated with robustness-driven optimization methods, malicious attacks on the hardware show distinct behaviors from noises, which requires a customized protection method tailored to optical hardware. In this work, we rethink the role of conventionally undesired non-idealities in photonic accelerators and claim their surprising effects on defending against weight attacks. Inspired by the protection effects from DNN quantization and pruning, we propose a synergistic defense framework tailored for optical AI hardware that proactively protects sensitive weights via pre-attack unary weight encoding and post-attack vulnerability-aware weight locking. Efficiency-reliability trade-offs are formulated as constrained optimization problems and efficiently solved offline without model re-training costs. Extensive evaluation of various DNN benchmarks with a multi-core photonic accelerator shows that our framework maintains near-ideal inference accuracy under adversarial bit-flip attacks with merely <3% memory overhead. Our codes are open-sourced at link. Haotian Lu 0002, Ziang Yin, Partho Bhoumik, Sanmitra Banerjee, Krishnendu Chakrabarty, Jiaqi Gu 0002 |
ASP-DAC | 5 |
| 2025 | AGILE: A Multi-Task Contrastive Learning Framework with Adversarial Gradient Iterative Learning for Bio-Signal Anonymization
Tamonash Bhattacharyya, Farshad Firouzi, Amir-Mohammad Rahmani, Sanaz R. Mousavi, Krishnendu Chakrabarty |
BSN | 5 |
| 2025 | Identifying System-on-Chip Security Assets with Structure-Based AnalysisabstractIn the evolving field of hardware design, ensuring the security of System-on-Chips (SoCs) has become increasingly vital. As SoCs grow in complexity, integrating components from various sources, the identification and protection of security assets are crucial to prevent vulnerabilities. Traditional methods of identifying these assets are manual and time-intensive. To address this challenge, automated tools for security asset identification are essential, enabling faster and more accurate detection of critical assets early in the design process. In this paper, we propose a framework for the automated identification of security assets within SoCs. By transforming register-transfer level (RTL) code into graphs and leveraging deep neural networks (DNNs) to classify assets based on their structural patterns, our approach can effectively differentiate between security and non-security assets. Experimental results show that the proposed method achieves high classification accuracy, with the model reaching up to 99% accuracy in identifying security assets, significantly reducing the need for manual intervention. Wei-Kai Liu, Benjamin Tan 0001, Krishnendu Chakrabarty |
DAC | 3 |
| 2025 | DEAR: Dependable 3D Architecture for Robust DNN TrainingabstractReRAM-based compute-in-memory (CiM) architectures present an attractive design choice for accelerating deep neural network (DNN) training. However, these architectures are susceptible to stuck-at faults (SAFs) in ReRAM cells, which arise from manufacturing defects and cell wearout over time, particularly due to the continuous weight updates during DNN training. These faults significantly degrade accuracy and compromise dependability. To address this issue, we propose DEAR: dependable 3D architecture for robust DNN training. DEAR introduces a novel online compensation method that employs a digital compensation unit to correct SAF-induced errors dynamically during both forward and backward propagation. This approach mitigates errors induced by SAFs during both the forward and backward phases of DNN training. Additionally, DEAR leverages an HBM-based 3D memory structure to store fault-related error information efficiently. Experimental results show that DEAR limits inferencing accuracy loss to under 2% even when up to 10% of cells are faulty with uniformly distributed faults, and under 2% for up to 5% faulty cells in clustered distributions. This high fault tolerance is achieved with an area overhead of 11.5% and energy overhead of less than 6% for VGG networks and less than 12% for ResNet networks. Ashish Reddy Bommana, Farshad Firouzi, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 7 |
| 2025 | Runtime Security Analysis of Monolithic 3D Embedded DRAM with Oxide-Channel TransistorabstractWe present the first security and disturbance study of monolithic 3D (M3D) embedded DRAM (eDRAM) with 2T gain cell using oxide-channel transistors. We explore the Rowhammer/Rowpress vulnerabilities on amorphous indium tungsten oxide (IWO) transistors for eDRAM with standalone 2D integration and memory-on-memory M3D integration. In addition, We examine M3D-specific electrical disturbances from memory-on-logic M3D integration. We evaluate IWO eDRAM's susceptibility to these vulnerabilities/disturbances and discuss the potential impact on M3D integration. We examine physical design and architecture strategies for M3D integration of IWO eDRAM. We provide systematic recommendations to inform security strategies for M3D integration and security of IWO eDRAM. Our results show that limiting the minimum vertical interlayer distance to 300 nm reduces vertical disturbances in memory-on-memory M3D integration. In addition, for memory-on-logic M3D integration, we observed that IWO eDRAM's read bitline is sensitive to crosstalk from high-speed switching logic circuits. In conjunction, we show that IWO eDRAM standalone 2D integration is 30× more resilient to Rowhammer than current state-of-the-art memory because the IWO transistor's$I_{ON}/I_{OFF}$ratio is roughly three orders of magnitude greater than standard memory access transistors. Eduardo Ortega, Jungyoun Kwak, Shimeng Yu, Krishnendu Chakrabarty |
DATE | 4 |
| 2025 | On the Impact of Warpage on BEOL Geometry and Path Delays in Fan-out Wafer-Level PackagingabstractWarpage is a major concern in fan-out wafer-level packaging (FOWLP) due to the complex thermal processing steps involved in manufacturing. These steps include curing, electroplating, and deposition, which induce residual stresses through differential thermal expansion and contraction of materials. This effect is further amplified by mismatches in the coefficients of thermal expansion (CTE) between different materials. In particular, high-density interconnects in the back-end of line (BEOL), redistribution layers (RDLs), and through-mold vias (TMVs) are susceptible to warpage-induced stress, strain, and deformation. This work conducts structural simulations to analyze warpage in the BEOL stack induced by FOWLP. Our results indicate that the impact of warpage is non-uniform across the entire BEOL geometry of a die, hence it impacts different metal layers differently, and different coordinates within one metal layer differently. We leverage this warpage analysis to calculate parasitics and evaluate the resulting changes in path delays. Dhruv Thapar, Arjun Chaudhuri, Ravi Mahajan, Krishnendu Chakrabarty |
DATE | 5 |
| 2025 | FLARE: Fault Attack Leveraging Address Reconfiguration Exploits in Multi-Tenant FPGAs
Jayeeta Chaudhuri, Hassan Nassar, Dennis Gnad, Jörg Henkel, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty |
ETS | 6 |
| 2025 | Prompt, Fab, Flex: Agentic LLMs for Flexible Electronics DesignabstractFlexible Electronics (FE) have emerged as a promising platform for extreme edge applications that demand attributes tailored to the application domain, such as ultra-low cost, low power consumption, mechanical flexibility, biocompatibility, and environmental sustainability. While advances in printed and flexible device technologies have demonstrated the feasibility of sensing, computing, and communication on deformable substrates, the design and implementation of FE-based systems remain limited by traditional Electronic Design Automation (EDA) workflows, which are complex, time-intensive, and largely inaccessible to non-experts. In parallel, recent progress in Large Language Models (LLMs) has enabled automation across multiple stages of integrated circuit design; however, existing approaches exclusively target conventional silicon technologies and are not designed to address the unique constraints of FE. This work introduces the first LLM-driven framework for end-to-end hardware design automation in flexible electronics. The proposed methodology supports Register-Transfer Level (RTL) generation, logic synthesis, and cross-layer Power–Performance–Area (PPA) Design Space Exploration (DSE) for bespoke Machine Learning (ML) classifiers. Experimental results demonstrate the feasibility and effectiveness of the approach in generating resource-efficient hardware designs optimized for FE, thereby lowering barriers to adoption and accelerating the development of personalized, application-specific FEs. Farshad Firouzi, Bahareh J. Farahani, Polykarpos Vergos, Deepesh Sahoo, Nathaniel Bleier, Krishnendu Chakrabarty |
ICCAD | 6 |
| 2025 | STAMP-2.5D: Structural and Thermal Aware Methodology for Placement in 2.5D IntegrationabstractChiplet-based architectures and advanced packaging have emerged as transformative approaches in semiconductor design. While conventional physical design for 2.5D heterogeneous systems typically prioritizes wirelength reduction through tight chiplet packing, this strategy introduces thermal bottlenecks and intensifies coefficient of thermal expansion (CTE) mismatches, compromising long-term reliability. Addressing these challenges requires holistic consideration of thermal performance, mechanical stress, and interconnect efficiency. We introduce STAMP-2.5D: Structural and Thermal Aware Methodology for Placement in 2.5D integration, the first automated placement methodology that simultaneously optimizes these critical factors. Our approach employs finite element analysis to simulate temperature distributions and stress profiles across chiplet configurations while minimizing interconnect wirelength. Experimental results demonstrate that a thermal and structurally aware automated placement approach reduces overall stress by 11 %, maintains excellent thermal performance with a negligible$\mathbf{0. 5 \%}$temperature increase and simultaneously reduces total wirelength by 11 % compared to temperature-only optimization. Additionally, we conduct an exploratory study on the effects of temperature gradients on structural integrity, providing crucial insights for reliability-conscious chiplet design. Varun Darshana Parekh, Zachary Wyatt Hazenstab, Srivatsa Rangachar Srinivasa, Krishnendu Chakrabarty, Kai Ni 0004, Narayanan Vijaykrishnan |
ICCD | 4 |
| 2025 | Taming Sparse Giants: Deploying Mixture-of-Experts on 3D Heterogeneous Compute-in-Memory SystemsabstractThe deployment of large Mixture-of-Experts (MoE) models on 3D heterogeneous integrated (3D-HI) Compute-in-Memory (CiM) architectures presents unique challenges, requiring joint optimization of area, energy, latency, and perplexity (PPL). We first introduce a detailed 3D-stacked CiM architecture model, incorporating both SRAM and ReRAM tiers with thermal and device-level considerations. Building on this foundation, we present OPTIMEX, a multi-objective optimization framework that efficiently maps MoE expert projections onto heterogeneous tiers. Our evaluation demonstrates substantial benefits: up to$\text{6 0. 9 \%}$area and$\text{5 4. 7 \%}$energy reduction versus an all-SRAM baseline, while lowering PPL by as much as 98.4% compared to all-ReRAM configurations. Furthermore, OPTIMEX outperforms common heuristics, delivering improvements of up to 73.1% in area, 67.0% in energy, 96.1% in PPL, and 74.9% in latency. Together, these contributions highlight a path toward scalable, energy-efficient, and reliable MoE deployment on advanced CiM platforms. Ashish Reddy Bommana, Farshad Firouzi, Krishnendu Chakrabarty |
ICCD | 4 |
| 2025 | MALLS: Multi-Agent LLMs for Synthetic Hardware Vulnerability Generation and DetectionabstractLLMs have demonstrated promising capabilities in generating RTL code from high-level functional descriptions of hardware modules. However, their effectiveness is constrained by the lack of high-quality, diverse datasets particularly for applications in IP design, verification, and security analysis. To address this limitation, we introduce MALLS, a multi-agent framework in which specialized LLM agents namely, a generator and a discriminator collaborate in an adversarial yet cooperative setting to improve the quality and correctness of RTL designs and curate a high quality synthetic hardware vulnerability dataset. In this architecture, the generator agent is responsible for producing RTL implementations from initial seed examples through in-context learning, while the discriminator agent assesses the generator's output for functional correctness and the presence of security vulnerabilities. This interaction creates a dynamic feedback loop, enabling both agents to iteratively improve through each other's responses leading to self-supervised learning. A hard bank of examples is maintained in a database which include instances that were difficult to generate or detect by either agents. By generating paired positive (correct) and negative (buggy) examples, the system learns to distinguish subtle design flaws and generalize across diverse RTL patterns while generating high quality synthetic examples of hardware vulnerabilities. Experimental results show that using the hard bank of examples produced by the adversarial multi-LLM setup improves both vulnerability generation and detection performance. Jonti Talukdar, Agastya Seth, Sanmitra Banerjee, Farshad Firouzi, Krishnendu Chakrabarty |
ICCD | 5 |
| 2025 | Exploration of Low-Power Flexible Stress Monitoring Classifiers for Conformal WearablesabstractConventional stress monitoring relies on episodic, symptom-focused interventions, missing the need for continuous, accessible, and cost-efficient solutions. State-of-the-art approaches use rigid, silicon-based wearables, which, though capable of multitasking, are not optimized for lightweight, flexible wear, limiting their practicality for continuous monitoring. In contrast, flexible electronics (FE) offer flexibility and low manufacturing costs, enabling real-time stress monitoring circuits. However, implementing complex circuits like machine learning (ML) classifiers in FE is challenging due to integration and power constraints. Previous research has explored flexible biosensors and ADCs, but classifier design for stress detection remains underexplored. This work presents the first comprehensive design space exploration of low-power, flexible stress classifiers. We cover various ML classifiers, feature selection, and neural simplification algorithms, with over 1200 flexible classifiers. To optimize hardware efficiency, fully customized circuits with low-precision arithmetic are designed in each case. Our exploration provides insights into designing real-time stress classifiers that offer higher accuracy than current methods, while being low-cost, conformable, and ensuring low power and compact size. Florentia Afentaki, Sri Sai Rakesh Nakkilla, Konstantinos Balaskas, Paula L. Duarte, Shiyi Jiang, Georgios Zervakis 0001, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ISLPED | 8 |
| 2025 | Fault Tolerance in RRAM-based AI Accelerator with Guided Randomized ActivationabstractResistive Random Access Memory (RRAM)-based analog in-memory computing (IMC) AI accelerators offer significant advantages over digital accelerators, including lower power consumption, reduced data movement, and higher computational efficiency. However, their deployment in safety-critical and edge applications is challenging due to their hardware non-idealities, such as programming error, conductance drift, and read noise, which degrade the inferencing accuracy of the implemented neural networks (NNs). Existing methods, including noise injection during training and activation function modifications, provide limited fault-tolerance in realistic scenarios with non-idealities. We propose a fault-tolerant activation function with architectural optimization that enhances robustness against hardware-induced variations with minimal hardware and NN architectural changes. During training, the proposed activation function features a stochastic negative region, which inherently injects noise into the negative region of the activation. During inferencing, the proposed activation function operates deterministically, ensuring compatibility with existing hardware while maintaining computational efficiency. Extensive evaluations with benchmark datasets demonstrate that the proposed approach significantly improves inferencing accuracy by up to 60% under varying noise levels, outperforming conventional activation functions as well as existing fault-tolerant activation functions. By enhancing fault-tolerance to hardware-induced errors, the proposed method enables reliable and energy-efficient RRAM-based analog IMC. Soyed Tuhin Ahmed, Eduardo Ortega, Ryan Depsey, T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty |
ITC | 8 |
| 2025 | Fault Modeling and Testing of Chiplet-to-Chiplet Interconnects in Fan-out Wafer-Level Packaging*abstractAdvanced packaging technologies are reshaping; semiconductor integration, offering unprecedented improvements in performance, efficiency, and miniaturization. Among these Fan-Out Wafer-Level Packaging (FOWLP) has emerged as a transformative solution, pushing the limits of chiplet integration through its superior electrical performance, reduced footprint, and enhanced thermal management However, FOWLP presents several challenges, including coefficient of thermal expansion mismatch, warpage, die shift, and post-molding protrusion, all of which can lead to misalignment and defects. Moreover, thermo-mechanical stresses in the organic package induce warpage and delamination, further exacerbating weak defects during runtime operation. To address these challenges, we propose a comprehensive defect analysis and testing framework for FOWLP interconnects. Defects are mapped to equivalent electrical circuit models, allowing for precise fault characterization. A built-in self-test (BIST) architecture is introduced to detect open and bridging faults while accurately diagnosing fault types and localizing them. An embedded ring oscillator within the BIST network detects weak opens, bridging and coupling faults, quantifying the size of the defects. We demonstrate how this test framework can be utilized at different supply voltage corners and for various transistor sizes to reduce fault-effect aliasing due to process variations, thereby ensuring robust diagnostics, yield learning, and silicon lifecycle management. The effectiveness of the ring oscillator circuit is validated through HSPICE simulations using equivalent faulty circuit models for a 7 nm CMOS technology. Partho Bhoumik, Arjun Chaudhuri, Sandeep Kumar Goel, Krishnendu Chakrabarty |
ITC | 4 |
| 2025 | SMART: Scalable and Modular Architecture for Routing-Aware Testing of Fan-out Wafer-Level Packages*abstractFan-out wafer-level packaging enables heterogeneous chiplet integration via Cu pillars and redistribution layers (RDLs). As FOWLP technology evolves, the focus is shifting towards many-chiplet designs, necessitating multi-layer RDL structures to route interconnects between these chiplets. However, defects such as opens, shorts, and coupling are a challenge for RDL structures. High-density, multi-layer RDLs exacerbate these challenges, leading to intensified coupling, elevated switching activity, and shorts within metal segments. Moreover, the finer pitch of RDLs increases electromigration due to rising current densities. We propose a routing-aware testing framework that leverages the multi-layer RDL routing Information to target realistic shorts and coupling defects. By physically partitioning interconnects into regions, we enable test scheduling and leverage shared test-pattern generators for launching test patterns and capturing responses to reduce test time and area overhead without compromising fault coverage. The framework’s effectiveness is demonstrated on four many-chiplet package designs with varying configurations of chiplet-to-chiplet connectivity. Our results show that this method can achieve over 99.8 % fault coverage with an n-fold reduction in test area and test time through an n-way partitioning strategy. Partho Bhoumik, Dhruv Thapar, Arjun Chaudhuri, Krishnendu Chakrabarty |
ITC | 4 |
| 2025 | LLM-Aided In-Field Workload Generation for Detecting Silent Data Corruptions at ScaleabstractComputational integrity is crucial in large-scale data centers where Silent Data Corruptions (SDCs) pose a growing reliability challenge. SDCs can lead to incorrect computation results not captured by traditional error detection mechanisms, making their detection and mitigation essential. However, existing post-manufacturing and in-field testing methods, such as opportunistic and ripple testing, face significant scalability challenges due to high computational costs and test times. We propose an LLM-aided approach for generating targeted test cases to detect SDCs. As a case study, we focus on the functional blocks of a RISC-V CV32E40P processor core. Our method generates targeted test cases that maximize voltage droops in given hardware modules, such as functional units, increasing the likelihood of triggering SDCs in-field. Additionally, our approach is architecture-aware and layout-aware, enhancing fault activation and enabling automated optimization of generated test cases. Experimental evaluations demonstrate that the proposed method significantly improves SDC detection efficiency by reducing the number of required test cases while preserving high test coverage. Compared to commonly used test cases, the proposed approach increases average voltage droops by up to 38%. By integrating LLM-aided test case generation, the proposed approach achieves voltage droops of up to 9% relative to the supply voltage, improving the effectiveness of in-field SDC detection and mitigation strategies. Peter Domanski, Deepesh Sahoo, Eduardo Ortega, Farshad Firouzi, Krishnendu Chakrabarty |
ITC | 5 |
| 2025 | OCTANE: On-Chip Telemetry-based Anomaly Notification EngineabstractSilicon lifecycle management (SLM) is essential for ensuring the reliability and quality of silicon products. Traditional approaches primarily rely on off-chip solutions to detect malware, diagnose hardware bugs, and characterize silicon health metrics. However, these methods do not incorporate hardware/software co-design for SLM. This work introduces On- Chip Telemetry-based Anomaly Notification Engine (OCTANE), designed to monitor chip status using performance counters and sensors. OCTANE features a compute- and memory-efficient, unsupervised anomaly detection mechanism implemented on-chip (OCTANE-edge) using fixed-point arithmetic. Furthermore, it enhances on-chip anomaly detection through unsupervised feature ranking based on telemetry feature information entropy and compression index. This unsupervised feature ranking technique is workload-independent and provides the generalizability required for SLM. The proposed solution extends to an end-to-end anomaly-informed diagnosis model that leverages OCTANE-edge compacted anomaly telemetry signatures to diagnose chip security or safety incidents (OCTANE-cloud). All telemetry data is collected via model-specific register space using open-source Linux tools and the performance counter monitor. To validate our approach, we capture chip telemetry signatures from the PAMPAR benchmark suite under anomaly-inducing events such as security attacks (e.g., Rowhammer and Spectre) and voltage droops across two Intel platforms. OCTANE demonstrates highly effective unsupervised on-chip anomaly detection with minimal area and idle power overhead (1.2% and 2.6%). OCTANE provides anomaly detection/diagnosis with accuracy surpassing 0.96/0.98. Eduardo Ortega, Arjun Hati, Jonti Talukdar, Woohyun Paik, Rita Chattopadhyay, Krishnendu Chakrabarty |
ITC | 7 |
| 2025 | TIDE: Telemetry-Informed Delay Testing for Silent Data Corruption *abstractSilent Data Corruption (SDC) is caused by undetected errors that yield incorrect results without triggering system alerts or error logs. Existing test methodologies are inadequate for capturing dynamic voltage fluctuations that occur under realistic workload conditions, thereby limiting their effectiveness for detecting SDCs. To address these limitations, we introduce Telemetry-Informed Delay Testing (TIDE), a novel methodology that enhances SDC detection by leveraging telemetry sensors to monitor voltage fluctuations and their impact on timing integrity. By incorporating dynamic, workload-aware test generation, the proposed framework overcomes key limitations of traditional approaches and facilitates early detection of SDCs. The effectiveness of TIDE is demonstrated through case studies conducted on two RISC-V-based SoCs and multiple workloads. Deepesh Sahoo, Eduardo Ortega, Peter Domanski, Farshad Firouzi, Krishnendu Chakrabarty |
ITC | 5 |
| 2025 | NeuralTPG: GPU-Accelerated Neural Twin-Based Test Pattern Generation for Transition Delay Faults in Safety-Critical ApplicationsabstractSafety-critical applications such as autonomous driving demand rigorous functional safety assurance. We present a safety-guided test pattern generation framework called NeuralTPG for transition faults in integrated circuits (ICs) based on Launch-on-Capture (LOC) delay testing. We model the logical state transition behavior of standard cells using multilayer perceptrons (MLPs), referred to as Cell-Nets. The neural twin is constructed by converting standard cell instances into Cell-Nets and replacing inter-cell wires with neural connections. We leverage the end-to-end differentiability of the neural twin to compute input test-pattern pairs for transition faults through back-propagation. The neural twin enhances fault propagation to primary outputs (POs) and generates test-pattern pairs that maximize faults’ propagation capability. The framework supports non-binary, user-defined criticality-factor (CF) assignment across the circuit’s internal nets and POs, enabling CF-guided test-pattern pair generation to propagate transition faults to more critical POs. NeuralTPG employs concurrent test generation, fully leveraging GPU acceleration to generate test-pattern pairs for all transition faults simultaneously, thereby improving test efficiency. Experimental evaluations on six benchmarks across three CF configurations demonstrate that NeuralTPG can be used to achieve safety-guided fault propagation. Xuanyi Tan, Gitanjali Mukherjee, Dhruv Thapar, Arjun Chaudhuri, Sanmitra Banerjee, Rubin A. Parekhji, Krishnendu Chakrabarty |
ITC | 7 |
| 2025 | FeTest: Defect Analysis and March Test Solution for FeFETs *abstractFerroelectric Field-Effect Transistors (FeFETs) are emerging as promising candidates for non-volatile memory and in-memory computing due to their low power consumption, fast switching speeds, and high integration density. However, device-level defects and process variations pose significant challenges to their reliability, particularly for multi-level cell (MLC) FeFETs. Polarization defects in the ferroelectric layer and back-end-of-line (BEOL) defects lead to memory window reduction and erroneous computations in FeFET-based crossbars. We propose a March test solution tailored for MLC FeFETs that detects and diagnoses BEOL defects in the presence of underlying process variations and polarization defects. Dhruv Thapar, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ITC | 4 |
| 2025 | Defect Analysis and Built-In-Self-Test for Chiplet Interconnects in Fan-out Wafer-Level Packaging*abstractFan-out wafer-level packaging (FOWLP) addresses the demand for higher interconnect densities by offering reduced form factor, improved signal integrity, and enhanced performance. However, FOWLP faces manufacturing challenges such as coefficient of thermal expansion (CTE) mismatch, warpage, die shift, and post-molding protrusion, causing misalignment and bonding issues during redistribution layer (RDL) buildup. Moreover, the organic nature of the package exposes it to severe thermo-mechanical stresses during fabrication and operation. In order to address these challenges, we propose a comprehensive defect analysis and testing framework for FOWLP interconnects. We use Ansys Q3D to map defects to equivalent electrical circuit models and perform fault simulations to investigate the impacts of these defects on chiplet functionality. Additionally, we present a built-in self-test (BIST) architecture to detect stuck-at and bridging faults while accurately diagnosing the fault type and location. Our simulation results demonstrate the efficacy of the proposed BIST solution and provide critical insights for optimizing design decisions in packages, balancing fault detection and diagnosis with the cost of testability insertion. Partho Bhoumik, Krishnendu Chakrabarty |
VTS | 3 |
| 2025 | Silent Data Corruption: Advancing Detection, Diagnosis, and Mitigation StrategiesabstractSilent Data Corruptions (SDCs) pose a critical challenge to computer system reliability, arising from vulnerabilities across different layers of the computing stack. This paper addresses this challenge through three complementary contributions that systematically target SDCs from hardware manufacturing to application-level resilience. First, we analyze timing failures caused by random process variations in advanced technology nodes, revealing that extreme slow paths at lower voltages are dominated by single weak transistors—insights crucial for manufacturing and in-field testing. Second, we introduce an LLM-driven framework that generates targeted functional test programs to induce SDCs, demonstrating its effectiveness in stressing hardware, uncovering latent vulnerabilities, and increasing energy consumption in a given device under test (DUT), making it a valuable tool for in-field testing. Third, as machine learning continues to drive advancements across critical domains such as healthcare, finance, and autonomous systems, ensuring its reliability is paramount. However, the susceptibility of these applications to SDCs threatens their reliability and robustness. To address this, we propose Fidelity-Q, a novel fault injection methodology to evaluate the impact of SDCs on Quantized Neural Networks (QNNs), showing that lower-bit quantization increases error susceptibility. Collectively, these contributions provide a comprehensive approach to identifying, analyzing, and mitigating SDCs across the computing stack, from hardware testing to machine learning applications. Peter Domanski, Mukarram Ali Faridi, Gabriel Kaunang, Wilson Pradeep, Adit D. Singh, Alfian Amrizal, Yanjing Li, Farshad Firouzi, Krishnendu Chakrabarty |
VTS | 9 |
| 2025 | ChipMnd: LLMs for Agile Chip DesignabstractThe increasing complexity of semiconductor design, along with stringent performance, power, and time-to-market requirements, has outpaced the capabilities of traditional Electronic Design Automation (EDA) methodologies. Conventional design workflows rely on manual intervention for critical tasks such as hardware description, synthesis optimization, and verification, leading to inefficiencies and scalability limitations. Large Language Models (LLMs) present a transformative approach by automating key stages of the design pipeline, enabling intelligent synthesis tuning, test generation, and security analysis. This paper introduces ChipMind, an LLM-driven framework comprising specialized agents and modules for digital and analog chip design. ChipMind integrates AI-driven methodologies to enhance design efficiency, accelerate prototyping, and optimize key design trade-offs, thereby addressing fundamental challenges in modern semiconductor development. Farshad Firouzi, David Z. Pan, Jiaqi Gu 0002, Bahareh J. Farahani, Jayeeta Chaudhuri, Ziang Yin, Pingchuan Ma 0012, Peter Domanski, Krishnendu Chakrabarty |
VTS | 9 |
| 2025 | Discretized-Isolation Forest: Memory- and Compute-Efficient Unsupervised Anomaly Detection for Resource-Constrained Internet of Things Edge DevicesabstractMemory and compute constraints make anomaly detection model training infeasible on Internet of Things (IoT) resource-limited edge devices. Many solutions train anomaly detection models (e.g., deep neural networks or DNNs) on the cloud and deploy them on IoT edge devices for inferencing. However, cloud-based training does not address overall communication latency and potential data leakage. Moreover, because anomalies rarely occur, using labels to define anomalies is impractical. Hence, supervised learning mechanisms are unsuitable for anomaly detection. There is a need for effective unsupervised anomaly detection for resource-constrained edge devices. We present the discretized isolation Forest (DIF) to address memory- and compute-efficient unsupervised anomaly detection for resource-constrained edge devices. We also present a discretization function, based on information entropy, to inform the growth of the isolation Forest (IF) ensemble to create the DIF. The DIF reduces the training time (memory usage) of the original IF by$79.38\times (166.66\times)$. We test the DIF against general anomaly detection benchmarks and edge-anomaly detection benchmarks. The edge-anomaly detection benchmarks were curated from the built-in iPhone edge sensors. Across all the edge-sensor anomaly detection datasets and against all the other considered models, the discretized isolation resulted in the lowest training time, lowest memory usage, preserved anomaly detection performance, and highest inference speeds. In addition, across all general anomaly detection benchmarks and against all considered models, DIF incurs lower training time and memory usage while retaining competitive inferencing execution time and anomaly detection performance. Eduardo Ortega, Rita Chattopadhyay, Krishnendu Chakrabarty |
IEEE Internet Things J. | 4 |
| 2025 | Enhancing Analog IC Security Using Randomized Obfuscation CircuitsabstractWith advances in technology scaling and globalization of the semiconductor industry, the vulnerability of analog integrated circuits (ICs) to reverse-engineering-based attacks, intellectual property theft, and unauthorized access has increased. Prior state-of-the-art analog deobfuscation techniques, such as those using genetic algorithms (GAs) and the satisfiability modulo theory, require an Oracle (i.e., unlocked IC) to recover the correct key. However, in some scenarios, an attacker present in an untrusted foundry might not have access to the Oracle. We demonstrate an Oracle-less attack using Bayesian optimization (BO) to retrieve the key of locked analog designs. To thwart both Oracle-guided and Oracle-less attacks, we present an automated obfuscation circuit generation framework for securing analog ICs. By employing randomness in obfuscation circuit generation, the proposed analog key-based methodology safeguards the integrity and reliability of analog ICs. Experimental results and security analysis for several analog designs demonstrate the robustness of the proposed technique to optimization attacks based on a GA and BO. We further show that the probability of guessing the correct key through brute force attack for an obfuscated analog circuit is negligibly small$(4.83\times 10^{-18})$. The proposed obfuscation scheme incurs an area overhead of less than 1.3% and power overhead of less than 2.64% for a mixed-signal IC. Jayeeta Chaudhuri, Mayukh Bhattacharya, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Reinforcement-Learning-Based Test Point Insertion for Power-Safe Testing in Monolithic 3-D ICsabstractMonolithic 3D (M3D) integration for integrated circuits (ICs) offers the promise of higher performance and lower power consumption over stacked-3D ICs. However, M3D suffers from large power supply noise (PSN) in the power distribution network due to high current demand and long conduction paths from voltage sources to local receivers. Excessive switching activities during the capture cycles in at-speed delay testing exacerbate the PSN-induced voltage droop problem. Therefore, PSN reduction is necessary for M3D ICs during testing to prevent the failure of good chips on the tester (i.e., yield loss). In this paper, we first develop an analysis flow for M3D designs to compute the PSN-induced voltage droop. Based on the analysis results, we extract the test patterns that are likely to cause yield loss. Next, we propose a reinforcement learning (RL)-based framework to insert test points and generate low-switching patterns that help in mitigating PSN without degrading the test coverage. Simulation results for benchmark M3D designs demonstrate the effectiveness of the proposed power-safe testing approach, compared to baseline cases that utilize commercial tools. Shao-Chun Hung, Arjun Chaudhuri, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Optimization of Droplet Routing in Microfluidic Biochips Using Calibrated Droplet-Shape MorphingabstractAdvanced digital microfluidic biochips based on technologies such as micro-electrode-dot-array (MEDA) and active matrix (AM) provide enhanced functionality compared to conventional biochips. Owing to the larger ratio of droplet size to electrode size, these platforms allow finer control of droplets and diagonal movement. Additionally, they allow dynamic grouping of micro-electrodes to form subsystems that can perform fluidic operations. Shape morphing is a key feature of MEDA/AM biochips that results in faster fluidic operations, thereby improving the efficiency of bioassays. To establish the benefits of shape morphing, we first numerically simulate the shape morphing operation. We employ a simplified 2-D flow model incorporating interface tracking through a level-set method to numerically simulate shape morphing induced by micro-electrode actuation. We also validate our numerical results with COMSOL simulations and experiments performed on MEDA/AM biochips. The validated shape morphing operations are subsequently used to optimize droplet routing for benchmark bioassays. We propose an algorithm to significantly reduce the size of the routing problem and the time needed to solve it. With the help of this improved approach, we show that droplet morphing operations reduce the time needed to complete bioassays. Arun Sankar Eenhakkattu Mana, Navajit Singh Baban, Hanbin Ma, Tsung-Yi Ho, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | TaintLock: Hardware IP Protection Against Oracle-Guided and Oracle-Reconstruction AttacksabstractScan-obfuscation schemes used with logic locking lack the ability to perform scan authentication on a per-pattern basis. These methods are of limited effectiveness in obfuscating scan data and they remain vulnerable to SAT-based scan deobfuscation attacks. In addition, prior methods designed to perform scan-data authentication are not adequate under the strongest threat models used to assess logic locking. To alleviate these problems, we propose enhancements to TaintLock, a lightweight dynamic per-pattern authentication and encryption scheme that uses taint and signature bits embedded within each test pattern to provide authenticated scan access. To prevent IP theft through Oracle-free and Oracle-guided attacks, TaintLock is paired with truly random logic locking (TRLL). TaintLock cryptographically authenticates each test pattern using the embedded taint and signature bits and passing them through a substitution-permutation (SP) network. It further uses cryptographically generated keys to dynamically encrypt scan data for unauthenticated users. TaintLock, while offering a low overhead and nonintrusive secure scan solution may remain susceptible to a new class of Oracle-reconstruction attacks that use machine learning. Additionally, assuming test pattern security is compromised, it may be potentially vulnerable to a template-based SAT attack aimed at partial key recovery. We analyze the susceptibility of TaintLock against these threats and demonstrate its resilience. We also demonstrate that TaintLock can be easily integrated with popular test architectures, such as embedded deterministic test (EDT). Finally, we also discuss the reconfigurable nature of TaintLock’s architecture to support different levels of encryption and authentication. Jonti Talukdar, Arjun Chaudhuri, Eduardo Ortega, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Detection of Voltage Droop-Induced Timing Fault Attacks Due to Hardware TrojansabstractRecent breakthroughs in heterogeneous integration (HI) using 2.5-D/3-D packaging technology have led to several advances in the semiconductor industry, increasing yield while reducing overall cost and time-to-market. However, the diversification of the HI supply chain has led to several sources of distrust due to the use of black-boxed third-party intellectual property (IP), outsourced fabrication, assembly and test facilities during the design and manufacturing process. We demonstrate the susceptibility of chiplet IPs to timing failure due to voltage droop in the power distribution network (PDN) induced by the insertion of chiplet level ring-oscillator (RO)-based hardware Trojans. We present an end-to-end methodology for design, placement, and insertion of RO-based Trojans in chiplet designs followed by characterizing their contribution to the dynamic voltage droop induced within the on-chip PDN. We quantify this PDN impact on timing paths and develop a systematic method to rank the susceptibility of different data paths toward a voltage droop event. We utilize this presilicon security analysis framework to evaluate voltage droop-based attack susceptibility for a variety of IPs, including some from the CEP benchmark. We also develop a machine learning-guided time-series anomaly detection framework to detect voltage droop-based anomalies on functional workloads running on different benchmarks, demonstrating the effectiveness of an convolutional autoencoders in detecting voltage droop-induced timing anomalies. Jonti Talukdar, Akshay Vyas, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | GINA: Exploiting Graph Neural Network Layer Features for Energy Efficient Inferencing in NVM-based PIM AcceleratorsabstractGraph Neural Networks (GNNs) are made up of multiple layers, with each layer comprising of different compute kernels involving weight vectors and adjacency matrices of input graph dataset. These layers exhibit varying features such as sparsity, storage requirement, and impact on predictive accuracy. Non-volatile memory (NVM)-based 3D Processing-In-Memory (PIM) architectures offer a promising approach to accelerate GNN inferencing. However, NVM device-based crossbars suffer from various non-idealities that affect the overall predictive accuracy. In this work, we consider the problem of finding a suitable mapping of GNN layers to PIM-based processing elements (PEs) in a 3D manycore architecture such that the impact of crossbar non-idealities on predictive accuracy is minimized. We develop a framework called GINA, which leverages low-cost, approximate Hessian-based methodology to automatically determine the GNN layers that are critical for accuracy and find a suitable GNN layer to PE mapping. To tackle non-idealities and to exploit sparsity at the crossbar level, a subset of the full crossbar is activated in a cycle, referred to as Operation Unit (OU). However, OU configurations vary with the above-mentioned GNN layer features, time-dependent conductance drift, and input graph dataset. GINA learns to optimize the OU configuration for unseen datasets as a function of GNN layer features and time-dependent conductance drift. Our experimental results demonstrate that GINA-enabled 3D PIM architecture reduces the latency and energy by 7.4 imes and 13 imes on an average, respectively, compared to state-of-the-art PIM architectures without compromising the predictive accuracy. Finally, we demonstrate the applicability of GINA to Convolutional Neural Networks (CNNs) and Vision Transformers. Gaurav Narang, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Krishnendu Chakrabarty, Partha Pratim Pande |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | Test-Fleet Scheduling in Complex Validation and Production EnvironmentsabstractWe present a solution to the complex design-automation problem of scheduling test operations in a validation laboratory or production facility. Our goal is to maximize the utilization of a fleet of test stations and minimize the overall test time for a set of products. We consider the realistic scenario where tests can have dependency graphs, implying that some tests must be completed and passed before others can proceed. We also consider a mix of product types that require different kinds of tests and a mix of testers, which implies that each product can only be tested only on a specific set of testers. To ensure scalability and flexibility, we have formulated this scheduling problem as a “partially observable stochastic game”, a multi-agent extension of a partially observable Markov decision process. We have implemented multi-agent reinforcement learning agents to maximize parallelization in a manner that speeds up both training and inferencing. We present scheduling results for synthetic test cases as well as real-life data from a production facility. Aniruddha Datta, Bhanu Vikas Yaganti, Mate Palocska, Andrew Dove, Arik Peltz, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | Patchability-Driven Design Exploration for System-on-Chip Patching ArchitecturesabstractAs System-on-Chip (SoC) designs become increasingly complex, ensuring comprehensive verification has become more challenging, leading to overlooked hardware bugs that can be found in the field. Addressing hardware bugs post-deployment is difficult, as they typically cannot be easily fixed like software bugs. To tackle this issue, hardware-based patching mechanisms have emerged as a potential solution for providing in-field fixes. However, the lack of a standardized method to evaluate the ”patchability” of different designs complicates the integration of patching infrastructure into SoCs. In this article, we propose a fully parameterized Patch Support Block (PSB) architecture that can be tailored for various hardware designs, enabling post-deployment patching. We introduce a novel patchability score formulation that provides a quantifiable metric for evaluating the effectiveness of patching designs. Our approach considers both the observability and controllability of the patching hardware and provides a framework for system integrators to maximize patchability while managing resource constraints. Through experimentation with multiple design configurations, we demonstrate how our methodology can enhance patchability in hardware systems and provide security-related fixes for SoCs in real-world scenarios. Wei-Kai Liu, Benjamin Tan 0001, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | SPICED+: Syntactical Bug Pattern Identification and Correction of Trojans in A/MS Circuits Using LLM-Enhanced DetectionabstractAnalog and mixed-signal (A/MS) integrated circuits (ICs) are crucial in modern electronics, playing key roles in signal processing, amplification, sensing, and power management. Many IC companies outsource manufacturing to third-party foundries, creating security risks such as syntactical bugs and stealthy analog Trojans. Traditional Trojan detection methods, including embedding circuit watermarks and hardware-based monitoring, impose significant area and power overheads while failing to effectively identify and localize the Trojans. To overcome these shortcomings, we present SPICED+, a software-based framework designed for syntactical bug pattern identification and the correction of Trojans in A/MS circuits, leveraging large language model (LLM)-enhanced detection. It uses LLM-aided techniques to detect, localize, and iteratively correct analog Trojans in SPICE netlists, without requiring explicit model training, and thus incurs zero area overhead. The framework leverages chain-of-thought reasoning and few-shot learning to guide the LLMs in understanding and applying anomaly detection rules, enabling accurate identification and correction of Trojan-impacted nodes. With the proposed method, we achieve an average Trojan coverage of 93.3%, average Trojan correction rate of 91.2%, and an average false-positive rate of 1.4%. Jayeeta Chaudhuri, Dhruv Thapar, Arjun Chaudhuri, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Theoretical Patchability Quantification for IP-Level Hardware Patching DesignsabstractAs the complexity of System-on-Chip (SoC) designs continues to increase, ensuring thorough verification becomes a significant challenge for system integrators. The complexity of verification can result in undetected bugs. Unlike software or firmware bugs, hardware bugs are hard to fix after deployment and they require additional logic, i.e., patching logic integrated with the design in advance in order to patch. However, the absence of a standardized metric for defining “patchability” leaves system integrators relying on their understanding of each IP and security requirements to engineer ad hoc patching designs. In this paper, we propose a theoretical patchability quantification method to analyze designs at the Register Transfer Level (RTL) with provided patching options. Our quantification defines patchability as a combination of observability and controllability so that we can analyze and compare the patchability of IP variations. This quantification is a systematic approach to estimate each patching architecture’s ability to patch at run-time and complements existing patching works. In experiments, we compare several design options of the same patching architecture and discuss their differences in terms of theoretical patchability and how many potential weaknesses can be mitigated. Wei-Kai Liu, Benjamin Tan 0001, Jason M. Fung, Krishnendu Chakrabarty |
ASPDAC | 4 |
| 2024 | Carbon Quantum Dot Fluorescent Stickers for Biochip AuthenticationabstractMicrofluidic biochips are widely used in biomedical research, clinical diagnostics, and point-of-care testing. However, their complex supply chains make them vulnerable to counterfeiting, overbuilding, and intellectual property (IP) piracy. We present fluorescent carbon quantum dot (CQD) stickers1that can be integrated with the polydimethylsiloxane (PDMS) based biochips for authentication. The stickers can be plasma-bonded to biochips made of glass and silicon. A protective spin-coated PDMS layer makes them obscured and tamperproof. However, they are detectable under UV light and can be authenticated via spectral analysis. The scheme exhibits unique excitation-dependent responses associated with the variability of the CQD sizes. This makes it ideal for physical authentication. Reliability studies concerning mechanical, photonic, and thermal degradation have demonstrated highly stable results. The stability of CQDs within the PDMS, their robust excitation-based emission fluorescence response, and the use of waste polypropylene masks make this a sustainable and robust authenticator for biochips. Navajit Singh Baban, Mohammed Abdelhameed, Mahmoud Elbeh, Khalil Ramadi, Yong-Ak Song, Sukanta Bhattacharjee, Ramesh Karri, Krishnendu Chakrabarty |
ATS | 8 |
| 2024 | Hacking the Fabric: Targeting Partial Reconfiguration for Fault Injection in FPGA FabricsabstractFPGAs are now ubiquitous in cloud computing infrastructures and reconfigurable system-on-chip, particularly for AI acceleration. Major cloud service providers such as Amazon and Microsoft are increasingly incorporating FPGAs for specialized compute-intensive tasks within their data centers. The availability of FPGAs in cloud data centers has opened up new opportunities for users to improve application performance by implementing customizable hardware accelerators directly on the FPGA fabric. However, the virtualization and sharing of FPGA resources among multiple users open up new security risks and threats. We present a novel fault attack methodology capable of causing persistent fault injections in partial bitstreams during the process of FPGA reconfiguration. This attack leverages powerwasters and is timed to inject faults into bitstreams as they are being loaded onto the FPGA through the reconfiguration manager, without needing to remain active throughout the entire reconfiguration process. Our experiments, conducted on a Pynq FPGA setup, demonstrate the feasibility of this attack on various partial application bitstreams, such as a neural network accelerator unit and a signal processing accelerator unit. Jayeeta Chaudhuri, Hassan Nassar, Dennis Gnad, Jörg Henkel, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty |
ATS | 6 |
| 2024 | Effective Runtime Fault Detection for DNN AcceleratorsabstractSystolic arrays are a popular choice for accelerating deep neural networks (DNNs) due to their inherent parallelism and efficient data reuse. However, ensuring the reliability of these DNN accelerators is crucial, as hardware faults can significantly degrade inferencing accuracy. Because systolic arrays utilize a large number of processing elements (PEs) for parallel processing of matrix multiplication, dataflow involving faulty PEs is especially of concern. Error propagation through PEs can reduce inferencing accuracy for DNN workloads. Even though many algorithm-based fault tolerance (ABFT) algorithms have been proposed to detect and correct errors in matrix multiplication, these ABFT methods cannot detect many errors originating from the accelerator hardware. We propose a run-time based fault detection technique leveraging functional data to generate checksums on-the-fly, avoiding the requirement for test patterns. Experimental evaluation shows that the proposed fault detection architecture can achieve 100% test coverage while incurring an area overhead of less than 2% for a 256 × 256 systolic array. Wei-Kai Liu, Jonti Talukdar, Benjamin Tan 0001, Krishnendu Chakrabarty |
ATS | 4 |
| 2024 | Neural Architecture Search for Blood Glucose Prediction in Type-1 DiabeticsabstractFor subjects affected with type-1 diabetes mellitus, accurately predicting future blood glucose values helps regulate insulin delivery. This paper introduces a dual Q-network-based neural architecture search approach to develop and train per-sonalized BG prediction models for individuals affected with type-1 diabetes mellitus. Utilizing historical blood glucose data collected via body sensor networks, the proposed model forecasts future blood glucose levels. When evaluated on the OhioTlDM dataset, the proposed approach shows significant improvements over the state-of-the-art, achieving a 46.78% reduction in root mean square error and a 56.05% reduction in mean absolute error while predicting blood glucose values 5 minutes into the future. Anthony Liardo, Aritra Ray, Farshad Firouzi, Kyle J. Lafata, Krishnendu Chakrabarty |
BSN | 5 |
| 2024 | PathDriver-Wash: A Path-Driven Wash Optimization Method for Continuous-Flow Lab-on-a-Chip SystemsabstractRapid advances in microfluidics technologies have facilitated the emergence of highly integrated lab-on-a-chip (LoC) biochip systems. With such coin-sized biochips, complicated bioassay procedures can be executed efficiently without any human intervention. To ensure the correctness and precision of assay outcomes, however, cross-contamination among different fluid samples/reagents needs to be dealt with separately during assay execution. As a consequence, wash operations have to be introduced and a wash path network needs to be established on the chip to remove the residues left in flow channels. To realize optimized assay procedures with efficient wash operations, we propose PathDriver-Wash in this paper, a path-driven wash optimization method for continuous-flow LoC biochip systems. The proposed method includes the following three key techniques: 1) The necessity of contamination removals is analyzed systemically to avoid unnecessary wash operations, 2) wash operations are integrated with the regular removal of excess fluids, so that extra path occupations caused by wash can be minimized, and 3) optimized wash paths and time windows are computed and assigned to wash operations, so that the completion time of assays can be minimized. Experimental results demonstrate that the proposed method leads to highly efficient wash procedures as well as minimized assay completion times. Xing Huang 0001, Zhiwen Yu 0001, Bin Guo 0001, Tsung-Yi Ho, Ulf Schlichtmann, Krishnendu Chakrabarty |
DATE | 7 |
| 2024 | Detection of Stealthy Bitstreams in Cloud FPGAs using Graph Convolutional NetworksabstractFPGAs are frequently utilized in cloud computing environments for high performance computing and neural network accelerators. Furthermore, multi-tenancy allows multiple users to upload customized modules on the FPGA, while maintaining logical isolation. However, attackers can take advantage of the multi-tenant environment to launch voltage-based attacks and denial-of-service (DoS). An attacker might stealthily split power-wasting ring oscillators (ROs) across multiple windows within an FPGA configuration bitstream, making it challenging for traditional detection mechanisms to identify these dispersed components as part of a larger malicious circuit. We propose a methodology to detect these malicious bitstreams by transforming individual windows within an FPGA bitstream into a graph-based representation. Leveraging this graph structure, our method employs graph convolutional networks (GCNs) in the training phase to capture malicious patterns from the bitstreams. We use the classification accuracy, true-positive rate, and false-positive rate metrics to quantify the effectiveness of our method across diverse power-wasting circuits on multiple FPGA boards. Jayeeta Chaudhuri, Krishnendu Chakrabarty |
ETS | 2 |
| 2024 | Test-Fleet Optimization Using Machine LearningabstractWe present a solution to the complex problem of scheduling test operations in a validation lab or production facility. Our goal is to maximize the utilization of a fleet of test stations and minimize the overall test time for a set of products. We consider the realistic scenario where tests can have dependency graphs, implying that some tests must be completed and passed before others can proceed. We also consider a mix of product types that require different kinds of tests and a mix of testers, which implies that each product can only be tested only on a specific set of testers. To ensure scalability and flexibility, we have formulated this scheduling problem as a “partially observable stochastic game”, a multi-agent extension of a partially observable Markov decision process. We have implemented multi-agent reinforcement learning agents to maximize parallelization in a manner that speeds up both training and inferencing. We present scheduling results for synthetic test cases as well as real-life data from a production facility. Aniruddha Datta, Bhanu Vikas Yaganti, Andrew Dove, Arik Peltz, Krishnendu Chakrabarty |
ETS | 5 |
| 2024 | LLM-AID: Leveraging Large Language Models for Rapid Domain-Specific Accelerator DevelopmentabstractThe challenges posed by the Dark Silicon era, combined with the escalating computational demands of emerging applications, such as Deep Learning (DL), have strained the capabilities of traditional CPUs and GPUs, necessitating the development of Domain-Specific Accelerators (DSAs). Despite offering substantial enhancements in Power, Performance, and Area (PPA), DSAs encounter significant challenges, including the rapid evolution of applications that necessitate the frequent development of new architectures. This, coupled with the expertise-intensive nature of the design process, often leads to reduced flexibility and extended development cycles, ultimately hindering the broader adoption and efficient deployment of DSAs. To address these challenges, this paper introduces LLM-AID, an agile framework that streamlines the DSA design flow by transforming high-level abstract specifications into Hardware Description Language (HDL) code and facilitating backend Computer-Aided Design (CAD) tool operations. By synergistically combining Large Language Models (LLMs), High-Level Synthesis (HLS) tools, design exploration techniques, and symbolic AI, LLM-AID dramatically accelerates design iterations, optimizes hardware performance, and significantly reduces time-to-market. This innovative approach democratizes DSA development, empowering designers to achieve unprecedented productivity while delivering high-quality DSA solutions. Farshad Firouzi, Sri Sai Rakesh Nakkilla, Chenghao Fu, Sanmitra Banerjee, Jonti Talukdar, Krishnendu Chakrabarty |
ICCAD | 6 |
| 2024 | KG-Infused LLM for Virtual Health Assistant: Accelerated Inference and Enhanced PerformanceabstractVirtual Health Assistants (VHAs) represent a significant advancement in patient care, leveraging artificial intelligence to provide continuous support and interaction. Despite the integration of advanced Large Language Models (LLMs), which have greatly improved conversational capabilities, VHAs still struggle to fully replicate the nuanced expertise of human medical professionals. This limitation is primarily due to their reliance on broad training data, which can result in responses that lack the necessary reliability and contextual relevance required in health-care, where precision is paramount. To address this challenge, recent research has focused on integrating healthcare databases, such as the Unified Medical Language System (UMLS), a comprehensive biomedical knowledge source, to enhance the reasoning capabilities and reliability of LLMs. However, previous efforts have been impeded by high inference times and suboptimal performance on various evaluation metrics due to inefficient retrieval of pertinent information from extensive databases. In this study, we propose a novel methodology that involves constructing a Knowledge Graph (KG) within the Neo4j graph database using UMLS data to facilitate faster retrieval. This approach is augmented by employing named entity recognition techniques to accurately identify relevant entities and by applying learning-to-rank and semantic matching algorithms to effectively rank the retrieved information. We validated our approach using various LLMs, including GPT-3.5 Turbo, GPT-4, LLaMA-7b, and LLaMA-13b, across the BioASQ, MedicationQA, and ExpertQA datasets. Our experiments demonstrated a 23% improvement in ROUGE-L scores, an 8-10% improvement in BERTScores, and a 20-30% improvement in BLEU scores, achieving up to an 80% reduction in inference time. Siva Kumar Katta, Aritra Ray, Farshad Firouzi, Krishnendu Chakrabarty |
ICMLA | 4 |
| 2024 | Preserving Accuracy While Stealing Watermarked Deep Neural NetworksabstractThe deployment of Deep Neural Networks (DNNs) as cloud services has accelerated significantly over the years. Training an application-specific DNN for cloud deployment requires substantial computational resources and costs associated with hyper-parameter tuning and model selection. To preserve Intellectual Property (IP) rights, model owners embed watermarks into publicly deployed DNNs. These trigger inputs and labels are uniquely selected and embedded into the watermarked DNN by the model owner, remaining undisclosed during deployment. If a watermarked DNN (target classifier) is stolen via white-box access and re-deployed by an adversary (pirated classifier) without securing the IP rights from the model owner, the model owner can identify their IP by sending trigger inputs to retrieve trigger labels. Typically, adversaries tamper with the model weights of the target classifier prior to deployment, which in turn reduces the utility of the well-trained DNN. The authors proposes re-deploying the target classifier without altering the model weights to preserve model utility, and using a small sample of non-identical in-distribution inputs (used for training the target classifier) to train a Siamese neural network to evade detection, at inference stage. Experimental evaluations on standard benchmark datasets- MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100- using ResNet architectures with varying triggers demonstrate that the proposed method achieves zero false positive rate (fraction of clean testing input incorrectly labelled as trigger inputs) and false negative rate (fraction of trigger inputs incorrectly labelled as clean in-distribution inputs) in nearly all cases, proving its efficacy. Aritra Ray, Farshad Firouzi, Kyle J. Lafata, Krishnendu Chakrabarty |
ICMLA | 4 |
| 2024 | SEC-CiM: Selective Error Compensation for ReRAM-based Compute-in-Memory*abstractReRAM-based Compute-in-Memory (CiM) architectures offer an attractive design choice for accelerating Convolutional Neural Network (CNN) inferencing in edge computing environments. However, these architectures are susceptible to stuck-at-faults (SAFs) in ReRAM cells stemming from manufacturing defects and cell wearout over time, significantly degrading CNN inferencing accuracy. To address this challenge, we propose a technique called Selective Error Compensation for CiM (SEC-CiM). This technique strategically mitigates errors by leveraging the insight that compensating for errors in a limited number of selected columns in a crossbar is sufficient to maintain CNN inferencing accuracy. With this strategy, SEC-CiM achieves significantly lower overhead compared to previous work. Notably, it effectively addresses errors resulting from stuck-at intermediate levels, a critical aspect that was previously overlooked. We develop a theoretical framework to determine the minimum number of columns requiring error compensation. Simulation results demonstrate that SEC-CiM limits the drop in inferencing accuracy to 2% for the ResNet18 and VGG16 models, even when up to 30% of the ReRAM cells in the crossbar are faulty. Similarly, for the Densenet121 CNN, comparable accuracy results are obtained when up to 15% of the ReRAM cells are faulty. We achieve this high level of fault tolerance with moderate area and power consumption overhead of 12.2% and 10.2%, respectively. Ashish Reddy Bommana, Farshad Firouzi, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ITC | 7 |
| 2024 | E-SCOUT: Efficient-Spatial Clustering-based Outlier Detection through TelemetryabstractSilicon lifecycle management (SLM) is needed to ensure silicon-product reliability and quality. Prior methods utilize off-chip solutions to identify malware, diagnose bugs, and characterize silicon health metrics. These methods do not explore hardware/software codesign for SLM. In this work, we present a new method called Efficient-Spatial Clustering-based OUlier detection through Telemetry (E-SCOUT) to monitor a chip’s status through performance counters/sensors. E-SCOUT includes a compute- and memory-efficient unsupervised 32-bit floating point outlier detection mechanism implemented on-chip (E-SCOUT edge). In addition, it enhances on-chip outlier detection through unsupervised feature ranking based on the telemetry feature information entropy. We also provide microarchitectural recommendations to enable a hardware/software co-design of E-SCOUT edge. The proposed solution includes an end-to-end outlier-informed diagnosis model with real telemetry data (E-SCOUT cloud). All telemetry data is collected through the model-specific register space using open-source Linux tools and Intel’s performance counter monitor. We capture the chip telemetry signatures of the PAMPAR benchmark suite in the presence of outlier events such as security attacks (e.g., Rowhammer and SPECTRE) and voltage droops. E-SCOUT provides effective unsupervised on-chip outlier detection performance with high accuracy levels (over 0.9) and with low area and low power over-head (2.2% die area overhead and 1% idle power consumption). Outlier diagnosis can identify the chip’s status with classification accuracy and F1-scores that exceed 0.8. Eduardo Ortega, Jonti Talukdar, Woohyun Paik, Rita Chattopadhyay, Krishnendu Chakrabarty |
ITC | 6 |
| 2024 | Safety-Guided Test Generation for Structural FaultsabstractMany real-life safety-critical applications such as autonomous driving require functional safety. We present a framework for functional safety-guided test pattern generation. We incorporate the functional information of each standard cell into a multi-layer-perceptron (MLP), referred to as Cell-Net. Each Cell-Net is a pre-trained MLP that models the behavior of the corresponding standard cell. The design netlist is translated into its neural twin, where the standard cell instances are substituted by their corresponding Cell-Nets and the wires in the netlist translate to neural connections between these Cell-Nets. We leverage the neural twin-enabled back-propagation for gradient computation, and utilize these gradients to compute test patterns. The output of every Cell-Net is associated with a bias that represents a perturbation in the signal propagating through that Cell-Net. We manipulate these bias values to inject stuck-at faults at the output of Cell-Nets. We utilize the neural twin to enhance the propagation of faults to primary outputs (POs), and find the test patterns that maximize the propagation of faults to POs. The neural twin also enables the assignment of non-binary criticalityfactors (CFs) to different POs and perform a test-pattern search for each fault for a given CF configuration. Our results on five benchmark circuits across three different CF configurations show an increased fault propagation achieved by the neural twin as compared to Automatic Test Pattern Generation (ATPG). Xuanyi Tan, Dhruv Thapar, Deepesh Sahoo, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty, Rubin A. Parekhji |
ITC | 6 |
| 2024 | Defect Analysis for FeFETs using a Compact ModelabstractFerroelectric field-effect transistors (FeFETs) are a promising emerging non-volatile memory, but the impact of manufacturing imperfections on these devices has yet to be studied comprehensively. We extend a previous FeFET compact model to combine the Preisach ferroelectric capacitor model with the BSIM-SOI MOSFET model. We calibrate this new compact model with data from a technology CAD (TCAD) model that is calibrated against a fabricated metal-ferroelectric-metal capacitor. We analyze polarization defects in the ferroelectric layer using this compact model. We address two classes of defects and map them to stuck-at-fault models, referred to as neutral faults (SAP0), and stuck-at-plus and stuck-at-minus (SAP+and SAP−) faults. This framework obviates the need for computationally expensive TCAD simulations for each defect scenario. Dhruv Thapar, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ITC | 4 |
| 2024 | Testing and Fault Diagnosis for Multi-level Resistive Random-Access Memory in Monolithic 3D Integration
Shao-Chun Hung, Partho Bhoumik, Krishnendu Chakrabarty |
VTS | 3 |
| 2024 | Low-Overhead Clustered Federated Learning for Personalized Stress MonitoringabstractStress, recognized widely as a substantial health concern, adversely affects individuals by undermining both their physical and mental well being. Prior studies on stress monitoring and management utilize a centralized cloud-based approach that combines data from each client for modeling. However, such a centralized approach raises data privacy concerns. To preserve privacy, decentralized federated learning (FL) has been proposed as a potential alternative framework. Nevertheless, existing FL algorithms have to deal with data heterogeneity; data skewness in each participant can significantly degrade the overall model performance. To tackle this challenge, we present a personalized, low-overhead clustered FL algorithm for stress-level recognition. The proposed algorithm outperforms two state-of-the-art baseline algorithms by providing over 7% and 12% increase in accuracy, respectively. The proposed algorithm also obtains a reduction of 37.5% and 9.6% in the training runtime compared to the two baseline algorithms. We also present a novel cold-start algorithm for new clients who join the trained system. Our results suggest that this cold-start algorithm is robust in terms of individual classification accuracy and total training time. Shiyi Jiang, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Internet Things J. | 3 |
| 2024 | Neuron grouping and mapping methods for 2D-mesh NoC-based DNN accelerators
Furkan Nacar, Alperen Cakin, Selma Dilek, Suleyman Tosun, Krishnendu Chakrabarty |
J. Parallel Distributed Comput. | 5 |
| 2024 | DAWN: Efficient Trojan Detection in Analog Circuits Using Circuit Watermarking and Neural TwinsabstractAs the globalization of integrated circuits (ICs) continues to advance, the threat of hardware Trojans has emerged as a major concern in ensuring the security and reliability of analog circuits. While a considerable body of prior work has focused on detecting digital Trojans in digital circuits, the detection of analog Trojans in analog circuits has received significantly less attention. We present DAWN, a sensitivity analysis-based analog Trojan detection framework using neural networks to identify potential analog Trojan hotspots and prevent them from being exploited through unauthorized modifications. We incorporate circuit watermarks in these hotspots to provide an additional layer of security. With these watermarks, any malicious modification to the circuit is automatically detected with high accuracy. We target the detection of stealthy, large-delay Trojans that might be inserted either during the chip design or fabrication stages. Experimental results for analog benchmark circuits and two commonly studied analog Trojans demonstrate the effectiveness of the proposed framework. Jayeeta Chaudhuri, Mayukh Bhattacharya, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Mitigating Slow-to-Write Errors in Memristor-Mapped Graph Neural Networks Induced by Adversarial AttacksabstractGraph neural networks (GNNs) are becoming popular in various real-world applications. However, hardware-level security is a concern when GNN models are mapped to emerging neuromorphic computing architectures such as memristor-based crossbars. We identify a vulnerability of memristor-mapped GNNs and propose an attack mechanism based on the identified vulnerability. The proposed attack tampers memristor-mapped graph-structured data of a GNN by injecting adversarial edges to the graph and inducing slow-to-write errors in crossbars. We present a defense mechanism based on the write-verify (WV) scheme. We analyze the effectiveness of the WV-based defense and provide theoretical security guarantees. This analysis also provides guidance for selecting appropriate design parameters for the WV scheme to ensure its effectiveness in countering slow-to-write errors induced by attacks. Experimental results for the proposed attack show that there is a 5.72× increase in the success rate compared to a software-based baseline. We also demonstrate the efficacy of the WV-based defense in mitigating all slow-to-write errors induced by the proposed attack. Ching-Yuan Chen, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Control-Logic Synthesis of Fully Programmable Valve Array Using Reinforcement LearningabstractFully programmable valve array (FPVA) biochips have emerged as a promising alternative for traditional application-specific microfluidic platforms thanks to their advantages in terms of flexibility and reconfigurability. By regularly deploying microvalves along vertical and horizontal flow channels, microfluidic modules with different sizes and shapes can be constructed dynamically on the chip, thereby enabling the automatic execution of various assay procedures in biology and biochemistry. The above advantages, however, result largely from the large-scale integration of valves as well as accurate control of their switchings, leading to very complicated control-logic design of such chips. In this article, we propose an reinforcement learning (RL)-based synthesis flow for the control-logic design of fully programmable valve array (FPVA) biochips, taking multichannel switching and control-cost minimization into consideration simultaneously. By employing a double deep$Q$-network (DDQN) and two Boolean-logic simplification techniques, control logics with both high-switching efficiency and low-fabrication cost can be constructed automatically. Furthermore, the solution space of multichannel-switching combinations is reduced to improve the search efficiency of the proposed method. Experimental results on multiple benchmarks demonstrate that the proposed synthesis flow leads to better-design solutions compared with the state-of-the-art techniques. Xing Huang 0001, Huayang Cai, Wenzhong Guo, Genggeng Liu, Tsung-Yi Ho, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | HuNT: Exploiting Heterogeneous PIM Devices to Design a 3-D Manycore Architecture for DNN TrainingabstractProcessing-in-memory (PIM) architectures have emerged as an attractive computing paradigm for accelerating deep neural network (DNN) training and inferencing. However, a plethora of PIM devices, e.g., resistive random-access memory, ferroelectric field-effect transistor, phase change memory, MRAM, static random-access memory, exists and each of these devices offers advantages and drawbacks in terms of power, latency, area, and nonidealities. A heterogeneous architecture that combines the benefits of multiple devices in a single platform can enable energy-efficient and high-performance DNN training and inference. 3-D integration enables the design of such a heterogeneous architecture where multiple planar tiers consisting of different PIM devices can be integrated into a single platform. In this work, we propose the HuNT framework, which hunts for (finds) an optimal DNN neural layer mapping, and planar tier configurations for a 3-D heterogeneous architecture. Overall, our experimental results demonstrate that the HuNT-enabled 3-D heterogeneous architecture achieves up to$10 {\times }$and$3.5 {\times }$improvement with respect to the homogeneous and existing heterogeneous PIM-based architectures, respectively, in terms of energy-efficiency (TOPS/W). Similarly, the proposed HuNT-enabled architecture outperforms existing homogeneous and heterogeneous architectures by up to$8 {\times }$and$2.4\times $, respectively, in terms of compute-efficiency (TOPS/mm2) without compromising the final DNN accuracy. Chukwufumnanya Ogbogu, Gaurav Narang, Biresh Kumar Joardar, Janardhan Rao Doppa, Krishnendu Chakrabarty, Partha Pratim Pande |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Block-Wise Mixed-Precision Quantization: Enabling High Efficiency for Practical ReRAM-Based DNN AcceleratorsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) architectures have demonstrated great potential to accelerate Deep Neural Network (DNN) training/ inference. However, the computational accuracy of analog PIM is compromised due to the non-idealities, such as the conductance variation of ReRAM cells. The impact of these non-idealities worsens as the number of concurrently activated wordlines and bitlines increases. To guarantee computational accuracy, only a limited number of wordlines and bitlines of the crossbar array can be turned on concurrently, significantly reducing the achievable parallelism of the architecture. While the constraints on parallelism limit the efficiency of the accelerators, they also provide a new opportunity for finegrained mixed-precision quantization. To enable efficient DNN inference on practical ReRAM-based accelerators, we propose an algorithm-architecture co-design framework called Block-Wise mixed-precision Quantization (BWQ). At the algorithm level, BWQ-A introduces a mixed-precision quantization scheme at the block level, which achieves a high weight and activation compression ratio with negligible accuracy degradation. We also present the hardware architecture design BWQ-H, which leverages the low-bit-width models achieved by BWQ-A to perform high-efficiency DNN inference on ReRAM devices. BWQ-H also adopts a novel precision-aware weight mapping method to increase the ReRAM crossbars throughput. Our evaluation demonstrates the effectiveness of BWQ, which achieves a 6.08× speedup and a 17.47× energy saving on average compared to existing ReRAM-based architectures. Xueying Wu, Edward Hanson, Nansu Wang, Qilin Zheng, Xiaoxuan Yang 0001, Huanrui Yang, Shiyu Li 0001, Partha Pratim Pande, Janardhan Rao Doppa, Krishnendu Chakrabarty, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2024 | Dynamic Adaptation Using Deep Reinforcement Learning for Digital Microfluidic BiochipsabstractWe describe an exciting new application domain for deep reinforcement learning (RL): droplet routing on digital microfluidic biochips (DMFBs). A DMFB consists of a two-dimensional electrode array, and it manipulates droplets of liquid to automatically execute biochemical protocols for clinical chemistry. However, a major problem with DMFBs is that electrodes can degrade over time. The transportation of droplet transportation over these degraded electrodes can fail, thereby adversely impacting the integrity of the bioassay outcome. We demonstrated that the formulation of droplet transportation as an RL problem enables the training of deep neural network policies that can adapt to the underlying health conditions of electrodes and ensure reliable fluidic operations. We describe an RL-based droplet routing solution that can be used for various sizes of DMFBs. We highlight the reliable execution of an epigenetic bioassay with the RL droplet router on a fabricated DMFB. We show that the use of the RL approach on a simple micro-computer (Raspberry Pi 4) leads to acceptable performance for time-critical bioassays. We present a simulation environment based on the OpenAI Gym Interface for RL-guided droplet routing problems on DMFBs. We present results on our study of electrode degradation using fabricated DMFBs. The study supports the degradation model used in the simulator. Tung-Che Liang, Yi-Chen Chang, Zhanwei Zhong, Yaas Bigdeli, Tsung-Yi Ho, Krishnendu Chakrabarty, Richard B. Fair |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2024 | Data Pruning-enabled High Performance and Reliable Graph Neural Network Training on ReRAM-based Processing-in-Memory AcceleratorsabstractGraph Neural Networks (GNNs) have achieved remarkable accuracy in cognitive tasks such as predictive analytics on graph-structured data. Hence, they have become very popular in diverse real-world applications. However, GNN training with large real-world graph datasets in edge-computing scenarios is both memory- and compute-intensive. Traditional computing platforms such as CPUs and GPUs do not provide the energy efficiency and low latency required in edge intelligence applications due to their limited memory bandwidth. Resistive random-access memory (ReRAM)-based processing-in-memory (PIM) architectures have been proposed as suitable candidates for accelerating AI applications at the edge, including GNN training. However, ReRAM-based PIM architectures suffer from low reliability due to their limited endurance, and low performance when they are used for GNN training in real-world scenarios with large graphs. In this work, we propose a learning-for-data-pruning framework, which leverages a trained Binary Graph Classifier (BGC) to reduce the size of the input data graph by pruning subgraphs early in the training process to accelerate the GNN training process on ReRAM-based architectures. The proposed light-weight BGC model reduces the amount of redundant information in input graph(s) to speed up the overall training process, improves the reliability of the ReRAM-based PIM accelerator, and reduces the overall training cost. This enables fast, energy-efficient, and reliable GNN training on ReRAM-based architectures. Our experimental results demonstrate that using this learning for data pruning framework, we can accelerate GNN training and improve the reliability of ReRAM-based PIM architectures by up to 1.6×, and reduce the overall training cost by 100× compared to state-of-the-art data pruning techniques. Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Krishnendu Chakrabarty, Janardhan Rao Doppa, Partha Pratim Pande |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Root-Cause Analysis with Semi-Supervised Co-Training for Integrated SystemsabstractRoot-cause analysis for integrated systems has become increasingly challenging due to their growing complexity. To tackle these challenges, machine learning (ML) has been applied to enhance root-cause analysis. Nonetheless, ML-based root-cause analysis usually requires abundant training data with root causes labeled by human experts, which are difficult or even impossible to obtain. To overcome this drawback, a semi-supervised co-training method is proposed for root-cause analysis in this article, which only requires a small portion of labeled data. First, a random forest is trained with labeled data. Next, we propose a co-training technique to learn from unlabeled data with semi-supervised learning, which pre-labels a subset of these data automatically and then retrains each decision tree in the random forest. In addition, a robust framework is proposed to avoid over-fitting. We further apply initialization by clustering and feature selection to improve the diagnostic performance. With two case studies from industry, the proposed approach shows superior performance against other state-of-the-art methods by saving up to 67% of labeling efforts. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Fault Diagnosis for Resistive Random Access Memory and Monolithic Inter-Tier Vias in Monolithic 3-D IntegrationabstractResistive random access memory (RRAM) constitutes a promising technology for next-generation memory architectures due to its simple structure, high on/off ratio, and processing-in-memory ability. Its compatibility with emerging monolithic 3-D (M3D) integration enables extremely high density using monolithic inter-tier vias (MIVs). However, both RRAM and M3D are susceptible to high defect rates due to immature manufacturing processes and process variations. Research efforts have been devoted to RRAM testing, while existing test solutions predominantly focus on fault detection. Fault diagnosis for M3D-integrated RRAM and MIVs remains unexplored. In this work, we propose a diagnosis procedure to identify the fault origin when a chip fails the manufacturing test. We present a detailed characterization of RRAM faulty behaviors in the presence of concurrent process variations and manufacturing defects. Based on RRAM characteristics, we develop a diagnosis sequence by identifying appropriate reference resistance and applied voltages to efficiently distinguish fault origins. Experimental results show that the proposed solution is compatible with existing test algorithms to significantly improve diagnostic resolution. By appending the proposed sequence to test algorithms, over 90% diagnostic resolution is achieved for every type of fault considered in an M3D-integrated RRAM. Shao-Chun Hung, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Rowhammer Vulnerability of DRAMs in 3-D IntegrationabstractWe investigate the vulnerability of 3-D-integrated dynamic random access memorys (DRAMs) [i.e., typically connected with silicon via (TSV), monolithic interconnect via (MIV)] to Rowhammer attacks. We have developed a SPICE framework to characterize Rowhammer attacks for the scenarios described. We utilize OPENROAD ASAP7 PDK for our simulation. We investigate horizontal (within the same tier) and vertical (across multiple tiers) variants of Rowhammer attacks. We show that horizontal Rowhammer vulnerability may be reduced through DRAM bank partitioning. In addition, we show that vertical parasitic capacitance in TSV 3D-DRAM is unlikely to lead to vertical Rowhammer attacks. However, vertical parasitic capacitance in MIV 3D-DRAM can make vertical Rowhammer attacks feasible. Eduardo Ortega, Jonti Talukdar, Woohyun Paik, Tyler K. Bletsch, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | ALT-Lock: Logic and Timing Ambiguity-Based IP Obfuscation Against Reverse EngineeringabstractWe present a logic ambiguity-based intellectual property (IP) obfuscation method that replaces traditional key gates with key-controlled functionally ambiguous logic gates, called LGA gates. We also protect timing paths by developing timing-ambiguous sequential cells called TA cells. We call this locking scheme ambiguous logic and timing logic locking (referred to as ALT-Lock). ALT-Lock ensures a two-pronged system-level security scheme where the attacker is forced to unlock not only combinational logic obfuscation but also timing obfuscation. We show that a combination of logic and timing ambiguity (TA) provides security against oracle-guided attacks. This method is superior to other traditional IP protection schemes such as combinational or sequential locking as it guarantees security against both oracle-guided and oracle-free attacks, while ensuring low power, performance, and area (PPA) overhead. Jonti Talukdar, Woohyun Paik, Eduardo Ortega, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Detection and Classification of Malicious Bitstreams for FPGAs in Cloud ComputingabstractAs FPGAs are increasingly shared and remotely accessed by multiple users and third parties, they introduce significant security concerns. Modules running on an FPGA may include circuits that induce voltage-based fault attacks and denial-of-service (DoS). An attacker might configure some regions of the FPGA with bitstreams that implement malicious circuits. Attackers can also perform side-channel analysis and fault attacks to extract secret information (e.g., secret key of an AES encryption). In this paper, we present a convolutional neural network (CNN)-based defense to detect bitstreams of RO-based malicious circuits by analyzing the static features extracted from FPGA bitstreams. We further explore the criticality of RO-based circuits in order to detect malicious Trojans that are configured on the FPGA. Evaluation on Xilinx FPGAs demonstrates the effectiveness of the security solutions. Jayeeta Chaudhuri, Krishnendu Chakrabarty |
ASP-DAC | 2 |
| 2023 | Attacking ReRAM-based Architectures using Repeated WritesabstractResistive random-access memory (ReRAM) is a promising technology for both memory and for in-memory computing. However, these devices have security vulnerabilities that are yet to be adequately investigated. In this work, we identify one such vulnerability that arises from the write mechanism in ReRAMs. Whenever a cell/row is written, a constant bias is automatically applied to the remaining cells/rows to reduce sneak current. We develop a new attack (referred as WriteHammer) that exploits this process. By repeatedly exposing a subset of cells to this bias, WriteHammer can cause noticeable resistance drift in the victim ReRAM cells. Experimental results indicate that WriteHammer can cause up to 3.5X change in cell resistance by repeatedly writing to the ReRAM cells for a duration of 4 ms. Biresh Kumar Joardar, Krishnendu Chakrabarty |
DATE | 2 |
| 2023 | Securing Heterogeneous 2.5D ICs Against IP Theft through Dynamic Interposer ObfuscationabstractRecent breakthroughs in heterogeneous integration (HI) technologies using 2.5D and 3D ICs have been key to advances in the semiconductor industry. However, heterogeneous integration has also led to several sources of distrust due to the use of third-party IP, testing, and fabrication facilities in the design and manufacturing process. Recent work on 2.5D IC security has only focused on attacks that can be mounted through rogue chiplets integrated in the design. Thus, existing solutions implement inter-chip let communication protocols that prevent unauthorized data modification and interruption in a 2.5D system. However, none of the existing solutions offer inherent security against IP theft. We develop a comprehensive threat model for 2.5D systems indicating that such systems remain vulnerable to IP theft. We present a method that prevents IP theft by obfuscating the connectivity of chiplets on the interposer using reconfigurable interconnection networks. We also evaluate the PPA impact and security offered by our proposed scheme. Jonti Talukdar, Arjun Chaudhuri, Sung Kyu Lim, Krishnendu Chakrabarty |
DATE | 5 |
| 2023 | Dynamic Task Remapping for Reliable CNN Training on ReRAM CrossbarsabstractA ReRAM crossbar-based computing system (RCS) can accelerate CNN training. However, hardware faults due to manufacturing defects and limited endurance impede the widespread adoption of RCS. We propose a dynamic task remapping-based technique for reliable CNN training on faulty RCS. Experimental results demonstrate that the proposed low-overhead method incurs only 0.85% accuracy loss on average while training popular CNNs such as VGGs, ResNets, and SqueezeNet with the CIFAR-IO, CIFAR-100, and SVHN datasets in the presence of faults. Chung-Hsuan Tung, Biresh Kumar Joardar, Partha Pratim Pande, Janardhan Rao Doppa, Hai Li 0001, Krishnendu Chakrabarty |
DATE | 6 |
| 2023 | Criticality Analysis of Ring Oscillators in FPGA Bitstreams *abstractThe popularity of cloud computing has led to increasing demand for efficient and scalable hardware. Multitenant FPGAs are becoming popular because of their ability to provide high performance and flexibility, yet being cost-effective. While multiple tenants have the ability to configure the same FPGA with customized modules, several security vulnerabilities can be exploited by adversaries. Attackers can use an FPGA to perform malicious actions, such as injecting malicious bitstreams and launching denial-of-service attacks. We propose a two-tier machine learning framework that first detects malicious features from an FPGA bitstream and then performs criticality analysis to evaluate the severity of potentially malicious ring oscillators (ROs) configured by that bitstream. The latter step is crucial as it ensures the security of FPGAs from voltage and power-based attacks and also reduces the risk of inappropriately blocking benign RO-based circuits from FPGA configuration. The proposed framework is evaluated using a diverse set of real-world bitstreams. We achieve an accuracy of 100% in detecting malicious bitstreams and an accuracy of 96.55% in detecting malicious bitstreams that are critical. Jayeeta Chaudhuri, Krishnendu Chakrabarty |
ETS | 2 |
| 2023 | Attacking Memristor-Mapped Graph Neural Network by Inducing Slow-to-Write ErrorsabstractGraph neural networks (GNNs) are becoming popular in various real-world applications. However, hardware-level security is a concern when GNN models are mapped to emerging neuromorphic technologies such as memristor-based crossbars. These security issues can lead to malfunction of memristor-mapped GNNs. We identify a vulnerability of memristor-mapped GNNs and propose an attack mechanism based on the identified vulnerability. The proposed attack tampers memristor-mapped graph-structured data of a GNN by injecting adversarial edges to the graph and inducing slow-to-write errors in crossbars. We show that 10% adversarial edge injection induces 1.11× longer write latency, eventually leading to a 44.33% error in node classification. Experimental results for the proposed attack also show that there is a 5.72× increase in the success rate compared to a software-based baseline. Ching-Yuan Chen, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ETS | 5 |
| 2023 | Test-Point Insertion for Power-Safe Testing of Monolithic 3D ICs using Reinforcement Learning*abstractMonolithic 3D (M3D) integration for integrated circuits (ICs) offers the promise of higher performance and lower power consumption over stacked-3D ICs. However, M3D suffers from large power supply noise (PSN) in the power distribution network due to high current demand and long conduction paths from voltage sources to local receivers. Excessive switching activities during the capture cycles in at-speed delay testing exacerbate the PSN-induced voltage droop problem. Therefore, PSN reduction is necessary for M3D ICs during testing to prevent the failure of good chips on the tester (i.e., yield loss). In this paper, we first develop an analysis flow for M3D designs to compute the PSN-induced voltage droop. Based on the analysis results, we extract the test patterns that are likely to cause yield loss. Next, we propose a reinforcement learning (RL)-based framework to insert test points and generate low-switching patterns that help in mitigating PSN without degrading the test coverage. Simulation results for benchmark M3D designs demonstrate the effectiveness of the proposed power-safe testing approach, compared to baseline cases that utilize commercial tools. Shao-Chun Hung, Arjun Chaudhuri, Krishnendu Chakrabarty |
ETS | 3 |
| 2023 | Energy-Efficient ReRAM-Based ML Training via Mixed Pruning and Reconfigurable ADCabstractMachine learning (ML) models have gained prominence in solving real-world tasks. However, implementing ML models is both compute- and memory-intensive. Domain-specific architectures such as Resistive Random Access Memory (ReRAM)-based Processing-in-Memory (PIM) platforms have been proposed to efficiently accelerate ML training and inference. However, existing ML workloads require a high amount of area and power for training. A major contributor to the area and power overheads is the Analog-to-Digital Converter (ADC). In this work, we propose a mixed pruning technique along with a novel reconfigurable ADC design to improve the power consumption profile. Overall, the pruned model with the reconfigurable ADC achieves ~50% reduction in power for training compared to existing state-of-the-art ReRAM-based architectures. Chukwufumnanya Ogbogu, Soumen Mohapatra, Biresh Kumar Joardar, Janardhan Rao Doppa, Deuk Hyoun Heo, Krishnendu Chakrabarty, Partha Pratim Pande |
ISLPED | 6 |
| 2023 | Biochip-PUF: Physically Unclonable Function for Microfluidic BiochipsabstractFlow-based microfluidic biochips (FMBs) have microvalves as key components. The physical characteristics of the microvalves vary instance-to-instance due to the inherent variability of numerous fabrication parameters. In this work, we leverage this unclonable, unpredictable instance-specific behavior and propose physically unclonable functions (PUFs) for FMBs, namely Biochip-PUFs (Bio-PUFs in short). We utilize variability in the microvalve membrane deflection response associated with the actuation pressure challenge to be our Bio-PUF parameter. Based on the distributions of the parameters measured on actual FMBs, we complement our Bio-PUF measurements via simulations of the FMB's microvalves in Comsol Multiphysics. Furthermore, we present a scheme based on the transient response of the microvalve actuation to augment the Bio-PUF authentication. The major advantage of this scheme is that we do not need any additional hardware to generate/implement the PUF module. The biochip itself can act as PUF instances while continuing to operate in normal functioning mode. Navajit Singh Baban, Ajymurat Orozaliev, Yong-Ak Song, Urbi Chatterjee, Sankalp Bose, Sukanta Bhattacharjee, Ramesh Karri, Krishnendu Chakrabarty |
ITC | 8 |
| 2023 | Scan Cell Segmentation Based on Reinforcement Learning for Power-Safe Testing of Monolithic 3D ICsabstractAs Moore's Law approaches its physical limits, monolithic 3D (M3D) integration offers continued power, performance, and density improvements. However, M3D integration can lead to large power supply noise (PSN) in the power distribution network due to high current demand and long conduction paths, leading to PSN-induced voltage droop problems. The PSN-induced voltage droop is more severe for at-speed delay testing than for the functional mode. Power-safe testing is therefore essential to prevent good chips from failing on the tester (i.e., yield loss). We propose a scan cell segmentation framework to reduce power consumption during scan capture. We use reinforcement learning to insert scan cell segments that can minimize switching activity without any adverse impact on test coverage. Simulation results for benchmark M3D designs highlight the effectiveness of the proposed framework. Shao-Chun Hung, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty |
ITC | 4 |
| 2023 | Simply-Track-and-Refresh: Efficient and Scalable Rowhammer MitigationabstractRowhammer is a memory vulnerability that can compromise system-level security. Rowhammer occurs when a DRAM row is accessed repeatedly, potentially causing bit-flips for neighboring rows. The threshold for Rowhammer has decreased from 139K accesses in 2014 to 3.2K in 2022. This threshold is projected to decrease further. Many existing solutions are not scalable, incur high overhead, or fail to offer protection in realistic scenarios. We propose Simply-Track-And-Refresh (STAR) as an effective and scalable Rowhammer mitigation. We compare STAR's performance overhead to recent solutions, HYDRA and AQUA. At ultra-low thresholds (500), STAR introduces 9.5x/31.7x lower average execution time overhead than HYDRA/AQUA. In addition, STAR introduces up to 4.3x lower area overhead and up to 3.3x lower power consumption compared to HYDRA and AQUA. We present proof of correctness, area and power consumption results derived using CACTI, and evaluation results from the PARSEC, SPLASH-2, SPEC2006, SPEC2017, and PAMPAR benchmark suites. Eduardo Ortega, Tyler K. Bletsch, Biresh Kumar Joardar, Jonti Talukdar, Woohyun Paik, Krishnendu Chakrabarty |
ITC | 6 |
| 2023 | Analysis and Characterization of Defects in FeFETsabstractEmerging devices are susceptible to manufacturing defects due to immature fabrication processes. Ferroelectric field-effect transistors, referred to as FeFETs, are promising emerging devices, but the impact of manufacturing imperfections on these devices has yet to be studied. Thus, we combine a technology CAD (TCAD) model with a fault-injection technique to represent fabrication defects in a FeFET. The TCAD model is calibrated against a fabricated metal-ferroelectric-metal capacitor and uses a multi-domain ferroelectric-layer structure. We address two classes of defects in the ferroelectric layer and map them to stuck-at-fault models referred to as neutral faults (SAP°) and stuck-at-plus and stuck-at-minus (SAP+and SAP−) faults. We also develop a machine-learning (ML) framework to characterize these fault-injected FeFET devices. The ML framework provides a significant speedup in predicting the health of the FE layer as compared to computationally heavy TCAD simulations. Our study of defects in ferroelectric FET (FeFET), which is done for the first time, and the insights gained thereof can provide valuable feedback for the fabrication and yield learning of FeFET-based circuits. Dhruv Thapar, Simon Thomann, Arjun Chaudhuri, Hussam Amrouch, Krishnendu Chakrabarty |
ITC | 5 |
| 2023 | Functional Test Generation for AI Accelerators using Bayesian Optimization∗abstractWe propose a black-box optimization method to generate functional test patterns for AI inferencing accelerators. Functional testing is faster than structural testing as scan chains are not used for shifting in patterns and shifting out test responses. Moreover, functional testing reduces "over-testing" by targeting the detection of functionally critical faults for a given application workload. We use Bayesian Optimization for targeted test-image generation for stuck-at faults in a systolic array-based accelerator. Our framework supports test-pattern compaction and leverages various types of error regularization for enforcing functional-likeness of the generated test images. We achieve high fault coverage using a small set of test images for pin-level faults in 16-bit and 32-bit floating-point processing elements of the systolic array achieves high fault coverage with a small set of test images. Arjun Chaudhuri, Ching-Yuan Chen, Jonti Talukdar, Krishnendu Chakrabarty |
VTS | 4 |
| 2023 | Special Session: Using Graph Neural Networks for Tier-Level Fault Localization in Monolithic 3D ICs *abstractMonolithic 3D (M3D) integration leverages fine-grained monolithic inter-tier vias (MIVs) to achieve significant improvements in power, performance, and area compared to conventional 2D integrated circuits (ICs). However, immature M3D fabrication flows lead to the degradation of device performance and unreliable interconnects between tiers. To improve yield learning, it is essential to perform fault localization at the tier level, which enables targeted diagnosis and process optimization efforts. This paper presents a graph neural network-based (GNN-based) diagnosis framework that efficiently localizes faults to a device tier and susceptible MIVs. The proposed solution offers rapid feedback to the foundry and improves the quality of diagnosis reports. The transferability of the GNN models makes it possible to perform diagnosis on designs with various design configurations without performance degradation. Results for four M3D benchmarks highlight the effectiveness of the proposed framework. Shao-Chun Hung, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty |
VTS | 4 |
| 2023 | Fusion of IoT, AI, Edge-Fog-Cloud, and Blockchain: Challenges, Solutions, and a Case Study in Healthcare and MedicineabstractThe digital transformation is characterized by the convergence of technologies—from the Internet of Things (IoT) to edge–fog–cloud computing, artificial intelligence (AI), and Blockchain—in multiple dimensions, blurring the lines between the physical and digital worlds. Although these innovations have evolved independently over time, they are increasingly becoming more intertwined, driving the development of new business models. With more adaptation, embracement, and development, we are witnessing a steady convergence and fusion of these technologies resulting in an unprecedented paradigm shift that is expected to disrupt and reshape the next-generation systems in vertical domains in a way that the capabilities of the technologies are aligned in the best possible way to complement each other. Despite the fact that the convergence of the four technologies can potentially tackle the main shortcomings of the existing systems, its adoption is still in its infancy phase, suffering from several issues, such as the absence of consensus toward any reference models or best practices. This article provides a comprehensive insight into the fusions of these paradigms by discussing a blend of topics addressing all the importation aspects from design to deployment. We will begin this article by providing an in-depth discussion on the main requirements, state-of-the-art reference architectures, applications, and challenges. Following this, we will present a reference architecture and a case study on privacy-preserving stress monitoring and management to better elaborate on the corresponding details and considerations. Farshad Firouzi, Shiyi Jiang, Krishnendu Chakrabarty, Bahareh J. Farahani, Mahmoud Daneshmand, Jaeseung Song, Kunal Mankodiya |
IEEE Internet Things J. | 3 |
| 2023 | Stool Image Analysis for Digital Health Monitoring By Smart Toilets
Jin Zhou 0014, Jackson McNabb, Nick DeCapite, Jose R. Ruiz, Deborah A. Fisher, Sonia Grego, Krishnendu Chakrabarty |
IEEE Internet Things J. | 7 |
| 2023 | Diagnosis of Malicious Bitstreams in Cloud Computing FPGAsabstractMultitenant field-programmable gate arrays (FPGAs) are increasingly being used in cloud computing technologies. Users are able to access the FPGA fabric remotely to implement custom accelerators in the cloud. However, the sharing of FPGA resources by untrusted third parties can lead to serious security threats. Attackers can configure a portion of the FPGA with a malicious bitstream. Such malicious use of the FPGA fabric may lead to severe voltage fluctuations and denial of service. In this work, we consider FPGAs that support time-based multitenancy, i.e., a single user has access to the FPGA at a time. We propose a convolutional neural network (CNN)-based approach to detect malicious RO-like circuits that are configured on an FPGA by learning features from the data-series representation of the bitstreams of malicious circuits. We use the classification accuracy, true-positive rate, and false-positive rate metrics to quantify the effectiveness of CNN-based classification of malicious bitstreams. Our threat model includes a variety of power-wasting circuits that are used to configure FPGAs in the cloud. We propose a two-stage malicious bitstream detection framework for the classification and diagnosis of the type of malicious circuit implemented by a particular bitstream. We further propose a novel window-merging technique to improve model performance in the second stage of the detection framework. Experimental results on Xilinx FPGAs demonstrate the effectiveness of the proposed method. Jayeeta Chaudhuri, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Deep Reinforcement Learning-Based Approach for Efficient and Reliable Droplet Routing on MEDA BiochipsabstractThe micro-electrode-dot-array (MEDA) architecture provides precise droplet control and real-time sensing in digital microfluidic biochips. Previous work has shown that trapped charge under microelectrodes (MCs) leads to droplets being stuck and failures in fluidic operations. A recent approach utilizes real-time sensing of MC health status, and attempts to avoid degraded electrodes during droplet routing. However, the problem with this solution is that the computational complexity is unacceptable for MEDA biochips of realistic size. Consequently, in this work, we introduce a deep reinforcement learning (DRL)-based approach to bypass degraded electrodes and enhance the reliability of routing. The DRL model utilizes the information of health sensing in real time to proactively reduce the likelihood of charge trapping and avoid using degraded MCs. Simulation results show that our approach provides effective routing strategies for COVID-19 testing protocols. We also validate our DRL-based approach using fabricated prototype biochips. Experimental results show that the developed DRL model completed the routing tasks using a fewer number of clock cycles and shorter total execution time, compared with a baseline routing method. Moreover, our DRL-based approach provides reliable routing strategies even in the presence of degraded electrodes. Our experimental results show that the proposed DRL-based routing is robust to occurrences of electrode faults, as well as increases the lifetime and usability of microfluidic biochips compared to existing strategies. Mahmoud Elfar, Yi-Chen Chang, Harrison Hao-Yu Ku, Tung-Che Liang, Krishnendu Chakrabarty, Miroslav Pajic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Learning Malicious Circuits in FPGA BitstreamsabstractComputing platforms are integrating field-programmable gate arrays (FPGAs) to support domain-specific customization. Multiple tenants can share these FPGAs by configuring them at runtime. However, attackers can abuse this capability by programming the FPGAs with malicious functions. A malicious configuration bitstream can launch denial of service, overheat the FPGA, leak sensitive information via side channels, enable remote monitoring, and launch voltage and timing attacks. We consider time-based multitenancy, where multiple tenants use the FPGA at different time intervals and not at the same time. We propose a defense based on machine learning (ML) algorithms to detect bitstreams of malicious circuits and malicious circuits mixed with legitimate circuits by analyzing the static features extracted from FPGA bitstreams. The proposed approach can help detect malicious bitstreams without the need for reverse engineering of the bitstream or having access to the design netlist. Our results on Xilinx FPGAs indicate that supervised classifiers may identify malicious bitstreams representing ring-oscillator circuits with a true-positive rate (TPR) of 100% and a false-positive rate (FPR) of only 4%. In addition, for the extremely difficult problem of detecting malicious bitstreams embedded in legitimate bitstreams, a pipeline of a random forest and a support vector machine classifiers trained on subarrays of bitstreams can help detect bitstreams of malicious circuits embedded in legitimate designs with TPR of 95.5% and FPR of 30.4%. Rana Elnaggar, Jayeeta Chaudhuri, Ramesh Karri, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Transferable Graph Neural Network-Based Delay-Fault Localization for Monolithic 3-D ICsabstractMonolithic 3-D (M3D) integration is a promising technology for achieving high performance and low-power consumption. However, the limitations of current M3D fabrication flows lead to performance degradation of devices in the top tier and unreliable interconnects between tiers. Fault localization at the tier level is therefore necessary to enhance yield learning, For example, tier-level localization can enable targeted diagnosis and process optimization efforts. In this article, we develop a graph neural network-based diagnosis framework to efficiently localize faults to a device tier. The proposed framework can be used to provide rapid feedback to the foundry and help enhance the quality of diagnosis reports generated by commercial tools. Results for four M3D benchmarks, with and without response compaction, show that the proposed solution achieves up to 32.86% improvement in diagnostic resolution with less than 1% loss of accuracy, compared to results from commercial tools. The proposed framework has also been demonstrated to be transferable to perform diagnosis on various design configurations without performance degradation. Shao-Chun Hung, Sanmitra Banerjee, Arjun Chaudhuri, Sung Kyu Lim, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Machine Learning-Based Rowhammer MitigationabstractRowhammer is a security vulnerability that arises due to the undesirable electrical interaction between physically adjacent rows in DRAMs. Bit flips caused by Rowhammer can be exploited to craft many types of attacks in platforms ranging from edge devices to datacenter servers. Existing DRAM protections using error-correction codes and targeted row refresh are not adequate for defending against Rowhammer attacks. In this work, we propose a Rowhammer mitigation solution using machine learning (ML). We show that the ML-based technique can reliably detect and prevent bit flips for all the different types of Rowhammer attacks (including the recently proposed Half-double and Blacksmith attacks) considered in this work. Moreover, the ML model is associated with lower power and area overhead compared to recently proposed Rowhammer mitigation techniques, namely, Graphene and Blockhammer, for 40 different applications from the Parsec, Pampar, Splash-2, SPEC2006, and SPEC 2017 benchmark suites. Biresh Kumar Joardar, Tyler K. Bletsch, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Hardware-Supported Patching of Security Bugs in Hardware IP BlocksabstractTo satisfy various design requirements and application needs, designers integrate multiple intellectual property blocks (IPs) to produce a system on chip (SoC). For improved survivability, designers should be able to patch the SoC to mitigate potential security issues arising from hardware IPs; for increased flexibility, we propose adding programmable hardware-based support for monitoring and bug mitigation. However, it is a challenge to decide how much additional cost a designer should expend up front to deal with unknown, future issues. We propose an approach that guides designers toward maximizing the benefits of adding “patchability” to various IPs in the system, given a target resource overhead. We frame the design problem as an integer quadratic program and show that our approach achieves superior patchability compared to the naïve and baseline approaches for a given cost limit. Experimental results show that when we set a cost limit of 2% field-programmable gate array adaptive logic module usage, our solution can generate a viable patching infrastructure with six patching blocks offering patches for seven different services in our case study. Wei-Kai Liu, Benjamin Tan 0001, Jason M. Fung, Ramesh Karri, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Accelerating Graph Neural Network Training on ReRAM-Based PIM Architectures via Graph and Model PruningabstractGraph neural networks (GNNs) are used for predictive analytics on graph-structured data, and they have become very popular in diverse real-world applications. Resistive random-access memory (ReRAM)-based PIM architectures can accelerate GNN training. However, GNN training on ReRAM-based architectures is both compute- and data intensive in nature. In this work, we propose a framework calledSlimGNNthat synergistically combines both graph and model pruning to accelerate GNN training on ReRAM-based architectures. The proposed framework reduces the amount of redundant information in both the GNN model and input graph(s) to streamline the overall training process. This enables fast and energy-efficient GNN training on ReRAM-based architectures. Experimental results demonstrate that using this framework, we can accelerate GNN training by up to$ {4}. {5} {\times }$while using$ {6}. {6} {\times }$less energy compared to the unpruned counterparts. Chukwufumnanya Ogbogu, Aqeeb Iqbal Arka, Lukas Pfromm, Biresh Kumar Joardar, Janardhan Rao Doppa, Krishnendu Chakrabarty, Partha Pratim Pande |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Unsupervised Two-Stage Root-Cause Analysis With Transfer Learning for Integrated SystemsabstractThe growing complexity of integrated systems makes root-cause analysis increasingly difficult. To address this challenge, advances in machine learning (ML) have been leveraged in recent years to design ML-based techniques for root-cause analysis. However, most of these methods require root-cause labels for defective samples obtained based on the analysis by human experts. In this article, we propose a multialgorithm two-stage clustering method with transfer learning for unsupervised root-cause analysis. First, a two-stage clustering method is proposed by applying multiple clustering methods to accommodate both numerical and categorical data and leveraging Silhouette score for model selection. Next, a double-bootstrapping method is proposed for data selection, transferring valuable information from a source product to a target product with insufficient data. In the first bootstrapping step, a random forest model is built to select effective source data. In the second bootstrapping step, clustering ensemble is applied to two-stage clustering to further improve the accuracy for root-cause analysis. Two case studies based on network products demonstrate the superior performance of the proposed approach compared to other state-of-the-art methods. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | ESSENCE: Exploiting Structured Stochastic Gradient Pruning for Endurance-Aware ReRAM-Based In-Memory Training SystemsabstractProcessing-in-memory (PIM) enables energy-efficient deployment of convolutional neural networks (CNNs) from edge to cloud. Resistive random-access memory (ReRAM) is one of the most commonly used technologies for PIM architectures. One of the primary limitations of ReRAM-based PIM in neural network training arises from the limited write endurance due to the frequent weight updates. To make ReRAM-based architectures viable for CNN training, the write endurance issue needs to be addressed. This work aims to reduce the number of weight reprogrammings without compromising the final model accuracy. We propose the ESSENCE framework with an endurance-aware structured stochastic gradient pruning method, which dynamically adjusts the probability of gradient update based on the current update counts. Experimental results with multiple CNNs and datasets demonstrate that the proposed method can extend ReRAM’s life time for training. For instance, with the ResNet20 network and CIFAR-10 dataset, ESSENCE can save the mean update counts of up to$10.29\times $compared to the stochastic gradient descent method and effectively reduce the maximum update counts compared with the No Endurance method. Furthermore, an aggressive tuning method based on ESSENCE can boost the mean update count savings by up to$14.41\times $. Xiaoxuan Yang 0001, Huanrui Yang, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Enhanced Built-In Self-Diagnosis and Self-Repair Techniques for Daisy-Chain Design in MEDA Digital Microfluidic BiochipsabstractDigital microfluidic biochips have emerged as a promising alternative for various laboratory procedures in biochemistry, such as drug discovery and DNA sequencing. A recent generation of digital biochips uses a micro-electrode-dot-array (MEDA) architecture, which provides finer controllability of droplets and seamlessly integrates microelectronics and microfluidics. To simplify the wiring design of such biochips, all microelectrodes and their control registers are daisy-chained together. Therefore, the ability to both identify faults in the chain and tolerate them is required in MEDA biochips. In this study, a new daisy-chain design approach is proposed, integrating a built-in self-repair scheme that can automatically detect faults and correct them. Moreover, an efficient test generation method that requires only a small number of test vectors is proposed to achieve 100% fault coverage without degrading the electrodes. The proposed self-repair scheme can be used in both offline and online modes. Experimental results show that detection and repair can be carried out for various types of faults that can occur in daisy chains. Xing Huang 0001, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Built-In Self-Test of High-Density and Realistic ILV Layouts in Monolithic 3-D ICsabstractNanoscale interlayer vias (ILVs) in monolithic 3-D (M3D) ICs have enabled high-density vertical integration of logic and memory tiers. However, the sequential assembly of M3D tiers via wafer bonding is prone to variability in the immature fabrication process and manufacturing defects. The yield degradation due to ILV faults can be mitigated via dedicated test and diagnosis of ILVs using built-in self-test (BIST). Prior work has carried out fault localization for a regular 1-D placement of ILVs in the M3D layout where shorts are assumed to arise only between unidirectional ILVs. However, to minimize wirelength in M3D routing, ILVs may be irregularly placed by a place-and-route tool, and shorts can also occur between an up-going ILV and a down-going ILV. To test and localize faults in realistic ILV layouts, we present a new BIST framework that is optimized for test time and PPA overhead. We also present a graph-theoretic approach for representing potential fault sites in the ILVs and carry out inductive fault analysis to drop noncritical sites. We describe a procedure for optimally assigning ILVs to the BIST pins and determining the BIST configuration for test-cost minimization. Evaluation results for M3D benchmarks demonstrate the effectiveness of the proposed framework. Arjun Chaudhuri, Sanmitra Banerjee, Sung Kyu Lim, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | Adaptive Droplet Routing for MEDA Biochips via Deep Reinforcement LearningabstractDigital microfluidic biochips (DMFBs) based on a micro-electrode-dot-array (MEDA) architecture provide fine-grained control and sensing of droplets in real-time. However, excessive actuation of microelectrodes in MEDA biochips can lead to charge trapping during bioassay execution, causing the failure of microelectrodes and erroneous bioassay outcomes. A recently proposed enhancement to MEDA allows run-time measurement of microelectrode health information, thereby enabling synthesis of adaptive routing strategies for droplets. However, existing synthesis solutions are computationally infeasible for large MEDA biochips that have been commercialized. In this paper, we propose a synthesis framework for adaptive droplet routing in MEDA biochips via deep reinforcement learning (DRL). The framework utilizes the real-time microelectrode health feedback to synthesize droplet routes that proactively minimize the likelihood of charge trapping. We show how the adaptive routing strategies can be synthesized using DRL. We implement the DRL agent, the MEDA simulation environment, and the bioassay scheduler using the OpenAI Gym environment. Our framework obtains adaptive routing policies efficiently for COVID-19 testing protocols on large arrays that reflect the sizes of commercial MEDA biochips available in the marketplace, significantly increasing probabilities of successful bioassay completion compared to existing methods. Mahmoud Elfar, Tung-Che Liang, Krishnendu Chakrabarty, Miroslav Pajic |
DATE | 3 |
| 2022 | Graph Neural Network-based Delay-Fault Localization for Monolithic 3D ICsabstractMonolithic 3D (M3D) integration is a promising technology for achieving high performance and low power consumption. However, the limitations of current M3D fabrication flows lead to performance degradation of devices in the top tier and unreliable interconnects between tiers. Fault localization at the tier level is therefore necessary to enhance yield learning, For example, tier-level localization can enable targeted diagnosis and process optimization efforts. In this paper, we develop a graph neural network-based diagnosis framework to efficiently localize faults to a device tier. The proposed framework can be used to provide rapid feedback to the foundry and help enhance the quality of diagnosis reports generated by commercial tools. Results for four M3D benchmarks, with and without response compaction, show that the proposed solution achieves up to 39.19% improvement in diagnostic resolution with less than 1% loss of accuracy, compared to results from commercial tools. Shao-Chun Hung, Sanmitra Banerjee, Arjun Chaudhuri, Krishnendu Chakrabarty |
DATE | 4 |
| 2022 | Graph Neural Network-based Delay-Fault Localization for Monolithic 3D ICsabstractMonolithic 3D (M3D) integration is a promising technology for achieving high performance and low power consumption. However, the limitations of current M3D fabrication flows lead to performance degradation of devices in the top tier and unreliable interconnects between tiers. Fault localization at the tier level is therefore necessary to enhance yield learning, For example, tier-level localization can enable targeted diagnosis and process optimization efforts. In this paper, we develop a graph neural network-based diagnosis framework to efficiently localize faults to a device tier. The proposed framework can be used to provide rapid feedback to the foundry and help enhance the quality of diagnosis reports generated by commercial tools. Results for four M3D benchmarks, with and without response compaction, show that the proposed solution achieves up to 39.19% improvement in diagnostic resolution with less than 1% loss of accuracy, compared to results from commercial tools. Shao-Chun Hung, Sanmitra Banerjee, Arjun Chaudhuri, Krishnendu Chakrabarty |
DATE | 4 |
| 2022 | Learning to Mitigate Rowhammer AttacksabstractRowhammer is a vulnerability that arises due to the undesirable interaction between physically adjacent rows in DRAMs. Existing DRAM protections are not adequate to defend against Rowhammer attacks. We propose a Rowhammer mitigation solution using machine learning (ML). We show that the ML-based technique can reliably detect and prevent bit flips for all the different types of Rowhammer attacks considered here. Moreover, the ML model is associated with lower power and area overhead compared to recently proposed Rowhammer mitigation techniques for 26 different applications from the Parsec, Pampar, and Splash-2 benchmark suites. Biresh Kumar Joardar, Tyler K. Bletsch, Krishnendu Chakrabarty |
DATE | 3 |
| 2022 | Machine Learning for Test, Diagnosis, Post-Silicon Validation and Yield OptimizationabstractRecent breakthroughs in machine learning (ML) technology are shifting the boundaries of what is technologically possible in several areas of Computer Science and Engineering. This paper discusses ML in the context of test-related activities, including fault diagnosis, post-silicon validation and yield optimization. ML is by now an established scientific discipline, and a large number of successful ML techniques have been developed over the years. This paper focuses on how to adapt ML approaches that were originally developed with other applications in mind to test-related problems. We consider two specific applications of learning in more depth: delay fault diagnosis in three-dimensional integrated circuits and tuning performed during post-silicon validation. Moreover, we examine the emerging concept of brain-inspired hyperdimensional computing (HDC) and its potential for addressing test and reliability questions. Finally, we show how to integrate ML into actual industrial test and yield-optimization flows. Hussam Amrouch, Krishnendu Chakrabarty, Dirk Pflüger, Ilia Polian, Matthias Sauer 0002, Matteo Sonza Reorda |
ETS | 2 |
| 2022 | Detection of Malicious FPGA Bitstreams using CNN-Based LearningabstractMulti-tenant FPGAs are increasingly being used in cloud computing technologies. Users are able to access the FPGA fabric remotely to implement custom accelerators in the cloud. However, sharing FPGA resources by untrusted third-parties can lead to serious security threats. Attackers can configure a portion of the FPGA with a malicious bitstream. Such malicious use of the FPGA fabric may lead to severe voltage fluctuations and eventually crash the FPGA. Attackers can also use side-channel and fault attacks to extract secret information (e.g., secret key of an AES encryption module). We propose a convolutional neural network (CNN)-based defense mechanism to detect malicious circuits that are configured on an FPGA by learning features from the data-series representation of the bitstreams of malicious circuits. We use the classification accuracy, true-positive rate, and false-positive rate metrics to quantify the effectiveness of CNN-based classification of malicious bitstreams. Experimental results on Xilinx FPGAs demonstrate the effectiveness of the proposed method. Jayeeta Chaudhuri, Krishnendu Chakrabarty |
ETS | 2 |
| 2022 | TaintLock: Preventing IP Theft through Lightweight Dynamic Scan Encryption using Taint Bits*abstractWe propose TaintLock, a lightweight dynamic scan data authentication and encryption scheme that performs per-pattern authentication and encryption using taint and signature bits embedded within the test pattern. To prevent IP theft, we pair TaintLock with truly random logic locking (TRLL) to ensure resilience against both Oracle-guided and Oracle-free attacks, including scan deobfuscation attacks. TaintLock uses a substitution-permutation (SP) network to cryptographically authenticate each test pattern using embedded taint and signature bits. It further uses cryptographically generated keys to encrypt scan data for unauthenticated users dynamically. We show that it offers a low overhead, non-intrusive secure scan solution without impacting test coverage or test time while preventing IP theft. Jonti Talukdar, Arjun Chaudhuri, Krishnendu Chakrabarty |
ETS | 3 |
| 2022 | LoCI: An Analysis of the Impact of Optical Loss and Crosstalk Noise in Integrated Silicon-Photonic Neural NetworksabstractCompared to electronic accelerators, integrated silicon-photonic neural networks (SP-NNs) promise higher speed and energy efficiency for emerging artificial-intelligence applications. However, a hitherto overlooked problem in SP-NNs is that the underlying silicon photonic devices suffer from intrinsic optical loss and crosstalk noise, the impact of which accumulates as the network scales up. Leveraging precise device-level models, this paper presents the first comprehensive and systematic optical loss and crosstalk modeling framework for SP-NNs. For an SP-NN case study with two hidden layers and 1380 tunable parameters, we show a catastrophic ~84% drop in inferencing accuracy due to optical loss and crosstalk noise. Amin Shafiee, Sanmitra Banerjee, Krishnendu Chakrabarty, Sudeep Pasricha, Mahdi Nikdast |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Don't CWEAT It: Toward CWE Analysis Techniques in Early Stages of Hardware DesignabstractTo help prevent hardware security vulnerabilities from propagating to later design stages where fixes are costly, it is crucial to identify security concerns as early as possible, such as in RTL designs. In this work, we investigate the practical implications and feasibility of producing a set of security-specific scanners that operate on Verilog source files. The scanners indicate parts of code that might contain one of a set of MITRE's common weakness enumerations (CWEs). We explore the CWE database to characterize the scope and attributes of the CWEs and identify those that are amenable to static analysis. We prototype scanners and evaluate them on 11 open source designs - 4 system-on-chips (SoC) and 7 processor cores - and explore the nature of identified weaknesses. Our analysis reported 53 potential weaknesses in the OpenPiton SoC used in [email protected], 11 of which we confirmed as security concerns. Baleegh Ahmad, Wei-Kai Liu, Luca Collini, Hammond A. Pearce, Jason M. Fung, Jonathan Valamehr, Mohammad Bidmeshki, Piotr Sapiecha, Krishnendu Chakrabarty, Ramesh Karri, Benjamin Tan 0001 |
ICCAD | 10 |
| 2022 | Machine Learning for Testing Machine-Learning Hardware: A Virtuous CycleabstractThe ubiquitous application of deep neural networks (DNN) has led to a rise in demand for AI accelerators. DNN-specific functional criticality analysis identifies faults that cause measurable and significant deviations from acceptable requirements such as the inferencing accuracy. This paper examines the problem of classifying structural faults in the processing elements (PEs) of systolic-array accelerators. We first present a two-tier machine-learning (ML) based method to assess the functional criticality of faults. While supervised learning techniques can be used to accurately estimate fault criticality, it requires a considerable amount of ground truth for model training. We therefore describe a neural-twin framework for analyzing fault criticality with a negligible amount of ground-truth data. We further describe a topological and probabilistic framework to estimate the expected number of PE's primary outputs (POs) flipping in the presence of defects and use the PO-flip count as a surrogate for determining fault criticality. We demonstrate that the combination of PO-flip count and neural twin-enabled sensitivity analysis of internal nets can be used as additional features in existing ML-based criticality classifiers. Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty |
ICCAD | 3 |
| 2022 | Deep Neural Network Piration without Accuracy LossabstractA deep neural network (DNN) classifier is often viewed as the intellectual property of a model owner due to the huge resources required to train it. To protect intellectual property, the model owner can embed a watermark into the DNN classifier (called target classifier) such that it outputs pre-determined labels (called trigger labels) for pre-determined inputs (called trigger inputs). Given the black-box access to a suspect classifier, the model owner can verify whether the suspect classifier is pirated version of its classifier by first querying the suspect classifier for trigger inputs and then checking whether the predicted labels match with the trigger labels. Many studies showed that an attacker can pirate the target classifier (called pirated classifier) via retraining or fine-tuning the target classifier to remove its watermark. However, they sacrifice the accuracy of the pirated classifier, which is undesired for critical applications such as finance and healthcare. In our work, we propose a new attack without sacrificing the accuracy of the pirated classifier for in-distribution testing inputs while preventing the detection from the model owner. Our idea is that an attacker can detect the trigger inputs in the inference stage of the pirated classifier. In particular, given a testing input, we let the pirated classifier return a random label if the input is detected as a trigger input. Otherwise, the pirated classifier predicts the same label as the target classifier. We evaluate our attack on benchmark datasets and find that our attack can effectively identify the trigger inputs. Our attack reveals that the intellectual property of a model owner can be violated with existing watermarking techniques, highlighting the need for new techniques. Aritra Ray, Jinyuan Jia 0001, Sohini Saha, Jayeeta Chaudhuri, Neil Zhenqiang Gong, Krishnendu Chakrabarty |
ICMLA | 6 |
| 2022 | Structural Test Generation for AI Accelerators using Neural TwinsabstractWe present a neural twin-based structural test pattern generation method for stuck-at faults in systolic array-based AI inferencing accelerators. The neural twin is a neural representation of the gate-level netlist of a processing element and it provides a one-to-one topological correspondence with the PE netlist. We leverage neural twin-enabled backpropagation for gradient computation to determine an input pattern that sensitizes a fault in the netlist. Our framework also supports pattern compaction for a batch of faults. Consequently, GPU-accelerated test-pattern generation is achieved with the proposed framework that can potentially detect hard-to-detect and random-pattern-resistant faults in AI accelerators. Experimental results for 4-bit, 8-bit, and 16-bit fixed-point accelerator arrays show the effectiveness of the proposed method. Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty |
IOLTS | 3 |
| 2022 | Fault Diagnosis for Resistive Random-Access Memory and Monolithic Inter-tier Vias in Monolithic 3D IntegrationabstractResistive random-access memory (RRAM) constitutes a promising technology for next-generation memory architectures due to its simple structure, high on/off ratio, and processing-in-memory ability. Its compatibility with emerging monolithic 3D (M3D) integration enables extremely high density using monolithic inter-tier vias (MIVs). However, both RRAM and M3D are susceptible to high defect rates due to immature manufacturing processes and process variations. Fault diagnosis for M3D-integrated RRAM and MIVs is therefore necessary to facilitate yield learning. In this work, we present a detailed characterization of RRAM faulty behaviors in the presence of process variations and manufacturing defects. We develop a diagnosis procedure by identifying appropriate reference resistance based on RRAM characteristics to efficiently distinguish fault origins. Results show that the proposed solution is compatible with existing test algorithms to significantly improve diagnostic resolution without affecting fault coverage. Shao-Chun Hung, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty |
ITC | 4 |
| 2022 | Automatic Structural Test Generation for Analog Circuits using Neural TwinsabstractThe growing size of analog IPs has made targeted structural testing of such designs a challenging problem. We present a gradient-based automated test generation framework for analog circuits using neural twins, which are neural equivalents of the corresponding analog circuit. A neural twin is constructed by combining several FET-twins that lie in the paths between the circuit's inputs and observation points. Each FET-twin is a fully-connected neural network that models the IV characteristics of individual MOSFETs in the design. We train different variants of FET-twins that can predict both the output current and nodal voltage with more than 99% accuracy. We create an analog neural miter circuit, for which tests are generated using gradient ascent to maximize the loss between the faulty and fault-free versions of the neural twin. By computing gradients in a batchwise fashion for all the faults in the design, we develop a test compaction scheme that covers all faults with minimum number of test patterns. The neural twin-driven test generation method is interpretable, faster to simulate through GPUs, and guarantees convergence through backpropagation. We demonstrate the effectiveness of this framework by generating tests for structural defects in analog benchmark circuits. We show that our method outperforms an existing black-box optimization method that can be repurposed for test generation. Jonti Talukdar, Arjun Chaudhuri, Mayukh Bhattacharya, Krishnendu Chakrabarty |
ITC | 4 |
| 2022 | Special Session: Fault Criticality Assessment in AI AcceleratorsabstractThe ubiquitous application of deep neural networks (DNN) has led to a rise in demand for AI accelerators. DNN-specific functional criticality analysis identifies faults that cause measurable and significant deviations from acceptable requirements such as the inferencing accuracy. This paper examines the problem of classifying structural faults in the processing elements (PEs) of systolic-array accelerators. We first present a two-tier machine-learning (ML) based method to assess the functional criticality of faults. The problem of minimizing misclassification is addressed by utilizing generative adversarial networks (GANs). The two-tier ML/GAN-based criticality assessment method leads to less than 1% test escapes during functional criticality evaluation of structural faults. While supervised learning techniques can be used to accurately estimate fault criticality, it requires a considerable amount of ground truth for model training. We therefore describe a neural-twin framework for analyzing fault criticality with a negligible amount of ground-truth data. A recently proposed misclassification-driven training algorithm is used to sensitize and identify biases that are critical to the functioning of the accelerator for a given application workload. The proposed framework achieves up to 100% accuracy in fault-criticality classification in 16-bit and 32-bit PEs by using the criticality knowledge of only 2% of the total faults in a PE. Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty |
VTS | 3 |
| 2022 | Semi-Supervised Root-Cause Analysis with Co-Training for Integrated SystemsabstractThe increasing complexity of integrated systems has exacerbated the challenges associated with system diagnosis. To tackle these challenges, intelligent root-cause-analysis facilitated by machine learning has been proposed in recent years. However, most of these methods rely on a large amount of data with root-cause labels, which are often either not available or difficult to obtain. In this paper, we propose a semi-supervised root-cause-analysis method with co-training, where only a small set of labeled data is required. Using random forest as the learning kernel, a co-training technique is proposed to leverage the unlabeled data by automatically pre-labeling a subset of them and retraining each decision tree. In addition, several novel techniques are proposed to avoid over-fitting and determine hyper-parameters. Two case studies based on industrial designs demonstrate that the proposed approach significantly outperforms state-of-the-art methods by saving up to 43% of labeling efforts by human experts. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
VTS | 3 |
| 2022 | A Resilient and Hierarchical IoT-Based Solution for Stress Monitoring in Everyday SettingsabstractThe conventional mental healthcare regime often follows a symptom-focused and episodic approach in a noncontinuous manner, wherein the individual discretely records their biomarker levels or vital signs for a short period prior to a subsequent doctor’s visit. Recognizing that each individual is unique and requires continuous stress monitoring and personally tailored treatment, we propose a holistic hybrid edge–cloud Wearable Internet of Things (WIoT)-based online stress monitoring solution to address the above needs. To eliminate the latency associated with cloud access, appropriate edge models—spiking neural network (SNN), Conditionally Parameterized Convolutions (CondConv), and support vector machine (SVM)—are trained, enabling low-energy real-time stress assessment near the subjects on the spot. This work leverages design-space exploration for the purpose of optimizing the performance and energy efficiency of machine learning inference at the edge. The cloud exploits a novel multimodal matching network model that outperforms six state-of-the-art stress recognition algorithms by 2%–7% in terms of accuracy. An offloading decision process is formulated to strike the right balance between accuracy, latency, and energy. By addressing the interplay of edge–cloud, the proposed hierarchical solution leads to a reduction of 77.89% in response time and 78.56% in energy consumption with only a 7.6% drop in accuracy compared to the Internet of Things (IoT)–Cloud scheme, and it achieves a 5.8% increase in accuracy on average compared to the IoT-Edge scheme. Shiyi Jiang, Farshad Firouzi, Krishnendu Chakrabarty, Eric B. Elbogen |
IEEE Internet Things J. | 3 |
| 2022 | Built-in Self-Test and Fault Localization for Inter-Layer Vias in Monolithic 3D ICsabstractMonolithic 3D (M3D) integration provides massive vertical integration through the use of nanoscale inter-layer vias (ILVs). However, high integration density and aggressive scaling of the inter-layer dielectric make ILVs especially prone to defects. We present a low-cost built-in self-test (BIST) method that requires only two test patterns to detect opens, stuck-at faults, and bridging faults (shorts) in ILVs. We also propose an extended BIST architecture for fault detection, called Dual-BIST, to guarantee zero ILV fault masking due to single BIST faults and negligible ILV fault masking due to multiple BIST faults. We analyze the impact of coupling between adjacent ILVs arranged in a 1D array in block-level partitioned designs. Based on this analysis, we present a novel test architecture called Shared-BIST with the added functionality of localizing single and multiple faults, including coupling-induced faults. We introduce a systematic clustering-based method for designing and integrating a delay bank with the Shared-BIST architecture for testing small-delay defects in ILVs with minimal yield loss. Simulation results for four two-tier M3D benchmark designs highlight the effectiveness of the proposed BIST framework. Arjun Chaudhuri, Sanmitra Banerjee, Heechun Park, Bon Woong Ku, Sukeshwar Kannan, Krishnendu Chakrabarty, Sung Kyu Lim |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2022 | Design Automation and Test Solutions for Monolithic 3D ICsabstractMonolithic 3D (M3D) is an emerging heterogeneous integration technology that overcomes the limitations of the conventional through-silicon-via (TSV) and provides significant performance uplift and power reduction. However, the ultra-dense 3D interconnects impose significant challenges during physical design on how to best utilize them. Besides, the unique low-temperature fabrication process of M3D requires dedicated design-for-test mechanisms to verify the reliability of the chip. In this article, we provide an in-depth analysis on these design and test challenges in M3D. We also provide a comprehensive survey of the state-of-the-art solutions presented in the literature. This article encompasses all key steps on M3D physical design, including partitioning, placement, clock routing, and thermal analysis and optimization. In addition, we provide an in-depth analysis of various fault mechanisms, including M3D manufacturing defects, delay faults, and MIV (monolithic inter-tier via) faults. Our design-for-test solutions include test pattern generation for pre/post-bond testing, built-in-self-test, and test access architectures targeting M3D. Lingjun Zhu, Arjun Chaudhuri, Sanmitra Banerjee, Gauthaman Murali, Pruek Vanna-Iampikul, Krishnendu Chakrabarty, Sung Kyu Lim |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2022 | C-Testing and Efficient Fault Localization for AI AcceleratorsabstractAccelerators for machine learning [artificial intelligence (AI)] inferencing applications are homogeneous designs composed of identical cores. Each core or processing element (PE) contains multiply-and-accumulate units, control logic, and registers for storing and forwarding weights and activations. Testing homogeneous array-based AI accelerator chips by running automatic test pattern generation (ATPG) at the array level results in a high CPU time and pattern count. We propose a constant-testable (C-testable) method for test generation at the PE level such that the ATPG effort does not increase with the number of PEs. Our results show that compared to the traditional array-level testing, the proposed method achieves up to$4.2\times $($3.5\times $),$1530\times $($2388\times $), and$170\times $($142\times $) reduction in the test pattern count, ATPG runtime, and test cycle count, respectively, for stuck-at (transition) faults in a$256\times 256$array, while preserving the test coverage. A reconfigurable scan architecture is introduced to enable the proposed C-testable solution for the entire accelerator array. The design-space exploration of a hierarchical test-compaction framework is presented. We also describe four debug solutions for fault localization and diagnosis. Arjun Chaudhuri, Chunsheng Liu 0002, Xiaoxin Fan, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Functional Criticality Analysis of Structural Faults in AI AcceleratorsabstractThe ubiquitous application of deep neural networks (DNNs) has led to a rise in demand for artificial intelligence (AI) accelerators. For example, the tensor processing unit from Google–based on a systolic array–and its variants are of considerable interest for DNN inferencing using AI accelerators. This article studies the problem of classifying structural faults in such an accelerator based on their functional criticality. We first analyze pin-level faults in the processing elements (PEs) of a systolic array. Simulation results for the LeNet network with 8-bit fixed-point, 16-bit floating-point (FP), and 32-bit FP data paths applied to the MNIST dataset show that over 93% of the pin-level structural faults in a PE are functionally benign. We present a greedy iterative framework for determining the criticality of stuck-at faults in a PE netlist and analyze the limitations of criticality analysis methods based on repeated fault simulations. We next present a scalable two-tier machine-learning (ML)-based method to assess the functional criticality of stuck-at faults in a computationally efficient manner. We address the problem of minimizing misclassification by utilizing generative adversarial networks (GANs). Two-tier ML/GAN-based criticality assessment leads to less than 1% test escapes during functional criticality evaluation of structural faults. Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Efficient Identification of Critical Faults in Memristor-Based Inferencing AcceleratorsabstractDeep neural networks (DNNs) are becoming ubiquitous, but hardware-level reliability is a concern when DNN models are mapped to emerging neuromorphic technologies such as memristor-based crossbars. As DNN architectures are inherently fault tolerant and many faults do not affect inferencing accuracy, careful analysis must be carried out to identify faults that are critical for a given application. We present a misclassification-driven training (MDT) algorithm to efficiently identify critical faults (FCFs) in the crossbar. Our results for three DNNs on the CIFAR-10 data set show that MDT can rapidly and accurately identify a large number of FCFs—up to$20\times $faster than a baseline method of forward inferencing with randomly injected faults. We use the set of FCFs obtained using MDT and the set of benign faults obtained using forward inferencing to train a machine learning (ML) model to efficiently classify all the crossbar faults in terms of their criticality. Using the ground truth generated using MDT and forward inferencing, we show that the ML models can classify millions of faults within minutes with a remarkably high classification accuracy of up to 99%. We also show that the ML model trained using CIFAR-10 provides high accuracy when it is used to carry out fault classification for the ImageNet data set. We present a fault-tolerance solution that exploits this high degree of criticality-classification accuracy, leading to a 92.5% reduction in the redundancy needed for fault tolerance. Ching-Yuan Chen, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Formal Synthesis of Adaptive Droplet Routing for MEDA BiochipsabstractA digital microfluidic biochip (DMFB) enables the miniaturization of immunoassays, point-of-care clinical diagnostics, and DNA sequencing. A recent generation of DMFBs uses a microelectrode-dot-array (MEDA) architecture, which provides fine-grained control of droplets and real-time droplet sensing using CMOS technology. However, microelectrodes in a MEDA biochip can degrade due to charge trapping when they are repeatedly charged and discharged during bioassay execution; such degradation leads to the failure of microelectrodes and erroneous bioassay outcomes. To address this problem, we first introduce a new microelectrode-cell design such that we can obtain the health status of all the microelectrodes in a MEDA biochip by employing the inherent sensing mechanism. Next, we present a stochastic game-based model for droplet manipulation, and a formal synthesis method for droplet routing that can dynamically change droplet transportation routes. This adaptation is based on the real-time health information obtained from microelectrodes. Comprehensive simulation results for four real-life bioassays show that our method increases the likelihood of successful bioassay completion with negligible impact on time-to-results. Mahmoud Elfar, Tung-Che Liang, Krishnendu Chakrabarty, Miroslav Pajic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Runtime Malware Detection Using Embedded Trace BuffersabstractAnti-virus software (AVS) tools are used to detect malware in a system. However, AVS are vulnerable to attacks. A malicious entity can exploit these vulnerabilities to subvert the AVS. Recently, hardware components such as hardware performance counters have been used for malware detection. In this article, we propose preempts malware by examining embedded processor traces (PREEMPT), a zero overhead, high-accuracy, low-latency technique to detect malware by repurposing embedded trace buffer (ETB), a debug hardware component available in most modern processors. The ETB is used for postsilicon validation and debug and allows us to control and monitor the internal activities of a chip, beyond what is provided by the input/output pins. PREEMPT combines these hardware-level observations with machine learning-based classifiers to preempt malware before it causes damage. The benefits of reusing ETB for malware detection include the increased robustness against attacks and no performance penalties. PREEMPT can detect malware on an OpenSPARC T1 core running Linux operating system with a F1-score of 96.6%. Rana Elnaggar, Kanad Basu, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Securing SoCs With FPGAs Against Rowhammer AttacksabstractHeterogeneous SoCs integrate FPGAs and microprocessor cores on the same fabric to accelerate applications, such as cryptography and deep learning. Since FPGAs share resources with the microprocessor cores, they can launch noncacheable synchronous DRAM (SDRAM) transactions through direct FPGA-to-microprocessor SDRAM interface. Therefore, if the FPGA 3rd party IPs (3PIPs) are malicious, they can launch rowhammer attacks on the SDRAM. Today’s countermeasures based on performance counters cannot detect these attacks because memory transactions from FPGAs do not pass through the cache. In addition, today’s countermeasures that count the frequency of activation of memory rows cannot identify the intellectual property (IP) that launches the attack from the FPGA. We present a security solution that monitors the SDRAM transactions from IPs on the FPGA to each bank of the microprocessor SDRAM through the FPGA-to-microprocessor SDRAM interface. The proposed monitor is implemented on the FPGA fabric. It can detect attempts to launch a rowhammer attack before it causes bit flips in the SDRAM. It utilizes 6.3% of the adaptive logic modules (ALMs) available in an Intel Cyclone V FPGA, when multiple IPs are monitored. Rana Elnaggar, Peilin Song, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Accurate and Robust Malware Detection: Running XGBoost on Runtime Data From Performance CountersabstractMalware applications are one of the major threats that computing systems face today. While security researchers develop new defense mechanisms to detect malware, attackers continue to release new malware families that evade detection. New defense mechanisms must therefore be developed to effectively counter malware. Hardware performance counters (HPCs) have been recently proposed as a means to detect malware. However, recent work has also shown that malware detection is not effective when performance counters are sampled in realistic scenarios. We show how proper data preprocessing and the use of the XGBoost classifier can be used to improve the performance of malware detection using HPCs by at least 15%. We also show that the proposed method can detect malware early (shortly after its launch) by classifying HPC datastreams at short time intervals. In addition, we propose a multitemporal classification model that ensures the early detection of a high percentage of malware while maintaining overall low false positive rates. Finally, we show that through robust training, the XGBoost classifier shows up to 50x less vulnerability to adversarial attacks that are intended to undermine its malware detection performance. Rana Elnaggar, Lorenzo Servadei, Shubham Mathur, Robert Wille, Wolfgang Ecker, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Efficient Regulation of Synthetic Biocircuits Using Droplet-Aliquot Operations on MEDA BiochipsabstractMicrofluidic platforms have recently emerged as an invaluable component for studying synthetic biology as they are capable of emulating complex molecular networks of biological pathways (biocircuits) on a chip. A special type of biochemical assays, known as biocircuit-regulatory scanning (BRS) assays, is employed to regulate gene expression, enabling comprehensive exploration of related biocircuit parameters. Prior work has provided high-level design methodologies for implementing BRS; however, most of these methods are abstract and cannot be used in practice as they overlook the dynamics of interactions between the samples and the biochip. In this article, we address this limitation by providing a comprehensive framework that implements BRS assays. The proposed framework, named BioScan, includes: 1) a statistical method that selects suitable volumetric ratios of biochemicals used to execute a BRS assay; 2) a high-level synthesis method that generates the specifications of the target BRS assay; 3) a translation technique enabling implementation of BRS on a microelectrode-dot array (MEDA) biochip; and 4) a Dirichlet-regressor that constructs the parameter space of the associated biocircuit. Simulation results show that the proposed framework can efficiently perform parameter-space exploration (PSE) while significantly reducing completion time and reagent cost. Mohamed Ibrahim 0002, Zhanwei Zhong, Bhargab B. Bhattacharya, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | High-Throughput Training of Deep CNNs on ReRAM-Based Heterogeneous Architectures via Optimized Normalization LayersabstractResistive random-access memory (ReRAM)-based architectures can be used to accelerate convolutional neural network (CNN) training. However, existing architectures either do not support normalization at all or they support only a limited version of it. Moreover, it is common practice for CNNs to add normalization layers after every convolution layer. In this work, we show that while normalization layers are necessary to train deep CNNs, only a few such layers are sufficient for effective training. A large number of normalization layers do not improve prediction accuracy; it necessitates additional hardware and gives rise to performance bottlenecks. To address this problem, we proposeDeepTrain, a heterogeneous architecture enabled by a Bayesian optimization (BO) methodology; together, they provide adequate hardware and software support for normalization operations. The proposed BO methodology determines the minimum number of normalization operations necessary for a given CNN. Experimental evaluation indicates that the BO-enabledDeepTrainarchitecture achieves up to$15\times $speedup compared to a conventional GPU for training CNNs with no accuracy loss while utilizing only a few normalization layers. Biresh Kumar Joardar, Aryan Deshwal, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Mixing Models as Integer Factorization: A Key to Sample Preparation With Microfluidic BiochipsabstractMicrofluidic biochips have recently emerged with significant promise and versatility in automating a variety of biochemical protocols on a tiny chip. Sample preparation, which involves the mixing of fluids with a specified target ratio in the minuscule scale, is an essential component of these protocols. Algorithms that optimize on-chip sample-preparation cost and time are closely intertwined with the underlying mixing model, mixing sequence, and fluidic architecture. Although numerous mixing models have been studied in the literature, their impact on the dynamics of mixing steps is hitherto not fully understood. In this article, we show that various mixing models can be envisaged in the light of prime factorization of integers thus establishing a connection among mixing algorithms, chip architectures, and performance. This insight has led to the development of the proposed factorization-based dilution algorithm (FacDA) considering a generalized mixing model suitable for micro-electrode-dot-array (MEDA) biochips. It further leads to target volume oriented dilution algorithm (TVODA) to cater to user’s demand for an output with a given volume. We formulate the optimization problem on the fabric of the satisfiability modulo theory (SMT) while determining mixing sequences. Simulation results on a large number of test-cases reveal thatFacDAandTVODAoutperform the state-of-the-art dilution algorithms for MEDA biochips with respect to reactant cost, mixing time, and waste production. Debraj Kundu, Sudip Roy 0001, Sukanta Bhattacharjee, Sohini Saha, Krishnendu Chakrabarty, P. P. Chakrabarti 0001, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Knowledge Transfer in Board-Level Functional Fault Diagnosis Enabled by Domain AdaptationabstractHigh integration densities and design complexity make board-level functional fault diagnosis extremely difficult. Machine-learning techniques can identify functional faults with high accuracy, but they require a large volume of data to achieve high-prediction accuracy. This drawback limits the effectiveness of traditional machine-learning algorithms for training a model in the early stage of manufacturing, when only a limited amount of fail data and repair records are available. We propose a board-level diagnosis workflow that utilizes domain adaptation (DA) to transfer the knowledge learned from mature boards to a new board in the ramp-up phase. First, based on the requirement of fault diagnosis, we select an appropriate domain-adaptation method to reduce differences between mature boards and the new board. Second, these DA methods utilize information from both the mature and the new boards with carefully designed domain-alignment rules and train a functional fault diagnosis classifier. Experimental results using three complex boards in volume production and one new board in the ramp-up phase show that, with the help of DA and the proposed workflow, the diagnosis accuracy is improved. Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Accelerating Large-Scale Graph Neural Network Training on Crossbar DietabstractResistive random-access memory (ReRAM)-based manycore architectures enable acceleration of graph neural network (GNN) inference and training. GNNs exhibit characteristics of both DNNs and graph analytics. Hence, GNN training/inferencing on ReRAM-based manycore architectures give rise to both computation and on-chip communication challenges. In this work, we leverage model pruning and efficient graph storage to reduce the computation and communication bottlenecks associated with GNN training on ReRAM-based manycore accelerators. However, traditional pruning techniques are either targeted for inferencing only, or they are not crossbar-aware. In this work, we propose a GNN pruning technique called DietGNN. DietGNN is a crossbar-aware pruning technique that achieves high accuracy training and enables energy, area, and storage efficient computing on ReRAM-based manycore platforms. The DietGNN pruned model can be trained from scratch without any noticeable accuracy loss. Our experimental results show that when mapped on to a ReRAM-based manycore architecture, DietGNN can reduce the number of crossbars by over 90% and accelerate GNN training by${\sim }{2}.{7}{\times }$compared to its unpruned counterpart. In addition, DietGNN reduces energy consumption by more than${\sim }{3}.{5}{\times }$compared to the unpruned counterpart. Chukwufumnanya Ogbogu, Aqeeb Iqbal Arka, Biresh Kumar Joardar, Janardhan Rao Doppa, Hai Li 0001, Krishnendu Chakrabarty, Partha Pratim Pande |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Unsupervised Two-Stage Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems have placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligence and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. We propose a two-stage unsupervised root-cause-analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to cluster the data in a coarse-grained manner. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. The proposed method can accommodate both numerical and categorical test items. A combination of the L-method, cross validation, and Silhouette score enables us to automatically determine all hyperparameters. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause-analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Online Fault Detection in ReRAM-Based Computing Systems for InferencingabstractA resistive switching random access memory (ReRAM)-based computing system (RCS) provides an energy-efficient hardware implementation of vector–matrix multiplication for machine-learning hardware. However, it is susceptible to faults due to the immature resistive switching random access memory (ReRAM) fabrication process. We propose an efficient online fault-detection method for RCS. This method monitors the dynamic power consumption of each ReRAM crossbar and determines the occurrence of faults when a changepoint is detected in the monitored power-consumption time series. To estimate the percentage of faulty cells in a faulty ReRAM crossbar, we compute statistical features immediately before and after the changepoint and use them as independent variables; we use the percentage of faulty cells as dependent variables to train a predictive model using machine learning. In this way, the computationally expensive fault localization and error-recovery steps are carried out only when a high fault rate is estimated. Simulation results show that, with the fault-detection method and the predictive model, the test time is significantly reduced, the hardware overhead is negligible, and high classification accuracy for the MNIST and CIFAR-10 datasets using RCS can still be ensured. Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Pruning of Deep Neural Networks for Fault-Tolerant Memristor-based AcceleratorsabstractHardware-level reliability is a major concern when deep neural network (DNN) models are mapped to neuromorphic accelerators such as memristor-based crossbars. Manufacturing defects and variations lead to hardware faults in the crossbar. Although memristor-based DNNs are inherently tolerant to these faults and many faults are benign for a given inferencing application, there is still a non-negligible number of critical faults (CFs) in the memristor crossbars that can lead to misclassification. It is therefore important to efficiently identify these CFs so that fault-tolerance solutions can focus on them. In this paper, we present an efficient technique based on machine learning to identify these CFs; CFs can be identified with over 98% accuracy and at a rate that is 20 times faster than a baseline using random fault injection. We next present a fault-tolerance technique that iteratively prunes a DNN by targeting weights that are mapped to CFs in the memristor crossbars. Our results for the CIFAR-10 data set and several benchmark DNNs show that the proposed pruning technique eliminates up to 95% of the CFs with less than 1% DNN inferencing accuracy loss. This reduction in the total number of CFs leads to a 99% savings in the hardware redundancy required for fault tolerance. Ching-Yuan Chen, Krishnendu Chakrabarty |
DAC | 2 |
| 2021 | ReGraphX: NoC-enabled 3D Heterogeneous ReRAM Architecture for Training Graph Neural NetworksabstractGraph Neural Network (GNN) is a variant of Deep Neural Networks (DNNs) operating on graphs. However, GNNs are more complex compared to traditional DNNs as they simultaneously exhibit features of both DNN and graph applications. As a result, architectures specifically optimized for either DNNs or graph applications are not suited for GNN training. In this work, we propose a 3D heterogeneous manycore architecture for on-chip GNN training to address this problem. The proposed architecture, ReGraphX, involves heterogeneous ReRAM crossbars to fulfill the disparate requirements of both DNN and graph computations simultaneously. The ReRAM-based architecture is complemented with a multicast-enabled 3D NoC to improve the overall achievable performance. We demonstrate that ReGraphX outperforms conventional GPUs by up to 3.5X (on an average 3X) in terms of execution time, while reducing energy consumption by as much as 11X. Aqeeb Iqbal Arka, Janardhan Rao Doppa, Partha Pratim Pande, Biresh Kumar Joardar, Krishnendu Chakrabarty |
DATE | 5 |
| 2021 | Advances in Testing and Design-for-Test Solutions for M3D Integrated CircuitsabstractMonolithic 3D (M3D) integration has the potential to achieve significantly higher device density compared to TSV-based 3D stacking. Sequential integration of transistor layers enables high-density vertical interconnects, known as inter-layer vias (ILVs), However, high integration density and aggressive scaling of the inter-layer dielectric make M3D integrated circuits especially prone to process variations and manufacturing defects. We explore the impact of these fabrication imperfections on chip-performance and present the associated test challenges. We introduce two M3D-specific design-for-test solutions - a low-cost built-in self-test architecture for the defect-prone ILVs and a tier-level fault localization method for yield learning. We describe the impact of defects on the efficiency of delay fault testing and highlight solutions for test generation under constraints imposed by the 3D power distribution network. Sanmitra Banerjee, Arjun Chaudhuri, Shao-Chun Hung, Krishnendu Chakrabarty |
DATE | 4 |
| 2021 | Modeling Silicon-Photonic Neural Networks under UncertaintiesabstractSilicon-photonic neural networks (SPNNs) offer substantial improvements in computing speed and energy efficiency compared to their digital electronic counterparts. However, the energy efficiency and accuracy of SPNNs are highly impacted by uncertainties that arise from fabrication-process and thermal variations. In this paper, we present the first comprehensive and hierarchical study on the impact of random uncertainties on the classification accuracy of a Mach-Zehnder Interferometer (MZI)-based SPNN. We show that such impact can vary based on both the location and characteristics (e.g., tuned phase angles) of a non-ideal silicon-photonic device. Simulation results show that in an SPNN with two hidden layers and 1374 tunable-thermal-phase shifters, random uncertainties even in mature fabrication processes can lead to a catastrophic 70% accuracy loss. Sanmitra Banerjee, Mahdi Nikdast, Krishnendu Chakrabarty |
DATE | 3 |
| 2021 | Fault-Criticality Assessment for AI Accelerators using Graph Convolutional NetworksabstractOwing to the inherent fault tolerance of deep neural networks (DNNs), many structural faults in DNN accelerators tend to be functionally benign. In order to identify functionally critical faults, we analyze the functional impact of stuck-at faults in the processing elements of a 128×128 systolic-array accelerator that performs inferencing on the MNIST dataset. We present a 2-tier machine-learning framework that leverages graph convolutional networks (GCNs) for quick assessment of the functional criticality of structural faults. We describe a computationally efficient methodology for data sampling and feature engineering to train the GCN-based framework. The proposed framework achieves up to 90% classification accuracy with negligible misclassification of critical faults. Arjun Chaudhuri, Jonti Talukdar, Jinwook Jung, Gi-Joon Nam, Krishnendu Chakrabarty |
DATE | 5 |
| 2021 | Efficient Identification of Critical Faults in Memristor Crossbars for Deep Neural NetworksabstractDeep neural networks (DNNs) are becoming ubiquitous, but hardware-level reliability is a concern when DNN models are mapped to emerging neuromorphic technologies such as memristor-based crossbars. As DNN architectures are inherently fault-tolerant and many faults do not affect inferencing accuracy, careful analysis must be carried out to identify faults that are critical for a given application. We present a misclassification-driven training (MDT) algorithm to efficiently identify critical faults (CFs) in the crossbar. Our results for two DNNs on the CIFAR-10 data set show that MDT can rapidly and accurately identify a large number of CFs-up to 20× faster than a baseline method of forward inferencing with randomly injected faults. We use the set of CFs obtained using MDT and the set of benign faults obtained using forward inferencing to train a machine learning (ML) model to efficiently classify all the crossbar faults in terms of their criticality. We show that the ML model can classify millions of faults within minutes with a remarkably high classification accuracy of over 99%. We present a fault-tolerance solution that exploits this high degree of criticality-classification accuracy, leading to a 93% reduction in the redundancy needed for fault tolerance. Ching-Yuan Chen, Krishnendu Chakrabarty |
DATE | 2 |
| 2021 | Formal Synthesis of Adaptive Droplet Routing for MEDA BiochipsabstractA digital microfluidic biochip (DMFB) enables the miniaturization of immunoassays, point-of-care clinical diagnostics, and DNA sequencing. A recent generation of DMFBs uses a micro-electrode-dot-array (MEDA) architecture, which provides fine-grained control of droplets and real-time droplet sensing using CMOS technology. However, microelectrodes in a MEDA biochip can degrade due to charge trapping when they are repeatedly charged and discharged during bioassay execution; such degradation leads to the failure of microelectrodes and erroneous bioassay outcomes. To address this problem, we first introduce a new microelectrode-cell design such that we can obtain the health status of all the microelectrodes in a MEDA biochip by employing the inherent sensing mechanism. Next, we present a stochastic game-based model for droplet manipulation, and a formal synthesis method for droplet routing that can dynamically change droplet transportation routes. This adaptation is based on the real-time health information obtained from microelectrodes. Comprehensive simulation results for four real-life bioassays show that our method increases the likelihood of successful bioassay completion with negligible impact on time-to-results. Mahmoud Elfar, Tung-Che Liang, Krishnendu Chakrabarty, Miroslav Pajic |
DATE | 3 |
| 2021 | Perspectives on Emerging Computation-in-Memory ParadigmsabstractThe traditional Von-Neumann architecture is reaching its limits and finding it difficult to cope up with the ever-increasing demands of modern workloads like artificial intelligence. This demand has fueled the search of technologies that can mimic human brain to efficiently combine both memory and computation within a single device. In this work, we present the state-of-the-art research in the domain of computation-in-memory. In particular, we take a look at memristors and its widespread application in neuromorphic computation. We introduce ReRAMs in terms of their novel computing paradigms and present ReRAM-specific design flows. We address the various circuit opportunities and challenges related to reliability and fault tolerance associated with them. Another high-potential candidate to leverage memory and computation from a single device is Ferroelectric Field-effect Transistor (FeFET). Here we present a co-integration of such FeFETs with another emerging nanotechnology concept, called Reconfigurable Field Effect Transistor (RFET) and discuss the impact of the higher amount of states provided by this combination. Shubham Rai, Anteneh Gebregiorgis, Debjyoti Bhattacharjee, Krishnendu Chakrabarty, Said Hamdioui, Anupam Chattopadhyay, Jens Trommer, Akash Kumar 0001 |
DATE | 5 |
| 2021 | DARe: DropLayer-Aware Manycore ReRAM architecture for Training Graph Neural NetworksabstractGraph Neural Networks (GNNs) are a variant of Deep Neural Networks (DNNs) operating on graphs. GNNs have attributes of both DNNs and graph computation. However, training GNNs on manycore architectures is a challenging task because it involves heavy communication that bottlenecks performance. DropEdge and Dropout, which we collectively refer to as DropLayer, are regularization techniques that can improve the predictive accuracy of GNNs. Moreover, when implemented on a manycore architecture, DropEdge and Dropout are capable of reducing the on-chip traffic. In this paper, we present a ReRAM-based 3D manycore architecture called DARe, tailored for accelerating on-chip training of GNNs. The key component of the DARe architecture is a Network-on-Chip (NoC) that reduces the amount of communication using DropLayer. The reduced traffic prevents communication hotspots and leads to better performance. We demonstrate that DARe outperforms conventional GPUs by up to 6.7X (5.6X on average) in terms of execution time, while being up to 30X (23X on average) more energy efficient for GNN training. Aqeeb Iqbal Arka, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ICCAD | 5 |
| 2021 | ParaMitE: Mitigating Parasitic CNFETs in the Presence of Unetched CNTsabstractCarbon nanotube FETs (CNFETs) are emerging as an alternative to silicon devices for next-generation computing systems. However, imperfect carbon nanotube deposition during CNFET fabrication can lead to the formation of difficult-to-etch CNT aggregates in the active layer. These CNT aggregates can form parasitic CNFETs (para-FETs) that are modulated by adjoining gate contacts or back-end-of-line metal layers, thereby forming conditional shorts and stuck-at faults. We show that even weak (parametric) para-FETs can lead to a degraded static noise margin in CNFET-based design. We propose ParaMitE, a layout optimization method that horizontally flips selected standard cells in situ to minimize the number of para-FETs that can arise due to unetched CNTs. As we modify only the cell orientation (and not the cell placement), the impact on the power, timing, and wire length of the CNFET-based design is negligible. Simulation results for several benchmarks show that the proposed method can mitigate up to 60% of the possible para-FET locations (90% of the most critical locations) with only a 3% increase in the total wire length. ParaMitE can enable yield ramp-up at the foundry by providing guidance on which para-FETs can be avoided by design, and conversely, which CNT aggregates must be removed through processing steps. Sanmitra Banerjee, Arjun Chaudhuri, Gauthaman Murali, Mark Nelson 0005, Sung Kyu Lim, Krishnendu Chakrabarty |
ICCAD | 7 |
| 2021 | Heterogeneous Manycore Architectures Enabled by Processing-in-Memory for Deep Learning: From CNNs to GNNs: (ICCAD Special Session Paper)abstractResistive random-access memory (ReRAM)-based processing-in-memory (PIM) architectures have recently become a popular architectural choice for deep-learning applications. ReRAM-based architectures can accelerate inferencing and training of deep learning algorithms and are more energy efficient compared to traditional GPUs. However, these architectures have various limitations that affect the model accuracy and performance. Moreover, the choice of the deep-learning application also imposes new design challenges that must be addressed to achieve high performance. In this paper, we present the advantages and challenges associated with ReRAM-based PIM architectures by considering Convolutional Neural Networks (CNNs) and Graph Neural Networks (GNNs) as important application domains. We also outline methods that can be used to address these challenges. Biresh Kumar Joardar, Aqeeb Iqbal Arka, Janardhan Rao Doppa, Partha Pratim Pande, Hai Li 0001, Krishnendu Chakrabarty |
ICCAD | 6 |
| 2021 | Multi-Objective Optimization of ReRAM Crossbars for Robust DNN Inferencing under Stochastic NoiseabstractResistive random-access memory (ReRAM) is a promising technology for designing hardware accelerators for deep neural network (DNN) inferencing. However, stochastic noise in ReRAM crossbars can degrade the DNN inferencing accuracy. We propose the design and optimization of a high-performance, area-and energy-efficient ReRAM-based hardware accelerator to achieve robust DNN inferencing in the presence of stochastic noise. We make two key technical contributions. First, we propose a stochastic-noise-aware training method, referred to as ReSNA, to improve the accuracy of DNN inferencing on ReRAM crossbars with stochastic noise. Second, we propose an information-theoretic algorithm, referred to as CF-MESMO, to identify the Pareto set of solutions to trade-off multiple objectives, including inferencing accuracy, area overhead, execution time, and energy consumption. The main challenge in this context is that executing the ReSNA method to evaluate each candidate ReRAM design is prohibitive. To address this challenge, we utilize the continuous-fidelity evaluation of ReRAM designs associated with prohibitive high computation cost by varying the number of training epochs to trade-off accuracy and cost. CF-MESMO iteratively selects the candidate ReRAM design and fidelity pair that maximizes the information gained per unit computation cost about the optimal Pareto front. Our experiments on benchmark DNNs show that the proposed algorithms efficiently uncover high-quality Pareto fronts. On average, ReSNA achieves 2.57% inferencing accuracy improvement for ResNet20 on the CIFAR-10 dataset with respect to the baseline configuration. Moreover, CF-MESMO algorithm achieves 90.91% reduction in computation cost compared to the popular multi-objective optimization algorithm NSGA-II to reach the best solution from NSGA-II. Xiaoxuan Yang 0001, Syrine Belakaria, Biresh Kumar Joardar, Huanrui Yang, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty, Hai Li 0001 |
ICCAD | 7 |
| 2021 | Parallel Droplet Control in MEDA Biochips using Multi-Agent Reinforcement LearningabstractMicrofluidic biochips are being utilized for clinical diagnostics, including COVID-19 testing, because of they provide sample-to-result turnaround at low cost. Recently, microelectrode-dot-array (MEDA) biochips have been proposed to advance microfluidics technology. A MEDA biochip manipulates droplets of nano/picoliter volumes to automatically execute biochemical protocols. During bioassay execution, droplets are transported in parallel to achieve high-throughput outcomes. However, a major concern associated with the use of MEDA biochips is microelectrode degradation over time. Recent work has shown that formulating droplet transportation as a reinforcement-learning (RL) problem enables the training of policies to capture the underlying health conditions of microelectrodes and ensure reliable fluidic operations. However, the above RL-based approach suffers from two key limitations: 1) it cannot be used for concurrent transportation of multiple droplets; 2) it requires the availability of CCD cameras for monitoring droplet movement. To overcome these problems, we present a multi-agent reinforcement learning (MARL) droplet-routing solution that can be used for various sizes of MEDA biochips with integrated sensors, and we demonstrate the reliable execution of a serial-dilution bioassay with the MARL droplet router on a fabricated MEDA biochip. To facilitate further research, we also present a simulation environment based on the PettingZoo Gym Interface for MARL-guided droplet-routing problems on MEDA biochips. Tung-Che Liang, Jin Zhou 0014, Yun-Sheng Chan, Tsung-Yi Ho, Krishnendu Chakrabarty, Cy Lee |
ICML | 5 |
| 2021 | Testing and Fault-Localization Solutions for Monolithic 3D ICs*abstractMonolithic 3D(M3D) integration has the potential to achieve significantly higher device density compared to TSV-based 3D stacking. Sequential integration of transistor layers enables high-density vertical interconnects, known as inter-layer vias (ILVs). However, high integration density and aggressive scaling of the inter-layer dielectric make M3D integrated circuits especially prone to process variations and manufacturing defects. Furthermore, the sequential assembly of M3D tiers and immature fabrication process are prone to inter-tier coupling and performance variations. In view of the impact of these fabrication imperfections on chip performance and the associated test challenges, we describe two M3D-specific design-for-test solutions – a low-cost built-in self-test architecture for the defect-prone ILVs and a tier-level fault localization method for yield learning. Arjun Chaudhuri, Krishnendu Chakrabarty |
ITC-Asia | 2 |
| 2021 | Efficient Fault-Criticality Analysis for AI Accelerators using a Neural Twin∗abstractOwing to the inherent fault tolerance of deep neural network (DNN) models used for classification, many structural faults in the processing elements (PEs) of a systolic array-based AI accelerator are functionally benign. Brute-force fault simulation for determining fault criticality is computationally expensive due to many potential fault sites in the accelerator array and the dependence of criticality characterization of PEs on the functional input data. Supervised learning techniques can be used to accurately estimate fault criticality but it requires ground truth for model training. The ground-truth collection involves extensive and computationally expensive fault simulations. We present a framework for analyzing fault criticality with a negligible amount of ground-truth data. We incorporate the gate-level structural and functional information of the PEs in their "neural twins", referred to as "PE-Nets". The PE netlist is translated into a trainable PE-Net, where the standard-cell instances are substituted by their corresponding "Cell-Nets" and the wires translate to neural connections. Each Cell-Net is a pre-trained DNN that models the Boolean-logic behavior of the corresponding standard cell. In the PE-Net, every neural connection is associated with a bias that represents a perturbation in the signal propagated by that connection. We utilize a recently proposed misclassification-driven training algorithm to sensitize and identify biases that are critical to the functioning of the accelerator for a given application workload. The proposed framework achieves up to 100% accuracy in fault-criticality classification in 16-bit and 32-bit PEs by using the criticality knowledge of only 2% of the total faults in a PE. Arjun Chaudhuri, Ching-Yuan Chen, Jonti Talukdar, Siddarth Madala, Abhishek Kumar Dubey, Krishnendu Chakrabarty |
ITC | 6 |
| 2021 | On-line Functional Testing of Memristor-mapped Deep Neural Networks using Backdoored ChecksumsabstractDeep learning (DL) applications are becoming in- creasingly ubiquitous. However, recent research has highlighted a number of reliability concerns associated with deep neural networks (DNNs) used for DL. In particular, hardware-level reliability of DNNs is of concern when DL models are mapped to specialized neuromorphic hardware such as memristor-based crossbars. Faults in the crossbars can deviate the corresponding DNN model weights from their trained values. It is therefore desirable to have an on-device "checksum" function to indicate if model weights are deviated. We present a backdooring technique that fine-tunes DNN weights to implement the checksum function. The backdoored checksum function is triggered only when inferencing is carried out using a special set of data points with watermarks. We show that backdooring, i.e., fine-tuning of DNN weights, has no impact on the inferencing accuracy of the original DNN model. Moreover, the implemented checksum functions for AlexNet and VGG-16 remarkably outperform baseline approaches. Based on the proposed on-line functional testing solution, we present a computing framework that can efficiently recover the inferencing accuracy of a memristor-mapped DNN from weight deviations. Compared to related recent work, the proposed framework achieves 5.6 × speed-up in time-to-recovery and reduces the on-chip test data volume by 99.99%. Ching-Yuan Chen, Krishnendu Chakrabarty |
ITC | 2 |
| 2021 | Adaptive Methods for Machine Learning-Based Testing of Integrated Circuits and BoardsabstractThe relentless growth in information technology and artificial intelligence (AI) is placing demands on integrated circuits and boards for high performance, added functionality, and low power consumption. However, these new trends lead to high test cost and challenges associated with test planning. Machine learning (ML) provides an opportunity to overcome the challenges associated with the testing of complex systems. Taking the advantages of ML techniques, useful information can be extracted from test data logs, and this information helps facilitate the testing process for both chips and boards. In addition, the ever-growing need to achieve test-cost reduction with no test-quality degradation is driving the adoption of ML-based adaptive methods for testing. Adaptive test methods observe changes in the distribution of test data and dynamically adjust the testing process, thus reducing test cost. In this paper, we describe efficient solutions for adapting machine-learning techniques to testing and diagnosis. To reduce manufacturing cost, we select different test items for chips with different predicted quality levels. To avoid the periodic interruption of computing tasks in accelerators for AI, we describe an efficient online testing method that interrupts the regular computation only when a high defect rate is estimated. To identify board-level functional faults with high accuracy, we utilize online incremental learning and transfer learning to address the practical issues that arise when we deal with real-life test data in a high-volume production environment. Krishnendu Chakrabarty |
ITC | 2 |
| 2021 | A BIST-based Dynamic Obfuscation Scheme for Resilience against Removal and Oracle-guided Attacks*abstractBISTLock is a recently proposed logic-locking technique that integrates a barrier finite-state-machine (FSM) with the built-in self-test (BIST) controller. We demonstrate the vulnerability of BISTLock to removal/bypass attacks and develop countermeasures to make it resilient against not only removal attacks but any form of Oracle-guided attack. Removal resilience is achieved through the incorporation of an input-signal scrambler. We demonstrate the vulnerability of the standalone scrambler to the SAT attack and present a reconfigurable LFSR-based dynamic authenticator that achieves SAT resilience. The proposed solution provides dynamic obfuscation upon the application of an incorrect key and prevents Oracle access to the attacker. We also present a security analysis of the overall system against Oraclefree attacks such as BMC-based sequential SAT and the FSM reverse engineering attack. We evaluate the security strength of the proposed solution and show that hardware overhead is low for a broad set of benchmark circuits. Jonti Talukdar, Amitabh Das, Sohrab Aftabjahani, Peilin Song, Krishnendu Chakrabarty |
ITC | 6 |
| 2021 | Unsupervised Root-Cause Analysis with Transfer Learning for Integrated SystemsabstractThe increasing complexity of integrated systems has exacerbated the problems associated with root-cause analysis. Leveraging advances artificial intelligence, a large amount of intelligent root-cause-analysis methods have been proposed in recent years. However, most of these methods rely on root-cause labels from repair history for defective samples, which are often expensive to obtain. In this paper, we propose an unsupervised root-cause-analysis method that utilizes transfer learning. A two-stage clustering method is first developed by exploiting model selection based on the concept of Silhouette score. Next, a data-selection method based on ensemble learning is proposed to transfer valuable information from a source product to improve the root-cause-analysis accuracy on the target product with insufficient data. Two case studies based on industry designs demonstrate that the proposed approach significantly outperforms other state-of-the-art unsupervised root-cause-analysis methods. Renjian Pan, Xin Li 0001, Krishnendu Chakrabarty |
VTS | 3 |
| 2021 | Security Against Data-Sniffing and Alteration Attacks in IJTAGabstractThe IEEE Std. 1687 (IJTAG) facilitates access to on-chip instruments in complex system-on-chip designs. However, a major security vulnerability in IJTAG has yet to be addressed. IJTAG supports the integration of tapped and wrapped instruments at the IP provider with hidden test-data registers (TDRs). The instruments with hidden TDRs can alter and steal the data that is shifted through them. These attacks are called “data-alteration” and “data-sniffing” attacks, respectively. We propose the addition of shadow TDRs (STDRs) and information-flow tracking logic to protect the shifted in test data from illegitimate alteration and leakage by malicious third-party IPs. We present two security architectures for IJTAG. The first architecture secures the IJTAG against data alteration and incurs no timing overhead. However, it does not secure IJTAG against data-sniffing attacks (DS). The second architecture is an upgrade to the first architecture where we repurpose the use of the STDRs and information-tracking logic to secure the IJTAG against both data-alteration and DS. However, it incurs timing overhead. We present security proofs, simulation results, and the overheads associated with these countermeasures for various benchmarks. We also discuss the tradeoffs in security and overhead between the two proposed architectures. Rana Elnaggar, Ramesh Karri, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | AccuReD: High Accuracy Training of CNNs on ReRAM/GPU Heterogeneous 3-D ArchitectureabstractThe growing popularity of convolutional neural networks (CNNs) along with their complexity has led to the search for efficient computational platforms suitable for them. Resistive random-access memory (ReRAM)-based architectures offer a promising alternative to commonly used GPU-based platforms for training CNNs. However, due to their low-precision storage capability, these architectures cannot support all types of CNN layers and suffer from accuracy loss of the learned model. In addition, ReRAM behavior varies with temperature. High temperature reduces noise margin and introduces additional noise. This makes training of CNNs challenging as outputs can be misinterpreted at higher operating temperatures leading to accuracy loss. In this work, we propose an M3D-enabled heterogeneous architecture: AccuReD, that combines ReRAM arrays with GPU cores, to address these challenges and achieve high accuracy CNN training. AccuReD supports all types of CNN layers and achieve near-GPU accuracy even with low-precision and nonideal behavior of ReRAMs. In addition, to reduce temperature, we present a performance-thermal-aware mapping policy that maps CNN layers to the computing elements of AccuReD. Experimental evaluation indicates that AccuReD does not lose accuracy while accelerating CNN training by 12× on an average compared to conventional GPU-only platforms. Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Hai Li 0001, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Board-Level Functional Fault Identification Using Streaming DataabstractHigh integration densities and design complexity of printed-circuit boards make board-level functional fault identification extremely difficult. Machine learning provides an opportunity to identify functional faults with high accuracy and thereby reduce repair cost. However, the large volume of manufacturing data comes in a streaming format and exhibits time-dependent concept drift in a production environment. These drawbacks limit the effectiveness of traditional machine-learning algorithms. We propose a diagnosis workflow that utilizes online learning to train classifiers incrementally with a small chunk of data at each step. These online-learning algorithms adapt to concept drift quickly with carefully designed update rules. A hybrid algorithm is also proposed to handle the scenario that data for varying numbers of boards are collected at different times. This hybrid algorithm concurrently implements two basic models. For each data chunk, this algorithm chooses the better model with high probability. The experimental results using two boards in high-volume production show that, with the help of online learning and the proposed hybrid algorithm, the F1-score for diagnosis based on binary classifiers can be improved from 57.3% to 81.0%. The top-3 accuracy for diagnosis based on multiclass classifiers can be improved from 78.3% to 91.4%. Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Black-Box Test-Cost Reduction Based on Bayesian Network ModelsabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. In this article, we propose a novel black-box test selection method based on Bayesian networks (BNs), which extract the strong relationship among tests. First, the problem of reducing the black-box test cost is formulated as a constrained optimization problem. Next, multiple structure learning and transfer learning algorithms are implemented to construct BN models. Based on these BN models, we propose an iterative test selection method with a new metric, Bayesian index, for test-cost reduction. In addition, averaging strategies are applied to enhance the reduction performance. Finally, a robust model selection framework is proposed to select the optimal BN model for test-cost reduction. Two case studies with production test data demonstrate that when no prior information is provided, our proposed approach effectively reduces the test cost by up to 14.7%, compared to the state-of-the-art greedy algorithm. Moreover, our proposed approach further reduces the test cost by up to 7.1% when prior information is provided from similar products. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | How Secure Are Checkpoint-Based Defenses in Digital Microfluidic Biochips?abstractA digital microfluidic biochip (DMFB) is a miniaturized laboratory capable of implementing biochemical protocols. Fully integrated DMFBs consist of a hardware platform, controller, and network connectivity, making it a cyber-physical system (CPS). A DMFB CPS is being advocated for safety-critical applications, such as medical diagnosis, drug development, and personalized medicine. Hence, the security of a DMFB CPS is of immense importance to their successful deployment. Recent research has made progress in devising corresponding defense mechanisms by employing so-called checkpoints (CPs). Existing solutions either rely on probabilistic security analysis that does not consider all possible actions an attacker may use to overcome an applied CP mechanism or rely on exhaustive monitoring of DMFB at all time-steps during the assay execution. For devising a defense scheme that is guaranteed to be secure, an exact analysis of the security of a DMFB is needed. This is not available in the current state-of-the-art. In this article, we address this issue by developing an exact method, which uses the deductive power of satisfiability solvers to verify whether a CP-based defense thwarts the execution of an attack. We demonstrate the usefulness of the proposed method by showcasing two applications on practical bioassays: 1) security analysis of various checkpointing strategies and 2) derivation of a counterexample-guided fool-proof secure CP scheme. Mohammed Shayan, Sukanta Bhattacharjee, Robert Wille, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Toward Hardware-Based IP Vulnerability Detection and Post-Deployment Patching in Systems-on-ChipabstractSystem integrators create heterogeneous systems-on-chip (SoCs) by integrating numerous third-party intellectual property blocks (3PIPs) to achieve application-specific design goals. With increasing intellectual property (IP) complexity, 3PIPs can suffer from hardware bugs or they can inadvertently introduce other software-exploitable security threats to the SoC. To ensure the ongoing survivability of new SoCs, we need infrastructure for patching newly discovered IP issues after an SoC has been deployed. To address the increasing risks from 3PIPs, we explore the feasibility and limitations of implementing monitoring and mitigation capabilities in hardware. Our proposed monitoring and mitigation patch (MoP) blocks provide a defensive foundation against critical IP-centric issues, focusing on situations where a system integrator only has interface-level visibility of 3PIP designs. The MoPs are distributed throughout the SoC to monitor and mitigate issues directly in hardware and transparently for potentially compromised software-the MoPs are resilient against run-time compromised software and firmware. We ensure that these monitors are reconfigurable after deployment by implementing them using embedded-FPGAs or as a reprogrammable, fixed-design module. We perform a case study of numerous IP-types and model a selection of security-relevant issues and bugs in the IPs, exploring the relative complexity and potential resource overhead. Our study shows the utility of our proposed approach, with MoP blocks requiring less than ~1.5% of the adaptive logic modules (ALMs) in a Cyclone V FPGA for interface monitoring and issue mitigation per IP. Benjamin Tan 0001, Rana Elnaggar, Jason M. Fung, Ramesh Karri, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Enhancing the Reliability of MEDA Biochips Using IJTAG and Wear LevelingabstractA digital microfluidic biochip (DMFB) enables the miniaturization of immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. A recent generation of biochips uses a micro-electrode-dot-array (MEDA) architecture, which provides fine-grained control of droplets and seamlessly integrates microelectronics and microfluidics using CMOS technology and a TSMC fabrication process. To ensure that bioassays are carried out on MEDA biochips efficiently, high-level synthesis algorithms have recently been proposed. However, as in the case of conventional DMFBs, microelectrodes are likely to fail when they are heavily utilized, and previous methods fail to consider reliability issues. In this article, we first present a new microelectrode cell (MC) design such that the droplet-sensing operation can be enabled/disabled for individual MCs. Next, “partial update” and “partial sensing” operations are presented based on an IEEE Std. 1687 IJTAG network design. Finally, wear-leveling synthesis method is proposed to ensure uniform utilization of MCs on MEDA. A comprehensive set of simulation results demonstrate the effectiveness of the proposed hardware design and design automation methods. Zhanwei Zhong, Tung-Che Liang, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Access-Time Minimization for the IJTAG Network Using Data Broadcast and Hardware ParallelismabstractThe IEEE Std. 1687 facilitates flexible access to on-chip instruments through the JTAG test-access port. This flexibility enables the minimization of the overall access time (OAT), and a number of techniques have been proposed in the literature to achieve this goal. However, the OAT is still high for instruments that require a large amount of test data if this data is shifted through the scan chain serially. In order to further reduce the OAT, we present an efficient test-scheduling method that exploits broadcast and hardware parallelism for instrument access. A broadcast scheduling method is synergistically combined with three parallel IJTAG designs. We show that under different cost criteria, we can select the most efficient parallel IJTAG design such that the equivalent access time (EAT) is minimized. In addition, an interconnect fabric design and an integer-linear-programming method is used to balance the lengths of multiple scan chains. Two industry chip designs and three IJTAG benchmarks are used to evaluate the effectiveness of the proposed method. Zhanwei Zhong, Guoliang Li 0004, Qinfu Yang, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Fault Modeling and Efficient Testing of Memristor-Based MemoryabstractMemristor-based memory technology is one of the emerging memory technologies, which is a potential candidate to replace traditional memories. Efficient test solutions are required to enable the quality and reliability of such products. In previous works, fault models are caused by open, short and bridge defects and parametric variations during the fabrication. However, these fault models cannot describe the bridge defects that cause the state of the faulty cell to an undefined state. In this paper, we analyze the different effects of bridge defects and aggregate their faulty behavior into new fault models, undefined coupling fault and dynamic undefined coupling fault. In addition, an enhanced March algorithm is designed to detect all the modeled faults. In one resistor crossbar with$N$memristors, the enhanced March algorithm requires$8N$write and$7N$read operations with negligible hardware overhead. To reduce the test time, a March RC algorithm is proposed based on read operations with new reference currents, which requires$4N+2$write and$6N$read operations. Analytical results show that the proposed test algorithms can detect all the modeled faults outperforming all the previous methods. Subsequently, a Design-for-Testability scheme is proposed to implement March RC algorithm with a little area overhead. Peng Liu 0045, Zhiqiang You, Jigang Wu, Bosheng Liu, Yinhe Han 0001, Krishnendu Chakrabarty |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Learning to Train CNNs on Faulty ReRAM-based Manycore AcceleratorsabstractThe growing popularity of convolutional neural networks (CNNs) has led to the search for efficient computational platforms to accelerate CNN training. Resistive random-access memory (ReRAM)-based manycore architectures offer a promising alternative to commonly used GPU-based platforms for training CNNs. However, due to the immature fabrication process and limited write endurance, ReRAMs suffer from different types of faults. This makes training of CNNs challenging as weights are misrepresented when they are mapped to faulty ReRAM cells. This results in unstable training, leading to unacceptably low accuracy for the trained model. Due to the distributed nature of the mapping of the individual bits of a weight to different ReRAM cells, faulty weights often lead to exploding gradients. This in turn introduces a positive feedback in the training loop, resulting in extremely large and unstable weights. In this paper, we propose a lightweight and reliable CNN training methodology using weight clipping to prevent this phenomenon and enable training even in the presence of many faults. Weight clipping prevents large weights from destabilizing CNN training and provides the backpropagation algorithm with the opportunity to compensate for the weights mapped to faulty cells. The proposed methodology achieves near-GPU accuracy without introducing significant area or performance overheads. Experimental evaluation indicates that weight clipping enables the successful training of CNNs in the presence of faults, while also reducing training time by 4 X on average compared to a conventional GPU platform. Moreover, we also demonstrate that weight clipping outperforms a recently proposed error correction code (ECC)-based method when training is carried out using faulty ReRAMs. Biresh Kumar Joardar, Janardhan Rao Doppa, Hai Li 0001, Krishnendu Chakrabarty, Partha Pratim Pande |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | Thwarting Bio-IP Theft Through Dummy-Valve-Based ObfuscationabstractResearchers develop bioassays following rigorous experimentation in the lab that involves considerable fiscal and highly-skilled-person-hour investment. Previous work shows that a bioassay implementation can be reverse-engineered by using images or video and control signals of the biochip. Hence, techniques must be devised to protect the intellectual property (IP) rights of the bioassay developer. This study is the first step in this direction and it makes the following contributions: (1) it introduces the use of a dummy valve as a security primitive to obfuscate bioassay implementations; (2) it shows how dummy valves can be used to obscure biochip building blocks such as multiplexers and mixers; (3) it presents design rules and security metrics to design and measure obfuscation. In our preliminary work, we presented the concept through the use of sieve-valve as a dummy-valve. However, sieve-valves are difficult to fabricate. To overcome fabrication complexities, we propose a novel multi-height-valve as an obfuscation primitive. Moreover, we showcase the suitability of multi-height-valve for obfuscation through COMSOL simulations. We demonstrate the practicality of the proposal by fabricating an obfuscated biochip using multi-height valves. We assess the cost-security trade-offs associated with this solution and study the practical implications of dummy-valve based obfuscation on real-life biochips. Mohammed Shayan, Sukanta Bhattacharjee, Ajymurat Orozaliev, Yong-Ak Song, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Performance and Accuracy Tradeoffs for Training Graph Neural Networks on ReRAM-Based ArchitecturesabstractGraph neural network (GNN) is a variant of deep neural networks (DNNs) operating on graphs. However, GNNs are more complex compared with DNNs as they simultaneously exhibit attributes of both DNN and graph computations. In this work, we propose a ReRAM-based 3-D manycore processing-in-memory architecture called ReMaGN, tailored for on-chip training of GNNs. ReMaGN implements GNN training using reduced-precision representation to make the computation faster and reduce the load on the communication backbone. However, reduced precision can potentially compromise the accuracy of training. Hence, we undertake a study of performance and accuracy tradeoffs in such architectures. We demonstrate that ReMaGN outperforms conventional GPUs by up to$9.5\times $(on average$7.1\times $) in terms of execution time, while being up to$42\times $(on average$33.5\times $) more energy efficient without sacrificing accuracy. Aqeeb Iqbal Arka, Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2021 | Variation-Aware Delay Fault Testing for Carbon-Nanotube FET CircuitsabstractSensitivity to process variations and manufacturing defects are major showstoppers for the high-volume manufacturing of carbon nanotube field-effect transistors (CNFETs). These imperfections affect gate delay and may remain undetected when test patterns obtained using conventional test-generation techniques are used. We propose a new test generation method that takes CNFET-specific process variations into account and identifies multiple testable long paths through each node in a netlist. In contrast to state-of-the-art techniques, our method can also handle variations that have a nonlinear impact on the propagation delay. The generated test patterns ensure the detection of delay faults through the longest path, even under random CNFET process variations. The proposed method shows significant improvement in the statistical delay quality level (SDQL) compared with a state-of-the-art technique and a commercial ATPG tool for multiple benchmarks. We observed a minimum of 17.1% improvement in the SDQL offered by our patterns over a test set of the same size generated by the commercial tool. We also show that our method, when integrated with the conventional transition fault test flow, offers a significant improvement in the quality of test patterns under random variations. Moreover, the proposed method is flexible and can be easily extended to other emerging device technologies. Sanmitra Banerjee, Arjun Chaudhuri, August Ning, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Power Supply Noise-Aware At-Speed Delay Fault Testing of Monolithic 3-D ICsabstractMonolithic 3-D (M3-D) integration is an emerging technology that offers significant power, performance, and area benefits for an integrated circuit (IC) design. However, a problem with the 3-D power distribution network in such ICs is that it can lead to high power supply noise (PSN) during the capture cycles in at-speed scan testing for transition delay faults. Therefore, the failure of good chips (i.e., yield loss) resulting from the PSN-induced voltage droop is a major concern for M3-D designs. In this article, we first assess the PSN and voltage droop problems and their impact on path delays for at-speed testing of benchmark M3-D designs. Next, we present an analysis framework to identify test patterns that are most likely to lead to yield loss. We describe a test-pattern reshaping solution based on integer linear programming to make appropriate changes to the test patterns that cause yield loss. Simulation results for four M3-D benchmarks highlight the effectiveness of the proposed solution. Shao-Chun Hung, Yi-Chen Lu, Sung Kyu Lim, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | Reliability-Oriented IEEE Std. 1687 Network Design and Block-Aware High-Level Synthesis for MEDA BiochipsabstractA digital microfluidic biochip (DMFB) enables miniaturization of immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. A recent generation of biochips uses a microelectrode-dot-array (MEDA) architecture, which provides fine-grained control of droplets and seamlessly integrates microelectronics and microfluidics using CMOS technology. To ensure that bioassays are carried out on MEDA biochips efficiently, high-level synthesis algorithms have recently been proposed. However, as in the case of conventional DMFBs, microelectrodes are likely to fail when they are heavily utilized, and previous methods fail to consider reliability issues. In this paper, we present the design of an IEEE Std. 1687 (IJTAG) network and a block-aware high-level synthesis method that can effectively alleviate reliability problems in MEDA biochips. A comprehensive set of simulation results demonstrate the effectiveness of the proposed method. Zhanwei Zhong, Tung-Che Liang, Krishnendu Chakrabarty |
ASP-DAC | 3 |
| 2020 | NodeRank: Observation-Point Insertion for Fault Localization in Monolithic 3D ICs∗abstractMonolithic 3D (M3D) ICs have emerged as a promising technology with significant improvement in power, performance, and area (PPA) over conventional 3D-stacked ICs. However, the sequential assembly of M3D tiers and immature fabrication process are prone to manufacturing defects and intertier process variations. Tier-level fault localization is therefore essential for yield ramp-up and diagnosis. Due to overhead concerns, only a limited number of observation points (OPs) can be inserted on the outgoing inter-layer vias (ILVs) of a tier to enable fault localization. We propose the computationally efficient NodeRank algorithm for observation-point insertion (OPI) on a small subset of outgoing ILVs. An ATPG-independent heuristic is presented, which is several orders-of-magnitude faster than ATPG fault simulation-based OPI. We introduce a metric called degree of fault localization to quantify the effectiveness of OPs. Evaluation results for two-tier M3D benchmark circuits show the effectiveness of the proposed method. Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty |
ATS | 3 |
| 2020 | C-Testing of AI Accelerators *abstractAccelerators for machine learning (AI) inferencing applications are homogeneous designs composed of identical cores. Each core, or processing element (PE), contains multiply-and-accumulate units, control logic, and registers for storing and forwarding weights and activations. Testing homogeneous array-based AI accelerator chips by running automatic test pattern generation (ATPG) at the array level results in a high CPU time and pattern count. We propose a constant-testable (C-testable) method for test generation at the PE level such that the ATPG effort does not increase with the number of PEs. Our results show that, compared to the traditional array-level testing, the proposed method achieves up to 4.2× (3.5 ×), 1530 × (2388 ×), and 170× (142×) reduction in the test pattern count, ATPG runtime, and test cycle count, respectively, for stuck-at (transition) faults in a 256 × 256 array, while preserving the test coverage. A reconfigurable scan architecture is introduced to enable C-testing for the entire accelerator array. Arjun Chaudhuri, Chunsheng Liu 0002, Xiaoxin Fan, Krishnendu Chakrabarty |
ATS | 4 |
| 2020 | Power Supply Noise-Aware Scan Test Pattern Reshaping for At-Speed Delay Fault Testing of Monolithic 3D ICs *abstractMonolithic 3D (M3D) integration is an emerging technology that offers significant power, performance, and area benefits for integrated circuit (IC) design. However, a problem with the 3D power distribution network in such ICs is that it can lead to high power supply noise (PSN) during the capture cycles in at-speed scan testing for transition delay faults. Therefore, the failure of good chips (i.e., yield loss) resulting from the PSN-induced voltage droop is a major concern for M3D designs. In this paper, we first assess the PSN and voltage droop problems, and their impact on path delays for at-speed testing of benchmark M3D designs. Next, we present an analysis framework to identify test patterns that are most likely to lead to yield loss. We describe a test-pattern reshaping solution based on integer linear programming to make appropriate changes to the test patterns that cause yield loss. Simulation results for four M3D benchmarks highlight the effectiveness of the proposed solution. Shao-Chun Hung, Yi-Chen Lu, Sung Kyu Lim, Krishnendu Chakrabarty |
ATS | 4 |
| 2020 | Exploring the Mysteries of System-Level TestabstractSystem-level test, or SLT, is an increasingly important process step in today's integrated circuit testing flows. Broadly speaking, SLT aims at executing functional workloads in operational modes. In this paper, we consolidate available knowledge about what SLT is precisely and why it is used despite its considerable costs and complexities. We discuss the types or failures covered by SLT, and outline approaches to quality assessment, test generation and root-cause diagnosis in the context of SLT. Observing that the theoretical understanding for all these questions has not yet reached the level of maturity of the more conventional structural and functional test methods, we outline new and promising directions for methodical developments leveraging on recent findings from software engineering. Ilia Polian, Jens Anders, Steffen Becker 0001, Paolo Bernardi 0002, Krishnendu Chakrabarty, Nourhan Elhamawy, Matthias Sauer 0002, Adit D. Singh, Matteo Sonza Reorda, Stefan Wagner 0001 |
ATS | 5 |
| 2020 | Design of a Reliable Power Delivery Network for Monolithic 3D ICs*abstractAs Moore’s law hits physical limits, monolithic 3D (M3D) integration based on fine-grained monolithic inter-tier vias is emerging as a promising technique to continue performance, power, and area improvements. However, the design of a reliable power delivery network (PDN) for M3D integrated circuits (ICs) is a formidable challenge due to higher power and current densities. In addition, compared to traditional designs, interconnects in M3D designs are more susceptible to electromigration and stress migration. Yield loss resulting from the power-supply noise (PSN) in functional and testing mode is also a major concern for M3D ICs. In this paper, we describe recent research efforts that provide solutions to mitigate these reliability concerns in M3D ICs. Shao-Chun Hung, Krishnendu Chakrabarty |
DATE | 2 |
| 2020 | GRAMARCH: A GPU-ReRAM based Heterogeneous Architecture for Neural Image SegmentationabstractDeep Neural Networks (DNNs) employed for image segmentation are computationally more expensive and complex compared to the ones used for classification. However, manycore architectures to accelerate the training of these DNNs are relatively unexplored. Resistive random-access memory (ReRAM)-based architectures offer a promising alternative to commonly used GPU-based platforms for training DNNs. However, due to their low-precision storage capability, these architectures cannot support all DNN layers and suffer from accuracy loss of the learned models. To address these challenges, we propose GRAMARCH, a heterogeneous architecture that combines the benefits of ReRAM and GPUs simultaneously by using a high-throughput 3D Network-on-Chip. Experimental results indicate that by suitably mapping DNN layers to processing elements, it is possible to achieve up to 53X better performance compared to conventional GPUs for image segmentation. Biresh Kumar Joardar, Nitthilan Kannappan Jayakodi, Janardhan Rao Doppa, Hai Li 0001, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 6 |
| 2020 | Microfluidic Trojan Design in Flow-based BiochipsabstractMicrofluidic technologies find application in various safety-critical fields such as medical diagnostics, drug research, and cell analysis. Recent work has focused on security threats to microfluidic-based cyberphysical systems and defenses. So far the threat analysis has been limited to the cases of tampering with control software/hardware, which is common to most cyberphysical control systems in general; in a sense, such an approach is not exclusive to microfluidics. In this paper, we present a stealthy attack paradigm that uses characteristics exclusive to the microfluidic devices - a microfluidic trojan. The proposed trojan payload is a valve whose height has been perturbed to vary its pressure response. This trojan can be triggered in multiple ways based on time or specific operations. These triggers can occur naturally in a bioassay or added into the controlling software. We showcase the trojan application in carrying out practical attacks -contamination, parameter-tampering and denial-of-service - on a real-life bioassay implementation. Further, we present guidelines to launch stealthy attacks and to counter them. Mohammed Shayan, Sukanta Bhattacharjee, Yong-Ak Song, Krishnendu Chakrabarty, Ramesh Karri |
DATE | 4 |
| 2020 | Functional-Like Transition Delay Fault Test-Pattern Generation using a Bayesian-Based Circuit ModelabstractFor high-performance integrated circuits with tight timing budgets, full-scan based transition delay fault (TDF) testing is mandatory to ensure high test quality. However, the discrepancy between the scan test mode and the functional mode is problematic. For example, the elevated switching activity during scan test application may degrade circuit performance and lead to overkill. In this paper, we address this problem by generating functional-like TDF test patterns. First, a Bayesian-based circuit model is constructed; the result is an enumeration of circuit states that closely mimics the functional mode. During test generation, the model guides the backtrace and fault propagation procedures more effectively than the conventional SCOAP or COP measures because reconvergent fanout is implicitly included in the model. Experimental results on processor benchmarks, including a MIPS32 and a RISC-V processor, show that the TDF test set generated using the Bayesian-based circuit model not only is more functional-like, but also achieves higher fault coverage. Ching-Yuan Chen, Ching-Hong Cheng, Jiun-Lang Huang, Krishnendu Chakrabarty |
ETS | 4 |
| 2020 | Detection of Rowhammer Attacks in SoCs with FPGAsabstractHeterogeneous SoCs integrate FPGAs and microprocessor cores on the same fabric to accelerate applications such as cryptography and deep learning. Since FPGAs share resources with the microprocessor cores, they can launch non-cacheable SDRAM transactions through direct FPGA-to-microprocessor SDRAM interface. Therefore, if the FPGA 3rd party IPs (3PIPs) are malicious, they can launch rowhammer attacks on the SDRAM. Today's countermeasures based on performance counters cannot detect these attacks because memory transactions from FPGAs do not pass through the cache. In addition, countermeasures that count the frequency of activation of memory rows require structural changes to the memory controller or DRAM chips. Moreover, today's countermeasures cannot identify the IP that launches the attack. We present a security solution that monitors the SDRAM transactions from IPs on the FPGA to each bank of the microprocessor SDRAM through the FPGA-to-microprocessor SDRAM interface. The proposed monitor is implemented on the FPGA fabric. It can detect attempts to launch a rowhammer attack before it causes bit flips in the SDRAM. It utilizes only 1% of the adaptive logic modules (ALMs) available in an Intel Cyclone V FPGA to monitor the transactions from one IP. Rana Elnaggar, Peilin Song, Krishnendu Chakrabarty |
ETS | 4 |
| 2020 | RTL-to-GDS Design Tools for Monolithic 3D ICsabstractIn this paper, we propose RTL-to-GDS design flow for monolithic 3D ICs (M3D) built with carbon nanotube field-effect transistors and resistive memory. Our tool flow is based on commercial 2D tools and smart ways to extend them to conduct M3D design and simulation. We provide a post-route optimization flow, which exploits the full potential of the underlying M3D process design kit (PDK) for power, performance and area (PPA) optimization. We also conduct IR-drop and thermal analysis on M3D designs to improve the reliability. To enhance the testability of our M3D designs, we develop design-for-test (DFT) methodologies and integrate a low-overhead built-in self-test module into our design for testing inter-layer vias (ILVs) as well as logic circuitries in the individual tiers. Our benchmark design is RISC-V Rocketcore, which is an open source processor. Our experiments show 8.1% of power, 19.6% of wirelength and 55.7% of area savings with M3D designs at iso-performance compared to its 2D counterpart. In addition, our IR-drop and thermal analyses indicate acceptable power and thermal integrity in our M3D design. Gauthaman Murali, Pruek Vanna-Iampikul, Dae Hyun Kim 0004, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty, Saibal Mukhopadhyay, Sung Kyu Lim |
ICCAD | 8 |
| 2020 | Adaptive Droplet Routing in Digital Microfluidic Biochips Using Deep Reinforcement LearningabstractWe present and investigate a novel application domain for deep reinforcement learning (RL): droplet routing on digital microfluidic biochips (DMFBs). A DMFB, composed of a two-dimensional electrode array, manipulates discrete fluid droplets to automatically execute biochemical protocols such as point-of-care clinical diagnosis. However, a major concern associated with the use of DMFBs is that electrodes in a biochip can degrade over time. Droplet-transportation operations associated with the degraded electrodes can fail, thereby compromising the integrity of the bioassay outcome. We show that casting droplet transportation as an RL problem enables the training of deep network policies to capture the underlying health conditions of electrodes and to provide reliable fluidic operations. We propose a new RL-based droplet-routing flow that can be used for various sizes of DMFBs, and demonstrate reliable execution of an epigenetic bioassay with the RL droplet router on a fabricated DMFB. To facilitate further research, we also present a simulation environment based on the OpenAI Gym Interface for RL-guided droplet-routing problems on DMFBs. Tung-Che Liang, Zhanwei Zhong, Yaas Bigdeli, Tsung-Yi Ho, Krishnendu Chakrabarty, Richard B. Fair |
ICML | 5 |
| 2020 | Functional Criticality Classification of Structural Faults in AI AcceleratorsabstractThe ubiquitous application of deep neural networks (DNNs) has led to a rise in demand for artificial intelligence (AI) accelerators. This paper studies the problem of classifying structural faults in such an accelerator based on their functional criticality. We analyze the impact of stuck-at faults in the processing elements (PEs) of a $128 \times 128$ systolic array designed to perform classification on the MNIST dataset using both 32-bit and 16-bit data paths. We present a two-tier machine-learning (ML) based method to assess the functional criticality of these faults. We address the problem of minimizing misclassification by utilizing generative adversarial networks (GANs). The two-tier ML/GAN-based criticality assessment method leads to less than 1% test escapes during functional criticality evaluation. Arjun Chaudhuri, Jonti Talukdar, Krishnendu Chakrabarty |
ITC | 4 |
| 2020 | BISTLock: Efficient IP Piracy Protection using BISTabstractThe globalization of IC manufacturing has increased the likelihood for IP providers to suffer financial and reputational loss from IP piracy. Logic locking prevents IP piracy by corrupting the functionality of an IP unless a correct secret key is inserted. However, existing logic-locking techniques can impose significant area overhead and performance impact (delay and power) on designs. In this work, we propose BISTLock, a logic-locking technique that utilizes built-in self-test (BIST) to isolate functional inputs when the circuit is locked. We also propose a set of security metrics and use the proposed metrics to quantify BISTLock's security strength for an open-source AES core. Our experimental results demonstrate that BISTLock is easy to implement and introduces an average of 0.74% area and no power or delay overhead across the set of benchmarks used for evaluation. Jinwook Jung, Peilin Song, Krishnendu Chakrabarty, Gi-Joon Nam |
ITC | 4 |
| 2020 | Online Fault Detection in ReRAM-Based Computing Systems by Monitoring Dynamic Power ConsumptionabstractA ReRAM-based computing system (RCS) provides an energy-efficient hardware implementation of vector-matrix multiplication for machine-learning hardware. However, it is vulnerable to faults due to the immature ReRAM fabrication process. We propose an efficient online fault-detection method for RCS; the proposed method monitors the dynamic power consumption of each ReRAM crossbar and determines the occurrence of faults when a changepoint is detected in the monitored power-consumption time series. In order to estimate the percentage of faulty cells in a faulty ReRAM crossbar, we compute statistical features before and after the changepoint and train a predictive model using machine-learning techniques. In this way, the computationally expensive fault localization and error-recovery steps are carried out only when a high fault rate is estimated. Simulation results show that, with the fault-detection method and the predictive model, the test time is significantly reduced while high classification accuracy for the MNIST and CIFAR-10 datasets using RCS can still be ensured. Krishnendu Chakrabarty |
ITC | 2 |
| 2020 | Unsupervised Root-Cause Analysis for Integrated SystemsabstractThe increasing complexity and high cost of integrated systems has placed immense pressure on root-cause analysis and diagnosis. In light of artificial intelligent and machine learning, a large amount of intelligent root-cause analysis methods have been proposed. However, most of them need historical test data with root-cause labels from repair history, which are often difficult and expensive to obtain. In this paper, we propose a two-stage unsupervised root-cause analysis method in which no repair history is needed. In the first stage, a decision-tree model is trained with system test information to roughly cluster the data. In the second stage, frequent-pattern mining is applied to extract frequent patterns in each decision-tree node to precisely cluster the data so that each cluster represents only a small number of root causes. In additional, L-method and cross validation are applied to automatically determine the hyper-parameters of our algorithm. Two industry case studies with system test data demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised root-cause analysis method. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 4 |
| 2020 | LSTM-based Analysis of Temporally- and Spatially-Correlated Signatures for Intermittent Fault DetectionabstractIntermittent faults are a critical reliability threat in deep submicron VLSI circuits. These faults occur non-deterministically due to unstable hardware and unpredictable operating conditions; they are activated/deactivated with changes in the runtime environment. Online fault prediction models are commonly used to predict soft errors and aging effects. A small set of flip-flops, whose states constitute the signature, conveys information about the fine-grained behavior of the circuit, and serves as the input to a machine-learning (ML) model. The nondeterministic failure mechanisms of intermittent faults, however, result in temporally- and spatially-correlated signatures (TSC-signatures). Moreover, the high-dimensional time-series features impede the use of traditional ML models for intermittent-fault detection. To cope with this challenge, we adapt the TSC-signatures to existing ML detection models. Moreover, we propose a novel detection model based on Recurrent Neural Network with Long Short-Term Memory (LSTM) that is inherently suitable for this problem. Simulation results for the ITC99 benchmark circuits highlight the effectiveness of the proposed model. Xingyi Wang, Li Jiang 0002, Krishnendu Chakrabarty |
VTS | 3 |
| 2020 | 3D-ReG: A 3D ReRAM-based Heterogeneous Architecture for Training Deep Neural NetworksabstractDeep neural network (DNN) models are being expanded to a broader range of applications. The computational capability of traditional hardware platforms cannot accommodate the growth of model complexity. Among recent technologies to accelerate DNN, resistive memory (ReRAM)-based processing-in-memory (PIM) emerged as a promising solution for DNN inference due to its high efficiency for matrix-based computation. We face two major technical challenges in extending the use of ReRAM-based accelerators for training: (1) full-precision data is essential in back-propagation; (2) the need to support both feed-forward and back-propagation aggravates the data-movement burden. We propose a heterogeneous architecture named as 3D-ReG, which leverages full-precision GPU to ensure training accuracy and low-overhead 3D integration to provide low-cost data movements. Moreover, we introduce conservative and aggressive task-mapping schemes, which partition the computation phases in different ways to balance execution efficiency and training accuracy. We evaluate 3D-ReG implemented with two 3D integration technologies, through-silicon vias (TSVs) and monolithic inter-tier vias (MIVs), and compare them with GPU-only and PIM-only counterparts. Various GPU-only platforms using two main-memory technologies (DRAM, ReRAM) and three interconnect technologies (2D, TSV, MIV) are evaluated as well. Experimental results show that 3D-ReG can achieve on average 5.64× training speedup and 3.56× higher energy efficiency compared with the GPU with DRAM as main memory, at the cost of 0.05%–3.39% accuracy drop. We define a new metric, gain-loss ratio (GLR), which quantitatively evaluates the capability of a DNN training hardware in terms of the model accuracy and hardware efficiency. The results of our comparison show that the aggressive task-mapping scheme on MIV-based 3D-ReG outperforms the other methods. Bing Li 0017, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty, Joe X. Qiu, Hai Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | Lotus: A New Topology for Large-scale Distributed Machine LearningabstractMachine learning is at the heart of many services provided by data centers. To improve the performance of machine learning, several parameter (gradient) synchronization methods have been proposed in the literature. These synchronization algorithms have different communication characteristics and accordingly place different demands on the network architecture. However, traditional data-center networks cannot easily meet these demands. Therefore, we analyze the communication profiles associated with several common synchronization algorithms and propose a machine learning--oriented network architecture to match their characteristics. The proposed design, named Lotus, because it looks like a lotus flower, is a hybrid optical/electrical architecture based on arrayed waveguide grating routers (AWGRs). In Lotus, a complete bipartite graph is used within the group to improve bisection bandwidth and scalability. Each pair of groups is connected by an optical link, and AWGRs between adjacent groups enhance path diversity and network reliability. We also present an efficient routing algorithm to make full use of the path diversity of Lotus, which leads to a further increase in network performance. Simulation results show that the network performance of Lotus is better than Dragonfly and 3D-Torus under realistic traffic patterns for different synchronization algorithms. Huaxi Gu, Xiaoshan Yu 0001, Krishnendu Chakrabarty |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | BioCyBig: A Cyberphysical System for Integrative Microfluidics-Driven Analysis of Genomic Association StudiesabstractThis paper presents a research vision to design a large-scale cyberphysical systems (CPS) experimental framework to enable collaborative and coordinated molecular biology studies. This framework will be based on the integration of CPS with microfluidic biochips and cloud computing. It has the potential to drastically advance personalized medicine through knowledge fusion among many research groups, and synchronization of research planning. This framework therefore leads to a better understanding of diseases such as cancer, and helps researchers in identifying effective treatments. A case study from cancer research is discussed to explain the significance of our framework in promoting coordinated genomic studies. Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Jun Zeng 0001 |
IEEE Trans. Big Data | 2 |
| 2020 | Timing-Driven Flow-Channel Network Construction for Continuous-Flow Microfluidic BiochipsabstractThe emergence of flow-based microfluidic biochips (FBMBs) has increased the automation level of biochemical procedures, and these lab-on-a-chip devices are now being used for enzyme-linked immunosorbent assay, point-of-care diagnosis, etc. As fabrication technology advances, the feature size of FBMBs keeps shrinking, thereby introducing a series of knotty challenges to the physical design of FBMBs. In particular, timing-sensitive bioassays, such as forensic DNA typing and chromatin immunoprecipitation, require a highly accurate time-control of fluids within a limited completion time. However, existing work does not consider the real-time requirements of these bioassays. In this paper, we formulate the first practical timing-driven flow-channel network construction problem for FBMBs and present a performance-driven placement and routing algorithm for solving this problem. Given the design specifications of a biochip and its biochemistry application, our goal is to construct a high-quality flow-channel network with minimized timing delay and total cost. The experimental results on 14 benchmarks confirm that our algorithm leads to better timing behavior and lower chip cost. Xing Huang 0001, Tsung-Yi Ho, Krishnendu Chakrabarty, Wenzhong Guo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Hierarchical Symbol-Based Health-Status Analysis Using Time-Series Data in a Core Router SystemabstractTo ensure high reliability and rapid error recovery in commercial core router systems, a health-status analyzer is essential to monitor the different features of core routers. However, traditional health analyzers need to store a large amount of historical data in order to identify health status. The storage requirement becomes prohibitively high when we attempt to carry out long-term health-status analysis for a large number of core routers. We describe the design of a symbol-based health status analyzer that first encodes, as a symbol sequence, the long-term complex time series collected from a number of core routers, and then utilizes the symbol sequence to do health analysis. The symbolic aggregation approximation (SAX), 1d-SAX, moving-average-based trend approximation, and nonparametric symbolic approximation representation methods are implemented to encode complex time series in a hierarchical way. Hierarchical agglomerative clustering and sequitur rule discovery are implemented to learn important global and local patterns. Three classification methods including a vector-space-model-based approach are then utilized to identify the health status of core routers. Data collected from a set of commercial core router systems are used to validate the proposed health-status analyzer. The experimental results show that our symbol-based health status analyzer requires much lower storage than traditional methods, but can still maintain comparable diagnosis accuracy. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Self-Learning and Efficient Health-Status Analysis for a Core Router SystemabstractThe health status of core router systems needs to be analyzed efficiently in order to ensure high reliability and timely error recovery. Although a large amount operational data is collected from core routers, due to high computational complexity and expensive labor cost, only a small part of this data is labeled by experts. The lack of labels is an impediment toward the adoption of supervised learning. We present an iterative self-learning procedure for assessing the health status of a core router. This procedure first computes a representative feature matrix to capture different characteristics of time-series data. Not only statistical-modeling-based features are computed from three general categories but also a recurrent neural network-based autoencoder is utilized to capture a wider range of hidden patterns. Moreover, both minimum-redundancy-maximum-relevance (mRMR) method and fully connected feedforward autoencoder are applied to further reduce dimensionality of extracted feature matrix. Hierarchical clustering is then utilized to infer labels for the unlabeled dataset. Finally, a classifier is built and iteratively updated using both labeled and unlabeled dataset. Field data collected from a set of commercial core routers are used to experimentally validate the proposed health-status analyzer. The experimental results show that the proposed feature-based self-learning health analyzer achieves higher precision and recall than the traditional supervised health analyzer as well the currently deployed rule-based health analyzer. Moreover, it achieves better performance than the three anomaly detection baseline methods under the transformed binary classification scenario. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | An Interlayer Interconnect BIST and Diagnosis Solution for Monolithic 3-D ICsabstractMonolithic 3-D (M3D) integration offers higher-density integration compared to 3-D integration based on through-silicon vias. Advances in testing are however needed to screen defects in M3D integration without significantly impacting manufacturing cost. We propose a built-in-self-test (BIST) solution to target shorts and opens in interlayer vias (ILVs). In the proposed solution, scan cells at the interface of two layers are stitched into a twisted-ring counter (TRC) using their functional outputs and ILVs. The interface-register cells launch and capture tests, and a test path consists of ILVs and a multiplexer. We map the problem of minimizing the length of the wires added to stitch the TRC to that of finding a minimum-cost Hamiltonian circuit in a weighted bipartite graph. Since the weighted Hamiltonian circuit problem is NP-Complete, we propose a heuristic algorithm for this problem. We also propose a framework based on artificial neural network to carry out diagnosis when a chip fails the proposed test. We show using simulations that the proposed BIST solution can detect all opens and shorts. We also show using simulations that the proposed diagnosis framework can accurately estimate the size of defects in ILVs. Abhishek Koneru, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Multitarget Sample Preparation Using MEDA BiochipsabstractSample preparation, as a key procedure in many biochemical protocols, mixes various samples, and/or reagents into solutions that contain the target concentrations. Digital microfluidic biochips (DMFBs) have been adopted as a platform for sample preparation because they provide automatic procedures that require less reactant consumption and reduce human-induced errors. However, the most existing methods only consider two-reactant sample preparation, and they cannot be used for many biochemical applications that involve multiple reactants. In addition, the existing methods that can be used for multiple-reactant sample preparation were proposed on traditional DMFBs where only the (1:1) mixing model is available. In the (1:1) mixing model, only two droplets of the same volume can be mixed at a time, which results in higher completion time and the wastage of valuable reactants. To overcome this limitation, the micro-electrode-dot-array (MEDA) architecture has been introduced; it provides the flexibility of mixing multiple droplets of different volumes in a single operation. In this article, we present a generic multiple-reactant sample preparation algorithm that exploits the novel fluidic operations on MEDA biochips. We also propose an enhanced algorithm that increases the operation-sharing opportunities when multiple target concentrations are needed, and therefore the usage of reactants can be further reduced. The simulated experiments show that the proposed method outperforms existing methods in terms of saving reactant cost, minimizing the number of operations, and reducing the amount of waste. Tung-Che Liang, Yun-Sheng Chan, Tsung-Yi Ho, Krishnendu Chakrabarty, Chen-Yi Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Extending the Lifetime of MEDA Biochips by Selective Sensing on MicroelectrodesabstractA digital microfluidic biochip (DMFB) enables miniaturization of immunoassays, point-of-care clinical diagnostics, and DNA sequencing. A recent generation of DMFBs uses a micro-electrode-dot-array (MEDA) architecture, which provides fine-grained control of droplets and real-time droplet sensing using the CMOS technology. However, microelectrodes in a MEDA biochip degrade when they are charged and discharged frequently during bioassay execution. In this article, we first make the key observation that the droplet-sensing operations contribute up to 94% of all microelectrode actuation in MEDA. Consequently, to reduce the number of droplet-sensing operations, we present a new microelectrode cell (MC) design as well as a selective-sensing method such that only a small fraction of microelectrodes perform droplet sensing during bioassay execution. The selection of microelectrodes that need to perform the droplet sensing is based on an analysis of experimental data. A comprehensive set of simulation results show that the total number of droplet-sensing operations is reduced to only 0.7%, which prolongs the lifespan of a MEDA biochip by 11× without any impact on bioassay time-to-response. Tung-Che Liang, Zhanwei Zhong, Miroslav Pajic, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Test Generation for Flow-Based Microfluidic Biochips With General ArchitecturesabstractFlow-based microfluidic biochips have become a promising platform for complex biochemical assays. As the integration of such chips is increasing, a flexible general reconfigurable platform, fully programmable valve array (FPVA), has emerged. Such a 2-D array comprises regularly arranged valves using which flow-networks with different geometry, size, and connectivity can be constructed dynamically. However, the test generation for such arrays becomes challenging due to the large number of potential flow-networks and transportation paths that can be configured on-chip. In this article, we propose a strategy to generate efficient test patterns for FPVAs based on the concepts of test paths and cuts. These patterns together can cover multiple faults in both flow and control layers. We also introduce the concept of test trees and multiple cuts for a test pattern to deal with faults in FPVAs with multiple ports. Moreover, the proposed method can be applied to generate test patterns for traditional flow-based biochips with predefined architectures. The simulation results demonstrate that defects in FPVAs can be detected reliably by a limited number of test patterns generated by the proposed method. For traditional biochips with predefined architectures, these patterns also exhibit an improved test efficiency. Bing Li 0005, Bhargab B. Bhattacharya, Krishnendu Chakrabarty, Tsung-Yi Ho, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | An Efficient Fault-Tolerant Valve-Based Microfluidic Routing Fabric for Droplet Barcoding in Single-Cell AnalysisabstractSingle-cell analysis is used to gain insights into diseases, such as cancer. Advances in microfluidic solutions have enabled the efficient classification and analysis of a heterogeneous population of cells. Recently, a hybrid microfluidic platform was proposed for concurrent single-cell analysis on thousands of heterogeneous cells. In this design, barcoding droplets are routed using a valve-based routing fabric to label the input cells. However, prior work overlooked defects that are likely to occur during chip fabrication and system integration and the fault tolerance of this routing fabric remains a major concern. We address the above limitation and introduce a low-overhead design technique for guaranteeing the tolerance of single faults, while maintaining the efficiency of the cell-analysis platform. We show that the proposed method is optimal in that it minimizes the overhead in terms of fabric size. Yasamin Moradi, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Toward Secure Checkpointing for Micro-Electrode-Dot-Array BiochipsabstractBiochemical experiments, such as diagnostics must be precise and trusted, and provide quick time to results. This has been enabled by automated digital microfluidics; however, it also exposes these experiments to security threats. Previous work has shown that the critical challenge in securing digital microfluidic devices is the lack of sensing resources. The micro-electrode-dot-array (MEDA) is a next-generation digital microfluidic biochip platform that supports fine-grained control and real-time sensing of droplet movements. These capabilities permit continuous monitoring and checkpoint (CP)-based validation of assay execution on MEDA. This article presents a class of “shadow attacks” that abuse the timing slack in the assay execution. State-of-the-art CP-based validation techniques cannot expose the shadow operations. We overcome this limitation by introducing extra CPs in the assay execution at time instances when the assay is prone to shadow attacks. We achieve this by identifying the conditions that enable shadow attacks. We use these conditions to minimize the number of CPs required to guarantee the correctness of bioassay implementation. Our simulation results confirm the effectiveness and practicality of the defense. Mohammed Shayan, Tung-Che Liang, Sukanta Bhattacharjee, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Synthesis of Tamper-Resistant Pin-Constrained Digital Microfluidic BiochipsabstractDigital microfluidic biochips (DMFBs) are an emerging technology that implements bioassays through manipulation of discrete fluid droplets. Recent results have shown that DMFBs are vulnerable to actuation tampering attacks, where a malicious adversary modifies control signals for the purposes of manipulating results or causing denial-of-service. Such attacks leverage the highly programmable nature of DMFBs. However, practical DMFBs often employ a technique called pin mapping to reduce control pin count while simultaneously reducing the degrees of freedom available for droplet manipulation. Attempts to control specific electrodes as part of an attack cannot be made without inadvertently actuating other electrodes on-chip, which makes the tampering evident. This paper explores this tamper resistance property of pin mapping in detail. We derive relevant security metrics, evaluate the tamper resistance of several existing pin mapping algorithms, and propose a new security-aware pin mapper. Further, we develop integer linear programming-based methodologies for inserting indicator droplets into a DMFB in order to boost tamper resistance. Experimental results show that the proposed techniques can significantly increase the difficulty for an attacker to make stealthy changes to the execution of a bioassay. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Analysis and Design of Tamper-Mitigating Microfluidic Routing FabricsabstractMicrofluidic routing fabrics are reconfigurable primitives that permit the dynamic redirection of fluids on a flow-based microfluidic biochip. Such primitives are bringing the benefits of rapid prototyping and on-the-fly reconfigurability from integrated circuits to the microfluidic domain. An unfortunate side effect of this increased flexibility is susceptibility to tampering. A malicious adversary can alter either the electronic control signals or the pneumatic control lines used to drive the routing fabric. In this paper, we provide a high-level security assessment of microfluidic systems utilizing routing fabrics, and analyze their security under actuation tampering attacks. We show that under reasonable assumptions, the permissible states of a routing fabric form a probability distribution. We provide methods for efficiently determining this distribution through a binary tree representation. We then show how to synthesize routings fabrics that exhibit well-defined behaviors. We call a routing fabric designed in such a way tamper-mitigating, as it makes the effects of tampering probabilistically less severe. We then show how the proposed methodology can be used to protect a forensic DNA barcoding application from attack. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Software-Based Self-Testing Using Bounded Model Checking for Out-of-Order Superscalar ProcessorsabstractGenerating functional tests for processors has been a challenging problem for decades in the very large-scale integration testing field. This paper presents a method that generates software-based self-tests by leveraging bounded model checking (BMC) techniques and targeting, for the first time, out-of-order [out-of-order execution (OOE)] superscalar processors. To combat the state-space explosion associated with BMC, the proposed method starts by combining module-level abstraction-refinement with slicing to reduce the size of the model under verification. Next, an off-the-shelf BMC solver is used on the obtained extended finite-state machines to generate the leading sequences that are necessary to excite internal processor functions. Finally, constrained automatic test-pattern generation is used to cover all structural faults within every function excited by the obtained leading sequences. Experimental results show that the proposed method leads to extremely high fault coverage on the critical components corresponding to OOE operations in functional mode. The method therefore helps in tackling the over-testing problem that is inherent to the full-scan test approach. Ying Zhang 0040, Krishnendu Chakrabarty, Zebo Peng, Ahmed Rezine, Huawei Li 0001, Petru Eles, Jianhui Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | IJTAG-Based Fault Recovery and Robust Microelectrode-Cell Design for MEDA BiochipsabstractA digital microfluidic biochip (DMFB) is an attractive platform for immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. A recent generation of biochips uses a micro-electrode-dot-array (MEDA) architecture, which provides fine-grained controllability of droplets and seamlessly integrates microelectronics and microfluidics using CMOS technology. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, chip testing, fault diagnosis, and fault recovery are critical for MEDA biochips. In this article, we present an effective fault-recovery solution based on the homogeneous structure of MEDA. Since the microelectrode cells (MCs) in an MEDA biochip are identical, we add multiplexers for reconfigurability, whereby an MC with faulty components can use the hardware resources in a neighboring MC. In addition, we use the IEEE 1687 (also known as IJTAG) network to reduce the number of control signals needed for the multiplexers, and to provide flexible subscan chain access for the fault-recovery control flow. A comprehensive set of simulation results demonstrates the effectiveness of the proposed fault-recovery solution for MEDA biochips. Zhanwei Zhong, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Molecular Barcoding as a Defense Against Benchtop Biochemical Attacks on DNA Fingerprinting and Information ForensicsabstractDNA fingerprinting can offer remarkable benefits, especially for point-of-care diagnostics, information forensics, and analysis. However, the pressure to drive down costs is likely to lead to cheap untrusted solutions and a multitude of unprecedented risks. These risks will especially emerge at the frontier between the cyberspace and DNA biology. To address these risks, we perform a forensic-security assessment of a typical DNA-fingerprinting flow. We demonstrate, for the first time, benchtop analysis of biochemical-level vulnerabilities in flows that are based on a standard quantification assay known as polymerase chain reaction (PCR). After identifying potential vulnerabilities, we realize attacks using benchtop techniques to demonstrate their catastrophic impact on the outcome of the DNA fingerprinting. We also propose a countermeasure, in which DNA samples are each uniquely barcoded (using synthesized DNA molecules) in advance of PCR analysis, thus demonstrating the feasibility of our approach using benchtop techniques. We discuss how molecular barcoding could be utilized within a cyber-biological framework to improve DNA-fingerprinting security against a wide range of threats, including sample forgery. We also present a security analysis of the DNA barcoding mechanism from a molecular biology perspective. Mohamed Ibrahim 0002, Tung-Che Liang, Kristin Scott, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Bio-chemical Assay Locking to Thwart Bio-IP TheftabstractIt is expected that as digital microfluidic biochips (DMFBs) mature, the hardware design flow will begin to resemble the current practice in the semiconductor industry: design teams send chip layouts to third-party foundries for fabrication. These foundries are untrusted and threaten to steal valuable intellectual property (IP). In a DMFB, the IP consists of not only hardware layouts but also of the biochemical assays (bioassays) that are intended to be executed on-chip. DMFB designers therefore must defend these protocols against theft. We propose to “lock” biochemical assays by inserting dummy mix-split operations. We experimentally evaluate the proposed locking mechanism, and show how a high level of protection can be achieved even on bioassays with low complexity. We also demonstrate a new class of attacks that exploit the side-channel information to launch sophisticated attacks on the locked bioassay. Sukanta Bhattacharjee, Jack Tang, Sudip Poddar, Mohamed Ibrahim 0002, Ramesh Karri, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2020 | Secure Assay Execution on MEDA Biochips to Thwart Attacks Using Real-Time SensingabstractDigital microfluidic biochips (DMFBs) have emerged as a promising platform for DNA sequencing, clinical chemistry, and point-of-care diagnostics. Recent research has shown that DMFBs are susceptible to various types of malicious attacks. Defenses proposed thus far only offer probabilistic guarantees of security due to the limitation of on-chip sensor resources. A micro-electrode-dot-array (MEDA) biochip is a next-generation DMFB that enables the real-time sensing of on-chip droplet locations, which are captured in the form of a droplet-location map. We propose a security mechanism that validates assay execution by reconstructing the sequencing graph (i.e., the assay specification) from the droplet-location maps and comparing it against the golden sequencing graph. We prove that there is a unique (one-to-one) mapping from the set of droplet-location maps (over the duration of the assay) to the set of possible sequencing graphs. Any deviation in the droplet-location maps due to an attack is detected by this countermeasure because the resulting derived sequencing graph is not isomorphic to the original sequencing graph. We highlight the strength of the security mechanism by simulating attacks on real-life bioassays. We also address the concern that the proposed mechanism may raise false alarms when some fluidic operations are executed on MEDA biochips. To avoid such false alarms, we propose an enhanced sensing technique that provides fine-grained sensing for the security mechanism. Tung-Che Liang, Mohammed Shayan, Krishnendu Chakrabarty, Ramesh Karri |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2020 | Fine-grained Adaptive Testing Based on Quality PredictionabstractThe ever-increasing complexity of integrated circuits inevitably leads to high test cost. Adaptive testing provides an effective solution for test-cost reduction; this testing framework selects the important test items for each set of chips. However, adaptive testing methods designed for digital circuits are coarse-grained, and they are targeted only at systematic defects. To incorporate fabrication variations and random defects in the testing framework, we propose a fine-grained adaptive testing method based on machine learning. We use the parametric test results from the previous stages of test to train a quality-prediction model for use in subsequent test stages. Next, we partition a given lot of chips into two groups based on their predicted quality. A test-selection method based on statistical learning is applied to the chips with high predicted quality. An ad hoc test-selection method is proposed and applied to the chips with low predicted quality. Experimental results using a large number of fabricated chips and the associated test data show that to achieve the same defect level as in prior work on adaptive testing, the fine-grained adaptive testing method reduces test cost by 90% for low-quality chips and up to 7% for all the chips in a lot. Renjian Pan, Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2020 | Algorithmic Fault Detection for RRAM-based Matrix OperationsabstractAn RRAM-based computing system (RCS) provides an energy-efficient hardware implementation of vector-matrix multiplication for machine-learning hardware. However, it is vulnerable to faults due to the immature RRAM fabrication process. We propose an efficient fault tolerance method for RCS; the proposed method, referred to as extended-ABFT (X-ABFT), is inspired by algorithm-based fault tolerance (ABFT). We utilize row checksums and test-input vectors to extract signatures for fault detection and error correction. We present a solution to alleviate the overflow problem caused by the limited number of voltage levels for the test-input signals. Simulation results show that for a Hopfield classifier with faults in 5% of its RRAM cells, X-ABFT allows us to achieve nearly the same classification accuracy as in the fault-free case. Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2020 | Runtime Identification of Hardware Trojans by Feature Analysis on Gate-Level Unstructured Data and Anomaly DetectionabstractAs the globalization of chip design and manufacturing process becomes popular, malicious hardware inclusions such as hardware Trojans pose a serious threat to the security of digital systems. Advanced Trojans can mask many architectural-level Trojan signatures and adapt against several detection mechanisms. Runtime Trojan detection techniques are considered as a last line of defense against Trojan inclusion and activation. In this article, we propose an offline analysis to select a subset of flip-flops as surrogates and build an anomaly detection model based on the activity profile of flip-flops. These flip-flops are monitored online, and the anomaly detection model implemented online analyzes the flip-flop data to detect any anomalous Trojan activity. The effectiveness of our approach has been tested on several Trojan-inserted designs of the Leon3 processor. Trojan activation is detected with an accuracy score of above 0.9 (ratio of the number of true predictions to total number of predictions) with no false positives by monitoring less than 0.5% of the total number of flip-flops. Arunkumar Vijayan, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2020 | Analysis of the Impact of Process Variations and Manufacturing Defects on the Performance of Carbon-Nanotube FETsabstractCarbon-nanotube FETs (CNFETs) are potential successors to CMOS transistors; these emerging devices have a low intrinsic delay due to near-ballistic transport in carbon nanotubes (CNTs). As CNFETs are evaluated for circuit/system design, it is important to analyze variations in CNT process parameters. In this article, we present a systematic approach to quantify the impact of these imperfections on the transistor- and gate-level performances of CNT-based circuits. Process variations are investigated to identify the critical device parameters that have maximum impact on the device on-current. We also present a model that predicts the realistic CNFET yield in the presence of process variations. Finally, the impact of manufacturing defects, such as pinholes in the gate dielectric and parasitic CNFETs formed due to imperfect etching, are modeled and evaluated using HSPICE. Sanmitra Banerjee, Arjun Chaudhuri, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Programmable Daisychaining of Microelectrodes to Secure Bioassay IP in MEDA BiochipsabstractAs digital microfluidic biochips (DMFBs) make the transition to the marketplace for commercial exploitation, security and intellectual property (IP) protection are emerging as important design considerations. Recent studies have shown that DMFBs are vulnerable to reverse engineering aimed at stealing biomolecular protocols (IP theft). The IP piracy of proprietary protocols may lead to significant losses for pharmaceutical and biotech companies. The microelectrode dot array (MEDA) is a next-generation DMFB platform that supports real-time sensing of droplets and has the added advantage of important security protection. However, real-time sensing offers opportunities to an attacker to steal the biochemical IP. We show that the daisychaining of microelectrodes and the use of one-time programmability in MEDA biochips provides effective bitstream scrambling of biochemical protocols. To examine the strength of this solution, we develop a Satisfiability (SAT)-based attack that can unscramble the bitstreams through repeated observations of bioassays executed on the MEDA platform. Based on insights gained from the SAT attack, we propose an advanced defense against IP theft. Simulation results using real-life biomolecular protocols confirm that while the SAT attack is effective for simple instances, our advanced defense can thwart it for realistic MEDA biochips and real-life protocols. Tung-Che Liang, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Sample preparation for multiple-reactant bioassays on micro-electrode-dot-array biochipsabstractSample preparation, as a key procedure in many biochemical protocols, mixes various samples and/or reagents into solutions that contain the target concentrations. Digital microfluidic biochips (DMFBs) have been adopted as a platform for sample preparation because they provide automatic procedures that require less reactant consumption and reduce human-induced errors. However, traditional DMFBs only utilize the (1:1) mixing model, i.e., only two droplets of the same volume can be mixed at a time, which results in higher completion time and the wastage of valuable reactants. To overcome this limitation, a next-generation micro-electrode-dot-array (MEDA) architecture that provides flexibility of mixing multiple droplets of different volumes in a single operation was proposed. In this paper, we present a generic multiple-reactant sample preparation algorithm that exploits the novel fluidic operations on MEDA biochips. Simulated experiments show that the proposed method outperforms existing methods in terms of saving reactant cost, minimizing the number of operations, and reducing the amount of waste. Tung-Che Liang, Yun-Sheng Chan, Tsung-Yi Ho, Krishnendu Chakrabarty, Chen-Yi Lee |
ASP-DAC | 4 |
| 2019 | Execution of provably secure assays on MEDA biochips to thwart attacksabstractDigital microfluidic biochips (DMFBs) have emerged as a promising platform for DNA sequencing, clinical chemistry, and point-of-care diagnostics. Recent research has shown that DMFBs are susceptible to various types of malicious attacks. Defenses proposed thus far only offer probabilistic guarantees of security due to the limitation of on-chip sensor resources. A micro-electrode-dot-array (MEDA) biochip is a next-generation DMFB that enables the sensing of on-chip droplet locations, which are captured in the form of a droplet-location map. We propose a security mechanism that validates assay execution by reconstructing the sequencing graph (i.e., the assay specification) from the droplet-location maps and comparing it against the golden sequencing graph. We prove that there is a unique (one-to-one) mapping from the set of droplet-location maps (over the duration of the assay) to the set of possible sequencing graphs. Any deviation in the droplet-location maps due to an attack is detected by this countermeasure because the resulting derived sequencing graph is not isomorphic to the original sequencing graph. We highlight the strength of the security mechanism by simulating attacks on real-life bioassays. Tung-Che Liang, Mohammed Shayan, Krishnendu Chakrabarty, Ramesh Karri |
ASP-DAC | 3 |
| 2019 | Fault tolerance in neuromorphic computing systemsabstractResistive Random Access Memory (RRAM) and RRAM-based computing systems (RCS) provide energy-efficient technology options for neuromorphic computing. However, the applicability of RCS is limited by reliability problems that arise from the immature fabrication process. In order to take advantage of RCS in practical applications, fault-tolerant design is a key challenge. We present a survey of fault-tolerant designs for RRAM-based neuromorphic computing systems. We first describe RRAM-based crossbars and training architectures in RCS. Following this, we classify fault models into different categories, and review post-fabrication testing methods. Subsequently, online testing methods are presented. Finally, we present various fault-tolerant techniques that were designed to tolerate different types of RRAM faults. The methods reviewed in this survey represent recent trends in fault-tolerant designs of RCS, and are expected motivate further research in this field. Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty |
ASP-DAC | 4 |
| 2019 | Factorization based dilution of biochemical fluids with micro-electrode-dot-array biochipsabstractSample preparation, an essential preprocessing step for biochemical protocols, is concerned with the generation of fluids satisfying specific target ratios and error-tolerance. Recent micro-electrode-dot-array (MEDA)-based DMF biochips provide the advantage of supporting both discrete and dynamic mixing models, the power of which has not yet been fully harnessed for implementing on-chip dilution and mixing of fluids. In this paper, we propose a novel factorization-based algorithm called FacDA for efficient and accurate dilution of sample fluid on a MEDA chip. Simulation results reveal that over a large number of test-cases with the mixing volume constraint in the range of 4--10 units, FacDA requires around 38% fewer mixing steps, 52% less sample units, and generates approximately 23% less wastage, all on average, compared to two prior dilution algorithms used for MEDA chips. Sohini Saha, Debraj Kundu, Sudip Roy 0001, Sukanta Bhattacharjee, Krishnendu Chakrabarty, P. P. Chakrabarti 0001, Bhargab B. Bhattacharya |
ASP-DAC | 5 |
| 2019 | Robust sample preparation on digital microfluidic biochipsabstractSample preparation is an important application for the digital microfluidic biochips (DMFBs) platform, and many methods have been developed to reduce the time and reagent usage associated with on-chip sample preparation. However, errors in fluidic operations can result in the concentration of the resulting droplet being outside the calibration range. Current error-recovery methods have the drawback that they need the use of on-chip sensors and further re-execution time. In this paper, we present two dilution-chain structures that can generate a droplet with a desired concentration even if volume variations occur during droplet splitting. Experimental results show the effectiveness of the proposed method compared to previous methods. Zhanwei Zhong, Robert Wille, Krishnendu Chakrabarty |
ASP-DAC | 3 |
| 2019 | PREEMPT: PReempting Malware by Examining Embedded Processor TracesabstractAnti-virus software (AVS) tools are used to detect Malware in a system. However, software-based AVS are vulnerable to attacks. A malicious entity can exploit these vulnerabilities to subvert the AVS. Recently, hardware components such as Hardware Performance Counters (HPC) have been used for Malware detection. In this paper, we propose PREEMPT, a zero overhead, high-accuracy and low-latency technique to detect Malware by re-purposing the embedded trace buffer (ETB), a debug hardware component available in most modern processors. The ETB is used for post-silicon validation and debug and allows us to control and monitor the internal activities of a chip, beyond what is provided by the Input/Output pins. PREEMPT combines these hardware-level observations with machine learning-based classifiers to preempt Malware before it can cause damage. There are many benefits of re-using the ETB for Malware detection. It is difficult to hack into hardware compared to software, and hence, PREEMPT is more robust against attacks than AVS. PREEMPT does not incur performance penalties. Finally, PREEMPT has a high True Positive value of 94% and maintains a low False Positive value of 2%. Kanad Basu, Rana Elnaggar, Krishnendu Chakrabarty, Ramesh Karri |
DAC | 3 |
| 2019 | RTL-to-GDS Tool Flow and Design-for-Test Solutions for Monolithic 3D ICsabstractMonolithic 3D IC overcomes the limitation of the existing through-silicon-via (TSV) based 3D IC by providing denser vertical connections with nano-scale inter-layer vias (ILVs). In this paper, we demonstrate a thorough RTL-to-GDS design flow for monolithic 3D IC, which is based on commercial 2D place-and-route (P&R) tools and clever ways to extend them to handle 3D IC designs and simulations. We also provide a low-cost built-in-self-test (BIST) method to detect various faults that can occur on ILVs. Lastly, we present a resistive random access memory (ReRAM) compiler that generates memory modules that are to be integrated in monolithic 3D ICs. Heechun Park, Kyungwook Chang, Bon Woong Ku, Daehyun Kim 0002, Arjun Chaudhuri, Sanmitra Banerjee, Saibal Mukhopadhyay, Krishnendu Chakrabarty, Sung Kyu Lim |
DAC | 10 |
| 2019 | System-level hardware failure prediction using deep learningabstractDisk and memory faults are the leading causes of server breakdown. A proactive solution is to predict such hardware failure at the runtime and then isolate the hardware at risk and backup the data. However, the current model-based predictors are incapable of using the discrete time-series data, such as the values of device attributes, which conveys high-level information of the device behavior. In this paper, we propose a novel deep-learning based prediction scheme for system-level hardware failure prediction. We normalize the distribution of samples' attributes from different vendors to make use of diverse training sets. We propose a temporal Convolution Neural Network based model that is insensitive to the noise in the time dimension. Finally, we design a loss function to train the model with extremely imbalanced samples effectively. Experimental results from an open S.M.A.R.T data set and an industrial data set show the effectiveness of the proposed scheme. Xiaoyi Sun, Krishnendu Chakrabarty, Ruirui Huang, Yiquan Chen, Hai Cao, Yinhe Han 0001, Xiaoyao Liang, Li Jiang 0002 |
DAC | 2 |
| 2019 | Multi-Tenant FPGA-based Reconfigurable Systems: Attacks and DefensesabstractPartial reconfiguration of FPGAs improves system performance, increases utilization of hardware resources, and enables run-time update of system capabilities. However, the sharing of FPGA resources among various tenants presents security risks that affect the privacy and reliability of tenant applications running in the FPGA-based system. In this study, we examine the security ramifications of co-tenancy with a focus on address-redirection and task-hiding attacks. We design a counter-measure that protects FPGA-based systems against such attacks and prove that it resists these attacks. We present simulation results and an experimental demonstration using a Xilinx FPGA board to highlight the effectiveness of the countermeasure. The proposed countermeasure incurs negligible cost in terms of the area utilization of FPGAs currently used in the cloud. Rana Elnaggar, Ramesh Karri, Krishnendu Chakrabarty |
DATE | 3 |
| 2019 | BioScan: Parameter-Space Exploration of Synthetic Biocircuits Using MEDA Biochips∗abstractRecent advances in microfluidic technology offer efficient platforms to emulate complex molecular networks of biological pathways (biocircuits) on a lab-on-chip. The behavior of biocircuits is governed by a number of gene-regulatory parameters. A fundamental challenge in synthesizing and verifying biocircuits is the lack of design tools that implement biocircuit-regulatory scanning (BRS) assays to explore the large parameter-space efficiently, while optimizing synthesis time and reagent cost. In this paper, we introduce an optimization flow named BioScan for systematic exploration of the parameter-space of a biocircuit. BioScan includes: (1) a statistical approach to determine a subset of mixing ratios of reagents that span the entire parameter space as densely as possible under cost constraints; (2) an ILP-based synthesis method that implements a BRS-assay on a micro-electrode dot-array biochip. Simulation results show that BioScan reduces reagent cost and enhances space-filling properties. Mohamed Ibrahim 0002, Bhargab B. Bhattacharya, Krishnendu Chakrabarty |
DATE | 3 |
| 2019 | REGENT: A Heterogeneous ReRAM/GPU-based Architecture Enabled by NoC for Training CNNsabstractThe growing popularity of Convolutional Neural Networks (CNNs) has led to the search for efficient computational platforms to enable these algorithms. Resistive random-access memory (ReRAM)-based architectures offer a promising alternative to commonly used GPU-based platforms for CNN training. However, backpropagation in CNNs is susceptible to the limited precision of ReRAMs. As a result, training CNNs on ReRAMs affects the final accuracy of learned model. In this work, we propose REGENT, a heterogeneous architecture that combines ReRAM arrays with GPU cores, and exploits the benefits provided by 3D integration along with a high-throughput yet energy efficient Network-on-Chip (NoC) for training CNNs. We also propose a bin-packing based framework that maps CNN layers and then optimize the placement of computing elements to meet the targeted design objectives. Experimental evaluations indicate that REGENT improves full-system EDP by 55.7% on average compared to conventional GPU-only platforms for training CNNs. Biresh Kumar Joardar, Bing Li 0017, Janardhan Rao Doppa, Hai Li 0001, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 6 |
| 2019 | Desieve the Attacker: Thwarting IP Theft in Sieve-Valve-based BiochipsabstractResearchers develop bioassays following rigorous experimentation in the lab that involves considerable fiscal and highly-skilled-person-hour investment. Previous work shows that a bioassay implementation can be reverse engineered by using images or video and control signals of the biochip. Hence, techniques must be devised to protect the intellectual property (IP) rights of the bioassay developer. This study is the first step in this direction and it makes the following contributions: (1) it introduces use of a sieve-valve as a security primitive to obfuscate bioassay implementations; (2) it shows how sieve-valves can be used to obscure biochip building blocks such as multiplexers and mixers; (3) it presents design rules and security metrics to design and measure obfuscated biochips. We assess the cost-security trade-offs associated with this solution and demonstrate practical sieve-valve based obfuscation on real-life biochips. Mohammed Shayan, Sukanta Bhattacharjee, Yong-Ak Song, Krishnendu Chakrabarty, Ramesh Karri |
DATE | 4 |
| 2019 | Built-in Self-Test for Inter-Layer Vias in Monolithic 3D ICsabstractMonolithic 3D integration provides massive vertical integration through the use of nanoscale inter-layer vias (ILVs). However, high integration density and aggressive scaling of the inter-layer dielectric make ILVs especially prone to defects. We present a low-cost built-in self-test (BIST) method to detect opens, stuck-at faults (SAFs), and bridging faults (shorts) in ILVs. Two test patterns-all-1s and all-0s-are applied to the input side of a set of ILVs (e.g., making up a bus between two tiers). On the adjacent tier (the output side of the ILVs), the test responses are compacted to a 2-bit signature through space compaction. We prove that this compaction solution does not introduce any fault aliasing. Simulations results using HSPICE and M3D benchmark designs show that the proposed BIST method requires low area overhead and test time, but provides effective fault localization and the detectability of a wide range of resistive faults. Arjun Chaudhuri, Sanmitra Banerjee, Heechun Park, Bon Woong Ku, Krishnendu Chakrabarty, Sung Kyu Lim |
ETS | 5 |
| 2019 | Machine Learning-based Prediction of Test PowerabstractWith the increase in circuit complexity, the gap between circuit development time and analysis time has widened. A large database is required in order to perform essential analysis tasks such as power, thermal, and IR-drop analysis, which, in turn, leads to long run times. This work focuses on test power analysis. Due to the large number of test patterns for modern designs and the excessive power analysis run time for each test, it is not feasible to obtain complete power profiles for all the tests. However, test power-safety is essential to produce reliable manufacturing test results and prevent yield loss and chip damage. Accurate power profiling can typically be done for a small subset of pre-selected tests only. An essential task is therefore to determine those tests, which potentially provide the worst-case scenarios with respect to test power. We propose machine learning-based power prediction for test selection. The prediction is applied in two different ways. First, we predict the activity of a test to identify tests with high power consumption. Second, the switching activity and the power information are related to the layout of the chip to identify local hot spots. Various machine learning-based algorithms are used to evaluate this approach. Additionally, the algorithms are compared against each other. The results indicate high prediction accuracy and effectiveness. This makes these algorithms well suited for worst-case test selection. Harshad Dhotre, Stephan Eggersglüß, Krishnendu Chakrabarty, Rolf Drechsler |
ETS | 3 |
| 2019 | Test and Design-for-Testability Solutions for Monolithic 3D Integrated CircuitsabstractM3D integration can result in reduced area and higher performance when compared to 3D die stacking. Due to the benefits of M3D integration, there is growing interest in industry towards the adoption of this technology. However, test challenges for M3D integration have remained largely unexplored. We present three key test challenges for M3D integration: (i) performance variations due to high-density integration, (ii) defect analysis and modeling, and (iii) defect isolation and yield enhancement. For each test challenge, we motivate the need to study its impact on an M3D IC, analyze the effectiveness of existing test solutions, and propose new solutions. Abhishek Koneru, Krishnendu Chakrabarty |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | The Internet of Microfluidic Things: Perspectives on System Architecture and Design Challenges: Invited PaperabstractThe integration of microfluidics and biosensor technology is transforming microbiology research by providing new capabilities for clinical diagnostics, cancer research, and pharmacology studies. This integration enables new approaches for biochemistry automation and cyber-physical adaptation. Similarly, recent years have witnessed the rapid growth of the Internet of Things (IoT) paradigm, where different types of real-world elements such as wearable sensors are connected and allowed to autonomously interact with each other. Combining the advances of both cyber-physical microfluidics and IoT domains can generate new opportunities for knowledge fusion by transforming distributed local microfluidic elements into a global network of coordinated microfluidic systems. This paper aims to streamline this transformation and it presents a research vision for enabling the Internet of Microfluidic Things (IoMT). To leverage advances in connected Microfluidic Things, we highlight new perspectives on system architecture, and describe technical challenges related to design automation, temporal flexibility, security, and service assignment. This vision is supported by case studies from cancer research and pharmacology studies to explain the significance of the proposed framework. Mohamed Ibrahim 0002, Maria Gorlatova, Krishnendu Chakrabarty |
ICCAD | 3 |
| 2019 | Can Multi-Layer Microfluidic Design Methods Aid Bio-Intellectual Property Protection?abstractResearchers develop bioassays by rigorously experimenting in the lab. This involves significant fiscal and skilled person-hour investment. A competitor can reverse engineer a bioassay implementation by imaging or taking a video of a biochip when in use. Thus, there is a need to protect the intellectual property (IP) rights of the bioassay developer. We introduce a novel 3D multilayer-based obfuscation to protect a biochip against reverse engineering. Mohammed Shayan, Sukanta Bhattacharjee, Yong-Ak Song, Krishnendu Chakrabarty, Ramesh Karri |
IOLTS | 4 |
| 2019 | Fault-Tolerant Neuromorphic Computing SystemsabstractThe emergence of non-volatile memories (NVM) such as resistive-oxide random access memory (RRAM), magnetoresistive random access memory (MRAM), and phase change memory (PCM) enables brain-inspired neuromorphic computing. However, due to immature fabrication process, NVMs are prone to process variations and manufacturing defects, which must be investigated for effective defect-to-fault mapping, high-coverage test generation, and diagnostics-driven yield learning. In this paper, we present a survey of research on fault modeling, test generation methodologies, and fault-tolerant design of neuromorphic computing systems based on RRAM and MRAM. Arjun Chaudhuri, Krishnendu Chakrabarty |
ITC | 3 |
| 2019 | Hardware Fault Tolerance for Binary RRAM CrossbarsabstractResistive random-access memory (RRAM)-based computing systems (RCS) are being advocated for neural network acceleration. The memristor is the unit cell of an RCS and it is susceptible to process variations and manufacturing defects. Therefore, it is essential to tolerate faulty memristors to ensure intended system operation. We present the architecture of a novel processing element to tolerate both stuck-at and undefined-state faults in binary RRAM cells. We also describe a 4T1R reconfigurable cell-based crossbar design with an ancillary 3T mesh to provide 100% hardware fault tolerance for random and clustered fault distributions for up to 50% fault density. The proposed 4T1R cell is 2.04× smaller than the state-of-the-art neuromorphic SRAM cell. Evaluation results for binary pattern-matching and digit recognition applications demonstrate the effectiveness of our fault tolerance methodology. Arjun Chaudhuri, Bonan Yan, Yiran Chen 0001, Krishnendu Chakrabarty |
ITC | 4 |
| 2019 | Programmable Daisychaining of Microelectrodes for IP Protection in MEDA BiochipsabstractAs digital microfluidic biochips (DMFBs) make the transition to the marketplace for commercial exploitation, security and intellectual property (IP) protection are emerging as important design considerations. Recent studies have shown that DMFBs are vulnerable to reverse engineering aimed at stealing biomolecular protocols (IP theft). The IP piracy of proprietary protocols may lead to significant losses for pharmaceutical and biotech companies. The micro-electrode-dot-array (MEDA) is a next-generation DMFB platform that supports real-time sensing of droplets and has the added advantage of important security protections. However, real-time sensing offers opportunities to an attacker to steal the biochemical IP. We show that the daisychaining of microelectrodes and the use of one-time-programmability in MEDA biochips provides effective bitstream scrambling of biochemical protocols. To examine the strength of this solution, we develop a SAT attack that can unscramble the bitstreams through repeated observations of bioassays executed on the MEDA platform. Based on insights gained from the SAT attack, we propose an advanced defense against IP theft. Simulation results using real-life biomolecular protocols confirm that while the SAT attack is effective for simple instances, our advanced defense can thwart it for realistic MEDA biochips and real-life protocols. Tung-Che Liang, Krishnendu Chakrabarty, Ramesh Karri |
ITC | 2 |
| 2019 | Knowledge Transfer in Board-Level Functional Fault Identification using Domain AdaptationabstractHigh integration densities and design complexity make board-level functional fault identification extremely difficult. Machine-learning techniques can identify functional faults with high accuracy, but they require a large volume of data to achieve high prediction accuracy. This drawback limits the effectiveness of traditional machine-learning algorithms for training a model in the early stage of manufacturing, when only a limited amount of fail data and repair records are available. We propose a board-level diagnosis workflow that utilizes domain adaptation to transfer the knowledge learned from a mature board to a new board in the ramp-up phase. First, a metric is designed to evaluate the similarity between products, and based on the calculated value of the similarity, either a homogeneous or a heterogeneous domain adaptation algorithm is selected. Second, these domain adaptation algorithms utilize information from both the mature and the new boards with carefully designed domain-alignment rules and train a functional fault identification classifier. Three complex boards in volume production and one new board in the ramp-up phase are used to validate the proposed domain-adaptation approach in terms of the diagnosis accuracy. Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 3 |
| 2019 | Fault Recovery in Micro-Electrode-Dot-Array Digital Microfluidic Biochips Using an IJTAG NetworkBehaviorsabstractA digital microfluidic biochip (DMFB) is an attractive platform for immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. A recent generation of biochips uses a micro-electrode-dot-array (MEDA) architecture, which provides fine-grained controllability of droplets and seamlessly integrates microelectronics and microfluidics using CMOS technology. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, chip testing, fault diagnosis and fault recovery are critical for MEDA biochips. In this paper, we present an effective fault- recovery solution based on the homogeneous structure of MEDA. Since the microelectrode cell (MCs) in a MEDA biochip are identical, we add multiplexers for reconfigurability, whereby an MC with faulty components can use the hardware resources in a neighboring MC. In addition, we use the IEEE 1687 (a.k.a. IJTAG) network to reduce the number of control signals need for the multiplexers, and to provide flexible sub-scan chain access for the fault-recovery control flow. A comprehensive set of simulation results demonstrates the effectiveness of the proposed fault-recovery solution for MEDA biochips. Zhanwei Zhong, Krishnendu Chakrabarty |
ITC | 2 |
| 2019 | Structural Test and Functional Test for Digital Acoustofluidic BiochipsabstractA digital microfluidic biochip (DMB) is an attractive platform for automating laboratory procedures in microbiology. However, a major problem associated with today's DMBs is the risk of cross-contamination due to undesirable fouling of the electrode surface, i.e., droplet materials stick to the surface. To overcome the above problem, a contactless liquid-handling biochip technology referred to as acoustofluidics has recently been proposed, and droplet manipulations on acoustofluidic biochips have also been experimentally demonstrated. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, acoustofluidic biochips must be adequately tested before they are used for bioassay execution. This paper presents the first approach for testing of an acoustofluidic biochip that includes an array of interdigital transducers (IDTs). We first present structural test techniques to evaluate the pass/fail status of each IDT, and identify the type of fault if it fails. In order to ensure correct operation of functional units, e.g., mixers and routers, we also present functional test techniques to address fundamental acoustofluidic operations such as droplet transportation and droplet mixing. We evaluate the proposed test methods using experiments on fabricated acoustofluidic biochips. Zhanwei Zhong, Haodong Zhu, Tony Jun Huang, Krishnendu Chakrabarty |
ITC | 5 |
| 2019 | Reliable Power Delivery and Analysis of Power-Supply Noise During Testing in Monolithic 3D ICsabstractMonolithic 3D (M3D) integration offers significant performance, power, and area benefits. However, the design of a reliable M3D power-delivery network (PDN) is challenging due to high power density and current demand per unit area. We propose a framework to design a reliable PDN for M3D ICs using accurate electrical and reliability models. We leverage genetic programming to explore the design space to optimize the PDN for M3D. We also analyze power-supply noise (PSN) during scan-based testing and compare it with that observed during functional operation. We quantify the impact of PSN during scan-based testing on yield loss. Our results show that the PDN obtained using the proposed approach significantly increases the reliability of at least 40% of the wire segments in the PDN. In addition, the proposed PDN design reduces the worst-case power-supply droop by 50.5% compared to a baseline PDN. The yield loss due to power-supply droop for the proposed design is also significantly lower compared to the baseline. Abhishek Koneru, Aida Todri, Krishnendu Chakrabarty |
VTS | 3 |
| 2019 | Board-Level Functional Fault Identification using Streaming DataabstractHigh integration densities and design complexity of printed-circuit boards make board-level functional fault identification extremely difficult. Machine learning provides an opportunity to identify functional faults with high accuracy and thereby reduce repair cost. However, the large volume of manufacturing data comes in a streaming format and exhibits time-dependent concept drift in a production environment. These drawbacks limit the effectiveness of traditional machine-learning algorithms. We propose a diagnosis workflow that utilizes online learning to train classifiers incrementally with a small chunk of data at each step. These online learning algorithms adapt to concept drift quickly with carefully designed update rules. A hybrid algorithm is also proposed to handle the scenario that data for varying numbers of boards are collected at different times. Experimental results using two boards in high-volume production show that, with the help of online learning and the proposed hybrid algorithm, the F1-score for diagnosis can be improved from 57.3% to 78.9%. Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 4 |
| 2019 | Black-Box Test-Coverage Analysis and Test-Cost Reduction Based on a Bayesian Network ModelabstractThe growing complexity of circuit boards makes manufacturing test increasingly expensive. In order to reduce test cost, a number of test selection methods have been proposed in the literature. However, only few of these methods can be applied to black-box test-cost reduction. The conventional greedy algorithm, which selects the most important tests by considering both strong and weak relationships among tests, suffers from overfitting. In order to overcome overfitting, we propose a novel black-box test selection method based on a Bayesian network model. First, the problem of reducing black-box test cost is formulated as a constrained optimization problem. Next, a score-based algorithm is implemented to construct the Bayesian network for black-box tests. Finally, we propose a Bayesian index with the property of Markov blankets, and then an iterative test selection method is developed based on our proposed Bayesian index. The proposed approach ensures that only the strong relationships among black-box tests are used for test selection so that this approach is more robust to overfitting. Two case studies with production test data demonstrate that the proposed approach effectively reduces test cost by up to 14.7%, compared to a conventional greedy algorithm. Renjian Pan, Zhaobo Zhang, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
VTS | 4 |
| 2019 | Test-Cost Reduction for 2.5D ICs Using Microspring Technology for Die Attachment and ReworkabstractInterposer-based 2.5D integrated circuits (ICs) are being increasingly adopted in the semiconductor industry for FPGAs and GPUs. However, the cost of testing is still a major concern for 2.5D ICs because if a faulty die is detected after it is bonded to the interposer, the entire 2.5D assembly has to be discarded. We consider 2.5D integration based on the use of microsprings for attaching dies to the interposer. A key advantage of microsprings is that they allow the 2.5D assembly to be reworkable. If a faulty die is detected during post-bond testing, we can replace the faulty die with a fault-free one. In order to quantify the benefit of the reworkable 2.5D assembly, we present a test-flow selection method for 2.5D ICs with microsprings. We compare the test cost of microspring-based integration with a baseline of test flows for microbump-only integration, with respect to some key parameters such as pre-bond test cost, fault coverage of tests, and microspring cost. For a large number of dies and a relatively low die yield, microsprings provide significant benefits over the baseline. Zhanwei Zhong, Tom B. Wrigglesworth, Eugene M. Chow, Krishnendu Chakrabarty |
VTS | 4 |
| 2019 | Efficient Generation of Dilution Gradients With Digital Microfluidic BiochipsabstractDigital microfluidic biochips (DMFBs) are now being extensively used to automate several biochemical laboratory protocols such as clinical analysis, point-of-care diagnostics, or DNA sequencing. In many biological assays, e.g., bacterial susceptibility tests and cellular response analysis, samples, or reagents are required in multiple concentration (or dilution) factors, satisfying certain gradient patterns such as linear, exponential, or parabolic. Dilution gradients are traditionally prepared using continuous-flow microfluidic devices. Unfortunately, most of them suffer from inflexibility and nonprogrammability, and they require large volumes of costly stock-solutions. DMFBs, on the other hand, are shown to produce, more efficiently, samples with multiple dilution factors. However, none of the existing DMFB-based algorithms utilize the properties of the gradient-profile while optimizing reactant-cost and sample-preparation time. In this paper, we explore the underlying combinatorial attributes of different gradients and harnessed them for efficient production of the desired concentration profile. For linear gradients, we present theoretical results concerning the number of mix-split operations and waste production, and prove an upper bound on on-chip storage requirement. A cost-effective method for generating a wide class of exponential gradients is also proposed. Finally, in order to handle a complex-shaped gradient, we posit a digital-geometric technique to approximate it with a sequence of linear gradients. Experimental results on various gradient-profiles are presented in support of the proposed method. Sukanta Bhattacharjee, Ansuman Banerjee, Tsung-Yi Ho, Krishnendu Chakrabarty, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Synthesis of a Cyberphysical Hybrid Microfluidic Platform for Single-Cell AnalysisabstractSingle-cell genomics is used to advance our understanding of diseases, such as cancer. Microfluidic solutions have recently been developed to classify cell types or perform single-cell biochemical analysis on preisolated types of cells. However, new techniques are needed to efficiently classify cells and conduct biochemical experiments on multiple cell types concurrently. Nondeterministic cell-type identification, system integration, and design automation are major challenges in this context. To overcome these challenges, we present a hybrid microfluidic platform that enables complete single-cell analysis on a heterogeneous pool of cells. We combine this architecture with an associated design-automation and optimization framework, referred to as co-synthesis (CoSyn). The proposed framework employs real-time resource allocation to coordinate the progression of concurrent cell analysis. Besides this framework, a probabilistic model based on a discrete-time Markov chain is also deployed to investigate protocol settings, where experimental conditions, such as sonication time, vary probabilistically among cell types. Simulation results show that CoSyn efficiently utilizes platform resources and outperforms baseline techniques. Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Synthesis of Reconfigurable Flow-Based Biochips for Scalable Single-Cell ScreeningabstractSingle-cell screening is used to sort a stream of cells into clusters (or types) based on prespecified biomarkers, thus supporting type-driven biochemical analysis. Reconfigurable flow-based microfluidic biochips (RFBs) can be utilized to screen hundreds of heterogeneous cells within a few minutes, but they are overburdened with the control of a large number of valves. To address this problem, we present a pin-constrained RFB design methodology for single-cell screening. The proposed design is analyzed using computational fluid dynamics simulations, mapped to an RC-lumped model, and combined with intervalve connectivity information to construct a high-level synthesis framework, referred to as cell sorter using multiplexed control (Sortex). Simulation results show that Sortex significantly reduces the number of control pins and fulfills the timing requirements of single-cell screening. Mohamed Ibrahim 0002, Aditya Sridhar, Krishnendu Chakrabarty, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Changepoint-Based Anomaly Detection for Prognostic Diagnosis in a Core Router SystemabstractPrognostic diagnosis is desirable for commercial core router systems to ensure early failure prediction and fast error recovery. The effectiveness of prognostic diagnosis depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the statistical properties of the monitored data change significantly as time proceeds. We describe the design of a changepoint (CP)-based anomaly detector that first detects CPs from collected time-series data, and then utilizes these CPs to detect anomalies. Different CP detection approaches are implemented to detect various types of CPs. A clustering method is then developed to identify normal/abnormal patterns from CP windows. Data collected from a set of commercial core router systems are used to validate the proposed anomaly detector. Experimental results show that our CP-based anomaly detector achieves better performance than traditional methods in terms of two metrics, namely success ratio and nonfalse-alarm ratio. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | A Design-for-Test Solution Based on Dedicated Test Layers and Test Scheduling for Monolithic 3-D Integrated CircuitsabstractMonolithic 3-D (M3D) integration has the potential to achieve significantly higher device density compared to 3-D integration based on through-silicon vias. We propose a test solution for M3D ICs based on dedicated test layers, which are inserted between functional layers. We evaluate the cost associated with the proposed design-for-test (DfT) solution and compare it with that for a potential DfT solution based on the IEEE Std. P1838. Our results show that the proposed DfT solution is more cost-efficient than the P1838-based DfT solution for a wide range of interlayer via density. We also present a test scheduling and optimization technique for wafer-level testing of M3D ICs. The proposed technique provides test schedules with minimum test time under power consumption and probe pad constraints. Abhishek Koneru, Sukeshwar Kannan, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Optimization of Multi-Target Sample Preparation On-Demand With Digital Microfluidic BiochipsabstractSample preparation is a fundamental preprocessing step needed in almost all biochemical assays and is conveniently automated on a microfluidic lab-on-chip. In digital microfluidics, it is accomplished by a sequence of droplet-mix-split steps on a biochip. Many real-life applications require a sample with multiple concentration factors (CFs). Existing algorithms, while producing multi-CF targets, attempt to share the mix-split steps in order to reduce reactant-cost and sample-preparation time. However, all prior approaches have two limitations: 1) sharing of intermediate droplets can be best effected only when all required target CFs are known a priori and 2) the processing time may vary depending on the allowable error-tolerance in target-CFs. In this paper, we present a cost-effective solution to multi-CF-dilution on-demand, by using only one (or two) mix-split step(s). In order to service dynamically arriving requests of multiple CFs quickly, we prepare dilutions of the sample with a few CFs in advance (called source-CFs), and fill on-chip reservoirs with these fluids. For minimizing the number of such preprocessed CFs, we present an integer linear programming-based method, an approximation algorithm, and a heuristic algorithm. The proposed methods also allow the users to tradeoff the number of on-chip reservoirs against service time for various applications. Simulation results for several target sets demonstrate the superiority of the proposed techniques over prior art in terms of the number of mix-split steps, waste droplets, and reactant usage when the on-chip reservoirs are preloaded with source-CFs using a customized droplet-streaming engine. Sudip Poddar, Sukanta Bhattacharjee, Subhas C. Nandy, Krishnendu Chakrabarty, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Predicting X-Sensitivity of Circuit-Inputs on Test-Coverage: A Machine-Learning ApproachabstractDigital circuits are often prone to suffer from uncertain timing, inadequate sensor feedback, limited controllability of past states or inability of initializing memory-banks, and erroneous behavior of analog-to-digital converters, which may produce an unknown (${X}$) logic value at various circuit nodes. Additionally, many design bugs that are identified during the post-silicon validation phase manifest themselves as${X}$-values. The presence of such${X}$-sources on certain primary or secondary inputs of a logic circuit may cause loss of fault-coverage of a test set, which, in turn, may impact its reliability and robustness. In this paper, we provide a mechanism for predicting the sensitivity of${X}$-sources in terms of loss of fault-coverage, on the basis of learning only a few structural features of the circuit that are easy to extract from the netlist. We show that the${X}$-sources can be graded satisfactorily according to their sensitivity using support vector regression, thereby obviating the need for costly explicit simulation. Experimental results on several benchmark circuits demonstrate the efficacy, speed, and accuracy of prediction. Manjari Pradhan, Bhaswar B. Bhattacharya, Krishnendu Chakrabarty, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Security Assessment of Micro-Electrode-Dot-Array BiochipsabstractDigital microfluidic biochips (DMFBs) are versatile, reconfigurable systems for manipulating discrete fluid droplets. Building on the success of DMFBs, platforms based on “sea-of-electrodes,” the micro-electrode-dot-array (MEDA), has been proposed to further increase scalability and reconfigurability. Research has shown that DMFBs are susceptible to actuation tampering attacks which alter control signals and result in fluid manipulation; such attacks have yet to be studied in the context of MEDA biochips. In this paper, we assess the security of MEDA biochips under such attacks, and further argue that it is inherently a more secure platform than traditional DMFBs. First, we identify a new class of actuation tampering attacks specific to MEDA biochips: the micro-droplet attack. We show that this new attack is stealthy as it produces a subtler difference in results compared to traditional DMFBs. We then illustrate our findings through a case study of an MEDA biochip implementing a glucose measurement assay. Second, we enumerate the system features required to secure an MEDA biochip against actuation tampering attacks and show that these features are naturally implemented in MEDA. Mohammed Shayan, Jack Tang, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Toward Secure and Trustworthy Cyberphysical Microfluidic BiochipsabstractTechnological shifts in the fields of microfluidics and security are now converging. New techniques in microfluidics increasingly rely on cyberphysical integration and concepts from computer-aided design automation to provide ease-of-use, reliability, and higher throughput. Meanwhile, security concerns are extending beyond traditional information technologies as low-cost computing and sensing proliferates into an ever-increasing number of devices. This keynote paper highlights recent findings and trends in these field to motivate research in the nascent field of cyberphysical microfluidic biochip security and trust. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Defect Clustering-Aware Spare-TSV Allocation in 3-D ICs for Yield EnhancementabstractThe manufacturing yield challenge of 3-D integrated circuit is one of the key obstacles in the industry adoption of 3-D integration based on through-silicon-vias (TSVs). The addition of spare TSVs to repair faulty functional TSVs (f-TSVs) is an effective method for yield and reliability enhancement, but this approach results in significant hardware cost and delay overhead. Most existing solutions are only suitable for a “dual-uniform” scenario in which both the placement and the defect probabilities of f-TSVs are assumed to be uniform. In this paper, we propose a design technique that is compatible with nonuniform TSV placement and it can repair faulty TSVs based on a realistic clustered defect-distribution model. The proposed solution is based on two consecutive stages, which utilize a greedy algorithm and an integer-linear-programming formulation, respectively. By considering the tradeoff between chip yield, hardware cost, and delay overhead, the proposed technique provides higher yield and reliability under a clustered defect distribution, and with minimum hardware cost and delay overhead, compared to the previous work. Shengcheng Wang, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Fault-Tolerant Training Enabled by On-Line Fault Detection for RRAM-Based Neural Computing SystemsabstractAn resistive random-access memory (RRAM)-based computing system (RCS) is an attractive hardware platform for implementing neural computing algorithms. On-line training for RCS enables hardware-based learning for a given application and reduces the additional error caused by device parameter variations. However, a high occurrence rate of hard faults due to immature fabrication processes and limited write endurance restrict the applicability of on-line training for RCS. We propose a fault-tolerant on-line training method that alternates between a fault-detection phase and a fault-tolerant training phase. In the fault-detection phase, a quiescent-voltage comparison method is utilized. In the training phase, a threshold-training method and a remapping scheme is proposed. Our results show that, compared to neural computing without fault tolerance, the recognition accuracy for the Cifar-10 dataset improves from 37% to 83% when using low-endurance RRAM cells, and from 63% to 76% when using RRAM cells with high endurance but a high percentage of initial faults. Lixue Xia, Xuefei Ning, Krishnendu Chakrabarty, Yu Wang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Synterface: Efficient Chip-to-World Interfacing for Flow-Based Microfluidic Biochips Using Pin-Count MinimizationabstractFlow-based microfluidic biochips can be used to perform bioassays by manipulating a large number of on-chip valves. These biochips are increasingly used today for biomolecular recognition, single-cell screening, and point-of-care disease diagnostics, and design-automation solutions for flow-based microfluidics enable the mapping and optimization of bimolecular protocols and software-based valve control. However, a key problem that has not received adequate attention is chip-to-world interfacing, which requires the use of off-chip control equipment to provide control signals for the on-chip valves. This problem is exacerbated by the increase in the number of valves as chips get more complex. To address the interfacing problem, we present an efficient pin-count minimization (synthesis) problem, referred to as Synterface, which uses on-chip microfluidic logic gates and optimization based on concepts from linear algebra. We present results to show that Synterface significantly reduces pin-count and simplifies the external interface for flow-based microfluidics. Aditya Sridhar, Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | Bio-Protocol Watermarking on Digital Microfluidic BiochipsabstractAdvancements in digital microfluidic biochip (DMFB) technologies are paving the way for low-cost and automated platforms for implementing bio-protocols. However, the deployment of DMFBs outside of controlled settings will make them vulnerable to intellectual property (IP) theft. Bio-protocol development requires large investments for cross-domain innovations in biochemical analysis, microfluidics, and cyberphysical systems. We propose a watermarking technique for bio-protocol IP protection-a first in microfluidics-that hierarchically embeds a secret signature across these domains. Such a signature can be exclusively attributed to the owner (like a hash). The proposed solution takes into account the inherent variability in domain-specific parameters such as mixing ratio, sensor calibration, and incubation time. We describe watermarking techniques of varying complexities for different bio-protocol steps. These include watermarking for bio-protocol synthesis parameters and the cyberphysical systems control path parameters. A watermarking scheme based on integer linear programming is proposed for the sample-preparation step of a bio-protocol. The practicality of our solution is demonstrated through case studies involving an immunoassay and several mixing ratios required in the sample-preparation process of bio-protocols. The effectiveness of this approach is evaluated through various security metrics: proof of ownership score, the probability of successful tampering of the watermark, and the probability of coincidence. We also analyze the integrity of the watermark against various possible attacks: brute force search, the insertion of a new watermark, and the watermarking of more parameters. Mohammed Shayan, Sukanta Bhattacharjee, Jack Tang, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | CAD-Base: An Attack Vector into the Electronics Supply ChainabstractFabless semiconductor companies design system-on-chips (SoC) by using third-party intellectual property (IP) cores and fabricate them in offshore, potentially untrustworthy foundries. Owing to the globally distributed electronics supply chain, security has emerged as a serious concern. In this article, we explore electronics computer-aided design (CAD) software as a threat vector that can be exploited to introduce vulnerabilities into the SoC. We show that all electronics CAD tools—high-level synthesis, logic synthesis, physical design, verification, test, and post-silicon validation—are potential threat vectors to different degrees. We have demonstrated CAD-based attacks on several benchmarks, including the commercial ARM Cortex M0 processor [1]. Kanad Basu, Samah Mohamed Saeed, Christian Pilato, Mohammed Ashraf, Mohammed Nabeel Thari Moopan, Krishnendu Chakrabarty, Ramesh Karri |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2019 | Impact of Electrostatic Coupling on Monolithic 3D-enabled Network on ChipabstractMonolithic-3D-integration (M3D) improves the performance and energy efficiency of 3D ICs over conventional through-silicon-vias-based counterparts. The smaller dimensions of monolithic inter-tier vias offer high-density integration, the flexibility of partitioning logic blocks across multiple tiers, and significantly reduced total wire-length enable high-performance and energy-efficiency. However, the performance of M3D ICs degrades due to the presence of electrostatic coupling when the inter-layer-dielectric thickness between two adjacent tiers is less than 50nm. In this work, we evaluate the performance of an M3D-enabled Network-on-chip (NoC) architecture in the presence of electrostatic coupling. Electrostatic coupling induces significant delay and energy overheads for the multi-tier NoC routers. This in turn results in considerable performance degradation if the NoC design methodology does not incorporate the effects of electrostatic coupling. We demonstrate that electrostatic coupling degrades the energy-delay-product of an M3D NoC by 18.1% averaged over eight different applications from SPLASH-2 and PARSEC benchmark suites. As a countermeasure, we advocate the adoption of electrostatic coupling-aware M3D NoC design methodology. Experimental results show that the coupling-aware M3D NoC reduces performance penalty by lowering the number of multi-tier routers significantly. Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2019 | Hardware Trojan Detection Using Changepoint-Based Anomaly Detection TechniquesabstractThere has been a growing trend in recent years to outsource various aspects of the semiconductor design and manufacturing flow to different parties spread across the globe. Such outsourcing increases the risk of adversaries adding malicious logic, referred to as hardware Trojans, to the original design. The increased complexity of modern microprocessors increases the difficulty in detecting hardware Trojans at early stages of design and manufacturing. Therefore, there is a need for run-time detection techniques to capture Trojans that escape detection at these stages. In this paper, we introduce a machine learning-based run-time hardware Trojan detection method for microprocessor cores. This approach uses changepoint-based anomaly detection algorithm to detect the activation of Trojans that introduce abnormal patterns in the data streams obtained from performance counters. It does not modify the original microprocessor design to integrate on-chip monitoring sensors. We evaluate our method by detecting the activation of Trojans that cause denial-of-service, the degradation of system performance, and change in functionality of a microprocessor core. Results obtained using the OpenSPARC T1 core and an field-programmable gate array (FPGA) prototyping framework show that the Trojan activation is detected with a true positive rate of above 99% and a false positive rate of 0% for most of the implemented Trojans. Rana Elnaggar, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Toward Secure Microfluidic Fully Programmable Valve Array BiochipsabstractThe fully programmable valve array (FPVA) is a general-purpose programmable flow-based microfluidic platform, akin to the VLSI field-programmable gate array (FPGA). FPVAs are dynamically reconfigurable and, hence, are suitable in a broad spectrum of applications involving immunoassays and cell analysis. Since these applications are safety critical, addressing security concerns is vital for the success and adoption of FPVAs. This study evaluates the security of FPVA biochips. We show that FPVAs are vulnerable to malicious operations similar to digital and flow-based microfluidic biochips. FPVAs are further prone to new classes of attacks-tunneling and deliberate aging. This study establishes security metrics and describes possible attacks on real-life bioassays. Furthermore, we study the use of machine learning (ML) techniques to detect and classify attacks based on the golden and real-time biochip state. In order to boost the classifier's performance, we propose a smart checkpointing mechanism. Experimental results are presented to showcase: 1) best-fit ML model classifier; 2) performance of different tradeoffs in checkpointing; and 3) effectiveness of the proposed smart checkpointing scheme. Mohammed Shayan, Sukanta Bhattacharjee, Yong-Ak Song, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Design-for-testability for continuous-flow microfluidic biochipsabstractFlow-based microfluidic biochips are gaining traction in the microfluidics community since they enable efficient and low-cost biochemical experiments. These highly integrated lab-on-a-chip systems, however, suffer from manufacturing defects, which cause some chips to malfunction. To test biochips after manufacturing, air pressure is applied to input ports of a chip and predetermined test vectors are used to change the states of microvalves in the chip. Pressure meters are connected to the output ports to measure pressure values, which are compared with expected values to detect errors. To reduce the cost of the test platform, the number of pressure sources and meters should be reduced. We propose a design-for-testability (DFT) technique that enables a test procedure with only a single pressure source and a single pressure meter. Furthermore, the valves inserted for DFT share control channels with valves in the original chip so that no additional control signals are required. Simulation results demonstrate that this technique can generate efficient chip architectures for single-source single-meter test in all experiment cases successfully to reduce test cost, while the performance of these chips in executing applications is still maintained. Bing Li 0005, Tsung-Yi Ho, Krishnendu Chakrabarty, Ulf Schlichtmann |
DAC | 4 |
| 2018 | Tamper-resistant pin-constrained digital microfluidic biochipsabstractDigital microfluidic biochips (DMFBs)---an emerging technology that implements bioassays through manipulation of discrete fluid droplets---are vulnerable to actuation tampering attacks, where a malicious adversary modifies control signals for the purposes of manipulating results or causing denial-of-service. Such attacks leverage the highly programmable nature of DMFBs. However, practical DMFBs often employ a technique called pin mapping to reduce control pin count while simultaneously reducing the degrees of freedom available for droplet manipulation. Attempts to control specific electrodes as part of an attack cannot be made without inadvertently actuating other electrodes on-chip, which makes the tampering evident. This paper explores this tamper-resistance property of pin mapping in detail. We derive relevant security metrics, evaluate the tamper-resistance of several existing pin mapping algorithms, and propose a new security-aware pin mapper with superior tamper-resistance as compared to prior work. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
DAC | 3 |
| 2018 | Pre-assembly testing of interconnects in embedded multi-die interconnect bridge (EMIB) diesabstractThe embedded multi-die interconnect bridge (EMIB) is an advanced packaging technology for 2.5D integration. This paper presents a bridge test architecture based on the proposed IEEE Std. P1838. The proposed test method enables access to interconnects at a pre-assembly stage by pairing the interconnects using metal shorts and probing on coarse-pitch C4 bumps. It can efficiently detect resistive-open and resistive-short defects in the bridge interconnects and micro-bumps. Simulation results are presented to evaluate the range of defects that can be detected by the proposed method. Sudipta Mondal, Krishnendu Chakrabarty |
DATE | 2 |
| 2018 | Fault-tolerant valve-based microfluidic routing fabric for droplet barcoding in single-cell analysisabstractHigh-throughput single-cell genomics is used to gain insights into diseases such as cancer. Motivated by this important application, microfluidics has emerged as a key technology for developing comprehensive biochemical procedures for studying DNA, RNA, proteins, and many other cellular components. Recently, a hybrid microfluidic platform has been proposed to efficiently automate the analysis of a heterogeneous sequence of cells. In this design, a valve-based routing fabric based on transposers is used to label/barcode the target cells. However, the design proposed in prior work overlooked defects that are likely to occur during chip fabrication and system integration. We address the above limitation by investigating the fault tolerance of the valve-based routing fabric. We develop a theory of failure assessment and introduce a design technique for achieving fault tolerance. Simulation results show that the proposed method leads to a slight increase in the fabric size and decrease in cell-analysis throughput, but this is only a small price to pay for the added assurance of fault tolerance in the new design. Yasamin Moradi, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
DATE | 3 |
| 2018 | On Designing All-Optical Multipliers Using Mach-Zender InterferometersabstractIn recent years, the design of all-optical circuits has received a great attention among the researchers due to high-speed and low-power characteristics and compatibility with CMOS technology. Some combinational logic circuits like adders, subtractors, multipliers, multiplexers, which are useful in optical communication network in data centers and high-perform computers, have been designed using optical components. There are two different design styles, called as Design1 (based on conventional truth-table based approach) and Design2 (based on binary decision diagram). In this paper, four different all-optical multipliers have been explored for array multiplier and carry save adder (CSA)-based multiplier based on these two design styles, using semiconductor optical amplifier (SOA) based Mach-Zender interferometers (MZIs). Simulation results confirm that MZI-based CSA multiplier Design1) has the lowest optical cost and delay compared to those of other three multiplier designs (CSA multiplier - Design2, array multiplier - Design1, array multiplier - Design2) with a precision of 2 or more bits. Further, the proposed all-optical CSA multiplier designs outperform in terms of both optical cost and delay compared to the state-of-the-art designs of all-optical multipliers. Sumit Sharma 0002, Krishnendu Chakrabarty, Sudip Roy 0001 |
DSD | 2 |
| 2018 | Locking of biochemical assays for digital microfluidic biochipsabstractIt is expected that as digital microfluidic biochips (DMFBs) mature, the hardware design flow will begin to resemble the current practice in the semiconductor industry: design teams send chip layouts to third party foundries for fabrication. These foundries are untrusted, and threaten to steal valuable intellectual property (IP). In a DMFB, the IP consists of not only hardware layouts, but also of the biochemical assays (bioassays) that are intended to be executed on-chip. DMFB designers therefore must defend these protocols against theft. We propose to “lock” biochemical assays through random insertion of dummy mix-split operations, subject to several design rules. We experimentally evaluate the proposed locking mechanism, and show how a high level of protection can be achieved even on bioassays with low complexity. We offer guidance on the number of dummy mixsplits required to secure a bioassay for the lifetime of a patent. Sukanta Bhattacharjee, Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
ETS | 4 |
| 2018 | Design of fault-tolerant neuromorphic computing systemsabstractNeuromorphic computing is rapidly becoming mainstream, and Resistive Random Access Memory (RRAM) and RRAM-based computing systems (RCS) provide a promising hardware implementation of neuromorphic computing. This emerging computing system helps us to realize vector-matrix multiplications in a time complexity of 0(1), and it improves energy efficiency dramatically. However, due to the immature fabrication process, RCS is susceptible to defects; the resulting errors lead to a significant accuracy drop in neuromorphic computing applications. In order to take advantage of RCS in practical applications, fault-tolerant design is necessary. We present a survey of fault-tolerant designs for RRAM-based neuromorphic computing systems. We first describe RRAM-based crossbars and their role in neuromorphic computing systems. Following this, we classify fault models into different categories, and review the test solutions. Subsequently, the framework of fault-tolerant design for RCS is presented, which contains an online testing phase and a fault-tolerant training phase. The techniques proposed for these two phases are classified and explained to highlight their similarities and differences. The methods reviewed in this survey represent recent trends in fault-tolerant designs of RCS, and are expected motivate further research in this field. Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty |
ETS | 4 |
| 2018 | An efficient fault-tolerant valve-based microfluidic routing fabric for single-cell analysisabstractSingle-cell analysis is used to gain insights into diseases such as cancer. Recently, a hybrid microfluidic platform was proposed for concurrent single-cell analysis on thousands of heterogeneous cells. In this design, barcoding droplets are routed using a valve-based routing fabric to label the input cells. The fault-tolerance of this routing fabric has also been studied and a design technique for implementing a fault-tolerant crossbar has been proposed. However, prior work leads to a significant increase in fabric size and a decrease in cell-analysis performance. We address the above drawbacks and introduce a low-overhead design technique for achieving fault-tolerance, while maintaining the efficiency of the cell-analysis platform. We show that the proposed method is optimal in that it minimizes the overhead in terms of fabric size. We also show that the new design outperforms the previous solution in terms of cell-analysis performance. Yasamin Moradi, Krishnendu Chakrabarty, Ulf Schlichtmann |
ETS | 2 |
| 2018 | Failure prediction based on anomaly detection for complex core routersabstractData-driven prognostic health management is essential to ensure high reliability and rapid error recovery in commercial core router systems. The effectiveness of prognostic health management depends on whether failures can be accurately predicted with sufficient lead time. This paper describes how time-series analysis and machine-learning techniques can be used to detect anomalies and predict failures in complex core router systems. First both a feature-categorization-based hybrid method and a changepoint-based method have been developed to detect anomalies in time-varying features with different statistical characteristics. Next, a SVM-based failure predictor is developed to predict both categories and lead time of system failures from collected anomalies. A comprehensive set of experimental results is presented for data collected during 30 days of field operation from over 20 core routers deployed by customers of a major telecom company. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ICCAD | 3 |
| 2018 | Shadow attacks on MEDA biochipsabstractThe Micro-electrode-dot-array (MEDA) is a next-generation digital microfluidic biochip (DMFB) platform that supports fine-grained control and real-time sensing of droplet movements. These capabilities permit continuous monitoring and checkpoint-based validation of assay execution on MEDA. This paper presents a class of “shadow attacks” that abuse the timing slack in the assay execution. State-of-the-art checkpoint-based validation techniques cannot expose the shadow operations. We develop a defense that introduces extra checkpoints in the assay execution at time instances when the assay is prone to shadow attacks. Experiments confirm the effectiveness and practicality of the defense. Mohammed Shayan, Sukanta Bhattacharjee, Tung-Che Liang, Jack Tang, Krishnendu Chakrabarty, Ramesh Karri |
ICCAD | 5 |
| 2018 | Analysis of Process Variations, Defects, and Design-Induced Coupling in MemristorsabstractEmerging devices are susceptible to process variations and manufacturing defects due to immature fabrication processes. Memristors constitute a promising emerging technology, but they are known to suffer from high defect rates that contribute to faulty behavior. It is therefore important to analyze memristor fault models and understand the root causes of defects and variations. We present a physics-based classification and analysis of memristor fault origins. These fault origins are systematically attributed to process variations and manufacturing defects. We also investigate coupling in dense memristor crossbars. This study of memristor fault origins and the resulting conclusions provides valuable feedback for the fabrication and the design of memristor-based circuits and systems. Arjun Chaudhuri, Krishnendu Chakrabarty |
ITC | 2 |
| 2018 | Self-Learning Health-Status Analysis for a Core Router SystemabstractThe health status of core router systems needs to be analyzed efficiently in order to ensure high reliability and timely error recovery. Although a large amount operational data is collected from core routers, only a small part of this data is labeled by experts. The lack of labels is an impediment towards the adoption of supervised learning. We present an iterative self-learning procedure for assessing the health status of a core router. This procedure first computes a representative feature matrix to capture different characteristics of time-series data. Hierarchical clustering is then utilized to infer labels for the unlabeled dataset. Finally, a classifier is built and iteratively updated using both labeled and unlabeled dataset. Field data collected from a set of commercial core routers are used to experimentally validate the proposed health-status analyzer. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 3 |
| 2018 | Fine-Grained Adaptive Testing Based on Quality PredictionabstractThe ever-increasing complexity of integrated circuits inevitably leads to high test cost. Adaptive testing provides an effective solution for test-cost reduction; this testing framework selects the important test items for each set of chips. However, adaptive testing methods designed for digital circuits are coarse-grained, and they are targeted only at systematic defects. In order to incorporate fabrication variations and random defects in the testing framework, we propose a fine-grained adaptive testing method based on machine learning. We use the parametric test results from the previous stages of test to train a quality-prediction model for use in subsequent test stages. Next, we partition a given lot of chips into two groups based on their predicted quality. A test-selection method based on statistical learning is applied to the chips with high predicted quality. An ad hoc test-selection method is proposed and applied to the chips with low predicted quality. Experimental results using a large number of fabricated chips and the associated test data show that to achieve the same defect level as in prior work on adaptive testing, the fine-grained adaptive testing method reduces test cost by 90% for low-quality chips, and up to 7% for all the chips in a lot. Renjian Pan, Fangming Ye, Xin Li 0001, Krishnendu Chakrabarty, Xinli Gu |
ITC | 5 |
| 2018 | Fault Tolerance for RRAM-Based Matrix OperationsabstractAn RRAM-based computing system (RCS) provides an energy efficient hardware implementation of vector-matrix multiplication for machine-learning hardware. However, it is vulnerable to faults due to the immature RRAM fabrication process. We propose an efficient fault tolerance method for RCS; the proposed method, referred to as extended-ABFT (X-ABFT), is inspired by algorithm-based fault tolerance (ABFT). We utilize row checksums and test-input vectors to extract signatures for fault detection and error correction. We present a solution to alleviate the overflow problem caused by the limited number of voltage levels for the test-input signals. Simulation results show that for a Hopfield classifier with faults in 5% of its RRAM cells, X-ABFT allows us to achieve nearly the same classification accuracy as in the fault-free case. Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty |
ITC | 4 |
| 2018 | Built-In Self-Diagnosis and Fault-Tolerant Daisy-Chain Design in MEDA BiochipsabstractA digital microfluidic biochip is an attractive platform for revolutionizing immunoassays, clinical diagnostics, drug discovery, DNA sequencing, and other laboratory procedures in biochemistry. A recent generation of biochips uses a microelectrode-dot-array (MEDA) architecture, which provides finer controllability of droplets and seamlessly integrates microelectronics and microfluidics using CMOS technology. In order to simplify the wiring problem for MEDA biochips, all microelectrodes and their control registers are daisy-chained together. Therefore, chain diagnosis and fault tolerance are critical for MEDA biochips. We propose the first daisy-chain design that can perform self-diagnosis and repair the detected faults. In our design, every microelectrode cell is fully controlled, and faults in the daisy chain can be detected and repaired in a timely manner. Moreover, the proposed design can be used in both online and off-line modes. Experimental results demonstrate the effectiveness of the proposed daisy-chain design. Krishnendu Chakrabarty |
ITC | 3 |
| 2018 | Access-Time Minimization in the IEEE 1687 Network Using Broadcast and Hardware ParallelismabstractThe IEEE Std. 1687 facilitates flexible access to on-chip instruments through the JTAG test-access port. This flexibility enables the minimization of the overall access time (OAT), and a number of techniques have been proposed in the literature to achieve this goal. However, the OAT is still high for instruments that require a large amount of test data if this data is shifted through the scan chain serially. In order to further reduce the OAT, we present an efficient test-scheduling method that exploits broadcast and hardware parallelism for instrument access. A broadcast scheduling method is synergistically combined with three parallel IJTAG designs. We show that under different cost criteria, we can select the most efficient parallel IJTAG design such that the equivalent access time (EAT) is minimized. Two industry chip designs and three IJTAG benchmarks are used to evaluate the effectiveness of the proposed method. Zhanwei Zhong, Guoliang Li 0004, Qinfu Yang, Krishnendu Chakrabarty |
ITC | 4 |
| 2018 | Abetting Planned Obsolescence by Aging 3D Networks-on-ChipabstractWe set up a security analysis framework by aging the Network-on-Chip (NoC) to study planned obsolescence by the original equipment manufacturer (OEM). An NoC is the communication backbone in a manycore System-on-Chip (SoC). Planned obsolescence may adopt any vulnerability in the NoC to cause the SoC to fail. We show how an OEM can craft workloads to generate electromigration-induced stress and crosstalk noise in TSV-based vertical links in the NoC to hasten failure. We analyzed three malicious workloads and confirm that a crafted workload that injects 3-10% more traffic on to a few selected critical vertical links can shorten the lifetime of the NoC by 11%-25% averaged over the benchmarks considered in this work. Sourav Das 0002, Kanad Basu, Janardhan Rao Doppa, Partha Pratim Pande, Ramesh Karri, Krishnendu Chakrabarty |
NOCS | 6 |
| 2018 | Special session on machine learning for test and diagnosisabstractThe special session focuses on using Machine Learning (ML) techniques on different applications in test and diagnosis. The first contribution discusses how to close the gap between working silicon and a working system by using ML. The second presentation then talks an alternative ML view and its various applications such as functional verification, Fmax prediction, and production yield optimization. The last presentation discusses using supervised ML on volume diagnosis to further improve the accuracy of identifying root causes. Krishnendu Chakrabarty, Li-C. Wang, Gaurav Veda, Yu Huang 0005 |
VTS | 1 |
| 2018 | Securing IJTAG against data-integrity attacksabstractThe IEEE Std. 1687 (IJTAG) facilitates access to on-chip instruments in complex system-on-chip designs. However, a major security vulnerability in IJTAG has yet to be addressed. IJTAG supports the integration of tapped and wrapped instruments at the IP provider with hidden test-data registers (TDRs). The instruments with hidden TDRs can manipulate the data that is shifted through them. We propose the addition of shadow test-data registers by the trusted IJTAG integrator to protect the shifted data from illegitimate manipulation by malicious third-party IPs. In addition, we use information-flow tracking to identify the modified bits during the attack and the attacking instruments in an IJTAG network. We present security proofs, simulation results and the overheads associated with these countermeasures for various benchmarks. Rana Elnaggar, Ramesh Karri, Krishnendu Chakrabarty |
VTS | 3 |
| 2018 | An inter-layer interconnect BIST solution for monolithic 3D ICsabstractMonolithic three-dimensional (M3D) integration offers higher-density integration compared to 3D integration based on through-silicon vias. Advances in testing are however needed to screen defects in M3D integration. We propose a built-in self-test solution to target shorts and opens in ILVs. In the proposed solution, scan cells at the interface of two layers are stitched into a twisted-ring counter (TRC) using their functional outputs and ILVs. The interface-register cells launch and capture tests, and a test path consists of ILVs and a multiplexer. We map the problem of minimizing the length of the wires added to stitch the TRC to that of finding a minimum-cost Hamiltonian circuit in a weighted bipartite graph. Since the weighted Hamiltonian circuit problem is NP-Complete, we propose a heuristic algorithm for this problem. Simulation results show that the proposed test solution can detect all opens and shorts. Abhishek Koneru, Krishnendu Chakrabarty |
VTS | 2 |
| 2018 | Broadcast-based minimization of the overall access time for the IEEE 1687 networkabstractThe IEEE Std. 1687 enables flexible access to on-chip instruments through the JTAG test-access port. This flexibility enables the minimization of the overall access time (OAT), and a number of techniques have been proposed to achieve this goal. However, these techniques do not utilize the broadcast feature in the 1687 network. In order to further reduce the OAT, we present an efficient test-scheduling method that exploits the broadcast feature for instrument access. Two optimization solutions are then proposed - the first solution minimizes the OAT without the retargeting time, while the second one reorders the configurations so as to minimize the overall retargeting time among configurations. Three industry test cases are used to evaluate the effectiveness of the proposed method. Zhanwei Zhong, Guoliang Li 0004, Qinfu Yang, Krishnendu Chakrabarty |
VTS | 5 |
| 2018 | Machine Learning for Hardware Security: Opportunities and Risks
Rana Elnaggar, Krishnendu Chakrabarty |
J. Electron. Test. | 2 |
| 2018 | Cyber-Physical Digital-Microfluidic Biochips: Bridging the Gap Between Microfluidics and MicrobiologyabstractDigital microfluidics is transforming microbiology research by providing new opportunities for high-throughput sample preparation and point-of-care diagnostics. Over the past decade, several design-automation (synthesis) techniques have been developed for on-chip droplet manipulation. However, these methods oversimplify the dynamics of biomolecular protocols and they have yet to make a significant impact in biochemistry/microbiology research, leading to a large gap between advances in biochip design and the adoption of biochips for running biomolecular protocols. In this paper, we bridge this gap by introducing a new paradigm for biochip design automation. By exploiting advances in the integration of sensing systems into a digital-microfluidic biochip, we present a number of synthesis solutions that use realistic models of biomolecular protocols to address real-world microbiology applications through cyber-physical adaptation. This paper also details a vision for continued research on design-automation and optimization methodologies for the realization of biomolecular protocols using microfluidic biochips. Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
Proc. IEEE | 2 |
| 2018 | Keynote Paper: From EDA to IoT eHealth: Promises, Challenges, and SolutionsabstractThe interaction between technology and healthcare has a long history. However, recent years have witnessed the rapid growth and adoption of the Internet of Things (IoT) paradigm, the advent of miniature wearable biosensors, and research advances in big data techniques for effective manipulation of large, multiscale, multimodal, distributed, and heterogeneous data sets. These advances have generated new opportunities for personalized precision eHealth and mHealth services. IoT heralds a paradigm shift in the healthcare horizon by providing many advantages, including availability and accessibility, ability to personalize and tailor content, and cost-effective delivery. Although IoT eHealth has vastly expanded the possibilities to fulfill a number of existing healthcare needs, many challenges must still be addressed in order to develop consistent, suitable, safe, flexible and power-efficient systems that are suitable fit for medical needs. To enable this transformation, it is necessary for a large number of significant technological advancements in the hardware and software communities to come together. This keynote paper addresses all these important aspects of novel IoT technologies for smart healthcare-wearable sensors, body area sensors, advanced pervasive healthcare systems, and big data analytics. It identifies new perspectives and highlights compelling research issues and challenges, such as scalability, interoperability, device-network-human interfaces, and security, with various case studies. In addition, with the help of examples, we show how knowledge from CAD areas, such as large scale analysis and optimization techniques can be applied to the important problems of eHealth. Farshad Firouzi, Bahareh J. Farahani, Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Testing 3D-SoCs Using 2-D Time-Division MultiplexingabstractThrough-silicon vias (TSVs) are used as high-speed vertical interconnects between dies in a 3-D system-on-a-chip (SoC). However, their speed cannot be exploited during test application due to inherent limitations of the scan-chains of the cores, which prevent the use of high shift frequencies during the scan-in/out operations. Moreover, due to their high area cost, only a limited number of TSVs can be utilized for test application. As a result, TSVs become the bottleneck for transferring the large volume of test-data to the various layers of the stack, and the time for testing the 3-D chip increases a lot. In this paper, we propose an efficient test-access mechanism (TAM) architecture that exploits the high speed of TSVs to minimize the time for testing 3-D SoCs. The proposed TAM architecture is based on a 2-D time-division-multiplexing approach, and by the means of a very effective test-scheduling method, it offers significant savings in test-time, TSV-count and TAM-cost under power and thermal constraints. Extensive experiments on two 3-D benchmark SoCs show the benefits of the proposed method. Panagiotis Georgiou, Fotis Vartziotis, Xrysovalantis Kavousianos, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Toward Predictive Fault Tolerance in a Core-Router System: Anomaly Detection Using Correlation-Based Time-Series AnalysisabstractFault tolerance is used in communication systems to ensure high reliability and rapid error recovery. The effectiveness of most proactive fault-tolerant mechanism depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the monitored data involves temporal measurements and exhibits significantly different statistical characteristics for its constituent features. We describe the design of an anomaly detector that monitors the time-series data of a complex core router system. Anomaly detection techniques are compared in terms of their effectiveness for detecting different types of anomalies. A feature-categorizing-based hybrid method is proposed to overcome the difficulty of detecting anomalies in features with different statistical characteristics. Furthermore, a correlation analyzer is implemented to remove irrelevant and redundant features. Three types of synthetic anomalies, generated using a small amount of real data for a commercial telecom system, are used to validate the proposed anomaly detector. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Efficient and Adaptive Error Recovery in a Micro-Electrode-Dot-Array Digital Microfluidic BiochipabstractA digital microfluidic biochip (DMFB) is an attractive technology platform for automating laboratory procedures in biochemistry. In recent years, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed. MEDA biochips can provide advantages of better capability of droplet manipulation and real-time sensing ability. However, errors are likely to occur due to defects, chip degradation, and the lack of precision inherent in biochemical experiments. Therefore, an efficient error-recovery strategy is essential to ensure the correctness of assays executed on MEDA biochips. By exploiting MEDA-specific advances in droplet sensing, we present a novel error-recovery technique to dynamically reconfigure the biochip using real-time data provided by on-chip sensors. Local recovery strategies based on probabilistic-timed-automata are presented for various types of errors. An online synthesis technique and a control flow are also proposed to connect local-recovery procedures with global error recovery for the complete bioassay. Moreover, an integer linear programming-based method is also proposed to select the optimal local-recovery time for each operation. Laboratory experiments using a fabricated MEDA chip are used to characterize the outcomes of key droplet operations. The PRISM model checker and three benchmarks are used for an extensive set of simulations. Our results highlight the effectiveness of the proposed error-recovery strategy. Kelvin Yi-Tse Lai, John McCrone, Po-Hsien Yu, Krishnendu Chakrabarty, Miroslav Pajic, Tsung-Yi Ho, Chen-Yi Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | Structural and Functional Test Methods for Micro-Electrode-Dot-Array Digital Microfluidic BiochipsabstractA digital microfluidic biochip (DMFB) is an attractive platform for immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. More recently, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed, and droplet manipulations on MEDA biochips have also been experimentally demonstrated. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, MEDA biochips must be adequately tested before they can be used for bioassay execution. This paper presents the first approach for testing of MEDA biochips that include both CMOS circuits and microfluidic components. We first present structural test techniques to evaluate the pass/fail status of each microcell (droplet actuation, droplet maintenance, and droplet sensing) and identify faulty microcells. In order to ensure correct operation of functional units, e.g., mixers and diluters, we also present functional test techniques to address fundamental MEDA operations, such as droplet dispensing, transportation, mixing, and splitting. We evaluate the proposed test methods using simulations as well as experiments for fabricated MEDA biochips. Kelvin Yi-Tse Lai, Po-Hsien Yu, Krishnendu Chakrabarty, Tsung-Yi Ho, Chen-Yi Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Secure Randomized Checkpointing for Digital Microfluidic BiochipsabstractDigital microfluidic biochips (DMFBs) integrated with processors and arrays of sensors form cyberphysical systems and consequently face a variety of unique, recently described security threats. It has been noted that techniques used for error recovery can provide some assurance of integrity when a cyberphysical DMFB is under attack. This paper proposes the use of such hardware for security purposes through the randomization of checkpoints in both space and time, and provides design guidelines for designers of such systems. We define security metrics and present techniques for improving performance through static checkpoint maps, and describe performance tradeoffs associated with static and random checkpoints. We also provide detailed classification of attack models and demonstrate the feasibility of our techniques with case studies on assays implemented in typical DMFB hardware. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Leakage Current Analysis for Diagnosis of Bridge Defects in Power-Gating DesignsabstractManufacturing defects that do not affect the functional operation of low power integrated circuits (ICs) can nevertheless impact their power saving capability. We show that stuck-ON faults on the power switches and resistive bridges between the power networks can impair the power saving capability of power-gating designs. For quantifying the impact of such faults on the power savings of power-gating designs, we propose a diagnosis technique that targets bridges between the power networks. The proposed technique is based on the static power analysis of a power-gating design in stand-by mode and it utilizes a novel on-chip signature generation unit, which is sensitive to the voltage level between power rails, the measurements of which are processed off-line for the diagnosis of bridges that can adversely affect power savings. We explore, through SPICE simulation of the largest IWLS’05 benchmarks synthesized using a 32 nm CMOS technology, the tradeoffs achieved by the proposed technique between diagnosis accuracy and area cost and we evaluate its robustness against process variation. The proposed technique achieves a diagnosis resolution that is higher than 98.6% and 97.9% for bridges of${R}~{\gtrsim }~{10~{ M}\Omega }$(weak bridges) and bridges of${R~\lesssim~10~{ M}\Omega }$(strong bridges), respectively, and a diagnosis accuracy higher than 94.5% for all the examined defects. The area overhead is small and scalable: it is found to be 1.8% and 0.3% for designs with 27 K and 157 K gate equivalents, respectively. Vasileios Tenentes, Daniele Rossi 0001, S. Saqib Khursheed, Bashir M. Al-Hashimi, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | Online Soft-Error Vulnerability Estimation for Memory Arrays and Logic CoresabstractRadiation-induced soft errors are a major reliability concern in circuits fabricated at advanced technology nodes. Online soft-error vulnerability estimation offers the flexibility of exploiting dynamic fault-tolerant mechanisms for cost-effective reliability enhancement. We propose a generic run-time method with low area and power overhead to predict the soft-error vulnerability of on-chip memory arrays as well as logic cores. The vulnerability prediction is based on signal probabilities (SPs) of a small set of flip-flops, chosen at design time, by studying the correlation between the soft-error vulnerability and the flip-flop SPs for representative workloads. We exploit machine learning to develop a predictive model that can be deployed in the system in software form. Simulation results on two processor designs show that the proposed technique can accurately estimate the soft-error vulnerability of on-chip logic core, such as sequential pipeline logic and functional units as well as memory arrays that constitute the instruction cache, the data cache, and the register file. Arunkumar Vijayan, Saman Kiamehr, Mojtaba Ebrahimi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Fine-Grained Aging-Induced Delay Prediction Based on the Monitoring of Run-Time StressabstractRun-time solutions based on online monitoring and adaptation are required for resilience in nanoscale integrated circuits, as design-time solutions and guard bands are no longer sufficient. Bias temperature instability-induced transistor aging, one of the major reliability threats in nanoscale very large scale integration, degrades path delay over time and may lead to timing failures. Chip health monitoring is, therefore, necessary to track delay changes on a per-chip basis over the chip lifetime operation. However, direct monitoring based on actual measurement of path delays can only track a coarse-grained aging trend in a reactive manner, not suitable for proactive fine-grain adaptations. In this paper, we propose a low cost and fine-grained workload-induced stress monitoring approach, based on machine learning techniques, to accurately predict aging-induced delay. We integrate space and time sampling of selective flip-flops into the runtime monitoring infrastructure in order to reduce the cost of monitoring the workload. The prediction model is trained offline using support-vector regression and implemented in software. This approach can leverage proactive adaptation techniques to mitigate further aging of the circuit by monitoring aging trends. Simulation results for realistic open-source benchmark circuits highlight the accuracy of the proposed approach. Arunkumar Vijayan, Abhishek Koneru, Saman Kiamehr, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Workload-Aware Static Aging Monitoring and Mitigation of Timing-Critical Flip-FlopsabstractIn advanced technology nodes, bias temperature instability (BTI) has emerged as a prominent reliability concern. The worst-case effects of BTI occur during specific workload phases in which flip-flops (FFs) on a critical path do not switch their logic values for a long duration. These inactive FFs in the circuit experience accelerated workload-dependent static-BTI (S-BTI) stress. The aging effect of S-BTI for a few hours has been shown to be equivalent to one year of aging due to dynamic BTI, which can eventually cause circuit failure. The techniques available to mitigate S-BTI stress during standby mode of circuits are pessimistic, thereby limiting the performance of the circuit. To address this problem, we propose a runtime monitoring method to raise a flag when a timing-critical FF experiences severe S-BTI stress. To reduce the monitoring costs, we select a small representative set of FFs offline based on workload-aware correlation analysis and these selected FFs are monitored online for static aging phases. Our experiments conducted on two processors show that less than 0.5% of the total number of FFs is required to be selected as representative FFs for S-BTI stress monitoring. We also propose a low-overhead mitigation scheme to relax critical FFs by executing a software subroutine that is designed to exercise critical FFs. Arunkumar Vijayan, Saman Kiamehr, Fabian Oboril, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Performance and Thermal Tradeoffs for Energy-Efficient Monolithic 3D Network-on-ChipabstractThree-dimensional (3D) integration enables the design of high-performance and energy-efficient network on chip (NoC) architectures as communication backbones for manycore chips. To exploit the benefits of the vertical dimension of 3D integration, through-silicon-via (TSV) has been predominantly used in state-of-the-art manycore chip design. However, for TSV-based systems, high power density and the resultant thermal hotspot remain major concerns from the perspectives of chip functionality and overall reliability. The power consumption and thermal profiles of 3D NoCs can be improved by incorporating a Voltage-Frequency-Island (VFI)-based power management strategy. However, due to inherent thermal constraints of a TSV-based 3D system, we are unable to fully exploit the benefits offered by the power management methodology. In this context, emergence of monolithic 3D (M3D) integration has opened up new possibility of designing ultra-low-power and high-performance circuits and systems. The smaller dimensions of the inter-layer dielectric (ILD) and monolithic inter-tier vias (MIVs) offer high-density integration, flexibility of partitioning logic blocks across multiple tiers, and significant reduction of total wire-length. In this work, we present the first-ever study of the performance-thermal tradeoffs for energy efficient monolithic 3D manycore chips. In particular, we present a comparative performance evaluation of M3D NoCs with respect to their conventional TSV-based counterparts. We demonstrate that the proposed M3D-based NoC architecture incorporating VFI-based power management achieves a maximum of 29.4% lower energy-delay-product (EDP) compared to the TSV-based designs for a large set of benchmarks. We also demonstrate that the M3D-based NoC shows up to 29.1% lower maximum temperature than the TSV-based counterpart for these benchmarks. Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2018 | Demand-Driven Single- and Multitarget Mixture Preparation Using Digital Microfluidic BiochipsabstractRecent studies in algorithmic microfluidics have led to the development of several techniques for automated solution preparation using droplet-based digital microfluidic (DMF) biochips. A major challenge in this direction is to produce a mixture of several reactants with a desired ratio while optimizing reactant cost and preparation time. The sequence of mix-split operations that are to be performed on the droplets is usually represented as a mixing tree (or graph). In this article, we present an efficient mixing algorithm, namely, Mixing Tree with Common Subtrees ( MTCS ), for preparing single-target mixtures. MTCS attempts to best utilize intermediate droplets, which were otherwise wasted, and uses morphing based on permutation of leaf nodes to further reduce the graph size. The technique can be generalized to produce multitarget ratios, and we present another algorithm, namely, Multiple Target Ratios ( MTR ). Additionally, in order to enhance the output load, we also propose an algorithm for droplet streaming called Multitarget Multidemand ( MTMD ). Simulation results on a large set of target ratios show that MTCS can reduce the mean values of the total number of mix-split steps ( T ms ) and waste droplets ( W ) by 16% and 29% over Min-Mix (Thies et al. 2008) and by 22% and 34% over RMA (Roy et al. 2015), respectively. Experimental results also suggest that MTR can reduce the average values of T ms and W by 23% and 44% over the repeated version of Min-Mix , by 30% and 49% over the repeated version of RMA , and by 9% and 22% over the repeated-version of MTCS , respectively. It is observed that MTMD can reduce the mean values of T ms and W by 64% and 85%, respectively, over MTR . Thus, the proposed multitarget techniques MTR and MTMD provide efficient solutions to multidemand, multitarget mixture preparationon a DMF platform. Shalu, Srijan Kumar, Ananya Singla, Sudip Roy 0001, Krishnendu Chakrabarty, P. P. Chakrabarti 0001, Bhargab B. Bhattacharya |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2018 | Multicast Testing of Interposer-Based 2.5D ICs: Test-Architecture Design and Test SchedulingabstractInterposer-based 2.5D integrated circuits (ICs) are seen today as a precursor to 3D ICs based on through-silicon vias (TSVs). All the dies in a 2.5D IC must be adequately tested for product qualification. However, due to the limited number of package pins, it is a major challenge to test 2.5D ICs using conventional methods. Moreover, due to higher integration levels, test-application time and test power consumption for 2.5D ICs are also increased compared to their 2D counterparts. Therefore, it is imperative to take these issues into account during 2.5D IC testing. In this article, we present an efficient multicast test architecture for targeting defects in dies, in which multiple dies can be tested simultaneously to reduce the test-application time under constraints on test power and fault coverage. We also propose a test scheduling and optimization technique that can be utilized with the multicast test architecture. By considering the trade-off between test-application time, test-power budget, and test quality, the proposed technique provides test schedules with minimum test-application time under constraints on power consumption and fault coverage. Compared to previous work, the proposed technique can reduce test-application time by up to 53.4 for benchmark designs while achieving higher fault coverage. Since the loss in fault coverage due to multicast testing is extremely small, we can use top-off patterns to achieve full fault coverage for the dies at negligible additional cost. Shengcheng Wang, Ran Wang 0002, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | Fault-Tolerant Unicast-Based Multicast for Reliable Network-on-Chip TestingabstractWe present a unified test technique that targets faults in links, routers, and cores of a network-on-chip design based on test sessions. We call an entire procedure, that delivers test packets to the subset of routers/cores, a test session. Test delivery for router/core testing is formulated as two fault-tolerant multicast algorithms. Test packet delivery for routers is implemented as a fault-tolerant unicast-based multicast scheme via the fault-free links and routers that were identified in the previous test sessions to avoid packet corruption. A new fault-tolerant routing algorithm is also proposed for the unicast-based multicast core test delivery in the whole network. Identical cores share the same test set, and they are tested within the same test session. Simulation results highlight the effectiveness of the proposed method in reducing test time. Krishnendu Chakrabarty, Hideo Fujiwara |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | EditorialabstractAs 2018 draws to an end, so does my second and final term as the Editor-in-Chief (EiC) of the IEEE Transactions on Very Large Scale Intergation (VLSI) Systems (TVLSI). It gives me great pleasure to announce that Prof. Massimo Alioto from the National University of Singapore will be the new EiC, effective from January 2019. Massimo is widely respected for his outstanding research in many areas of VLSI circuits and system design, and he served admirably as an Associate EiC of TVLSI during my two EiC terms. I am grateful to him for his help and constant support. Under Massimo’s able leadership, TVLSI will continue to be the leading journal and top-choice publication venue for VLSI system design. His biography and photograph are included below. As the incoming EiC, Massimo will select the new editorial board. However, the current AEs will continue to make decisions until all the papers assigned to them have received final decisions. This should enable a seamless transition from this TVLSI EiC term to the next. Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Computation-oriented fault-tolerance schemes for RRAM computing systemsabstractThe emerging metal-oxide resistive switching random-access memory (RRAM) devices and RRAM crossbar arrays have demonstrated their potential in enormously boosting the speed and energy-efficiency of analog matrix-vector multiplication. Unfortunately, due to the immature fabrication technology, commonly occurring Stuck-At-Faults (SAFs) seriously degrade the computational accuracy of RRAM crossbar based Computing System (RCS). In this paper, we propose a Mapping Algorithm with inner fault-tolerant ability (MAO) to convert matrix parameters into RRAM conductances in RCS by providing larger mapping space and fully exploring the available mapping space. Furthermore, we present two computation-oriented redundancy schemes — ‘Redundant Crossbars’ (RX) and ‘Independent Redundant Columns’ (IRC) to alleviate the loss of computational accuracy due to SAFs. RX adds redundant RRAM crossbar arrays and IRC introduces independent redundant RRAM columns to compensate the computational errors brought by SAFs. Wenqin Huangfu, Lixue Xia, Xiling Yin, Tianqi Tang 0001, Boxun Li, Krishnendu Chakrabarty, Yuan Xie 0001, Yu Wang 0002, Huazhong Yang |
ASP-DAC | 7 |
| 2017 | Exact routing for micro-electrode-dot-array digital microfluidic biochipsabstractDigital microfluidics is an emerging technology that provide fluidic-handling capabilities on a chip. One of the most important issues to be considered when conducting experiments on the corresponding biochips is the routing of droplets. A recent variant of biochips uses a micro-electrode-dot-array (MEDA) which yields a finer controllability of the droplets. Although this new technology allows for more advanced routing possibilities, it also poses new challenges to corresponding CAD methods. In contrast to conventional microfluidic biochips, droplets on MEDA biochips may move diagonally on the grid and are not bound to have the same shape during the entire experiment. In this work, we present an exact routing method that copes with these challenges while, at the same time, guarantees to find the minimal solution with respect to completion time. For the first time, this allows for evaluating the benefits of MEDA biochips compared to their conventional counterparts as well as a quality assessment of previously proposed routing methods in this domain. Oliver Keszöcze, Andreas Grimmer, Robert Wille, Krishnendu Chakrabarty, Rolf Drechsler |
ASP-DAC | 5 |
| 2017 | Workload-aware static aging monitoring of timing-critical flip-flopsabstractIn advanced technology nodes, Bias Temperature Instability (BTI) has emerged as a prominent reliability concern. The worst-case effects of BTI occur during specific workload phases in which flip-flops on a critical path do not switch their logic values for a long duration. These inactive flip-flops in the circuit experience accelerated workload-dependent static-BTI stress. The aging effect of static BTI for a few hours has been shown to be equivalent to one year of aging due to dynamic BTI, which can eventually cause circuit failure. The techniques available to mitigate static-BTI stress during standby mode of circuits are pessimistic, thereby limiting the performance of the circuit. To address this problem, we propose a runtime monitoring method to raise a flag when a timing-critical flip-flop experiences severe static-BTI stress. To reduce the monitoring costs, we select a small representative set of flip-flops offline based on workload-aware correlation analysis and these selected flip-flops are monitored online for static aging phases. Our experiments conducted on two processors show that, less than 0.5% of the total number of flip-flops is required to be selected as representative flip-flops for S-BTI stress monitoring. Arunkumar Vijayan, Saman Kiamehr, Fabian Oboril, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ASP-DAC | 4 |
| 2017 | Security Implications of Cyberphysical Flow-Based Microfluidic BiochipsabstractFlow-based microfluidic biochips are revolutionizing biochemical research by automating complex protocols and reducing sample and reagent consumption. Integration of these biochips with sensors, actuators, and intelligent control have compounded these benefits while increasing reliability. And, many flow-based platforms have successfully transitioned to the marketplace, demonstrating their utility through several recent scientific publications. However, these microfluidic technologies and platforms have unintended security and trust implications that threaten their continued success. We survey cyberphysical flow-based microfluidic platforms and perform a security assessment. We then describe an attack on digital polymerase chain reactions and how such attacks undermine research integrity. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
ATS | 3 |
| 2017 | Fault-Tolerant Training with On-Line Fault Detection for RRAM-Based Neural Computing SystemsabstractAn RRAM-based computing system (RCS) is an attractive hardware platform for implementing neural computing algorithms. Online training for RCS enables hardware-based learning for a given application and reduces the additional error caused by device parameter variations. However, a high occurrence rate of hard faults due to immature fabrication processes and limited write endurance restrict the applicability of on-line training for RCS. We propose a fault-tolerant on-line training method that alternates between a fault-detection phase and a fault-tolerant training phase. In the fault-detection phase, a quiescent-voltage comparison method is utilized. In the training phase, a threshold-training method and a re-mapping scheme is proposed. Our results show that, compared to neural computing without fault tolerance, the recognition accuracy for the Cifar-10 dataset improves from 37% to 83% when using low-endurance RRAM cells, and from 63% to 76% when using RRAM cells with high endurance but a high percentage of initial faults. Lixue Xia, Xuefei Ning, Krishnendu Chakrabarty, Yu Wang 0002 |
DAC | 4 |
| 2017 | Robust TSV-based 3D NoC design to counteract electromigration and crosstalk noiseabstractA 3D network-on-chip (3D NoC) is an enabler for the design of high-performance and energy-efficient manycore chips. Most popular 3D NoCs utilize the Through-Silicon-Via (TSV)-based vertical links (VLs) as the communication pillars between the planar dies. However, the TSVs in a 3D NoC may fail due to both workload-induced stress and crosstalk capacitance. This failure negatively affects the overall achievable performance of the 3D NoC. In this work, we analyze the joint effects of workload-induced stress and crosstalk on the TSV mean-time-to-failure (MTTF) and hence the 3D NoC lifetime. We demonstrate that if we only consider the effects of electromigration on the TSVs due to workload-induced stress then the estimated MTTF and the subsequently lifetime of 3D NoC are too optimistic. Due to the combined effects of workload and crosstalk noise, the lifetime of 3D NoC reduces significantly. Subsequently, we demonstrate that a spare TSV allocation methodology considering the joint effects of workload and crosstalk noise enhances the lifetime of the 3D NoC by a factor of 4.6 compared to when only the workload is considered for a given spare budget of 5%. Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 4 |
| 2017 | Optimization of retargeting for IEEE 1149.1 TAP controllers with embedded compressionabstractWe present a formal optimization technique that enables retargeting for codeword-based IEEE 1149.1-compliant TAP controllers. The proposed method addresses the problem of high test data volume and Test Application Time (TAT) for a system-on-chip design during board or in-field testing, as well as during debugging. This procedure determines an optimal set of codewords with respect to given hardware constraints, e.g., embedded dictionary size and the interface to the Test Data Register in the IEEe 1149.1 Std. A complete traversal of the spanned search space is possible through the use of formal methods. An optimal set of codewords can be determined, which is directly utilized for retargeting. The proposed method is evaluated using test data with high-entropy, which is known to be the least amenable to compression, as well as input data for debugging and Functional Verification (FV) test data. Our results show a compression ratio improvement of more than 30% and a reduction in TAT up to 20% compared to previous techniques. Sebastian Huhn 0001, Stephan Eggersglüß, Krishnendu Chakrabarty, Rolf Drechsler |
DATE | 3 |
| 2017 | Digital-microfluidic biochips for quantitative analysis: Bridging the Gap between microfluidics and microbiologyabstractDigital-microfluidics technology has shown considerable promise for advancing sample preparation and point-of-care diagnostics; therefore, it has the potential to transform microbiology and biochemistry research. Over the past decade, a number of microfluidics design-automation techniques have been developed for on-chip droplet manipulation. However, these methods overlook the myriad complexities of biomolecular protocols and they have yet to make a significant impact in biochemistry/microbiology research. A paradigm shift in biochip design automation and a “phase transition” in research are clearly needed to bridge this gap between microfluidics and microbiology. In this paper, we explain how researchers from design-automation and embedded systems can play a key role in this transition. We present a new synthesis flow that uses realistic models of biomolecular protocols and cyberphysical adaptation to address real-world microbiology applications. We also present a list of metrics that can be used for the assessment of design-automation techniques for microbiology applications. Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
DATE | 2 |
| 2017 | CoSyn: Efficient single-cell analysis using a hybrid microfluidic platformabstractSingle-cell genomics is used to advance our understanding of diseases such as cancer. Microfluidic solutions have recently been developed to classify cell types or perform single-cell biochemical analysis on pre-isolated types of cells. However, new techniques are needed to efficiently classify cells and conduct biochemical experiments on multiple cell types concurrently. System integration and design automation are major challenges in this context. To overcome these challenges, we present a hybrid microfluidic platform that enables complete single-cell analysis on a heterogeneous pool of cells. We combine this architecture with an associated design-automation and optimization framework, referred to as Co-Synthesis (CoSyn). The proposed framework employs real-time resource allocation to coordinate the progression of concurrent cell analysis. Simulation results show that CoSyn efficiently utilizes platform resources and outperforms baseline techniques. Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ulf Schlichtmann |
DATE | 2 |
| 2017 | Testing microfluidic Fully Programmable Valve Arrays (FPVAs)abstractFully Programmable Valve Array (FPVA) has emerged as a new architecture for the next-generation flow-based microfluidic biochips. This 2D-array consists of regularly-arranged valves, which can be dynamically configured by users to realize microfluidic devices of different shapes and sizes as well as interconnections. Additionally, the regularity of the underlying structure renders FPVAs easier to integrate on a tiny chip. However, these arrays may suffer from various manufacturing defects such as blockage and leakage in control and flow channels. Unfortunately, no efficient method is yet known for testing such a general-purpose architecture. In this paper, we present a novel formulation using the concept of flow paths and cut-sets, and describe an ILP-based hierarchical strategy for generating compact test sets that can detect multiple faults in FPVAs. Simulation results demonstrate the efficacy of the proposed method in detecting manufacturing faults with only a small number of test vectors. Bing Li 0005, Bhargab B. Bhattacharya, Krishnendu Chakrabarty, Tsung-Yi Ho, Ulf Schlichtmann |
DATE | 4 |
| 2017 | Design automation and testing of monolithic 3D ICs: Opportunities, challenges, and solutions: (Invited paper)abstractMonolithic 3D ICs (M3D) are fabricated using a sequential process that grows new device and interconnect tiers in a bottom-up fashion. This fabrication process is in contrast to through-silicon via (TSV) technology that aligns and bonds pre-built tiers. M3D offers key advantages over TSVs, including (1) orders-of-magnitude smaller inter-tier vias, (2) no need for high alignment accuracy, (3) finer-grained tier partitioning options, etc. Recent studies have shown power, performance, area, and reliability (PPAR) advantages of M3D over TSV. However, M3D also suffers from its own problems, including (1) device and interconnect performance mismatch between tiers, (2) lack of EDA solutions, (3) testing challenges, (4) cost, etc. Research efforts have also been made to model and mitigate the impact of these undesirable characteristics of M3D. We will provide a survey of work that address the above issues and conclude with future directions. Kyungwook Chang, Abhishek Koneru, Krishnendu Chakrabarty, Sung Kyu Lim |
ICCAD | 3 |
| 2017 | Sortex: Efficient timing-driven synthesis of reconfigurable flow-based biochips for scalable single-cell screeningabstractSingle-cell screening is used to sort a stream of cells into clusters (or types) based on pre-specified biomarkers, thus supporting type-driven biochemical analysis. Reconfigurable flow-based microfluidic biochips (RFBs) can be utilized to screen hundreds of heterogeneous cells within a few minutes, but they are overburdened with the control of a large number of valves. To address this problem, we present a pin-constrained RFB design methodology for single-cell screening. The proposed design is analyzed using computational fluid dynamics simulations, mapped to an RC-lumped model, and combined with a high-level synthesis framework, referred to as Sortex. Simulation results show that Sortex significantly reduces the number of control pins and fulfills the timing requirements of single-cell screening. Mohamed Ibrahim 0002, Aditya Sridhar, Krishnendu Chakrabarty, Ulf Schlichtmann |
ICCAD | 3 |
| 2017 | Adaptive error recovery in MEDA biochips based on droplet-aliquot operations and predictive analysisabstractDigital microfluidic biochips (DMFBs) are being increasingly used in biochemistry labs for automating bioassays. However, traditional DMFBs suffer from some key shortcomings: 1) inability to vary droplet volume in a flexible manner; 2) difficulty of integrating on-chip sensors; 3) the need for special fabrication processes. To overcome these problems, DMFBs based on micro-electrode-dot-array (MEDA) have recently be-en proposed. However, errors are likely to occur on a MEDA DMFB due to chip defects and the unpredictability inherent to biochemical experiments. We present fine-grained error-recovery solutions for MEDA by exploiting real-time sensing and advanced MEDA-specific droplet operations. The proposed methods rely on adaptive droplet-aliquot operations and predictive analysis of mixing. Experimental results on three representative benchmarks demonstrate the efficiency of the proposed error-recovery strategy. Zhanwei Zhong, Krishnendu Chakrabarty |
ICCAD | 3 |
| 2017 | Monolithic 3D-Enabled High Performance and Energy Efficient Network-on-ChipabstractEmergence of monolithic 3D (M3D) integration has opened up the possibility of designing the ultra-low-power and high-performance circuits and systems. The smaller dimensions of monolithic inter-tier vias (MIVs) offer high density integration, the flexibility of partitioning logic blocks across multiple tiers, and significantly reduced total wire-length. In this work, we explore the design space of M3D-enabled energy-efficient NoC architectures and present a comparative performance evaluation with TSV-based counterparts. We describe the optimization of the link and router placements of the M3D-enabled NoC to ensure maximum achievable performance. The placement of M3D-enabled routers and links are explored using a machine-learning-inspired optimization algorithm. The proposed M3D-enabled NoC architecture achieves 32% lower energy-delay-product (EDP) compared to the conventional mesh-based counterpart. We also demonstrate that for the diverse set of benchmarks considered in this work, the M3D-enabled NoC, on an average, achieves 28% lower EDP than the TSV-based counterpart. Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ICCD | 4 |
| 2017 | A Design-for-Test Solution for Monolithic 3D Integrated CircuitsabstractMonolithic three-dimensional (M3D) integration has the potential to achieve significantly higher device density compared to 3D integration based on through-silicon vias (TSVs). We propose a test solution for M3D ICs based on dedicated test layers that are inserted between functional layers. We evaluate the cost associated with the proposed design-for-test (DfT) solution and compare it with that for a potential DfT solution based on the IEEE Std. P1838. Our results show that the proposed solution is more cost-efficient than the P1838-based solution for a wide range of inter-layer via (ILV) density, ILV yield, and defect density. Abhishek Koneru, Sukeshwar Kannan, Krishnendu Chakrabarty |
ICCD | 3 |
| 2017 | Security Trade-Offs in Microfluidic Routing FabricsabstractMicrofluidic routing fabrics, or crossbars, based on transposer primitives provide benefits in manufacturability, performance, and on-the-fly reconfigurability. Many applications in microfluidics, such as DNA barcoding for single-cell analysis, are expected to benefit from these new devices. However, the control of these critical devices poses new security questions that may impact the functional integrity of a microbiology application. This paper explores the many security implications of microfluidic crossbars that directly result from their structure, programmability and use in critical applications. We analyze security performance using new metrics describing how fluids can be "scattered" to incorrect locations under fault-injection attacks, and from these derive a probability model describing the likelihood of a successful attack. We present a case study of a recently described routing fabric proposed for use in a hybrid DNA barcoding platform, and discuss how fabric designers can improve security through architectural choices. Jack Tang, Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Ramesh Karri |
ICCD | 3 |
| 2017 | Run-time hardware trojan detection using performance countersabstractThere has been a growing trend in recent years to outsource various aspects of the semiconductor design and manufacturing flow to different parties spread across the globe. Such outsourcing increases the risk of adversaries adding malicious logic, referred to as hardware Trojans, to the original design. In this paper, we introduce a run-time hardware Trojan detection method for microprocessor cores. This approach uses Half-space trees to detect the activation of Trojans that introduce abnormal patterns in the data streams obtained from performance counters. It does not require any additional hardware or the monitoring of a large number of internal signals. We evaluate our method by detecting the activation of Trojans that cause denial-of-service, the degradation of system performance, and change in functionality of a microprocessor core. Results obtained using the OpenSPARC T1 core and an FPGA prototyping framework show that Trojan activation is detected with true positive ratio of above 0.9 and a false positive ratio of below 0.1 for most of the implemented Trojans. Rana Elnaggar, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ITC | 2 |
| 2017 | Changepoint-based anomaly detection in a core router systemabstractPrognostic diagnosis is desirable for commercial core router systems to ensure early failure prediction and fast error recovery. The effectiveness of prognostic diagnosis depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the statistical properties of the monitored data change significantly as time proceeds. We describe the design of a changepoint-based anomaly detector that first detects changepoints from collected time-series data, and then utilizes these changepoints to detect anomalies. Two approaches based on maximum-likelihood estimation are implemented to detect different types of changepoints. A clustering method is then developed to identify a wide range of normal/abnormal patterns from changepoint windows. Data collected from a set of commercial core router systems are used to validate the proposed anomaly detector. Experimental results show that our changepoint-based anomaly detector achieves better performance than traditional methods. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 3 |
| 2017 | Symbol-based health-status analysis in a core router systemabstractTo ensure high reliability and rapid error recovery in commercial core router systems, a health-status analyzer is essential to monitor the different features of core routers. However, traditional health analyzers need to store a large amount of historical data in order to identify health status. The storage requirement becomes prohibitively high when we attempt to carry out long-term health-status analysis for a large number of core routers. We describe the design of a symbol-based health status analyzer that first encodes, as a symbol sequence, the long-term complex time series collected from a number of core routers, and then utilizes the symbol sequence to do health analysis. The symbolic aggregation approximation (SAX) and moving-average-based trend approximation methods are implemented to encode complex time series in a hierarchical way. Hierarchical agglomerative clustering and sequitur rule discovery are implemented to learn important global and local patterns. Two classification methods are then utilized to identify the health status of core routers. Data collected from a set of commercial core router systems are used to validate the proposed health-status analyzer. The experimental results show that our symbol-based health status analyzer requires much lower storage than traditional methods, but can still maintain comparable diagnosis accuracy. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 3 |
| 2017 | Software-based online self-testing of network-on-chip using bounded model checkingabstractOnline testing is critical to ensure reliable operation of manycore systems based on a network-on-chip (NoC) interconnection fabric. We present a software-based online NoC self-testing solution based on bounded model checking (BMC). The proposed method first implements BMC on a sliced extended finite-state machine, and extracts the leading sequences necessary to excite NoC functions. Next, it targets the structural faults within every function excited by the leading sequence through constrained ATPG. Finally, a test protocol is developed to make the test responses observable. Experimental results show that the proposed method achieves high fault coverage in functional mode and outperforms previously proposed solutions. In addition, the fault coverage is very close to that of full-scan testing, but without any area overhead. Ying Zhang 0040, Krishnendu Chakrabarty, Huawei Li 0001, Jianhui Jiang |
ITC | 2 |
| 2017 | Test-cost optimization in a scan-compression architecture using support-vector regressionabstractScan compression is widely used in high-volume testing of complex integrated circuits. With an increase in design complexity, the increased density of unknown (X) values from output responses reduces compression efficiency. In order to effectively block X values and maximize the effectiveness of test compression, a scan-compression architecture has recently been proposed, in which deterministic test patterns can be loaded into selected scan cells by controlling the initial state of the pseudo-random pattern generator (PRPG). A careful selection of the PRPG length is however essential to reduce test cost. We propose an optimization method based on support-vector regression to determine the PRPG length for test-cost reduction in a given scan-compression architecture. A correlation-based feature selection methodology is also proposed to reduce the amount of data needed for the accurate selection of the PRPG length. Experimental results on industrial designs highlight the effectiveness of the proposed method. Jonathon E. Colburn, Vinod Pagalone, Kaushik Narayanun, Krishnendu Chakrabarty |
VTS | 5 |
| 2017 | Offline Error Detection in MEDA-Based Digital Microfluidic Biochips Using Oscillation-Based Testing Methodology
Vineeta Shukla, Fawnizu Azmadi Hussin, Nor Hisham Hamid, Noohul Basheer Zain Ali, Krishnendu Chakrabarty |
J. Electron. Test. | 5 |
| 2017 | Impact of Electrostatic Coupling and Wafer-Bonding Defects on Delay Testing of Monolithic 3D Integrated CircuitsabstractMonolithic three-dimensional (M3D) integration is gaining momentum, as it has the potential to achieve significantly higher device density compared to 3D integration based on through-silicon vias. M3D integration uses several techniques that are not used in the fabrication of conventional integrated circuits (ICs). Therefore, a detailed analysis of the M3D fabrication process is required to understand the impact of defects that are likely to occur during chip fabrication. In this article, we first analyze electrostatic coupling in M3D ICs, which arises due to the aggressive scaling of the interlayer dielectric (ILD) thickness. We then analyze defects that arise due to voids created during wafer bonding, a key step in most M3D fabrication processes. We quantify the impact of these defects on the threshold voltage of a top-layer transistor in an M3D IC. We also show that wafer-bonding defects can lead to a change in the resistance of interlayer vias (ILVs), and in some cases lead to an open in an ILV or a short between two ILVs. We then analyze the impact of these defects on path delays using HSpice simulations. We study their impact on the effectiveness of delay-test patterns for multiple instances of IWLS 2005 benchmarks in which these defects were randomly injected. Our results show that the timing characteristics of an M3D IC can be significantly altered due to coupling and wafer-bonding defects if the thickness of its ILD is less than 100nm. Therefore, for such M3D ICs, test-generation methods must be enhanced to take M3D fabrication defects into account. Abhishek Koneru, Sukeshwar Kannan, Krishnendu Chakrabarty |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2017 | Adaptation of Biochemical Protocols to Handle Technology-Change for Digital MicrofluidicsabstractAdvances in digital microfluidic (DMF) technologies offer a promising platform for a variety of biochemical applications, ranging from massively parallel DNA analysis and computational drug discovery to toxicity monitoring and medical diagnosis. In this paper, we address the migration problem that arises when the technology undergoes a change in the context of DMFs. Given a biochemical reaction synthesized for actuation on a given DMF architecture, we discuss how the same biochemical reaction can be ported seamlessly to an enhanced architecture, with possible modifications to the architectural parameters (e.g., clock frequency, mixer size, and mixing time) or geometric changes (e.g., change in reservoir locations or mixer positions, inclusion of new sensors or other physical resources). Complete resynthesis of the protocol for the new architecture may often become either inefficient or even infeasible due to scalability, proprietary, security, or cost issues. We propose an adaptation method for handling such technology-changes by modifying the existing actuation sequence through an incremental procedure. The foundation of our method lies in symbolic encoding and satisfiability-solvers, enriched with pertinent graph-theoretic and geometric techniques. This enables us to generate functionally correct solutions for the new target architecture without necessitating a complete resynthesis step, thereby enabling the utilization of these chips by users in biology who are not familiar with the on-chip synthesis tool-flow. We highlight the benefits of the proposed approach through extensive simulations on assay benchmarks. Sukanta Bhattacharjee, Sharbatanu Chatterjee, Ansuman Banerjee, Tsung-Yi Ho, Krishnendu Chakrabarty, Bhargab B. Bhattacharya |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Design-Space Exploration and Optimization of an Energy-Efficient and Reliable 3-D Small-World Network-on-ChipabstractA 3-D network-on-chip (NoC) enables the design of high performance and low power many-core chips. Existing 3-D NoCs are inadequate for meeting the ever-increasing performance requirements of many-core processors since they are simple extensions of regular 2-D architectures and they do not fully exploit the advantages provided by 3-D integration. Moreover, the anticipated performance gain of a 3-D NoC-enabled many-core chip may be compromised due to the potential failures of through-silicon-vias that are predominantly used as vertical interconnects in a 3-D IC. To address these problems, we propose a machine-learning-inspired predictive design methodology for energy-efficient and reliable many-core architectures enabled by 3-D integration. We demonstrate that a small-world network-based 3-D NoC (3-D SWNoC) performs significantly better than its 3-D MESH-based counterparts. On average, the 3-D SWNoC shows 35% energy-delay-product improvement over 3-D MESH for the PARSEC and SPLASH2 benchmarks considered in this paper. To improve the reliability of 3-D NoC, we propose a computationally efficient spare-vertical link (sVL) allocation algorithm based on a state-space search formulation. Our results show that the proposed sVL allocation algorithm can significantly improve the reliability as well as the lifetime of 3-D SWNoC. Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Control-Layer Routing and Control-Pin Minimization for Flow-Based Microfluidic BiochipsabstractRecent advances in flow-based microfluidic biochips have enabled the emergence of lab-on-a-chip devices for bimolecular recognition and point-of-care disease diagnostics. However, the adoption of flow-based biochips is hampered today by the lack of computer-aided design tools. Manual design procedures not only delay product development but they also inhibit the exploitation of the design complexity that is possible with current fabrication techniques. In this paper, we present the first practical problem formulation for automated control-layer design in flow-based microfluidic very large-scale integration (mVLSI) biochips and propose a systematic approach for solving this problem. Our goal is to find an efficient routing solution for control-layer design with a minimum number of control pins. The pressure-propagation delay, an intrinsic physical phenomenon in mVLSI biochips, is minimized in order to reduce the response time for valves, decrease the pattern set-up time, and synchronize valve actuation. Two fabricated flow-based devices and six synthetic benchmarks are used to evaluate the proposed optimization method. Compared with manual control-layer design and a baseline approach, the proposed approach leads to fewer control pins, better timing behavior, and shorter channel length in the control layer. Kai Hu 0003, Trung Anh Dinh, Tsung-Yi Ho, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Synthesis of Cyberphysical Digital-Microfluidic Biochips for Real-Time Quantitative AnalysisabstractConsiderable effort has recently been directed toward the implementation of molecular bioassays on digital-microfluidic biochips (DMFBs). However, today's solutions suffer from the drawback that multiple sample pathways are not supported and on-chip reconfigurable devices are not efficiently exploited. As a result, impractical manual intervention is needed to process protocols for gene-expression analysis. To overcome this problem, we first describe our benchtop experimental studies to understand gene-expression analysis and its relationship to the biochip design specification. We then introduce an integrated framework for quantitative gene-expression analysis using DMFBs. The proposed framework includes: 1) a spatial-reconfiguration technique that incorporates resource-sharing specifications into the synthesis flow; 2) an interactive firmware that collects and analyzes sensor data based on quantitative polymerase chain reaction; and 3) a real-time resource-allocation scheme that responds promptly to decisions about the protocol flow received from the firmware layer. This framework is combined with cyberphysical integration to develop the first design-automation framework for quantitative gene expression. Simulation results show that our adaptive framework efficiently utilizes on-chip resources to reduce time-to-result without sacrificing the chip's lifetime. Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Kristin Scott |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | A Branch-&-Bound Test-Access-Mechanism Optimization Method for Multi-Vdd SoCsabstractThe use of multiple voltage levels introduces new challenges for testing multi-Vddsystems-on-chip (SoCs). Timedivision-multiplexing (TDM) tackles many of these challenges and offers very effective test-schedules. However, the effectiveness of TDM for minimizing test time depends on the test-access-mechanism (TAM) in the SoC. Single-VddTAM optimization techniques consider neither the highly constrained test environment of multi-VddSoCs nor the benefits provided by TDM, therefore they are not suitable for multi-VddSoCs. In this paper, we propose the first TAM optimization technique for multi-VddSoCs. The proposed method exploits unique scheduling opportunities and flexibility offered by TDM, and by the means of a branch-&-bound approach, it quickly identifies the most effective TAM configurations. Experiments using large benchmark SoCs as well as SoCs from industry highlight the benefits of the proposed technique on multi-Vdddesigns, for both single-site and multisite test applications. Fotis Vartziotis, Xrysovalantis Kavousianos, Panagiotis Georgiou, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Prebond Testing and Test-Path Design for the Silicon Interposer in 2.5-D ICsabstractIn interposer-based 2.5-D integrated circuits, the passive silicon interposer is the least expensive component in the chip. Thus, it is desirable to test the interposer before bonding to ensure that more expensive and defect-free dies are not stacked on a faulty interposer. We present an efficient method to locate defects in a passive interposer before stacking. The proposed test architecture uses e-fuses that can be programmed to connect or disconnect functional paths inside the interposer. The concept of die footprint is utilized for interconnect testing, and the overall assembly and test flow is described. Moreover, the concept of weighted critical area is defined and utilized to reduce test time. In order to fully determine the location of each e-fuse and the order of functional interconnects in a test path, we also present a test-path design algorithm. The proposed algorithm can generate all test paths for interconnect testing. We present HSPICE simulation results to demonstrate the effectiveness of the prebond test solution. Test-path designs are also presented to highlight the efficiency of the test-path design algorithm. The benefit of using weighted critical area is demonstrated using a commercial interposer from industry. Ran Wang 0002, Sukeshwar Kannan, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | ExTest Scheduling and Optimization for 2.5-D SoCs With Wrapped TilesabstractInterposer-based 2.5-D integrated circuits (ICs) enable high-density interconnects, but introduce new challenges for the testing of a system-on-chip (SoC) die on an interposer. This paper presents two efficient ExTest scheduling strategies that implements interconnect testing between tiles inside an SoC die while satisfying the practical constraint that the number of required test pins cannot exceed the number of available pins at the chip level. These strategies target two different ways in which SoC dies are wrapped in 2.5-D ICs. The first scheduling approach is aimed at an extremely large SoC in which the wrapper design requires concurrent testing of the interconnects driving the tile under test. The second scheduling approach is applicable to more general wrapper designs that provide more flexibility in terms of the manner in which these interconnects can be tested. In both test strategies, the tiles in the SoC die are divided into groups based on the manner in which they are interconnected. In order to minimize the test time, two optimization solutions are introduced. The first solution minimizes the number of input test pins, and the second solution minimizes the number of output test pins. In addition, two subgroup configuration methods are further proposed to generate subgroups inside each test group. To highlight the effectiveness of the proposed test strategies, we present scheduling and optimization results for two SoC dies for 2.5-D ICs currently in production. Ran Wang 0002, Guoliang Li 0004, Rui Li 0084, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Synthesis of Error-Recovery Protocols for Micro-Electrode-Dot-Array Digital Microfluidic BiochipsabstractA digital microfluidic biochip (DMFB) is an attractive technology platform for various biomedical applications. However, a conventional DMFB is limited by: (i) the number of electrical connections that can be practically realized, (ii) constraints on droplet size and volume, and (iii) the need for special fabrication processes and the associated reliability/yield concerns. To overcome the above challenges, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed and fabricated recently. Error recovery is of key interest for MEDA biochips due to the need for system reliability. Errors are likely to occur during droplet manipulation due to defects, chip degradation, and the uncertainty inherent in biochemical experiments. In this paper, we first formalize error-recovery objectives, and then synthesize optimal error-recovery protocols using a model based on Stochastic Multiplayer Games (SMGs). We also present a global error-recovery technique that can update the schedule of fluidic operations in an adaptive manner. Using three representative real-life bioassays, we show that the proposed approach can effectively reduce the bioassay completion time and increase the probability of success for error recovery. Mahmoud Elfar, Zhanwei Zhong, Krishnendu Chakrabarty, Miroslav Pajic |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | VFI-Based Power Management to Enhance the Lifetime of High-Performance 3D NoCsabstractThe emergence of 3D network-on-chip (NoC) has revolutionized the design of high-performance and energy-efficient manycore chips. However, the anticipated performance gain can be compromised due to the degradation and failure of vertical links (VLs). The Through-Silicon-Via (TSV)-enabled VLs may fail due to workload-induced stress; the failure of a VL can affect the neighboring VLs, thereby causing a cascade of failures and reducing the lifetime of the chip. To enhance the reliability of 3D NoC-enabled manycore chips, we propose to incorporate a voltage-frequency island (VFI)-based power management strategy that helps to reduce the energy consumption and hence, the workload-induced stress of the highly utilized VLs. The adopted power-management strategy relies on control decisions about the voltage/frequency (V/F) levels on VLs. We demonstrate that compared to the well-known spare TSV allocation and adaptive routing strategies, power management is more effective in enhancing the reliability of a 3D NoC. VFI-based power management improves the reliability of the 3D NoC by one order of magnitude compared to both adaptive routing and spare allocation while running popular SPLASH-2 and PARSEC benchmarks. The principal benefit of power management is that it is capable of reducing the operating temperature of the system, which in turn enhances the Mean-Time-To-Failure (MTTF) of the VLs and reliability of the overall 3D NoC. Sourav Das 0002, Wonje Choi 0001, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | EditorialabstractWe are pleased to announce this year’s winners of the IEEE Transactions on Very Large Scale integrated (VLSI) Systems (TVLSI) Circuits and Systems (CAS) Society Best Reviewer and Associate Editor Awards. We had a number of qualified candidates for each award, but after a thorough evaluation process including input from the Associate Editors and Selection Committee, these five individuals stood out among all the candidates: Krishnendu Chakrabarty, Massimo Alioto, Rajiv V. Joshi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Test and Reliability Issues in 2.5D and 3D IntegrationabstractIncreasing wire delay and higher interconnect power consumption are major concerns for nanoscale CMOS ICs. Three-dimensional integrated circuits (3D ICs) based on through-silicon-vias (TSVs) appear to be a promising solution to overcome bottleneck in CMOS scaling. However, volume production and commercial exploitation of 3D ICs are not feasible before pressing concerns about heat dissipation and test cost, as well as manufacturing yield and resiliency challenges are adequately addressed. At present, interposer-based 2.5D ICs are being advocated as a precursor to 3D ICs. All the dies in a 2.5D or 3D IC must be adequately tested for product qualification. Moreover, the introduction of TSVs for both signal routing across multiple dies as well as power delivery network (PDN), imposes new challenges in terms of manufacturing yield and resiliency issues which should be addressed in both design and test flows. The purpose of this special session, consisting of a set of talks given by experts from the US, Asia and Europe, is to present the test and resiliency challenges faced for the 2.5D and 3D integrated circuits, and discuss the path to overcome such challenges. Mehdi Baradaran Tahoori, Krishnendu Chakrabarty |
ATS | 2 |
| 2016 | Testing of Interposer-Based 2.5D Integrated Circuits: Challenges and SolutionsabstractInterposer-based 2.5D integrated circuits (ICs) are seen today as a precursor to 3D ICs based on through-silicon vias. This paper describes some of the major challenges related to testing of 2.5D ICs and presents some solutions to these problems. We first describe a test architecture using e-fuses for pre-bond interposer testing. We next present an efficient built-in self-test (BIST) technique that targets the dies and the interposer interconnects. Finally, we present a programmable method for shift-clock stagger assignment to reduce power supply noise during SoC die testing in 2.5D ICs. Ran Wang 0002, Krishnendu Chakrabarty |
ATS | 2 |
| 2016 | Multicast Test Architecture and Test Scheduling for Interposer-Based 2.5D ICsabstractInterposer-based 2.5D integrated circuits (ICs) are seen today as a precursor to 3D ICs based on through-silicon vias (TSVs). All the dies in a 2.5D IC must be adequately tested for product qualification. However, due to the limited number of package pins, it is a a major challenge to test 2.5 ICs using conventional methods. Moreover, due to higher integration levels, test-application time and test power consumption for 2.5D ICs are also increased compared to their 2D counterparts. Therefore, it is imperative to take these issues into account during 2.5D IC testing. In this work, we present an efficient multicast test architecture for targeting defects in dies, in which multiple dies can be tested simultaneously to reduce the test-application time under constraints on test power and fault coverage. We also propose a test scheduling and optimization technique that can be utilized with the multicast test architecture. Compared to previous work, the proposed technique can reduce testapplication time by 53:4% for benchmark designs while achieving higher fault coverage. Shengcheng Wang, Ran Wang 0002, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ATS | 3 |
| 2016 | A real-time digital-microfluidic platform for epigeneticsabstractAdvances in digital-microfluidic biochips have led to miniaturized platforms that can implement biomolecular assays. However, these designs are not adequate for running multiple sample pathways because they consider unrealistic static schedules; hence runtime adaptation based on assay outcomes is not supported and only a rigid path of bioassays can be run on the chip. We present a design framework that performs fluidic task assignment, scheduling, and dynamic decision-making for quantitative epigenetics. We first describe our benchtop experimental studies to understand the relevance of chromatin structure on the regulation of gene function and its relationship to biochip design specifications. The proposed method models biochip design in terms of real-time multiprocessor scheduling and utilizes a heuristic algorithm to solve this NP-hard problem. Simulation results show that the proposed algorithm is computationally efficient and it generates effective solutions for multiple sample pathways on a resource-limited biochip. We also present experimental results using an embedded microcontroller as a testbed. Mohamed Ibrahim 0002, Craig Boswell, Krishnendu Chakrabarty, Kristin Scott, Miroslav Pajic |
CASES | 3 |
| 2016 | High-level synthesis for micro-electrode-dot-array digital microfluidic biochipsabstractA digital microfluidic biochip (DMFB) is an attractive technology platform for automating laboratory procedures in biochemistry. However, today's DMFBs suffer from several limitations: (i) constraints on droplet size and the inability to vary droplet volume in a fine-grained manner; (ii) the lack of integrated sensors for real-time detection; (iii) the need for special fabrication processes and reliability/yield concerns. To overcome the above problems, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have recently been demonstrated. However, due to the inherent differences between today's DMFBs and MEDA, existing synthesis solutions cannot be utilized for MEDA-based biochips. We present the first biochip synthesis approach that can be used for MEDA. The proposed synthesis method targets operation scheduling, module placement, routing of droplets of various sizes, and diagonal movement of droplets in a two-dimensional array. Simulation results using benchmarks and experimental results using a fabricated MEDA biochip demonstrate the effectiveness of the proposed co-optimization technique. Kelvin Yi-Tse Lai, Po-Hsien Yu, Tsung-Yi Ho, Krishnendu Chakrabarty, Chen-Yi Lee |
DAC | 5 |
| 2016 | Reliability and performance trade-offs for 3D NoC-enabled multicore chips
Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
DATE | 4 |
| 2016 | Integrated and real-time quantitative analysis using cyberphysical digital-microfluidic biochips
Mohamed Ibrahim 0002, Krishnendu Chakrabarty, Kristin Scott |
DATE | 2 |
| 2016 | Pre-bond testing of the silicon interposer in 2.5D ICs
Ran Wang 0002, Sukeshwar Kannan, Krishnendu Chakrabarty |
DATE | 4 |
| 2016 | Thermal-aware TSV repair for electromigration in 3D ICs
Shengcheng Wang, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty |
DATE | 3 |
| 2016 | Two-dimensional time-division multiplexing for 3D-SoCsabstractThrough-silicon vias (TSVs) are used as high-speed interconnects between dies in a 3D System-on-Chip (SoC). However, their speed cannot be utilized during test application due to inherent limitations of the scan chains of the cores, which prevent the use of high shift frequencies. Moreover, due to their high area cost, only a limited number of TSVs can be used for test application. As a result, the time needed for transferring test data to the cores in multiple dies can be considerable. We propose an efficient test-access mechanism (TAM) architecture, which exploits the high speed of TSVs to minimize the time for testing 3D SoCs. By the means of time-division multiplexing and an effective test scheduling method, the proposed TAM architecture offers significant savings in test time, TSV count and TAM cost. Panagiotis Georgiou, Fotis Vartziotis, Xrysovalantis Kavousianos, Krishnendu Chakrabarty |
ETS | 4 |
| 2016 | Analysis of electrostatic coupling in monolithic 3D integrated circuits and its impact on delay testingabstractMonolithic 3D (M3D) integration is a promising technology that offers considerable performance and area benefits. A number of techniques have been proposed in the literature for the design and fabrication of M3D integrated circuits (ICs). Despite these advances, test challenges have remained unexplored. As a first step towards the development of test solutions, we analyze electrostatic coupling in M3D ICs and quantify its impact on delay testing. We carry out a detailed study of coupling between device layers, and quantify the change in threshold voltage of transistors in the top layers. Such variations in threshold voltage can significantly impact circuit timing. Next, we analyze the impact of coupling on the effectiveness of delay-test patterns using the statistical delay quality level (SDQL) as a metric. Our results show that significant rethinking in test generation is needed to effectively screen delay defects in M3D ICs. Abhishek Koneru, Krishnendu Chakrabarty |
ETS | 2 |
| 2016 | A design-for-test solution for monolithic 3D integrated circuitsabstractMonolithic three-dimensional integrated circuits (M3D ICs) are being advocated as the next generation of 3D integration beyond 3D ICs based on through-silicon-vias. Testing of the bottom layer of an M3D IC is necessary to target defects arising from the layered manufacturing process. We present an efficient design-for-test (DfT) method for the bottom layer by isolating it from the top layer. A bypass structure based on e-fuses is proposed to connect pairs of inter-layer vias (ILVs). In order to minimize the wire length between paired ILVs, an ILV-pairing problem is formulated and then solved using a technique based on maximum-weighted bipartite matching. The independent ILVs, i.e., those that are not paired, are made controllable and observable using four different types of DfT structures. A cost-optimization problem is solved to minimize the DfT cost. We present ILV-pairing and cost-optimization results for designs based on the ITC'02 benchmarks as well as for an industry design. We also present HSpice simulation results to show that testing using e-fuses is feasible. Ran Wang 0002, Krishnendu Chakrabarty |
ETS | 2 |
| 2016 | Energy-efficient and reliable 3D network-on-chip (NoC): architectures and optimization algorithmsabstractThe Network-on-Chip (NoC) paradigm has emerged as an enabler for integrating a large number of embedded cores in a single die. Three-dimensional (3D) integration, a breakthrough technology to achieve “More Moore and More Than Moore,” provides numerous benefits e.g., better performance, lower power consumption, and higher bandwidth, by utilizing vertical interconnects and 3D stacking. Energy-efficient and high-bandwidth vertical interconnects enable the design of an energy efficient 3D NoC for massive manycore platforms. Existing 3D NoCs are deficient for meeting ever-increasing performance requirements of manycore processors since they are simple extension of regular 2D architectures and they do not fully exploit the advantages provided by 3D integration. Moreover, the anticipated performance gain of a 3D NoC-enabled manycore chip will be compromised due to the potential failures of through-silicon-vias (TSVs) that are predominantly used as vertical interconnects in a 3D IC. In this paper, we present the various challenges and possible solutions for designing energy-efficient and reliable manycore chips enabled by the 3D integration. Sourav Das 0002, Janardhan Rao Doppa, Partha Pratim Pande, Krishnendu Chakrabarty |
ICCAD | 4 |
| 2016 | Error recovery in a micro-electrode-dot-array digital microfluidic biochip?abstractA digital microfluidic biochip (DMFB) is an attractive technology platform for automating laboratory procedures in biochemistry. However, today's DMFBs suffer from several limitations: (i) constraints on droplet size and the inability to vary droplet volume in a fine-grained manner; (ii) the lack of integrated sensors for real-time detection; (iii) the need for special fabrication processes and the associated reliability/yield concerns. To overcome the above problems, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed recently, and droplet manipulation on these devices has been experimentally demonstrated. Errors are likely to occur due to defects, chip degradation, and the lack of precision inherent in biochemical experiments. Therefore, an efficient error-recovery strategy is essential to ensure the correctness of assays executed on MEDA biochips. By exploiting MEDA-specific advances in droplet sensing, we present a novel error-recovery technique to dynamically reconfigure the biochip using real-time data provided by on-chip sensors. Local recovery strategies based on probabilistic-timed-automata are presented for various types of errors. A control flow is also proposed to connect local recovery procedures with global error recovery for the complete bioassay. Laboratory experiments using a fabricated MEDA chip are used to characterize the outcomes of key droplet operations. The PRISM model checker and three analytical chemistry benchmarks are used for an extensive set of simulations. Our results highlight the effectiveness of the proposed error-recovery strategy. Kelvin Yi-Tse Lai, Po-Hsien Yu, Krishnendu Chakrabarty, Miroslav Pajic, Tsung-Yi Ho, Chen-Yi Lee |
ICCAD | 4 |
| 2016 | The hype, myths, and realities of testing 3D integrated circuitsabstractThree-dimensional (3D) integration using through-silicon vias (TSVs) promises higher integration levels in a single package, keeping pace with Moore's law. Despite the promise and benefits offered by 3D integration, testing remains a major obstacle that hinders its widespread adoption. This paper examines the hype, myths, and realities of 3D IC testing. We describe a number of testing and DfT challenges, and present some solutions being advocated for the challenges of “What to Test”, “How to Test”, and “When to Test”. Techniques highlighted in this paper include: (i) testing of the silicon interposer; (ii) pre-bond TSV testing; (iii) cost modeling and test-flow selection; (iv) a reconfigurable built-in self-test infrastructure. Ran Wang 0002, Sergej Deutsch, Mukesh Agrawal 0001, Krishnendu Chakrabarty |
ICCAD | 4 |
| 2016 | Accurate anomaly detection using correlation-based time-series analysis in a core router systemabstractFault tolerance is used in communication systems to ensure high reliability and rapid error recovery. The effectiveness of most proactive fault-tolerant mechanism depends on whether anomalies can be accurately detected before a failure occurs. However, traditional anomaly detection techniques fail to detect “outliers” when the monitored data involves temporal measurements and exhibits significantly different statistical characteristics for its constituent features. We describe the design of an anomaly detector that monitors the time-series data of a complex core router system. Anomaly detection techniques are compared in terms of their effectiveness for detecting different types of anomalies. A feature-categorizing-based hybrid method is proposed to overcome the difficulty of detecting anomalies in features with different statistical characteristics. Furthermore, a correlation analyzer is implemented to remove irrelevant and redundant features. Three types of synthetic anomalies, generated using a small amount of real data for a commercial telecom system, are used to validate the proposed anomaly detector. Shi Jin 0001, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
ITC | 3 |
| 2016 | Supply-voltage optimization to account for process variations in high-volume manufacturing testingabstractIn a high-volume manufacturing environment, it is difficult to set up test conditions, including supply-voltage levels, that can take into account the wide range of interdie process variations on a wafer. Test conditions that do not take into consideration these process variations can lead to either yield loss or poor quality control. In this work, we propose a method to identify supply-voltage levels to test semiconductor chips based on the process variations experienced by them, while also adapting these supply-voltage levels based on the chip locations on the wafer. For this purpose, we first identified various process zones on the wafer, based on frequencies of on-chip ring oscillators. Next we modeled the ring oscillator (RO) performance in terms of the variations in SPICE parameters using the design of experiments (DoE) method. The equation corresponding to the DoE model was solved for each process zone on the wafer in order to fit a set of independent SPICE parameters set for the corresponding zone. These independent SPICE parameters were then used to create an updated SPICE model for the ring oscillator corresponding to each zone, and these SPICE models were used to derive an appropriate supply-voltage level for each zone. The results were used to test 250 devices to identify chips that exhibit significant performance deviation compared to other chips from the same zone. Among these 250 devices, 89 belonged to the slow zone, 87 belonged to the medium-fast zone, and 74 belonged to the fast process zone. A total of seven devices from the medium-fast zone and nine devices from the fast zone showed performance deviations and they were successfully screened. Gurunath Kadam, Markus Rudack, Krishnendu Chakrabarty, Juergen Alt |
ITC | 3 |
| 2016 | Defect tolerance for CNFET-based SRAMsabstractSRAMs based on carbon nanotube field-effect transistors (CNFETs) offer a promising alternative to conventional SRAMs due to their high energy efficiency and low leakage. However, the imperfect CNT fabrication process introduces high defect rates and a unique defect distribution; these problems may offset the power/performance benefits of CNFET-based SRAMs and lead to yield degradation. We propose a redundancy architecture with asymmetrically partitioned column blocks and the sharing of spares among column blocks. We also present a analytical model to characterize the distribution of faults, which can guide the design exploration of the proposed redundancy architecture. Simulation results highlight the accuracy of the proposed model, as well as the efficiency and effectiveness of the redundancy architecture. Tianjian Li, Li Jiang 0002, Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty |
ITC | 5 |
| 2016 | Built-in self-test for micro-electrode-dot-array digital microfluidic biochipsabstractA digital microfluidic biochip (DMFB) is an attractive platform for immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. However, today's DMFBs suffer from several limitations, including (i) the lack of integrated sensors for real-time detection, (ii) constraints on droplet size and the inability to vary droplet volume in a fine-grained manner, and (iii) the need for special fabrication processes and the associated reliability/yield concerns. To overcome the above limitations, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed recently. Droplet manipulation on MEDA biochips has also been experimentally demonstrated. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, MEDA biochips must be adequately tested before they can be used for bioassay execution. We present an efficient built-in self-test (BIST) architecture for MEDA biochips. The proposed BIST architecture can effectively detect defects in a MEDA biochip, and faulty microcells can be identified. Simulation results based on HSPICE and experiments using fabricated MEDA biochips highlight the effectiveness of the proposed BIST architecture. Kelvin Yi-Tse Lai, Po-Hsien Yu, Krishnendu Chakrabarty, Tsung-Yi Ho, Chen-Yi Lee |
ITC | 4 |
| 2016 | Securing digital microfluidic biochips by randomizing checkpointsabstractMuch progress has been made in digital microfluidic biochips (DMFB), with a great body of literature addressing low-cost, high-performance, and reliable operation. Despite this progress, security of DMFBs has not been adequately addressed. We present an analysis of a DMFB system prone to malicious modification of routes and propose a DMFB defense based on spatio-temporal randomized checkpoints using CCD cameras. Absent the knowledge of the time- and space-randomized checkpoints, an attacker cannot navigate the DMFB without alerting the system. We present an algorithm to guide the placement and timing of the checkpoints such that the probability that an attack can evade detection is minimized. The efficacy of the defense mechanism is illustrated with a case study under stealthy malicious modifications. Jack Tang, Ramesh Karri, Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
ITC | 4 |
| 2016 | Testing of interposer-based 2.5D integrated circuitsabstractInterposer-based 2.5D integrated circuits (ICs) are seen today as a precursor to 3D ICs based on through-silicon vias (TSVs). All the dies and the interposer in a 2.5D IC must be adequately tested for product qualification. This work provides solutions to new challenges related to testing of 2.5D ICs. We propose a test architecture using e-fuses for pre-bond interposer testing. We design a test architecture that is fully compatible with the IEEE 1149.1 standard and relies on an enhancement of the standard test access port (TAP) controller. We present an efficient built-in self-test (BIST) technique that targets the dies and the interposer interconnects. We next describe two efficient ExTest scheduling strategies that implement interconnect testing between tiles within a system on chip (SoC) die on the interposer. Finally, we present a programmable method for shift-clock stagger assignment to reduce power supply noise during SoC die testing in 2.5D ICs. Ran Wang 0002, Krishnendu Chakrabarty |
ITC | 2 |
| 2016 | A unified test and fault-tolerant multicast solution for network-on-chip designsabstractWe present a unified test technique that targets all the components of a network-on-chip design. The proposed technique targets faults in links, routers, and cores. Link faults are first located using built-in self-test hardware inserted in the routers. Test packets for routers are delivered to the routers via the fault-free links and routers identified in the previous steps. A test packet can be corrupted by faulty links or routers, therefore, it is delivered across only previously identified fault-free routers/links. Test packet delivery for routers is implemented as a fault-tolerant unicast-based multicast scheme within the tested part of the network-on-chip. After all faulty routers are identified, a new fault-tolerant unicast-based multicast routing technique is proposed to deliver test packets for the cores. Identical cores share the same test set, and they are tested within the same test session. Simulation results highlight the effectiveness of the proposed method in reducing test time. Krishnendu Chakrabarty, Hideo Fujiwara |
ITC | 2 |
| 2016 | ForewordabstractOn behalf of the Organizing and Program Committees, it is our great pleasure to welcome you to the 24th Annual IFIP/IEEE International Conference on Very Large Scale Integration, VLSI-SoC'16, in Tallinn/Estonia. The conference is held in the Radisson Park Inn Meriton Conference & Spa Hotel, just a couple of footsteps away from the beautiful Old Town of the city. VLSI-SoC 2016 is the 24th in a series of international conferences sponsored by the IFIP TC 10 Working Group 10.5, IEEE CEDA and IEEE CASS, which explores the state-of-the-art in the areas that surround Ultra Large Scale Integration (ULSI) and System-on-Chip (SoC) design and test as well as mixed-technology devices. Jaan Raik, Ian O'Connor, Thomas Hollstein, Krishnendu Chakrabarty |
VLSI-SoC | 4 |
| 2016 | Optimization of the IEEE 1687 access network for hybrid access schedulesabstractThe IEEE 1687 Standard specifies an access network and a description language for embedded instruments. In this paper, we present an optimization technique to minimize the segment insertion bit (SIB) programming overhead for IEEE 1687-compliant access architectures. We first present an optimal solution based on dynamic programming for concurrent access schedules. This technique is then utilized to minimize the SIB programming overhead for more general hybrid access schedules. The proposed optimization technique is computationally efficient and it leads to significant reductions (as large as 97%) in the SIB programming overhead for hybrid access schedules. Srinivasa Shashank Nuthakki, Rajit Karmakar, Santanu Chattopadhyay, Krishnendu Chakrabarty |
VTS | 4 |
| 2016 | Online soft-error vulnerability estimation for memory arraysabstractRadiation-induced soft errors are a major reliability concern in circuits fabricated at advanced technology nodes. Online soft-error vulnerability estimation offers the flexibility of exploiting dynamic fault-tolerant mechanisms for cost-effective reliability enhancement. We propose a generic run-time method with low area and power overhead to predict the soft-error vulnerability of on-chip memory arrays. The vulnerability prediction is based on signal probabilities (SPs) of a small set of flip-flops, chosen at design time, by studying the correlation between the soft-error vulnerability and the flip-flop SPs for representative workloads. We exploit machine learning to develop a predictive model that can be deployed in the system in software form. Simulation results on two processor designs show that the proposed technique can accurately estimate the soft-error vulnerability of on-chip memory arrays that constitute the instruction cache, the data cache, and the register file. Arunkumar Vijayan, Abhishek Koneru, Mojtaba Ebrahimi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
VTS | 4 |
| 2016 | A programmable method for low-power scan shift in SoC integrated circuitsabstractWe present a programmable method for shift-clock stagger assignment to reduce power supply noise during system-on-chip (SoC) testing. An SoC design is typically composed of several blocks and two neighboring blocks that share the same power rails should not be toggled at the same time during shift. Therefore, the proposed programmable method does not assign the same stagger value to neighboring blocks. The positions of all blocks are first analyzed and the shared boundary length between blocks is then calculated. Based on the position relationships between the blocks, a mathematical model is presented to derive optimal result for small-to-medium sized problems. For larger designs, a heuristic algorithm is proposed and evaluated. We present assignment results as well as power-analysis results and silicon data for industry designs to highlight the effectiveness of the proposed method. Ran Wang 0002, Bonita Bhaskaran, Karthikeyan Natarajan, Ayub Abdollahian, Kaushik Narayanun, Krishnendu Chakrabarty, Amit Sanghani |
VTS | 6 |
| 2016 | Multicast-Based Testing and Thermal-Aware Test Scheduling for 3D ICs with a Stacked Network-on-ChipabstractA 3D stacked network-on-chip (NOC) promises the integration of a large number of cores in a many-core system-on-chip (SOC). The NOC can be used to test the embedded cores in such SOCs, whereby the added cost of dedicated test-access hardware can be avoided. However, a potential problem associated with 3D NOC-based test access is the emergence of hotspots due to stacking and the high toggle rates associated with structural test patterns used for manufacturing test. High temperatures and hotspots can lead to the failure of good parts, resulting in yield loss. We describe a unicast-based multicast approach and a thermal-driven test scheduling method to avoid hotspots, whereby the full NOC bandwidth is used to deliver test packets. Test delivery is carried out using a new unicast-based multicast scheme. Experimental results highlight the effectiveness of the proposed method in reducing test time under thermal constraints. Krishnendu Chakrabarty, Hideo Fujiwara |
IEEE Trans. Computers | 2 |
| 2016 | A Distributed, Reconfigurable, and Reusable BIST Infrastructure for Test and Diagnosis of 3-D-Stacked ICsabstractWe present an end-to-end design of a built-in self-test (BIST) and BIST-based diagnosis infrastructure for 3-D-stacked integrated circuits (ICs) that facilitates the use of BIST at multiple stages of 3-D integration. The proposed BIST design is distributed, reusable, and reconfigurable, hence it is attractive for both prebond and post-bond testing. We also provide support for translating a static BIST schedule into a set of BIST control instructions. The BIST design is validated using detailed simulations of the various operating modes. A framework for fault diagnosis using the BIST infrastructure for 3-D-stacked ICs is also proposed. We present results on synthetic stacks created from ITC'99 and Open-Core benchmark circuits and assess the impact of inserting BIST in these designs in terms of area, timing, and power overhead. Results show that the overhead due to BIST is negligible. We also formulate a test-scheduling problem that aims at minimizing test time under BIST-resource and power constraints, and use two algorithms based on bin packing for solving the problem. Mukesh Agrawal 0001, Krishnendu Chakrabarty, Bill Eklow |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Fault Diagnosis for Leakage and Blockage Defects in Flow-Based Microfluidic BiochipsabstractAdvances in flow-based microfluidics now allow an efficient implementation of biochemistry on-a-chip for DNA sequencing, drug discovery, and point-of-care disease diagnosis. However, the adoption of flow-based biochips is hampered by defects that frequently occur in chips fabricated using soft lithography techniques. Recently published work has shown how we can automate the testing of flow-based biochips; diagnosis methods are now needed to identify the flaws in the fabrication process and to facilitate the use of partially defective chips. Since disposable biochips are being targeted for a highly competitive and low-cost market segment, such diagnosis methods need to be inexpensive, quick, and effective. In this paper, we present the first approach for the automated diagnosis of leakage and blockage defects in flow-based microfluidic biochips. The proposed method targets the identification of fault types and their locations based on test outcomes. It reduces the number of possible fault sites significantly while identifying their exact locations. We use a graph representation of flow paths and a formulation based on hitting sets for the analysis of observed error syndromes. The diagnosis technique is evaluated on three fabricated biochips, and the localization of faults and their classification are achieved correctly in all cases. Kai Hu 0003, Bhargab B. Bhattacharya, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Wash Optimization and Analysis for Cross-Contamination Removal Under Physical Constraints in Flow-Based Microfluidic BiochipsabstractRecent advances in flow-based microfluidics have enabled the emergence of biochemistry-on-a-chip as a new paradigm in drug discovery, point-of-care disease diagnosis, and biomolecular recognition. However, these applications in biology and biochemistry require high precision to avoid erroneous assay outcomes and, therefore, are vulnerable to contamination between two fluidic flows with different biochemistries. Moreover, to wash contaminated sites, the buffer solution in flow-based biochips has to be guided along pre-etched channel networks. In this paper, we propose the first approach for automated wash optimization for contamination removal in flow-based microfluidic biochips. The proposed approach targets the generation of washing pathways to clean all contaminated microchannels with minimum execution time under physical constraints. Two representative and fabricated biochips are used to evaluate the proposed washing method. Compared with a baseline approach, the proposed approach leads to more efficient washing in all cases. Kai Hu 0003, Tsung-Yi Ho, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Efficient Board-Level Functional Fault Diagnosis With Missing SyndromesabstractFunctional fault diagnosis is widely used in board manufacturing to ensure product quality and improve product yield. Advanced machine-learning techniques have recently been advocated for reasoning-based diagnosis; these techniques are based on the historical record of successfully repaired boards. However, traditional diagnosis systems fail to provide appropriate repair suggestions when the diagnostic logs are fragmented and some error outcomes, or syndromes, are not available during diagnosis. We describe the design of a diagnosis system that can handle missing syndromes and can be applied to four widely used machine-learning techniques. Several imputation methods are discussed and compared in terms of their effectiveness for addressing missing syndromes. Moreover, a syndrome-selection technique based on the minimum-redundancy-maximum-relevance criteria is also incorporated to further improve the efficiency of the proposed methods. Two large-scale synthetic data sets generated from the log information of complex industrial boards in volume production are used to validate the proposed diagnosis system in terms of diagnosis accuracy and training time. Shi Jin 0001, Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | A Novel Test Method for Metallic CNTs in CNFET-Based SRAMsabstractStatic random access memories (SRAMs) built on carbon nanotube field effect transistors (CNFETs) are promising alternatives to conventional CMOS-based SRAMs, due to their advantages in terms of power consumption and noise immunity. However, the nonideal carbon nanotube (CNT) fabrication process generates metallic-CNTs (m-CNTs) along with semiconductor-CNTs, leading to correlated faulty cells along the growth direction of the m-CNTs. In this paper, we propose a novel low-cost test solution to detect such faults. Instead of using conventional March test to test each and every SRAM cell, we selectively test certain SRAM cells and judiciously skip testing other SRAM cells between the selected cells. To ensure high fault coverage, we propose three jump test algorithms for different CNFET-SRAM layouts. Moreover, we model m-CNT-induced SRAM faults and characterize their distribution in the SRAM array. Experimental results show that the proposed solutions are able to achieve high fault coverage with low test cost. Tianjian Li, Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty, Naifeng Jing, Li Jiang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | On-Chip Droop-Induced Circuit Delay Prediction Based on Support-Vector MachinesabstractVoltage droop is a major reliability concern in nano-scale very large-scale integration designs. Undesirable voltage droop is often a result of excessive IR drop. On the other hand, Ldi/dt-induced droop occurs when logic gates in the circuit draw high-switching current from the on-chip power supply network, and this problem is exacerbated at high-clock frequencies and smaller technology nodes. A consequence of voltage droop is usually an increase in path delays and the occurrence of intermittent faults during circuit operation. The addition of conservative timing margins, also known as guardbands, is a common practice to tackle the problem of voltage droop. However, such static and pessimistic guardbands, which are calculated at design time based on worst-case conditions, lead to significant performance loss. Dynamic frequency scaling is an alternative approach that enables the dynamic adjustment of clock frequency based on the actual voltage droop seen during runtime. For dynamic voltage-frequency to be effective, accurate and real-time prediction of voltage droop is essential. We propose a support-vector machine (SVM)-based regression method to predict voltage droop due to pattern-dependent IR drop based on inputs to the chip at runtime. Moreover, we reduce the amount of data needed for accurate prediction by using correlation-based feature selection. Several benchmarks from ITC'99 and International Work on Logic and Synthesis'05 highlight the effectiveness of the proposed method in terms of delay-prediction accuracy. Since real-time droop prediction requires hardware implementation of the predictor, we present the hardware design and synthesis results to demonstrate that the hardware overhead for the SVM predictor is negligible for large circuits. Fangming Ye, Farshad Firouzi, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Adaptive Board-Level Functional Fault Diagnosis Using Incremental Decision TreesabstractBoard-level functional fault diagnosis is needed for high-volume production to improve product yield. However, to ensure diagnosis accuracy and effective board repair, a large number of syndromes must be used. Therefore, the diagnosis cost can be prohibitively high due to the increase in diagnosis time and the complexity of test execution and analysis. We propose an adaptive diagnosis method based on incremental decision trees (DTs). Faulty components are classified according to the discriminative ability of the syndromes in DT training. The diagnosis procedure is constructed as a binary tree, with the most discriminative syndrome as the root and final repair suggestions are available as the leaf nodes of the tree. The syndrome to be used in the next step is determined based on the observation of syndromes thus far in the diagnosis procedure. The number of syndromes required for diagnosis can be significantly reduced compared to the total number of syndromes used for system training. Moreover, online learning is facilitated in the proposed diagnosis system using an incremental version of DTs, so as to bridge the knowledge obtained at test-design stage with the knowledge gained during volume production. The diagnosis system can thus adapt to occurrences of new error scenarios on-the-fly. Diagnosis results for three complex boards from industry, currently in volume production, highlight the effectiveness of the proposed approach. Fangming Ye, Zhaobo Zhang, Krishnendu Chakrabarty, Xinli Gu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Security Assessment of Cyberphysical Digital Microfluidic BiochipsabstractA digital microfluidic biochip (DMFB) is an emerging technology that enables miniaturized analysis systems for point-of-care clinical diagnostics, DNA sequencing, and environmental monitoring. A DMFB reduces the rate of sample and reagent consumption, and automates the analysis of assays. In this paper, we provide the first assessment of the security vulnerabilities of DMFBs. We identify result-manipulation attacks on a DMFB that maliciously alter the assay outcomes. Two practical result-manipulation attacks are shown on a DMFB platform performing enzymatic glucose assay on serum. In the first attack, the attacker adjusts the concentration of the glucose sample and thereby modifies the final result. In the second attack, the attacker tampers with the calibration curve of the assay operation. We then identify denial-of-service attacks, where the attacker can disrupt the assay operation by tampering either with the droplet-routing algorithm or with the actuation sequence. We demonstrate these attacks using a digital microfluidic synthesis simulator. The results show that the attacks are easy to implement and hard to detect. Therefore, this work highlights the need for effective protections against malicious modifications in DMFBs. Subidh Ali, Mohamed Ibrahim 0002, Ozgur Sinanoglu, Krishnendu Chakrabarty, Ramesh Karri |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | Optimization of 3D Digital Microfluidic Biochips for the Multiplexed Polymerase Chain ReactionabstractA digital microfluidic biochip (DMFB) is an attractive technology platform for revolutionizing immunoassays, clinical diagnostics, drug discovery, DNA sequencing, and other laboratory procedures in biochemistry. In most of these applications, real-time polymerase chain reaction (PCR) is an indispensable step for amplifying specific DNA segments. To reduce the reaction time to meet the requirement of “real-time” applications, multiplexed PCR is widely utilized. In recent years, three-dimensional (3D) DMFBs that integrate photodetectors (i.e., cyberphysical DMFBs) have been developed, which offer the benefits of smaller size, higher sensitivity, and faster result generations. However, current DMFB design methods target optimization in only two dimensions, thus ignoring the 3D two-layer structure of a DMFB. Furthermore, these techniques ignore practical constraints related to the interference between on-chip device pairs, the performance-critical PCR thermal loop, and the physical size of devices. Moreover, some practical issues in real scenarios are not stressed (e.g., the avoidance of the cross-contamination for multiplexed PCR). In this article, we describe an optimization solution for a 3D DMFB and present a three-stage algorithm to realize a compact 3D PCR chip layout, which includes: (i) PCR thermal-loop optimization, (ii) 3D global placement based on Strong-Push-Weak-Pull (SPWP) model, and (iii) constraint-aware legalization. To avoid cross-contamination between different DNA samples, we also propose a Minimum-Cost-Maximum-Flow-based (MCMF-based) method for reservoir assignment. Simulation results for four laboratory protocols demonstrate that the proposed approach is effective for the design and optimization of a 3D chip for multiplexed real-time PCR. Tsung-Yi Ho, Krishnendu Chakrabarty |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Error-Correcting Sample Preparation with Cyberphysical Digital Microfluidic Lab-on-ChipabstractDigital (droplet-based) microfluidic technology offers an attractive platform for implementing a wide variety of biochemical laboratory protocols, such as point-of-care diagnosis, DNA analysis, target detection, and drug discovery. A digital microfluidic biochip consists of a patterned array of electrodes on which tiny fluid droplets are manipulated by electrical actuation sequences to perform various fluidic operations, for example, dispense, transport, mix, or split. However, because of the inherent uncertainty of fluidic operations, the outcome of biochemical experiments performed on-chip can be erroneous even if the chip is tested a priori and deemed to be defect-free. In this article, we address an important error recoverability problem in the context of sample preparation. We assume a cyberphysical environment, in which the physical errors, when detected online at selected checkpoints with integrated sensors, can be corrected through recovery techniques. However, almost all prior work on error recoverability used checkpointing-based rollback approach, that is, re-execution of certain portions of the protocol starting from the previous checkpoint. Unfortunately, such techniques are expensive both in terms of assay completion time and reagent cost, and can never ensure full error-recovery in deterministic sense. We consider imprecise droplet mix-split operations and present a novel roll-forward approach where the erroneous droplets, thus produced, are used in the error-recovery process, instead of being discarded or remixed. All erroneous droplets participate in the dilution process and they mutually cancel or reduce the concentration-error when the target droplet is reached. We also present a rigorous analysis that reveals the role of volumetric-error on the concentration of a sample to be prepared, and we describe the layout of a lab-on-chip that can execute the proposed cyberphysical dilution algorithm. Our analysis reveals that fluidic errors caused by unbalanced droplet splitting can be classified as being either critical or non-critical , and only those of the former type require correction to achieve error-free sample dilution. Simulation experiments on various sample preparation test cases demonstrate the effectiveness of the proposed method. Sudip Poddar, Sarmishtha Ghoshal, Krishnendu Chakrabarty, Bhargab B. Bhattacharya |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2016 | Editorial First TVLSI Best AE and Reviewer AwardsabstractWe are pleased to announce the first IEEE Transactions on VLSI Systems Circuits and Systems (CAS) Society Best Reviewer and Associate Editor (AE) Awards. With the support of the IEEE CAS Society, we have been able to establish an award that recognizes the contributions of our top AEs and Reviewers from both academia and industry who are CAS members. Each year we would like to award two Best AE and three Best Reviewer Awards to those whose efforts and contributions help us to meet our mission of performing an expeditious selection of very high-quality submissions. Krishnendu Chakrabarty, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Software-based test and diagnosis of SoCs using embedded and wide-I/O DRAMabstractModern CMOS technology enables the integration of billions of transistors on a single chip. Emerging three-dimensional (3D) stacking techniques using through-silicon vias (TSVs) promise even higher integration by combining multiple dies in a single package. In order to keep the test cost low and enhance field reliability, there is a need to re-think conventional test practices, such as test-data compression and online testing, as well as test-application techniques and fault diagnosis. Traditional hardware-based on-chip decompression solutions are limited to compression techniques that do not require large hardware overhead for decompression. However, today's system-on-chip designs (SoCs) offer resources, such as embedded processors and large amounts of fast embedded memories, that can be exploited for efficient on-chip test application, online testing, and diagnosis using software-based compression. Examples of such systems are 3D ICs with wide-I/O DRAM or traditional ICs with embedded DRAM (eDRAM). We propose a test and diagnosis solution that makes use of software-based decompression of deterministic scan-test pattern and allows for test application from on-chip DRAM to the logic die, extending traditional hardware-based methods and allowing for online scan-based test and diagnosis. This solution therefore targets SoCs that contain, in addition to a microprocessor, multiple digital-logic cores and glue logic, all of which need to be tested using scan test patterns. Simulation results for benchmarks show that we can achieve high test-data compression, comparable with what is obtained using commercial tools, as well as high-resolution on-chip diagnosis with negligible hardware and test-time overhead. Sergej Deutsch, Krishnendu Chakrabarty |
ASP-DAC | 2 |
| 2015 | Design and optimization of 3D digital microfluidic biochips for the polymerase chain reactionabstractA digital microfluidic biochip (DMFB) is an attractive technology platform for revolutionizing immunoassays, clinical diagnostics, drug discovery, DNA sequencing, and other laboratory procedures in biochemistry. In most of these applications, real-time polymerase chain reaction (PCR) is an indispensable step for amplifying specific DNA segments. In recent years, three-dimensional (3D) DMFBs that integrate photodetectors (i.e., cyberphysical DMFBs) have been developed. They offer the benefits of smaller size, higher sensitivity and quicker time-to-results. However, current DMFB design methods target optimization in only two dimensions, hence they ignore the 3D two-layer structure of a DMFB. Moreover, these techniques ignore practical constraints related to the interference between on-chip device pairs, the performance-critical PCR thermal loop, and the physical size of devices. In this paper, we describe an optimization solution for a 3D DMFB, and present a three-stage algorithm to realize a compact 3D PCR chip layout, which includes: (i) PCR thermal-loop optimization; (ii) 3D global placement based on Strong-Push-Weak-Pull (SPWP) model; (iii) constraint-aware legalization. Simulation results for four laboratory protocols demonstrate that the proposed approach is effective for the design and optimization of a 3D chip for real-time PCR. Tsung-Yi Ho, Krishnendu Chakrabarty |
ASP-DAC | 3 |
| 2015 | Self-learning and adaptive board-level functional fault diagnosisabstractFunctional fault diagnosis is necessary for board-level product qualification. However, ambiguous diagnosis results can lead to long debug times and wrong repair actions, which significantly increase repair cost and adversely impact yield. A state-of-the-art functional fault diagnosis system involves several key components: (1) design of functional test programs, (2) collection of functional-failure syndromes, (3) building of the diagnosis engine, (4) isolation of root causes, and (5) evaluation of the diagnosis engine. Advances in each of these components can pave the way for a more effective diagnosis system, thus improving diagnosis accuracy and reducing diagnosis time. Machine-learning and data analysis techniques offer an unprecedented opportunity to develop an automated and adaptive diagnosis system to increase diagnosis accuracy and reduce diagnosis time. This paper describes how all the above components of an advanced diagnosis system can benefit from machine learning and information theory. Topics discussed include incremental learning, decision trees, root-cause analysis and evaluation metrics, data acquisition, and knowledge transfer. Fangming Ye, Krishnendu Chakrabarty, Zhaobo Zhang, Xinli Gu |
ASP-DAC | 2 |
| 2015 | Jump test for metallic CNTs in CNFET-based SRAMabstractSRAMs built on Carbon Nanotube Field Transistors (CNFET) are promising alternatives to conventional CMOS-based SRAMs, due to their advantages in terms of both power consumption and noise margin. However, non-ideal Carbon Nanotube (CNT) fabrication process generates metallic-CNTs (m-CNTs) along with semiconductor-CNTs (s-CNTs), rendering correlated faulty cells along the growth direction of the m-CNTs. Based on this phenomenon, we propose a novel testing algorithm for detecting m-CNTs, wherein consecutive write and read operations jump over multiple cells rather than marching through each and every cell, thereby significantly reducing the testing cost. The proposed jump test can be invoked before the march test to screen out those CNFET-SRAMs doomed to failure, and this can reduce the subsequent test overhead. Experimental results show that the proposed solution is able to achieve a high fault coverage with much less testing cost. Xiaoyao Liang, Qiang Xu 0001, Krishnendu Chakrabarty, Naifeng Jing, Li Jiang 0002 |
DAC | 4 |
| 2015 | Error recovery in digital microfluidics for personalized medicine
Mohamed Ibrahim 0002, Krishnendu Chakrabarty |
DATE | 2 |
| 2015 | An online thermal-constrained task scheduler for 3D multi-core processors
Chien-Hui Liao, Charles H.-P. Wen, Krishnendu Chakrabarty |
DATE | 3 |
| 2015 | Microfluidic very large-scale integration for biochips: Technology, testing and fault-tolerant designabstractMicrofluidic biochips are replacing the conventional biochemical analyzers by integrating all the necessary functions for biochemical analysis using microfluidics. Biochips are used in many application areas, such as, in vitro diagnostics, drug discovery, biotech and ecology. The focus of this paper is on continuous-flow biochips, where the basic building block is a microvalve. By combining these microvalves, more complex units such as mixers, switches, multiplexers can be built, hence the name of the technology, “microfluidic Very Large-Scale Integration” (mVLSI). A roadblock in the deployment of microfluidic biochips is their low reliability and lack of test techniques to screen defective devices before they are used for biochemical analysis. Defective chips lead to repetition of experiments, which is undesirable due to high reagent cost and limited availability of samples. This paper presents the state-of-the-art in the mVLSI platforms and emerging research challenges in the area of continuous-flow microfluidics, focusing on testing techniques and fault-tolerant design. Ismail Emre Araci, Paul Pop, Krishnendu Chakrabarty |
ETS | 3 |
| 2015 | Testing of digital microfluidic biochips with arbitrary layoutsabstractAs in the case of VLSI circuits, digital microfluidic biochips must be adequately tested after manufacturing to guarantee the correctness of the biomedical experiments. In this work, we propose an efficient test method for digital microfluidic biochips. In contrast to related prior work, the proposed test method is not only able to cover all chip defects but also applicable to arbitrary chip layouts. Experiments demonstrate that using the proposed test method, the test-application time can be reduced significantly compared to related prior work. Trung Anh Dinh, Shigeru Yamashita, Tsung-Yi Ho, Krishnendu Chakrabarty |
ETS | 4 |
| 2015 | Re-using BIST for circuit aging monitoringabstractBias Temperature Instability (BTI)-induced transistor aging degrades path delay over time and may eventually induce circuit failure due to timing violations. Chip health monitoring is therefore necessary to track delay changes on a per-chip basis. We propose a method to accurately predict the fine-grained circuit-delay degradation with minimal area and performance overhead. It re-uses on-chip design-for-test (DfT) infrastructure to track the severity of run-time stress by periodiclly capturing system state and compacting it using a multiple input signature register (MISR). The captured stress information is fed to a software-based prediction model in realtime. The prediction model is trained offline using support vector regression. Aging prediction based on run-time stress monitoring can be used to proactively activate aging mitigation techniques. Experimental results for benchmark circuits highlight the accuracy of the proposed approach. Farshad Firouzi, Fangming Ye, Arunkumar Vijayan, Abhishek Koneru, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ETS | 5 |
| 2015 | A branch-&-bound algorithm for TAM optimization in multi-Vdd SoCsabstractIn this paper, we present the first TAM optimization technique for multi-VddSoCs. The proposed method exploits unique scheduling opportunities and flexibility offered by TDM, and by the means of a very efficient Branch-&-Bound approach it quickly identifies the most effective TAM configurations. Experiments upon an industrial SoC highlight the benefits of the proposed technique on multi-Vdddesigns. Fotis Vartziotis, Xrysovalantis Kavousianos, Krishnendu Chakrabarty |
ETS | 3 |
| 2015 | Digital Microfluidic Biochips: Towards Functional Diversity, More than Moore, and Cyberphysical IntegrationabstractAdvances in droplet-based "digital" microfluidics have led to the emergence of biochip devices for automating laboratory procedures in biochemistry and molecular biology. These devices enable the precise control of nanoliter-volume droplets of biochemical samples and reagents. Therefore, integrated circuit (IC) technology can be used to transport and transport "chemical payload" in the form of micro/nanofluidic droplets. As a result, non-traditional biomedical applications and markets (e.g., high-throughout DNA sequencing, portable and point-of-care clinical diagnostics, protein crystallization for drug discovery), and fundamentally new uses are opening up for ICs and systems. Krishnendu Chakrabarty |
ACM Great Lakes Symposium on VLSI | 1 |
| 2015 | Optimizing 3D NoC Design for Energy Efficiency: A Machine Learning ApproachabstractThree-dimensional (3D) Network-on-Chip (NoC) is an emerging technology that has the potential to achieve high performance with low power consumption for multicore chips. However, to fully realize their potential, we need to consider novel 3D NoC architectures. In this paper, inspired by the inherent advantages of small-world (SW) 2D NoCs, we explore the design space of SW network-based 3D NoC architectures. We leverage machine learning to intelligently explore the design space to optimize the placement of both planar and vertical communication links for energy efficiency. We demonstrate that the optimized 3D SW NoC designs perform significantly better than their 3D MESH counterparts. On an average, the 3D SW NoC shows 35% energy-delay-product (EDP) improvement over 3D MESH for the nine PARSEC and SPLASH2 benchmarks considered in this work. The highest performance improvement of 43% was achieved for RADIX. Interestingly, even after reducing the number of vertical links by 50%, the optimized 3D SW NoC performs 25% better than the fully connected 3D MESH, which is a strong indication of the effectiveness of our optimization methodology. Sourav Das 0002, Janardhan Rao Doppa, Dae Hyun Kim 0004, Partha Pratim Pande, Krishnendu Chakrabarty |
ICCAD | 5 |
| 2015 | A General and Exact Routing Methodology for Digital Microfluidic BiochipsabstractAdvances in microfluidic technologies have led to the emergence of Digital Microfluidic Biochips (DMFBs), which are capable of automating laboratory procedures in biochemistry and molecular biology. During the design and use of these devices, droplet routing represents a particularly critical challenge. Here, various design tasks have to be addressed for which, depending on the corresponding scenario, different solutions are available. However, all these developments eventually result in an “inflation” of different design approaches for routing of DMFBs - many of them addressing a very dedicated routing task only. In this work, we propose a comprehensive routing methodology which (1) provides one (generic) solution capable of addressing a variety of different design tasks, (2) employs a “push-button”-scheme that requires no (manual) composition of partial results, and (3) guarantees minimality e.g., with respect to the number of timesteps or the number of required control pins. Experimental evaluations demonstrate the benefits of the solution, i.e., the applicability for a wide range of design tasks as well as improvements compared to specialized solutions presented in the past. Oliver Keszöcze, Robert Wille, Krishnendu Chakrabarty, Rolf Drechsler |
ICCAD | 3 |
| 2015 | Fine-Grained Aging Prediction Based on the Monitoring of Run-Time Stress Using DfT InfrastructureabstractRun-time solutions based on real-time monitoring and adaptation are required for resilience in nanoscale integrated circuits as design-time solutions and guard bands are no longer sufficient. Bias Temperature Instability (BTI)-induced transistor aging, one of the major reliability threats in nanoscale VLSI, degrades path delay over time and may eventually induce circuit failure due to timing violations. Chip health monitoring is, therefore, necessary to track delay changes on a per-chip basis. Chip-monitoring techniques based on actual measurement of path delays can only track a coarse-grained aging trend in a reactive manner. In this paper, we show how the on-chip design for test (DfT) infrastructure can be reused in order to perform fine-grain workload-induced stress monitoring for accurate aging prediction. The captured stress information is fed to a prediction model in real-time. The prediction model is trained offline using support-vector regression and implemented in software. This approach can leverage proactive adaptation techniques to mitigate further aging of the circuit by monitoring aging trends. Simulation results for realistic open-source benchmark circuits highlight the accuracy of the proposed approach. Abhishek Koneru, Arunkumar Vijayan, Krishnendu Chakrabarty, Mehdi Baradaran Tahoori |
ICCAD | 3 |
| 2015 | Defect Clustering-Aware Spare-TSV Allocation for 3D ICsabstractThe manufacturing yield challenge of three-dimensional integrated circuit (3D ICs) is one of the key obstacles in the industry adoption of 3D integration based on through-silicon-vias (TSVs). The addition of spare TSVs to repair faulty functional TSVs is an effective method for yield and reliability enhancement, but this approach results in significant hardware cost and delay overhead. Most existing solutions are only suitable for a “dual-uniform” scenario in which both the placement and the defect probabilities of functional TSVs are assumed to be uniform. In this paper, we propose a design technique that is compatible with non-uniform TSV placement and it can repair faulty TSVs based on a realistic clustered defect-distribution model. The proposed solution is based on two consecutive stages, which utilize a greedy algorithm and an integer-linear-programming formulation, respectively. By considering the trade-off between chip yield, hardware cost, and delay overhead, the proposed technique provides higher yield and reliability under a clustered defect distribution, and with minimum hardware cost and delay overhead, compared to the previous work. Shengcheng Wang, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty |
ICCAD | 3 |
| 2015 | Security implications of cyberphysical digital microfluidic biochipsabstractA digital microfluidic biochip (DMFB) is an emerging technology that enables miniaturized analysis systems for point-of-care clinical diagnostics, DNA sequencing, and environmental monitoring. A DMFB reduces the rate of sample and reagent consumption, and automates the analysis of assays. In this paper, we highlight the security vulnerabilities of DMFBs by identifying two potential attacks on a DMFB that performs enzymatic glucose assay on serum. In the first attack, the attacker adjusts the concentration of the glucose sample and thereby modifies the final result. In the second attack, the calibration curve of the assay operation is maliciously modified in order to make it deviate from the nominal/golden calibration curve. We demonstrate these attacks using a digital microluidics synthesis simulator. The results show that the attacks are stealthy as they do not result in any noticeable change in the DMFB synthesis. Subidh Ali, Mohamed Ibrahim 0002, Ozgur Sinanoglu, Krishnendu Chakrabarty, Ramesh Karri |
ICCD | 4 |
| 2015 | Cyber-physical integration in programmable microfluidic biochipsabstractMicrofluidic biochip technology integrates miniaturized components into a chip that can perform traditional biochemical laboratory procedures. Commercial impact is highlighted by the recent acquisition of Advanced Liquid Logic by Illumina Inc., a leader in DNA sequencing and biomolecular analysis. Due to the inherent variability involved in many biochemical processes, uncertainties manifest themselves in many ways in microluidics. Cyber-physical integration of on-chip sensors permits feedback-driven monitoring in-real time to detect and correct errors, along with other benefits such as adaptive control and dynamic re-synthesis. This paper overviews low-based and digital (droplet-based) microluidic biopchips, and discusses the state-of-the-art in microluidic device fabrication, the interplay between sensor feedback and adaptive control software, and practical experiences relating to biochip cyber-physical integration. It demonstrates the connections between the many fundamental principles of chip design and engineering, and the needs of the biochip community. Tsung-Yi Ho, William H. Grover, Shiyan Hu 0001, Krishnendu Chakrabarty |
ICCD | 4 |
| 2015 | Self-awareness and self-learning for resiliency in real-time systemsabstractWhile the notion of self-awareness has a long history in biology, psychology, medicine, engineering and (more recently) computing, we are seeing the emerging need for self-awareness in the context of complex Systems-on-Chip that must address the often conflicting requirements of performance, resiliency, energy, cost, etc. in the face of highly dynamic operational behaviors coupled with process, environment, and workload variabilities. Unlike traditional Systems-on-Chip (SoCs), self-aware SoCs must deploy an intelligent co-design of the control, communication, and computing infrastructure that interacts with the physical environment in real-time in order to modify the systems behavior so as to adaptively achieve desired objectives and Quality-of-Service (QoS). Self-aware SoCs require a combination of ubiquitous sensing and actuation, health-monitoring, and self-learning to enable the SoCs adaptation over time and space. This special session targets self-learning and self-awareness in two domains. The first one is a self-learning runtime reliability prediction approach by reusing Design-for-Test (DfT) infrastructure. The other one discusses real-time systems and applications to wireless communication, signal processing and control. Mehdi Baradaran Tahoori, Abhijit Chatterjee, Krishnendu Chakrabarty, Abhishek Koneru, Arunkumar Vijayan, Debashis Banerjee |
IOLTS | 3 |
| 2015 | Contactless pre-bond TSV fault diagnosis using duty-cycle detectors and ring oscillatorsabstractDefects in TSVs due to fabrication steps decrease the yield and reliability of 3D stacked ICs, hence these defects need to be screened early in the manufacturing flow. We propose a non-invasive method for pre-bond TSV test and diagnosis that does not require TSV probing. We use open TSVs as capacitive loads of their driving gates and measure the propagation delay by means of ring oscillators. Defects in TSVs cause variations in their RC parameters and therefore lead to variations in the propagation delay. By measuring these variations, we can detect resistive open and leakage faults. In addition, we use duty-cycle detectors to measure the duty cycle of the oscillation signal. These measurements provide additional information for fault analysis and hence increase the diagnosis accuracy. We exploit different voltage levels to increase the sensitivity of the test and its robustness against random process variations. We also present a method to create a regression model based on artificial neural networks to predict the fault size. As input, this model uses both the oscillation period and the duty cycle measured at multiple different voltage levels. The model classifies the type of the fault and predicts its size. Moreover, the regression model can effectively determine whether a TSV has both leakage and resistive-open defects. Results on fault-diagnosis effectiveness are presented through HSPICE simulations using realistic models for a 45nm CMOS technology. The estimated DfT area cost of our method is negligible for dies of realistic size. Sergej Deutsch, Krishnendu Chakrabarty |
ITC | 2 |
| 2015 | Test and debug solutions for 3D-stacked integrated circuitsabstractThree-dimensional (3D) stacking using through-silicon vias (TSVs) promises higher integration levels in a single package, keeping pace with Moore's law. Testing has been identified as a showstopper for volume manufacturing of 3D-stacked integrated circuits (3D ICs). This work provides solutions to new challenges related to 3D test content, test access, diagnosis and debug. We analyze the the impact of thermo-mechanical stress due to TSV fabrication process on test quality. We propose a test-generation flow that takes TSV-induced stress into account by using stress-aware circuit models. Pre-bond TSV test is a challenge due to limited accessibility of TSV at the pre-bond stage. We develop a non-invasive method for TSV test and diagnosis using ring oscillators, duty-cycle detectors, and a regression model based on artificial neural networks. In order to efficiently deliver test content, 3D design-for-test (DfT) architectures are needed. We propose an optimization approach that takes uncertainties in input parameters into account and provides a solution that is efficient in the presence of input-parameter variations and minimizes test time. Finally, post-silicon debug is a major challenge due to continuously increasing design complexity. We develop a low-cost debug architecture for massive signal tracing in 3D-stacked ICs with wide-I/O DRAM dies that significantly increases the observation window compared to traditional methods that use trace buffers. Sergej Deutsch, Krishnendu Chakrabarty |
ITC | 2 |
| 2015 | A general testing method for digital microfluidic biochips under physical constraintsabstractDigital microfluidics is viewed as one of the most promising technologies for biomedical experiments. Digital microfluidic biochips are often used today for applications such as point-of-care health assessment, drug discovery, and air-quality monitoring. Therefore, such devices must be adequately tested after manufacturing to guarantee the correctness of the biomedical experiments. Previous test methods for digital microfluidic biochips are either unable to cover all chip defects or inapplicable for application-specific biochips with arbitrary layouts. Furthermore, previous methods also ignore the fluidic constraints required for droplet routing, which makes the test droplet routing problem much more challenging in realistic test-application scenarios. In this paper, we propose the first test method for digital microfluidic biochips that is not only able to cover all chip defects, but is also applicable for arbitrary chip layouts. Moreover, we propose an optimization technique to route test droplets with minimum test-application time. A polynomial-time scheduling algorithm is also presented to solve the optimization problem in an efficient manner. Experiments demonstrate that the proposed test method requires significantly less test-application time compared to related previous work. Trung Anh Dinh, Shigeru Yamashita, Tsung-Yi Ho, Krishnendu Chakrabarty |
ITC | 4 |