EDBT 2026 Demo / reviewers in the wild / expert
Alex Orailoglu
dblp:o/AlexOrailoglu
· DBLP profile ↗
214ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0002-6104-3923ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 202 · 9 first-author · 15 since 2021Software engineering, systems software and programming languages · 29 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 3Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PrometheusFree: Concurrent Detection of Laser Fault Injection Attacks in Optical Neural Networks
Kota Nishida, Yoshihiro Midoh, Noriyuki Miura, Satoshi Kawakami, Alex Orailoglu, Jun Shiomi |
ASP-DAC | 5 |
| 2026 | Corrections to "AdaTrust: Combinational Hardware Trojan Detection Through Adaptive Test Pattern Construction"abstractIn the above article [1], the errors were addressed. Due to a production error, several reference citations were mistakenly typeset as “[?].” The correct references are as follows. 1)To improve Trojan isolation, the approach proposed in [11] reorders the circuit’s scan chains to create physically segregated regions that can be more easily activated individually.2) The most straightforward of the metrics is Trojan-to-circuit activity ( $ TCA $ ), defined in [11] as $TCA(TP_{i}) =\frac{TA_{i}}{CA_{i}}$ , where $TP_{i}$ denotes the test pattern under evaluation, $TA_{i}$ denotes the number of active Trojan gates for $TP_{i}$ , and $CA_{i}$ denotes the number of active non-Trojan gates for $TP_{i}$ .3)We performed experiments on Trust-Hub benchmarks with combinational Trojans inserted within sequential designs [20], [21] to confirm the proposed test pattern generation as a viable approach.4)To evaluate the effectiveness of this adaptive flow, the proposed methodology is executed on five Trojan-inserted ISCAS benchmarks made available on the Trust-Hub website [20], [21].5)2) Considering Process Variation: To introduce process variation, we model interdie and intradie variation as normal distributions with a mean of 1 and respective magnitudes of $15\% = 3\sigma_{inter}$ and $ 5\% = 3\sigma_{intra}$ (similar to values used in [11] and [12]).6)An alternate design-based approach [11] changes the ordering of scan chains to improve the ability to isolate a given region of the design. Chris Nigh, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | White-box logic obfuscation: A Transparent Solution to Hardware Piracy and Reverse EngineeringabstractReverse engineering the functional specification from a netlist is a challenging task that enables IP piracy and tampering. Traditional logic locking techniques, which depend on external activation with secrets stored in tamper-proof locations to thwart reverse engineering, have repeatedly been compromised by key recovery attacks. Their vulnerability highlights the flawed assumption of relying on tamper-proof secrets for building hardware security solutions, especially when activated devices are deployed in open environments where they are exposed to attackers for functional queries and probing. This paper presents white-box logic obfuscation (WBLO) as a novel solution to safeguard control logic from functional reverse engineering, even when attackers have full visibility of the operational netlist. WBLO eliminates the need for post-manufacturing activation and reliance on tamper-proof key storage by securing the design with keys that are autonomously updated internally through the legitimate sequential execution of the device. The proposed approach invalidates functional analysis under arbitrary probing and combinational queries. We examine the implementation challenges inherent in the WBLO process and identify critical design considerations that enhance security and efficiency. Building on these insights, we suggest a prioritization for various design transformation and synthesis rules that achieve robust security in the white-box attack model while minimizing implementation overheads. Leon Li, Alex Orailoglu |
ASP-DAC | 2 |
| 2025 | Breaking the One-Shot Barrier: Progressive Error Detection in FSM via Key-Driven CorruptionabstractFinite State Machines (FSMs) are central to control logic in digital systems as errors can result in critical system failures. Conventional error detection embeds Error Detection Codes (EDCs) into state encodings, but surpassing the 99.5% detection probability for arbitrary-magnitude errors, as required in mission-critical systems, demands significant redundancy and consequently results in prohibitive overhead in resource-constrained environments. In this work, we introduce a new FSM protection methodology that breaks the one-shot detection barrier of traditional EDCs. By leveraging white-box logic obfuscation (WBLO) to synthesize a randomized corruption mode and integrating loworder nonlinear parity constraints into the expanded state representation, we convert transient faults into persistent anomalies detectable across subsequent cycles. This progressive detection model substantially propels the Pareto frontier of implementation overhead versus detection probability, while maintaining sub-cycle average detection latency. Leon Li, Alex Orailoglu |
ATS | 2 |
| 2024 | Transcoders: A Better Alternative To Denoising AutoencodersabstractImage denoising is a popular technique that is used to remove noise incurred due to hardware faults or noise carefully crafted by an attacker. Autoencoders are some of the most popular denoisers. Their ability to learn a distribution’s latent space helps them achieve this property, and they are generally good at it. However, they are known to fumble in a white-box threat model where an attacker knows everything about the victim classifier and its denoiser network - including its architecture and hyperparameters. We show that this problem stems from the autoencoder’s learning goal. In this paper, we augment an autoencoder’s learning goal to conceive what we call transcoders. This modification forces the transcoder network to learn a function that is more adept at denoising a given image. Our results, evaluated on two datasets - MNIST and CIFAR10, a slew of attacks, and two threat models - gray-box and white-box, help us argue the following: given a denoising tool built using an autoencoder, one can update the learning goal of the autoencoder to that of a transcoder, and achieve a transcoder-based denoiser that is significantly better at handling both fault-induced and attack-induced noise. Pushpak Raj Gautam, Alex Orailoglu |
ETS | 2 |
| 2024 | Locked-by-Design: Enhancing White-box Logic Obfuscation with Effective Key MutationabstractThis paper proposes an obfuscation technique that thwarts functional reverse engineering despite the operational netlist being completely visible to the attacker. Central to this approach is the recognition that reverse engineering necessitates not only the recovery of an operational netlist but also the extraction of functional understandings from it to drive specific illegitimate applications. The proposed approach applies self-generated and mutating keys to obfuscate Finite State Machines (FSMs), forcing attackers to perform complex sequential analysis to learn even the simplest aspects of a design’s functionality. This work examines the impact of key mutation operations on the effectiveness and efficiency of obfuscation. It suggests an enhanced key mutation scheme capable of significantly reducing the implementation overheads without compromising attack resilience. The experimental results show that the proposed obfuscation algorithm leads to drastic overhead improvements and demonstrates strong resilience to sequential SAT attacks and functional reverse engineering. Leon Li, Alex Orailoglu |
ITC | 2 |
| 2023 | ClearLock: Deterring Hardware Reverse Engineering Attacks in a White-BoxabstractLogic obfuscation is a popular method for safeguarding semiconductor intellectual properties from reverse engineering threats. As key recovery attacks continue to advance, the once widely accepted notion of key secrecy has become increasingly untenable. This research proposes a novel method to thwart effective reverse engineering methods even when the attacker is armed with complete control of a fully-functional netlist, i.e., in a white-box. The proposed obfuscation technique derives mutating secrets through setting up an inherently hard problem for sequential designs to lock the netlist at design time and perform self-activation at runtime. The mutating secrets render any reverse engineering shortcuts ineffective, condemning the reverse engineering attacker to full sequential analysis at an intimidating complexity. The obfuscation procedure incorporates four design transformation techniques to ensure secure activation while minimizing overhead. The practicality and security of the proposed white-box obfuscation solution are validated through experiments on MCNC benchmarks. Leon Li, Alex Orailoglu |
ATS | 2 |
| 2023 | Thwarting Reverse Engineering Attacks through Keyless Logic ObfuscationabstractLogic obfuscation protects semiconductor IPs against reverse engineering threats by concealing IP implementation details using a tamper-proof key. With the continuous evolution of key recovery attacks exploiting functional, structural, and physical key exposures, the typical assumption of key secrecy becomes increasingly untenable. This work aims to end the tug of war between key-based defenses and key recovery attacks by delivering reverse engineering resilience through a novel keyless obfuscation approach that demands no external secret. The proposed solution locks the full functionality of a design using internally-generated and constantly-changing secrets that can be only extracted from the FSM transition history. The intrinsic secrets are secure against reverse engineering attempts due to the hardness of identifying valid transition paths to a target state from the obfuscated gate-level netlist. We develop an algorithm to synthesize the obfuscated FSM logic and the dynamic key update logic which jointly activate the design for all valid sequential queries so as to deliver unimpeded functionality for legal users. The algorithm enforces key consistency through equivalence-preserving FSM transformation and constraint-based state encoding to handle complex reconverging transition paths. Experimental results on MCNC benchmarks confirm the practicality and security of the proposed keyless logic obfuscation methodology. Leon Li, Alex Orailoglu |
VTS | 2 |
| 2023 | Redundancy Attack: Breaking Logic Locking Through Oracleless Rationality AnalysisabstractDuring the last decade, logic locking has been proposed to protect integrated circuits against piracy and reverse engineering threats. Functional pruning attacks, such as the Boolean satisfiability (SAT) attack, have greatly challenged the security of logic locking techniques, yet they require access to an oracle and suffer from muted efficacy on inherently SAT-hard circuits. In this article, we present a novel oracleless attack on both XOR and MUX-based logic locking by analyzing the redundancy level deviation under different keys. Our fundamental insight is that incorrect keys would produce circuits that violate universally followed design principles, prominent among which stands the minimization of logic redundancy. We leverage the redundancy-level deviation produced by individual and pairs of key bits to recover key bit values and establish pairwise equivalence, leading to an efficient nearly linear time attack algorithm. We experimentally verify that the proposed attack quickly unveils more than half of the key bits with high accuracy on 21 ISCAS’85 and MCNC circuits locked with a variety of locking techniques. To fortify logic locking in the face of this successful redundancy attack, we supplement this article with a key gate insertion methodology that smooths the deviation in redundancy level. We achieve this goal by selecting key gates that exhibit pairwise dissociation among a set of individually secure locations, which effectively breaks down correlations in arbitrary sets of key gates and, thus, imposes a consistent redundancy level throughout the key space. Experimental results confirm the strong redundancy attack resilience of the proposed defense strategy. Leon Li, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Unleashing the Potential of Sparse DNNs Through Synergistic Hardware-Sparsity Co-DesignabstractSparsity, a widely recognized path to curbing the computational needs of deep neural networks (DNNs), still suffers a number of roadblocks in practice, despite a decade of intensive research on sparse neural networks. The extant structured sparsity patterns often fail to attain significant model compression, while the hardware challenges posed by unstructured sparsity are yet to be fully overcome. As algorithmic and hardware innovations individually deliver limited benefits, a synergistic approach is necessary to unleash the potential of sparse DNNs. This work proposes a tightly integrated design methodology for the sparsity patterns and associated hardware platforms to reach the highest model compression goals while simultaneously facilitating efficient hardware processing. We demonstrate that novel complementary sparsity patterns can offer utmost expressiveness levels with inherent hardware exploitable regularity. Our novel dynamic training method converts the expressiveness of such sparsity configurations into highly accurate and compact sparse neural networks. Complementary sparsity is represented in a dense format, and when synergistically coupled with minimal yet strategic hardware modifications, can be processed in close concordance with the conventional dataflow of the dense matrix operations. We thus demonstrate that there is ample room for innovation beyond conventional techniques and immense practical potential for sparse neural networks through the synergistic design of sparsity patterns and hardware architectures. Elbruz Ozen, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | JANUS-HD: Exploiting FSM Sequentiality and Synthesis Flexibility in Logic Obfuscation to Thwart SAT Attack While Offering Strong CorruptionabstractLogic obfuscation has been proposed as a counter-measure towards chip counterfeiting and IP piracy by obfuscating circuit designs with a key-controlled locking mechanism. However, the extensive output corruption of early key gate based logic obfuscation techniques has exposed them to effective SAT attacks. While current SAT resilient logic obfuscation techniques succeed in undermining the attack by offering near-trivial output corruption, they do so at the expense of a drastic reduction in functional and structural protection scope. In this work, we present JANUS-HD based on novel insights that succeed to deliver the heretofore elusive goal of simultaneously boosting corruptibility and foiling SAT attacks. JANUS-HD obfuscates an FSM through diverse FF configurations for different transitions with the overall configuration setting as the obfuscation secret. A key-controlled Hamming distance comparator controls the obfuscation status at the minimized number of entrance states identified through a custom graph partitioning algorithm. Reliance on the inherent state transition patterns extends the obfuscation benefits to non-entrance states without exposing any additional key space pruning trace. We leverage the flexibility of state encoding and equivalence-based FSM transformations to generate an obfuscated netlist at low overhead using standard synthesis tools. Finally, we present a scan chain crippling mechanism that delivers unfettered scan chain access while eradicating any key trace leakage in the scan mode, thus thwarting chosen-input attacks aimed at the Hamming distance comparator. We illustrate through experiments that JANUS-HD delivers obfuscation scope improvements of up to 45.5x over the state-of-the-art, establishing the first cost-effective solution to offer a broad yet attack-resilient obfuscation scope against supply chain threats. Leon Li, Alex Orailoglu |
DATE | 2 |
| 2022 | Architecting Decentralization and Customizability in DNN Accelerators for Hardware Defect AdaptationabstractThe efficiency of machine intelligence techniques has improved noticeably in the embedded application domains thanks to the dedicated hardware accelerators for deep neural networks (DNNs). Despite the economic criticality of yield and reliability problems in advanced semiconductor nodes, these concerns have attracted limited attention in the context of embedded machine intelligence devices. The micro-architectural features of deep learning accelerators, when paired with the algorithmic characteristics of DNNs, unlock novel opportunities to tackle semiconductor reliability problems in embedded deep learning devices. While the fine-grained bypassing of the faulty processing elements reins the computational impact of hardware defects, a one-time training of DNNs with Hardware-Aware Dropout/Dropconnect techniques boosts model decentralization and facilitates accurate neural network inference in the degraded computational fabrics. Furthermore, on-device calibration methods can improve resilience even further without necessitating expensive defect compensation methods such as device-specific training. Our work confirms the potential for improving the yield, reliability, and operational lifetime of embedded machine intelligence devices through a highly practical co-design of DNNs and configurable hardware architectures. Elbruz Ozen, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Evolving Complementary Sparsity Patterns for Hardware-Friendly Inference of Sparse DNNsabstractSparse deep learning models are known to be more accurate than their dense counterparts for equal parameter and computational budgets. Unstructured model pruning can deliver dramatic compression rates, yet the consequent irregular sparsity patterns lead to severe computational challenges for modern computational hardware. Our work introduces a set of complementary sparsity patterns to construct both highly expressive and inherently regular sparse neural network layers. We propose a novel training approach to evolve inherently regular sparsity configurations and transform the expressive power of the proposed layers into a competitive classification accuracy even under extreme sparsity constraints. The structure of the introduced sparsity patterns engenders optimal compression of the layer parameters into a dense representation. Moreover, the constructed layers can be processed in the compressed format with full-hardware utilization in minimally modified non-sparse computational hardware. The experimental results demonstrate superior compression rates and remarkable performance improvements in sparse neural network inference in systolic arrays. Elbruz Ozen, Alex Orailoglu |
ICCAD | 2 |
| 2021 | SNR: Squeezing Numerical Range Defuses Bit Error Vulnerability Surface in Deep Neural NetworksabstractAs deep learning algorithms are widely adopted, an increasing number of them are positioned in embedded application domains with strict reliability constraints. The expenditure of significant resources to satisfy performance requirements in deep neural network accelerators has thinned out the margins for delivering safety in embedded deep learning applications, thus precluding the adoption of conventional fault tolerance methods. The potential of exploiting the inherent resilience characteristics of deep neural networks remains though unexplored, offering a promising low-cost path towards safety in embedded deep learning applications. This work demonstrates the possibility of such exploitation by juxtaposing the reduction of the vulnerability surface through the proper design of the quantization schemes with shaping the parameter distributions at each layer through the guidance offered by appropriate training methods, thus delivering deep neural networks of high resilience merely through algorithmic modifications. Unequaled error resilience characteristics can be thus injected into safety-critical deep learning applications to tolerate bit error rates of up to at absolutely zero hardware, energy, and performance costs while improving the error-free model accuracy even further. Elbruz Ozen, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | AdaTrust: Combinational Hardware Trojan Detection Through Adaptive Test Pattern ConstructionabstractAs society becomes increasingly reliant on products and systems that make use of integrated circuits, the defense against potential hardware Trojan attacks by an untrusted foundry becomes an important part of any certification flow for critical components. The slew of recent proposals notwithstanding, a satisfactory solution is still wanting as the solutions offered heretofore either require impractical design/test pattern cost or deliver insufficient detection capabilities, primarily challenged by the noise induced by process variation. The methodology put forth by this proposal aims to remedy this, leveraging an adaptive approach that applies superposition to perform a fine-grained circuit analysis and expose any extant Trojan circuitry. Iterative test pattern modifications, circuit response analysis, and adaptive decision-making are deployed, all embedded within the design-for-test and test pattern cost paradigms of a common industrial circuit. We demonstrate the efficacy of this technique on standard Trust-Hub benchmark circuits with combinational Trojans inserted in sequential designs, showing significant improvement over prior techniques. We also explore the potential cost-benefit tradeoffs that exist within such a methodology, with the intent to provide an efficient solution for an array of potential product markets. This methodology provides a reliable and effective means for Trojan detection, addressing an important piece of the overall circuit certification puzzle. Chris Nigh, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Concurrent Monitoring of Operational Health in Neural Networks Through Balanced Output PartitionsabstractThe abundant usage of deep neural networks in safety-critical domains such as autonomous driving raises concerns regarding the impact of hardware-level faults on deep neural network computations. As a failure can prove to be disastrous, low-cost safety mechanisms are needed to check the integrity of the deep neural network computations. We embed safety checksums into deep neural networks by introducing a custom regularization term in the network training. We partition the outputs of each network layer into two groups and guide the network to balance the summation of these groups through an additional penalty term in the cost function. The proposed approach delivers twin benefits. While the embedded checksums deliver low-cost detection of computation errors upon violations of the trained equilibrium during network inference, the regularization term enables the network to generalize better during training by preventing overfitting, thus leading to significantly higher network accuracy. Elbruz Ozen, Alex Orailoglu |
ASP-DAC | 2 |
| 2020 | Hunting Sybils in Participatory Mobile Consensus-Based NetworksabstractWe focus on detecting adversarial non-existent nodes, Sybils, in anonymized participatory mobile networks where nodes support both a node-to-server and peer-to-peer connection capabilities. As data-driven decisions within such networks typically rely on local consensuses, they are susceptible to adversarial injection attacks which impersonate honest nodes and overpower local data through forgery. Nickolai Verchok, Alex Orailoglu |
AsiaCCS | 2 |
| 2020 | Test Pattern Superposition to Detect Hardware TrojansabstractCurrent methods for the detection of hardware Trojans inserted by an untrusted foundry are either accompanied by unreasonable costs in design/test pattern overhead, or return results that fail to provide confident trustability. The challenges faced by these side-channel techniques are primarily a result of process variation, which renders pre-silicon expectations nearly meaningless in predicting the behavior of a manufactured IC. To overcome this hindrance in a cost-effective manner, we propose an easy-to-implement test pattern-based approach that is self-referential in nature, capable of dissecting and understanding the characteristics of a given manufactured IC to hone in on aberrant measurements that are demonstrative of malicious Trojan hardware. By leveraging the superposition principle to cancel out non-Trojan noise, we can isolate and magnify Trojan circuit effects, all within a regime considerate of practical test and design-for-test infrastructures. Experimental results performed on Trust-Hub benchmarks demonstrate the proposed method provides a clear and significant boost in our ability to confidently certify manufactured ICs over similar state-of-the-art techniques. Chris Nigh, Alex Orailoglu |
DATE | 2 |
| 2020 | Just Say Zero: Containing Critical Bit-Error Propagation in Deep Neural Networks With Anomalous Feature SuppressionabstractDNNs are abundantly employed in a variety of applications, including real-time systems with strict safety constraints. The consequences of errors prove disastrous in safety-critical systems, such as autonomous driving, healthcare, and industrial applications. DNNs are resilient to limited numerical perturbations yet fragile under large deviations in weights and activations. The traditional error tolerance measures fail to meet the tight design constraints of DNN processing systems due to extensive overheads or limited advantages in abundant error conditions. The algorithmic particularities of DNNs though create novel opportunities to deal with errors more effectively and economically. We revisit the two fundamental tasks in fault-tolerant system design, namely, error detection and correction, and demonstrate that the precise versions of these operations could be replaced by approximated counterparts in DNNs to deliver an extensive bit-error resilience even at high error rates while necessitating no information redundancy. We first maintain DNN accuracy even under extreme error rates by suppressing the numerical contributions of anomalous activations, eliminating any reliance on precise error correction. We tackle the problem of no redundancy error detection by establishing in training numerical associations among activations, and employing them for anomaly detection. Anomalous feature detection and suppression, performed efficiently at inference with minimal resources in a DNN accelerator, is shown to deliver significant resilience boosts while imposing neither information redundancy nor perceptible overheads. Elbruz Ozen, Alex Orailoglu |
ICCAD | 2 |
| 2020 | A Crowd-Based Explosive Detection System with Two-Level Feedback Sensor CalibrationabstractLarge, open, public events, such as marathons and festivals, have always presented a unique safety challenge. These sprawling events, which can take up entire city blocks or stretch for many miles, can draw tens to hundreds of thousands of spectators and in some cases have open admission. As it is impracticable to guarantee the subjection of every event-goer to a security screening, we propose a crowd-based explosive detection system that uses a multitude of low-cost ChemFET sensors which are distributed to attendees. As the sensors offer limited accuracy, we further propose a server-based decision-making framework that utilizes a two-level feedback loop between the sensors and the server and explores spatial and temporal locality of the collected data to overcome the inherent low-accuracy of individual sensors. We thoroughly explore two distinct detection schemes, stressing their performance under a myriad of conditions, thus showing that such a crowd-based detection system comprised of low-cost and low-accuracy sensors can deliver high detection accuracy with minimal false positives. Chengmo Yang, Patrick Cronin, Agamyrat Agambayev, Sule Ozev, A. Enis Çetin, Alex Orailoglu |
ICCAD | 6 |
| 2020 | Squeezing Correlated Neurons for Resource-Efficient Deep Neural Networks
Elbruz Ozen, Alex Orailoglu |
ECML/PKDD (2) | 2 |
| 2020 | Taming Combinational Trojan Detection Challenges with Self-Referencing Adaptive Test PatternsabstractWhile many side-channel methods have been proposed for detecting hardware Trojans inserted by an untrusted foundry, they are challenged in the face of process variation noise. The impacts of process variation have forced researchers to propose costly design enhancements to improve detection as a counter to the deficiency of current easy-to-implement test pattern-based methods. To overcome process variation noise with no design cost, we propose a novel self-referencing adaptive approach based on test pattern construction, which learns from and conforms to device characteristics to maximally magnify the Trojan signal. Through iterative test pattern modifications, response analyses, and decision-making, we can pursue suspicious behaviors and increase the likelihood of Trojan detection. Experiments on Trust-Hub Trojan circuit benchmarks show the efficacy of this technique, magnifying an equivocal starting signal 22 to 130 to deliver crisp resolution to the question of Trojan existence. Chris Nigh, Alex Orailoglu |
VTS | 2 |
| 2020 | Low-Cost Error Detection in Deep Neural Network Accelerators with Linear Algorithmic Checksums
Elbruz Ozen, Alex Orailoglu |
J. Electron. Test. | 2 |
| 2020 | Boosting Bit-Error Resilience of DNN Accelerators Through Median Feature SelectionabstractDeep learning techniques have enjoyed wide adoption in real life, including in various safety-critical embedded applications. While neural network computations require protection against hardware errors, the substantial overheads of conventional error-tolerance techniques limit their use on embedded platforms, which carry out demanding deep neural network computations with limited resources. The utilization of conventional techniques is further constrained in high error rate scenarios, increasingly prevalent under aggressive energy and performance optimizations. To resolve this conundrum, we introduce a novel median feature selection technique to filter the impact of bit errors prior to the execution of each layer. While our technique can be deemed as a fine-grained modular redundancy scheme, its construction purely out of the inherent redundancy of the network necessitates neither additional parameters nor extra multiply-accumulate operations, squashing the inordinate overheads typically associated with such techniques. Median feature selection can be efficiently performed in hardware and seamlessly integrated into embedded deep learning accelerators as a modular plug-in. Deep learning models can be trained with standard tools and techniques to ensure a graceful operational interface with the feature selection stages. The proposed technique allows the system to perform accurately even at high error rates by improving its resilience up to four orders of magnitude, yet incurs negligible 0.19%–0.48% area and 0.07%–0.19% power overheads for the required operations. Elbruz Ozen, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Sanity-Check: Boosting the Reliability of Safety-Critical Deep Neural Network ApplicationsabstractThe widespread usage of deep neural networks in autonomous driving necessitates a consideration of the safety arguments against hardware-level faults. This study confirms the possible catastrophic impact of hardware-level faults on DNN accuracy; the consequent need for low-cost fault tolerance methods can be met through a rigorous exploration of the mathematical properties of the associated computations. We propose Sanity-Check, which makes use of the linearity property and employs spatial and temporal checksums to protect fully-connected and convolutional layers in deep neural networks. Sanity-Check can be purely implemented on software and deployed on different execution platforms with no additional modification. We also propose Sanity-Check hardware which integrates seamlessly with modern DNN accelerators and neutralizes the small performance overhead in pure software implementations. Sanity-Check delivers perfect error-caused misprediction coverage in our experiments, which makes it a promising candidate for boosting the reliability of safety-critical deep neural network applications. Elbruz Ozen, Alex Orailoglu |
ATS | 2 |
| 2019 | Piercing Logic Locking Keys through Redundancy IdentificationabstractThe globalization of the IC supply chain witnesses the emergence of hardware attacks such as reverse engineering, hardware Trojans, IP piracy and counterfeiting. The consequent losses sum to billions of dollars for the IC industry. One way to defend against these threats is to lock the circuit by inserting additional key-controlled logic such that correct outputs are produced only when the correct key is applied. The viability of logic locking techniques in precluding IP piracy has been tested by researchers who have identified extensive weaknesses when access to a functional IC is guaranteed.In this paper, we uncover weaknesses of logic locking techniques when the attacker has no access to an activated IC, thus exposing vulnerabilities at the earliest stage even for applications that seek refuge from attacks through functional opaqueness. We develop an attack algorithm that prunes out the incorrect value of each key bit when it introduces a significant level of logic redundancy. Throughout our experiments on ISCAS-85 and ISCAS-89 benchmark circuits, the attack deciphers more than half of the key bits on average with a high accuracy. Leon Li, Alex Orailoglu |
DATE | 2 |
| 2019 | Detecting Gas Vapor Leaks through Uncalibrated Sensor Based CPSabstractWhile Volatile Organic Compounds (VOC) and ammonia have a place in our daily lives, their leakage into the environment is harmful to human health. In order to prevent and detect gaseous leaks of harmful VOCs, a cyber-physical system (CPS) comprised of ordinary people or first responders is proposed. This CPS uses small, low-cost sensors coupled to smart phones or mobile devices with the necessary computation and communication capabilities. The efficacy of such a CPS hinges on its ability to address technical challenges stemming from the fact that identically produced sensors may produce different results under the same conditions due to sensor drift, noise, or resolution errors. The proposed system makes use of time-varying signals produced by sensors to detect gas leaks. Sensors sample the gas vapor level in a continuous manner and time-varying sensor data is processed using deep neural networks. One of the neural networks (NN) is an energy efficient Additive Neural Network (AddNet) which can be implemented in host devices. The second NN is the discriminator of a GAN and the third a regular convolutional NN. AddNet produces comparable VOC gas leak detection results to regular convolutional networks while reducing area requirements by two thirds. Diaa Badawi, Sule Ozev, Jennifer Blain Christen, Chengmo Yang, Alex Orailoglu, A. Enis Çetin |
ICASSP | 5 |
| 2019 | Shielding Logic Locking from Redundancy AttacksabstractThe security of logic locking has been extensively examined under the threat model that assumes the availability of an activated IC. Recently, structural attacks such as ones based on redundancy analysis have challenged the viability of logic locking even when stringent measures are taken to preclude access to an activated IC. In this paper, we propose a gate selection based logic locking technique to identify key gate insertion sites such that the redundancy level deviates minimally under all key assignments. The proposed logic locking technique is evaluated on a set of benchmark circuits to confirm its resistance against redundancy analysis based attacks. Leon Li, Alex Orailoglu |
VTS | 2 |
| 2018 | Variation-Aware Hardware Trojan Detection through Power Side-channelabstractA hardware Trojan (HT) denotes the malicious addition or modification of circuit elements. The purpose of this work is to improve the HT detection sensitivity in ICs using power side-channel analysis. This paper presents three detection techniques in power based side-channel analysis by increasing Trojan-to-circuit power consumption and reducing the variation effect in the detection threshold. Incorporating the three proposed methods has demonstrated that a realistic fine-grain circuit partitioning and an improved pattern set to increase HT activation chances can magnify Trojan detectability. Fakir Sharif Hossain, Michihiro Shintani, Michiko Inoue, Alex Orailoglu |
ITC | 4 |
| 2017 | Ensuring system security through proximity based authenticationabstractAs Internet of Things applications using embedded systems enter wider markets, securing systems against attacks becomes necessary. In many applications, securely determining transmitter location helps maintaining system security. A relatively new RF-based localization technique called Received Signal Strength Ratio (RSSR), has potential utility in securing ad-hoc networks and body-area networks. We describe a novel attack on the security of such systems, and discuss a set of mitigation strategies that restore the effectiveness of RSSR. Joshua Marxen, Alex Orailoglu |
ASP-DAC | 2 |
| 2017 | Intra-Die-Variation-Aware Side Channel Analysis for Hardware Trojan DetectionabstractHigh detection sensitivity in the presence of process variation is a key challenge for hardware Trojan detection through side channel analysis. In this work, we present an efficient Trojan detection approach in the presence of elevated process variations. The detection sensitivity is sharpened by 1) comparing power levels from neighboring regions within the same chip so that the two measured values exhibit a common trend in terms of process variation, and 2) generating test patterns that toggle each cell multiple times to increase Trojan activation probability. Detection sensitivity is analyzed and its effectiveness demonstrated by means of RPD (relative power difference). We evaluate our approach on ISCAS'89 and ITC'99 benchmarks and the AES-128 circuit for both combinational and sequential type Trojans. High detection sensitivity is demonstrated by analysis on RPD under a variety of process variation levels and experiments for Trojan inserted circuits. Fakir Sharif Hossain, Tomokazu Yoneda, Michihiro Shintani, Michiko Inoue, Alex Orailoglu |
ATS | 5 |
| 2017 | Detecting hardware Trojans without a Golden IC through clock-tree defined circuit partitionsabstractThe sensitive identification and detection of hardware Trojans in ICs without a golden reference constitutes a key challenge. Traditional circuit partitioning and side-channel analysis techniques fall short of perfect sensitivity and accuracy and rely on golden references. In this work, a novel layout-aware clock tree driven circuit partitioning is coupled with an algorithm that selects transition delay fault test patterns that will deliver equal power on partitions. The circuit partitioning through the clock tree results in minimal hardware additions that can be effected through ECO. The selection of pairs of power uniform small regions results in reduction of inter-die variation effects, thus delivering increased detection sensitivity. The comparison for equal power of numerous pairs thoroughly perturbs the circuit under various activation conditions, resulting in elevated sensitivity and accuracy while obviating the need for reliance on Golden ICs. We evaluate our approach on ISCAS89 and ITC benchmarks and AES circuits for both combinational and sequential type Trojans. Experimental results show that our approach can effectively detect Trojans with ratios to circuit area as low as 0.023%. Fakir Sharif Hossain, Tomokazu Yoneda, Michiko Inoue, Alex Orailoglu |
ETS | 4 |
| 2016 | Power-Aware Delay Test Quality Optimization for Multiple Frequency DomainsabstractAs the number of frequency domains aggressively grows in today's systems-on-chip (SoCs), the delivery of high-delay test quality across numerous frequency domains while meeting test budgets assumes crucial importance. This paper proposes a method to explore the delay test quality tradeoffs across these domains, determining an optimal distribution of the test time budget across all domains while minimizing the overall SoC delay defect escape level. Satisfaction of this goal necessitates not only consideration of fault coverage but also of the distinct characteristics of each domain, such as frequency, path length distribution, scan length, and shift speed as well as full utilization of concurrent test support while remaining within the constraints of power thresholds to provide a reliable test environment. An optimization formulation as well as efficient test time allocation methods based on convexity and fast concurrent test planning algorithms are provided. Baris Arslan, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Aggressive Test Cost Reductions Through Continuous Test Effectiveness AssessmentabstractThe inclusion of various new test types in production test suites with the hope of keeping defect escape level in check continuously increases the test cost while the economics of the intensely competitive consumer marketplace dictates test strategies that are low cost yet still effective in defect detection. An adaptive test methodology is proposed in this paper to address the simultaneous demand for efficiency and effectiveness in testing through an adaptive prioritization of test vectors according to defect detection effectiveness, subsequently, leading to the identification of a compact yet effective test set. Experimental results confirm the effectiveness of the vector prioritization process and show substantial cost reduction levels. Baris Arslan, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Mobile ecosystem driven application-specific low-power control microarchitectureabstractA well-established technique for reducing embedded processor power is to leverage rigorous static code analysis to enable application-specific hardware customizations. Energy can be saved by leveraging compile-time information to eliminate certain hardware operations. One such research area reduced branch prediction hardware power by statically extracting application control flow information. However, the applicability of these static techniques has stalled with respect to modern smartphones. These high-performance mobile processors inhabit a unique mobile ecosystem domain space. Rather than an application consisting of a monolithic instruction sequence compiled on the target processor, mobile applications are stored in a centralized marketplace and consist of high-level object-oriented code that dynamically binds with device-specific foundation libraries. While this model helps reduce development time and enables applications to be downloaded and run on a variety of device models and operating system versions, it significantly hinders whole application static analysis and application-specific hardware optimizations. This paper addresses the unique challenges of the mobile ecosystem by enabling on-device application analysis to guide reconfiguration of the branch target buffer. Software and hardware customizations are intelligently combined to greatly reduce power dissipation while maintaining or even improving performance. Garo Bournoutian, Alex Orailoglu |
ICCD | 2 |
| 2015 | Joint Profit and Process Variation Aware High Level Synthesis With Speed BinningabstractAs integrated circuits continuously scale up, process variation plays an increasingly significant role in system design and semiconductor economic return. In this paper, we explore the potential of profit improvement under the inherent semiconductor variability based on the speed binning technique. We aim to develop a set of high level synthesis (HLS) solutions, for which purpose heuristic techniques, including allocation, scheduling, and resource binding, are proposed. The goal is to construct designs that maximize the number of chips that can be sold at the most advantageous price, leading to the maximization of the overall profit. In addition, a genetic algorithm-based formulation is constructed for HLS solutions. Then, we complement the HLS techniques with near-optimal bin placement strategies for further profit improvement. Experimental results confirm the superiority of the HLS results and the associated improvement in profit margins. Mengying Zhao, Alex Orailoglu, Chun Jason Xue |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | On-device objective-C application optimization framework for high-performance mobile processorsabstractSmartphones provide applications that are increasingly similar to those of interactive desktop programs, providing rich graphics and animations. To simplify the creation of these interactive applications, mobile operating systems employ highlevel object-oriented programming languages and shared libraries to manipulate the device's peripherals and provide common userinterface frameworks. The presence of dynamic dispatch and polymorphism allows for robust and extensible application coding. Unfortunately, the presence of dynamic dispatch also introduces significant overheads during method calls, which directly impact execution time. Furthermore, since these applications rely heavily on shared libraries and helper routines, the quantity of these method calls is higher than those found in typical desktop-based programs. Optimizing these method calls centrally before consumers download the application onto a given phone is exacerbated due to the large diversity of hardware and operating system versions that the application could run on. This paper proposes a methodology to tailor a given Objective-C application and its associated device-specific shared library codebase using on-device post-compilation code optimization and transformation. In doing so, many polymorphic sites can be resolved statically, improving the overall application performance. Garo Bournoutian, Alex Orailoglu |
DATE | 2 |
| 2014 | Sleep-aware variable partitioning for energy-efficient hybrid PRAM and DRAM main memoryabstractEnergy consumption of memories is always a significant issue for computing systems. Recently, hybrid PRAM and DRAM memory architectures have been proposed. It combines the advantages of DRAM and PRAM, such as low leakage power in PRAM and short write latency in DRAM. However, the leakage power in DRAM is still considerable in hybrid memories. The leakage power can only be reduced by turning DRAM into sleep state. In this paper, a novel proximity concept is proposed to guide the variable partitioning to maximize the possibility of turning DRAM into sleep mode. A novel Sleep-Aware Variable Partition Algorithm (SAVPA) is then proposed with the objective of maximizing the sleep time of DRAM while satisfying the performance and endurance constraints. The experiment results show that SAVPA reduces the energy consumption by 11.25% in average (up to 15.84%) compared to the state-of-art work with simple sleep technique. Chenchen Fu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
ISLPED | 4 |
| 2014 | Branch Prediction-Directed Dynamic Instruction Cache Locking for Embedded SystemsabstractCache locking is a cache management technique to preclude the replacement of locked cache contents. Cache locking is often adopted to improve cache access predictability in Worst-Case Execution Time (WCET) analysis. Static cache locking methods have been proposed recently to improve Average-Case Execution Time (ACET) performance. This article presents an approach, Branch Prediction-directed Dynamic Cache Locking (BPDCL), to improve system performance through cache conflict miss reduction. In the proposed approach, the control flow graph of a program is first partitioned into disjoint execution regions, then memory blocks worth locking are determined by calculating the locking profit for each region. These two steps are conducted during compilation time. At runtime, directed by branch predictions, locking routines are prefetched into a small high-speed buffer. The predetermined cache locking contents are loaded and locked at specific execution points during program execution. Experimental results show that the proposed BPDCL method exhibits an average improvement of 25.9%, 13.8%, and 8.0% on cache miss rate reduction in comparison to cases with no cache locking, the static locking method, and the dynamic locking method, respectively. Keni Qiu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2014 | Examining Timing Path Robustness Under Wide-Bandwidth Power Supply Noise Through Multi-Functional-Cycle Delay TestabstractCircuits designed and fabricated with nanometer-scale technology are increasingly sensitive to power ground noise across a wide frequency range, thus necessitating a strict examination of circuit robustness against noise during manufacturing tests. Conventional at-speed testing techniques possibly result in the escape of marginal timing failures, as they are unable to account for the impact of middle- and low-frequency noise on circuit timing. To address this challenge, we propose, in this paper, a novel multi-functional-cycle test scheme that targets the noise-induced failures on critical paths of the circuit. The proposed technique explores the noise profile of at-speed functional cycles and approximates it in delay testing through the application of multiple capture operations, thus maximally detecting the timing failures that potentially take place under the worst case functional mode noise. The noise impact of individual devices on the critical paths is characterized through simulations on the power mesh model extracted from the circuit layout. This enables a computationally efficient yet SPICE-accurate estimation of the compound noise profile of the test pattern through the linear superposition of individual ones. Guided by this noise estimation technique, a test pattern transformation flow is proposed to maximize the noise in pseudo-functional test operations. Simulation results show that the proposed scheme can examine the effect of wide-bandwidth noise and thus perform a much more rigorous testing on critical paths than conventional delay testing schemes, thereby significantly improving test quality. Mingjing Chen, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Full exploitation of process variation space for continuous delivery of optimal delay test qualityabstractThe increasing magnitude of process variations individualizes effectively each chip, necessitating distinct quantities of test resources for each in order to optimize overall delay test quality without exceeding set test budgets. This paper proposes an analytical framework that delivers the optimal test time assignment per chip in order to minimize the delay defect escape rate. Adjustment of the chip-specific test time in the continuous process variation space is attained through an adaptive test flow that utilizes process data measurements from the device under test. The results evince that a substantial improvement in the delay test quality can be obtained at no increase whatsoever to test time consumed by conventional test flows. Baris Arslan, Alex Orailoglu |
ASP-DAC | 2 |
| 2013 | Profit maximization through process variation aware high level synthesis with speed binningabstractAs integrated circuits continuously scale up, process variation plays an increasingly significant role in system design and semiconductor economic return. In this paper, we explore the potential of profit improvement under the inherent semiconductor variability based on the speed binning technique. We first accordingly propose a set of high level synthesis techniques, including allocation, scheduling and resource binding, thus essentially constructing designs that maximize the number of chips that can be sold at the most advantageous price, leading to the maximization of the overall profit. We explore subsequently the optimal bin placement strategy for further profit improvement. Experimental results confirm the superiority of the high level synthesis results and the associated improvement in profit margins. Mengying Zhao, Alex Orailoglu, Chun Jason Xue |
DATE | 2 |
| 2013 | Branch Prediction directed Dynamic instruction Cache Locking for embedded systemsabstractCache locking is a cache management technique to preclude the replacement of locked cache contents. Cache locking is often used to improve cache access predictability in Worst-Case Execution Time (WCET) analysis. Static cache locking methods have been proposed recently to improve average system performance. This paper presents an approach, Branch Prediction directed Dynamic Cache Locking (BPDCL), to improve average system performance through effective cache conflict miss reduction in different execution regions. In this proposed approach, the control flow graph of a program is partitioned into regions and memory blocks worth locking for each region are calculated during compilation time. At runtime, directed by branch predictions, locking routines are prefetched into a high-speed buffer. The pre-determined cache locking contents are loaded and locked at specific execution points during program execution. Experimental results show that the proposed BPDCL method exhibits an average improvement of 21.8% and 10.3% on cache miss rate reduction in comparison to the case with no cache locking and the static locking method respectively. Keni Qiu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
RTCSA | 4 |
| 2013 | Tracing the best test mix through multi-variate quality trackingabstractThe increasing multiplicity of defect types forces the inclusion of tests from a variety of fault models. The quest for test quality is checkmated though by the considerable and frequently unnecessary cost of the large number of tests, driven by the lack of a clear correspondence between defects and fault models. While the static derivation of the appropriate test mixes from a variety of fault models to deliver high test quality at low cost is a desirable goal, it is challenged by the frequent changes in defect characteristics. The consequent necessity for adaptivity is addressed in this paper through a test framework that utilizes the continuous stream of failing test data during production testing to track the varying test quality based on evolving defect characteristics and thus dynamically adjust the production test set to deliver a target defect escape level at minimal test cost. Baris Arslan, Alex Orailoglu |
VTS | 2 |
| 2013 | Towards a cost-effective hardware trojan detection methodologyabstractDue to the increasing globalization of integrated circuit fabrication, hardware security has emerged as a major issue, necessitating hardware trojan detection mechanisms. Numerous techniques exist, a subset of which we applied to 6 sets of combinational circuits and 2 sets of sequential circuits in order to effectively determine the presence of a hardware trojan. Utilizing only a single type of functional test and a single type of side-channel test, we were able to make determinations about all 6 sets of combinational circuits and 1 of the 2 sets of sequential circuits. Raymond Paseman, Alex Orailoglu |
VTS | 2 |
| 2013 | Application-aware adaptive cache architecture for power-sensitive mobile processorsabstractToday, mobile smartphones are expected to be able to run the same complex, algorithm-heavy, memory-intensive applications that were originally designed and coded for general-purpose processors. All the while, it is also expected that these mobile processors be power-conscientious as well as of minimal area impact. These devices pose unique usage demands of ultra-portability but also demand an always-on, continuous data access paradigm. As a result, this dichotomy of continuous execution versus long battery life poses a difficult challenge. This article explores a novel approach to mitigating mobile processor power consumption while abating any significant degradation in execution speed. The concept relies on efficiently leveraging both compile-time and runtime application memory behavior to intelligently target adjustments in the cache to significantly reduce overall processor power, taking into account both the dynamic and leakage power footprint of the cache subsystem. The simulation results show a significant reduction in power consumption of approximately 13% to 29%, while only incurring a nominal increase in execution time and area. Garo Bournoutian, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Register allocation for embedded systems to simultaneously reduce energy and temperature on registersabstractEnergy and thermal issues are two important concerns for embedded system design. Diminished energy dissipation leads to a longer battery life, while reduced temperature hotspots decelerate the physical failure mechanisms. The instruction fetch logic associated with register access has a significant contribution towards the total energy consumption. Meanwhile, the register file has also been previously shown to exhibit the highest temperature compared to the rest of the components in an embedded processor. Therefore, the optimization of energy and the resolution of the thermal issue for register accesses are of great significance. In this article, register allocation techniques are studied to simultaneously reduce energy consumption and heat buildup on register accesses for embedded systems. Contrary to prevailing intuition, we observe that optimizing energy and optimizing temperature on register accesses conflict with each other. We introduce a rotator hardware in the instruction decoder to facilitate a balanced solution for the two conflicting objectives. Algorithms for register allocation and refinement are proposed based on the access patterns and the effects of the rotator. Experimental results show that the proposed algorithms obtain notable improvements of energy and peak temperature for embedded applications. Tiantian Liu 0001, Alex Orailoglu, Chun Jason Xue, Minming Li |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Delay test resource allocation and scheduling for multiple frequency domainsabstractAs the number of frequency domains aggressively grows in today's SOCs, the delivery of high delay test quality across numerous frequency domains while meeting test budgets is crucial. This goal necessitates not only the consideration of fault coverage but also the distinct characteristics of each domain such as frequency and the distribution of path lengths and, additionally, the delay test quality tradeoffs across these domains. This paper proposes a method to identify the optimal test time allocation per domain based on the distinct characteristics of each in order to minimize overall delay defect escape level. The proposed method not only considers test time allocation but also concurrent scheduling of domains to optimize the delay test quality for SOCs that support the testing of multiple frequency domains in parallel. Baris Arslan, Alex Orailoglu |
VTS | 2 |
| 2012 | Small-delay defects detection under process variation using Inter-Path CorrelationabstractDetection of Small Delay Defects (SDDs) is a major concern in modern circuits using nanometer technologies. They are difficult to test and an important source of test escapes, and even when SDDs do not produce functional failures, they represent a reliability risk. The detection of these defects aggravates in the presence of process variations. In this paper, a methodology to detect SDDs in the presence of process variations using delay correlation information between paths of a circuit is proposed. This methodology exploits the concept that for two highly correlated paths, an important part of the delay variance in one path can be described by the delay variance in the second path. The methodology has been further extended to consider multiple path correlation thus improving the detection of SDDs. This methodology is able to distinguish delay defects from process variations. A metric is also proposed to quantify the SDD screenable variance that represents the percentage of variance where a defect can be detected. A statistical timing analysis framework has been developed and implemented to compute timing information and Inter-Path Correlation (IPC). Spatial and structural correlation, and random dopant fluctuations are considered. Simulation results in 74LS85 and ISCAS85 benchmark circuits evince the feasibility of the proposed methodology. Francisco J. Galarza-Medina, Jose Luis Garcia-Gervacio, Víctor H. Champac, Alex Orailoglu |
VTS | 4 |
| 2012 | On Diagnosis of Timing Failures in Scan ArchitectureabstractExcessive test mode power-ground noise in nanometer scale chips causes large delay uncertainties in scan chains, resulting in a highly elevated rate of timing failures. The hybrid timing violation types in scan chains, compounded by their possibly intermittent manifestations, invalidate the conventional assumptions in scan chain fault behavior, significantly increasing the ambiguity and difficulty in diagnosis. In this paper, we propose a novel methodology to identify the root cause of scan chain timing failures. The proposed work addresses the challenge of diagnosing multiple permanent or intermittent timing faults in scan chains and the associated clock trees, which closely approximate the realistic failure mechanisms observed in silicon. Instead of relying on fault simulation that is incapable of approximating the intermittent fault manifestation, the proposed technique characterizes the impact of timing faults by analyzing the phase movement of scan patterns. Extracting fault-sensitive statistical features of phase movement information provides strong signals for the precise identification of fault locations and types. The manifestation probability of each fault is furthermore computed through a mathematical transformation framework which accurately models the behavior of multiple faults as a Markov chain. The identification of failing scan cells enables a further examination of the possible delay defects in the scan clock buffers, which ascertains the possible root causes of the observed scan chain failures. The proposed scheme characterizes the timing impact of the defective clock buffers by extracting the change in the delay distribution of the clock paths, enabling the effective pruning of unrealistic fault hypotheses that would result in highly deviant timing behavior. Simulation results have confirmed that the proposed methodology can yield highly accurate diagnosis results for complex fault manifestations. Mingjing Chen, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Tackling Resource Variations Through Adaptive Multicore Execution FrameworksabstractMulticore architectures have been widely adopted to accommodate the rising performance demand in various application domains, ranging from high-end supercomputing to low-end consumer electronics. Yet due to the ever growing integration density and application complexity, such architectures suffer from increased level of core availability variations. At runtime, issues such as device failures, heat buildup, as well as resource competitions and preemptions can make computational resources unavailable, necessitating execution schedules capable of delivering diverse performance levels to match the varying resource allocations. The adaptive execution framework introduced in this paper delivers high-quality schedules capable of predictably reconfiguring execution and gracefully degrading performance in the face of resource unavailability. By adhering to a novel band structure, a set of possible execution schedules are compactly engendered in readiness at compile time, thus delivering predictable responses to runtime resource variations. More importantly, through the exploitation of an extra degree of freedom in the scheduling process, the scheduler can perform task assignments in such a way that adaptivity can be embedded within the preoptimized schedules at almost no cost. The efficacy of the proposed technique is confirmed by incorporating it into a conventional, widely adopted scheduling heuristic and experimentally verifying it in the context of single core degradations. Chengmo Yang, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Scan Power Reduction for Linear Test Compression Schemes Through Seed SelectionabstractXOR network-based on-chip test compression schemes have been widely employed in large industrial scan designs due to their high compression ratio and efficient decompression mechanism. Nevertheless, such a scheme necessitates high unspecified bit ratios in the original test cubes, resulting in quite significant difficulties in preprocessing test cubes for scan power reduction. The linear mapping from the original cubes to the compressed seeds typically provides extra degrees of flexibility as multiple seeds may reconstruct the test cube. Due to the highly divergent power impact of distinct seeds though, appreciable power reductions in the decompressed test data can be attained through the pinpointing of the power-optimal seeds during the compression phase. This work explores the aforementioned flexibility in the seed space, and outlines a mathematical and algorithmic framework for a power-aware linear test compression scheme. The proposed technique incurs no hardware overhead over the traditional linear compression scheme; it can be easily embedded furthermore into the industrial test compaction/compression flow. Experimental results confirm that the proposed technique delivers significant scan power reduction with negligible impact on the compression ratio. Mingjing Chen, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Adaptive Test Framework for Achieving Target Test Quality at Minimal CostabstractDefect characteristics and consequently test quality vary throughout the production life cycle. An optimal test methodology needs to adjust the test set based on the ever changing defect characteristics to deliver consistent test quality levels. In this paper, we propose an adaptive test framework that alters the test set continuously by utilizing the history of recent failures to track the instantaneous defect escape level, thus reaching the target test quality level at minimal test cost. Baris Arslan, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2011 | Diagnosing scan clock delay faults through statistical timing pruningabstractA novel methodology for diagnosing the delay faults in the scan clock tree is proposed. The proposed scheme characterizes the timing impact of the defective clock buffers by extracting the change in the delay distribution of the clock paths, enabling the effective pruning of unrealistic fault hypotheses that would result in highly deviant timing behavior. The proposed scheme models the statistical delay variation due to test mode power-ground noise, thus maximally approximating the actual failure behavior. Simulation results have confirmed that the proposed methodology can yield highly accurate diagnosis results for complex fault manifestations. Mingjing Chen, Alex Orailoglu |
DAC | 2 |
| 2011 | Adaptive test optimization through real time learning of test effectivenessabstractProduction test suites include a large number of redundant test patterns due to the inclusion of multiple test types with overlapping defect detection and the use of simple fault models for test generation. Identification and elimination of ineffective test patterns promises a significant reduction in test cost. This paper proposes a test framework that learns, without extensive data collection and at no additional test time, the effectiveness of individual test patterns during production testing by getting defect detection feedback from a dynamic test flow. The proposed technique is further capable of adapting to changes in the underlying defect mechanisms by tracking the defect detection trend of test patterns. Baris Arslan, Alex Orailoglu |
DATE | 2 |
| 2011 | Diagnosing scan chain timing faults through statistical feature analysis of scan imagesabstractExcessive test mode power-ground noise in nanometer scale chips causes large delay uncertainties in scan chains, resulting in a highly elevated rate of timing failures. The hybrid timing violation types in scan chains, plus their possibly intermittent manifestations, invalidate the traditional assumptions in scan chain fault behavior, significantly increasing the ambiguity and difficulty in diagnosis. In this paper, we propose a novel methodology to resolve the challenge of diagnosing multiple permanent or intermittent timing faults in scan chains. Instead of relying on fault simulation that is incapable of approximating the intermittent fault manifestation, the proposed technique characterizes the impact of timing faults by analyzing the phase movement of scan patterns. Extracting fault-sensitive statistical features of phase movement information provides strong signals for the precise identification of fault locations and types. The manifestation probability of each fault is furthermore computed through a mathematical transformation framework which accurately models the behavior of multiple faults as a Markov chain. The fault model utilized in the proposed scheme considers the effect of possibly asymmetric fault manifestation, thus maximally approximating the realistic failure behavior. Simulations on large benchmark circuits and two industrial designs have confirmed that the proposed methodology can yield highly accurate diagnosis results even for complicated fault manifestations such as multiple intermittent faults with mixed fault types. Mingjing Chen, Alex Orailoglu |
DATE | 2 |
| 2011 | Register allocation for simultaneous reduction of energy and peak temperature on registersabstractIn this paper, we focus on register allocation techniques to simultaneously reduce energy consumption and heat buildup of register accesses. The conflict between these two objectives is resolved through the introduction of a hardware rotator. A register allocation algorithm followed by a refinement method is proposed based on the access patterns and the effects of the rotator. Experimental results show that the proposed algorithms obtain notable improvements in energy consumption and temperature reduction for embedded applications. Tiantian Liu 0001, Alex Orailoglu, Chun Jason Xue, Minming Li |
DATE | 2 |
| 2011 | Frugal but flexible multicore topologies in support of resource variation-driven adaptivityabstractGiven the projected higher variations in the availability of computational resources, adaptive static schedules have been developed to attain high-speed execution reconfiguration with no reliance on any runtime rescheduling decisions. These schedules are able to deliver predictable execution despite the increased levels of device unreliability in future multicore systems. Yet the associated runtime reconfiguration overhead is largely determined by the underlying system topology. Fully connected architectures, although they can effectively hide the overhead in execution migration, become infeasible as the core count grows to hundreds in the near future. We exploit in this paper the high locality associated with adaptive static schedules, and outline a scalable and locally shareable system organization for multicore platforms. With the incorporation of a limited set of neighborhood-centered communication links, threads are allowed to be directly migrated among adjacent cores without physical data movement. At the architecture level, a set of 2-dimensional physical topologies with such a local sharing property embedded is furthermore proposed. The inherent regularity allows these topologies to be adopted as a fixed-silicon multicore platform that can be flexibly redefined according to the parallelism characteristics and resilience needs of each application. Chengmo Yang, Alex Orailoglu |
DATE | 2 |
| 2011 | Migration-aware adaptive MPSoC static schedules with dynamic reconfigurability
Chun Jason Xue, Chengmo Yang, Alex Orailoglu |
J. Parallel Distributed Comput. | 4 |
| 2011 | Full Fault Resilience and Relaxed Synchronization Requirements at the Cache-Memory InterfaceabstractWhile multicore platforms promise significant speedup for many current applications, they also suffer from increased reliability problems as a result of ever scaling device size. The projected elevation in fault rate, together with the diverse behavior of fault manifestation, argues for highly efficient solutions of full fault resilience. Traditional duplication and checkpointing strategies typically impose sizable overhead in checkpointing execution results, or in constantly synchronizing two threads for value checking. To reduce such overhead while at the same time delivering full fault resilience, we propose an integrated fault detection and checkpointing framework, wherein the comparison and checkpointing process is performed at the cache-memory interface. By sharing a single cache between two duplicated threads, execution results can be directly verified in the cache before being written back, thus strictly protecting the memory against execution faults. Meanwhile, as unconfirmed data are allowed to be written into the cache, one thread can run well ahead of the other, thus relaxing the straightjacket of the strict execution synchronization model. If a cache block is constantly updated, further synchronization relaxation can be achieved through extending the cache design to duplicate a cache block and skip the comparison of the intermediate values. Chengmo Yang, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Cost-effective IR-drop failure identification and yield recovery through a failure-adaptive test schemeabstractEver-increasing test mode IR-drop results in a significant amount of defect-free chips failing at-speed testing. The lack of a systematic IR-drop failure identification technique engenders a highly increased failure analysis time/cost and significant yield loss. In this paper, we propose a failure-adaptive test scheme that enables a fast differentiation of the IR-drop induced failure from the actual defects of the chip. The proposed technique debugs the failing chips using low IR-drop vectors that are custom-generated from the observed faulty response. Since these special vectors are designed in such a way that all the actual defects captured by the original vectors are still manifestable, their application can clearly pinpoint whether the root cause of failure is IR-drop or not, thus eliminating reliance on an intrusive debugging process that incurs quite a high cost. Such a test scheme further enables effective yield recovery from failing chips by passing the ones validated by the debugging vectors whose IR-drop level matches the functional mode. Experimental results show that the proposed scheme delivers a significant IR-drop reduction in the second test (debugging) phase, thus enabling a highly effective IR-drop failure identification and yield recovery at a slightly increased test cost. Mingjing Chen, Alex Orailoglu |
DATE | 2 |
| 2010 | Performance and energy efficient cache migrationapproach for thermal management in embedded systemsabstractIn this paper we propose an approach for performance and power aware warm start for the data cache during core level migration events that originate from overheating. We utilize the concept of reuse in the references to eliminate unnecessary information from being migrated. Furthermore, we exploit the temperature predictability to trigger the cache migration slightly before the actual thermal limit to allow sufficient time for the extraction of the reuse information and transfer of it to the destination cache while the execution is resuming normally. The suggested hardware not only is cost efficient but is also programmable so as to maintain flexibility in targeting application particularities. The experimental results we provide confirm the applicability of this approach. Raid Ayoub, Alex Orailoglu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Delay test quality maximization through process-aware selection of test set sizeabstractThe quality of a delay test set hinges not only on test patterns and the distribution of the delay defects but on the variations in process parameters as well. Process variations result in the same delay test set displaying differences from die to die in the detection of particular delay defects at the identical circuit node. The application of an identical test set to all devices independent of process variations consequently results in delivering inefficiencies in test time utilization. This paper proposes a delay test technique that adaptively changes the size of the test set based on the position of the device in the process variation space in order to maximize test quality within a given test time. Baris Arslan, Alex Orailoglu |
ICCD | 2 |
| 2010 | Fully adaptive multicore architectures through statically-directed dynamic execution reconfigurationsabstractAs a result of ever growing integration density and application complexity, future multicore architectures will suffer from increased levels of core availability variations. Full resource utilization in the face of various levels of resource availability necessitates techniques that compactly engender numerous schedules in readiness at compile time. Such schedules, each of which can make maximum utilization of the available resources, can be adaptively applied at runtime, thus enabling a toleration of up to an arbitrarily large amount of resource variations. Core binding permutations furthermore minimize the performance impact imposed by adaptivity on the pre-reconfiguration schedules while retaining all the concomitant benefits. The efficacy of the proposed technique is confirmed by incorporating it into a conventional, widely adopted scheduling heuristic and experimentally verifying it in the context of multiple core deallocations. This paper thus offers critical improvement over prior state-of-the-art, which targets solely single core failures, a subset of resource variation modes in future nanoscale MPSoCs which are projected to display elevated device failures, heat buildup, resource competition and preemptions. Chengmo Yang, Alex Orailoglu |
VLSI-SoC | 2 |
| 2010 | Fine-grained adaptive CMP cache sharing through access history exploitationabstractAdvances in semiconductor technologies have enabled the integration of multiple processor cores as well as varying sizes of L1 and L2 caches on a single chip. The ever growing complexity and diversity of the associated workloads impose a crucial challenge on the organization and management of the on-chip cache resources. As each core generates a varying amount of accesses to each cache line during execution, sharing a single L2 cache among all the cores can minimize off-chip misses. However, each access to a shared L2 cache imposes significant performance and power overhead, as the tags of all the blocks on a cache line need to be compared in parallel. To efficiently utilize cache resources while saving power, we present in this paper a fine-grained L2 cache management technique with minimum hardware overhead. Each core is allowed to set an ownership bit in an L2 cache block to directly signify the necessity of tag checking, thus reducing the latency and power consumption of each cache access. Joint block ownership approaches provide shareability, thus precluding costly data replication from which private L2 caches typically suffer. Meanwhile, through monitoring line-based access histories, a core that produces a large amount of misses is precluded from replacing blocks belonging to other cores, thus efficiently attaining fine-grained cache partitioning. Experimental results confirm that the proposed technique can effectively reduce the access latency and power consumption of traditional shared L2 caches, accompanied by additionally a slight reduction in the miss rate. Chengmo Yang, Chun Jason Xue, Alex Orailoglu |
VLSI-SoC | 3 |
| 2010 | VDDmin test optimization for overscreening minimization through adaptive scan chain maskingabstractVery-low-voltage (VDDmin) test assumes significant importance in detecting flaws in marginal chips as it can magnify the electrical impact of flaws. Yet such a test scheme is increasingly sensitive to test mode IR-drop. Exceedingly high test mode current and voltage surge may result in good chips failing VDDmin test, thus resulting in yield loss. A technique to counteract such yield loss, scan chain masking, has been widely incorporated in commercial DFT tools to address this issue. However, the masking of scan cells might lead to a reduced flaw coverage, necessitating a thorough investigation of its impact on test quality during the development of the optimal test plan. In this paper, we propose an analytical model to evaluate the cost of scan chain masking in terms of test escapes and overscreening effect. An adaptive scan chain masking flow guided by this model is then proposed to enhance the VDDmin test effectiveness in terms of the overall cost. The proposed flow adaptively identifies the IR-drop sensitive scan chains through the use of silicon debugging data of known-good parts and/or IR-drop simulation results, thus effectively avoiding the overmasking of scan chains that do not contribute to test mode IR-drop failures. This flow additionally delivers rapid convergence to a scan chain masking scheme that minimizes the overall escape/overscreening cost, thus significantly reducing engineering time in silicon debugging. Experimental results on real silicon confirm that the proposed methodology drastically reduces overscreening of good parts at negligible impact on test quality. Mingjing Chen, Alex Orailoglu |
VTS | 2 |
| 2010 | DiSC: A New Diagnosis Method for Multiple Scan Chain FailuresabstractIn scan-based testing environments, identifying the scan chain failures can be of significant help in guiding the failure analysis process for yield improvement. In this paper, we propose an efficient scan chain diagnosis method using a symbolic fault simulation to achieve high diagnostic resolution and small candidate list for single and multiple defects in scan chains. The main ideas of the proposed scan chain diagnosis method are twofold: 1) the reduction of the candidate scan cells through the analysis of the symbolic simulation responses, and 2) the identification of final candidate scan cells using the backward tracing method with the symbolic simulation responses. Experimental results show the effectiveness. Sunghoon Chun, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Filtering Global History: Power and Performance Efficient Branch PredictorabstractIn this paper we present an Application Customizable Branch Predictor, ACBP, that delivers efficiency in energy savings and performance without compromising prediction accuracy. The idea of our technique is to filter unnecessary global history information within the global history register to minimize the predictor size while maintaining prediction accuracy. We suggest in this work an efficient algorithm to capture the beneficial correlations. A cost-efficient and programmable hardware architecture is presented. Extensive experimental analysis confirms significant improvements in power savings and latency, ranging up to 84% and 30%,respectively. Raid Ayoub, Alex Orailoglu |
ASAP | 2 |
| 2009 | Reducing impact of cache miss stalls in embedded systems by extracting guaranteed independent instructionsabstractToday, embedded processors are expected to be able to run complex, algorithm-heavy, memory-intensive applications that were originally designed and coded for general-purpose processors. As such, the impact of memory latencies on the execution time increasingly becomes evident. All the while, it is also expected that embedded processors be power-conscientious as well as of minimal area impact. As a result, traditional methods for addressing performance and memory latencies, such as multiple issue, out-of-order execution and large, associative caches, are not aptly suited for the embedded domain due to the significant area and power overhead. This paper explores a novel approach to mitigating execution delays caused by memory latencies that would otherwise not be possible in a regular in-order, single-issue embedded processor without large, power-hungry constructs like a Reorder Buffer (ROB). The concept relies on both compile-time and run-time information to safely allow non-data-dependent instructions to continue executing while a memory stall has occurred. The simulation results show significant improvement in execution throughput of approximately 11%, while having a minimal impact on area overhead and power. Garo Bournoutian, Alex Orailoglu |
CASES | 2 |
| 2009 | Making DNA self-assembly error-proof: Attaining small growth error rates through embedded information redundancyabstractDNA self-assembly is emerging as the most promising technique for nanoscale self-assembly as it uses the simple, yet precise rules of DNA binding to create macroscale assemblies from nanoscale components. However, DNA self-assembly is also highly error-prone and requires the use of error-resilience techniques in order to unlock its potential. In this paper we propose a technique for error-resilience that is based on information redundancy but, in contrast to previous information redundancy schemes, can achieve much higher resilience to growth errors. By expanding the neighborhood from which redundant information is taken, we can extend the distance that errors are propagated and therefore increase the likelihood of the error being reversed. Given a growth error rate of ∈, we show that with a neighborhood of only 2 we can reduce the error rate to ∈3.64for arbitrary functions (as compared to ∈2.33previously achieved). Compared with spatial redundancy approaches, our technique allows for higher density nanostructures and has a greatly reduced assembly time. Saturnino Garcia, Alex Orailoglu |
DATE | 2 |
| 2009 | Towards no-cost adaptive MPSoC static schedules through exploitation of logical-to-physical core mapping latitudeabstractThe computing engines of many current applications are powered by MPSoC platforms, which promise significant speedup but induce increased reliability problems as a result of ever growing integration density and chip size. While static MPSoC execution schedules deliver predictable worst-case performance, the absence of dynamic variability unfortunately constrains their usefulness in such an unreliable execution environment. Adaptive static schedules with predictable responses to run-time resource variations have consequently been proposed, yet the extra constraints imposed by adaptivity on task assignment have resulted in schedule length increases. We propose to eradicate the associated performance degradation of such techniques while retaining all the concomitant benefits, by exploiting an inherent degree of freedom in task assignment regarding the logical to physical core mapping. The proposed technique relies on the use of core reordering and rotation through utilizing a graph representation model, which enables a direction translation of inter-core communication paths into order requirements between cores. The algorithmic implementation results confirm that the proposed technique can drastically reduce the schedule length overhead of both pre- and post-reconfiguration schedules. Chengmo Yang, Alex Orailoglu |
DATE | 2 |
| 2009 | Processor reliability enhancement through compiler-directed register file peak temperature reductionabstractEach semiconductor technology generation brings us closer to the imminent processor architecture heat wall, with all its associated adverse effects on system performance and reliability. Temperature hotspots not only accelerate the physical failure mechanisms such as electromigration and dielectric breakdown, but furthermore make the system more vulnerable to timing-related intermittent failures. Traditional thermal management techniques suffer from considerable performance overhead as the entire processor needs to be stalled or slowed down to preclude heat accumulation. Given the significant temporal and spatial variations of the chip-wide temperature, we propose in this paper a technique that directly targets one of the resources that is most likely to overheat in current processors, namely, the register files. Instead of duplicating or physically distributing the register file, we suggest to attain power density control through exploiting the extant spatial slack associated with register file accesses. Based on application-specific access profiles, a compiler-directed register shuffling strategy is proposed to deterministically construct the logical to physical register mapping in a rotating manner. Simulation results confirm that the proposed technique attains, within a limited hardware budget and negligible performance degradation, effective reduction in peak temperature and hence in the expected fault rates for the entire chip. Chengmo Yang, Alex Orailoglu |
DSN | 2 |
| 2009 | Deflecting crosstalk by routing reconsideration through refined signal correlation estimationabstractCrosstalk between interconnects has become one of the major factors that tamper with VLSI signal integrity as device feature size scales down to UDSM and nanometer level. Traditional techniques for crosstalk reduction focus on reducing the coupling capacitance between interconnect nets to maximize the crosstalk slack among the nets. Although effective in reducing the worst-case crosstalk, such a strategy is incapable of minimizing the average run-time crosstalk which is equivalently critical for the timing and functional correctness of nanoscale VLSI. The minimization of both worst-case and average crosstalk necessitates the consideration of signal correlation information determined by circuit logic during the layout optimization stage. In this paper, a post-global routing technique is proposed to reduce the run-time crosstalk risk without violating the worst-case crosstalk bound specified by traditional techniques. A measure is proposed to accurately capture the signal correlation and model the run-time behavior of net pairs. Adjustment in the routing track assignment is performed under the guidance of run-time information to reduce the chance of neighboring nets having crosstalk-generating signal transitions. Meanwhile, the coupling length of each net is still controlled to within a specific bound for the maximization of the minimal crosstalk slack.Experimental results show that, compared to the conventional approaches, the proposed technique achieves significant reduction in average crosstalk without exacerbating the minimal crosstalk slack of the circuit. Mingjing Chen, Alex Orailoglu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | Scan power reduction in linear test data compression schemeabstractXOR network-based on-chip test compression schemes have been widely employed in large industrial scan designs due to their high compression ratio and efficient decompression mechanism. Nevertheless, such a scheme necessitates high unspecified bit ratios in the original test cubes, resulting in quite significant difficulties in preprocessing test cubes for scan power reduction. The linear mapping from the original cubes to the compressed seeds typically provides extra degrees of flexibility as multiple seeds may reconstruct the test cube. Appreciable power reductions in the decompressed test data can be attained through the pinpointing of the power-optimal seeds during the compression phase. The proposed work explores the aforementioned flexibility in the seed space, and proposes the mathematical and algorithmic framework for a power-aware linear test compression scheme. The proposed technique incurs no hardware overhead over the traditional linear compression scheme; it can be easily embedded furthermore into the industrial test compaction/compression flow. Experimental results confirm that the proposed technique delivers significant scan power reduction with negligible impact on the compression ratio. Mingjing Chen, Alex Orailoglu |
ICCAD | 2 |
| 2009 | Scan Cell Positioning for Boosting the Compression of Fan-Out Networks
Ozgur Sinanoglu, Mohammed Al-Mulla, Noora A. Shunaiber, Alex Orailoglu |
J. Comput. Sci. Technol. | 4 |
| 2009 | Guest Editorial Special Section on the IEEE Symposium on Application Specific Processors 2008abstractThe four papers in this special section were originally presented at the IEEE Symposium on Application Specific Processors 2008. Alex Orailoglu, Laura Pozzi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Low-Power Scan Testing for Test Data Compression Using a Routing-Driven Scan ArchitectureabstractA new scan architecture is proposed to reduce peak test power and capture power. Only a subset of scan flip-flops is activated to shift test data or capture test responses in any clock cycle. This can effectively reduce the capture test power and peak test power. Two routing-driven schemes are proposed to reduce the routing overhead. Experimental results show that the proposed scan architecture can effectively reduce peak test power, capture power, test data volume, and test application cost. Dianwei Hu, Qiang Xu 0001, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2008 | A light-weight cache-based fault detection and checkpointing scheme for MPSoCs enabling relaxed execution synchronizationabstractWhile technology advances have made MPSoCs a standard architecture for embedded systems, their applicability is increasingly being challenged by dramatic increases in the amount of device failures that may occur during execution. Conventional fault tolerance techniques employ a duplication-and-comparison strategy to detect arbitrary execution faults, as well as a checkpointing-and-rollback strategy to recover from the faulty state. Comparison and checkpointing are performed either at task level, thus imposing a large amount of overhead in verifying and backing up memory pages, or at instruction level, thus necessitating a lock-step execution model which significantly limits the attainable performance. To overcome the shortcomings of both strategies, in this paper we propose a cache-based fault tolerance scheme wherein the comparison and checkpointing process is performed at the cache-memory interface. By allowing two processors that execute duplicated tasks to share a single data cache, the proposed scheme is able to verify execution results before writing them back into memory, thus protecting the memory from being polluted by execution faults. This in turn significantly reduces the checkpointing overhead. Meanwhile, since only the data written into memory are compared, the strict instruction-by-instruction synchronization model used in multithreading processors can be relaxed. The simulation results confirm that the proposed scheme only imposes a performance overhead ranging from 1.4% to 10.4%, while both fault detection and execution checkpointing can be effectively attained. Chengmo Yang, Alex Orailoglu |
CASES | 2 |
| 2008 | Miss reduction in embedded processors through dynamic, power-friendly cache designabstractToday, embedded processors are expected to be able to run complex, algorithm-heavy applications that were originally designed and coded for general-purpose processors. As a result, traditional methods for addressing performance and determinism become inadequate. This paper explores a new data cache design for use in modern high-performance em-bedded processors that will dynamically improve execution time, power efficiency, and determinism within the system. The simulation results show significant improvement in cache miss ratios and reduction in power consumption of approxi-mately 30 % and 15%, respectively. Garo Bournoutian, Alex Orailoglu |
DAC | 2 |
| 2008 | Towards fault tolerant parallel prefix adders in nanoelectronic systemsabstractFuture nanoelectronics based arithmetic components will enjoy abundant hardware, yet at the same time confront severe unreliability challenges. We focus on the fault tolerance of high performance parallel prefix adders (PPA), and exploit the inherent redundancy in PPAs to develop efficient fault tolerance approaches. We show that the internal invariant inherent in the parallel prefix adders provides support for online fault detection and fault masking. Furthermore, based on the particular regular structure of PPAs, an online diagnosis scheme can be developed, thus enabling the application of reconfigurability of nanoelectronics for the highly flexible online repair approaches. In contrast to traditional fault tolerance techniques that rely solely on significant external overhead, the proposed approach opens up a new genre of efficient fault tolerance techniques for arithmetic components in the nanoelectronic environment. Wenjing Rao, Alex Orailoglu |
DATE | 2 |
| 2008 | Test cost minimization through adaptive test developmentabstractThe ever-increasing complexity of mixed-signal circuits imposes an increasingly complicated and comprehensive parametric test requirement, resulting in a highly lengthened manufacturing test phase. Attaining parametric test cost reduction with no test quality degradation constitutes a critical challenge during test development. The capability of parametric test data to capture systematic process variations engenders a highly accurate prediction of the efficiency of each test for a particular lot of chips even on the basis of a small quantity of characterized data. The predicted test efficiency further enables the adjustment of the test set and test order, leading to an early detection of faults. We explore such an adaptive strategy, by introducing a technique that prunes the test set based on a test correlation analysis. A test selection algorithm is proposed to identify the minimum set of tests that delivers a satisfactory defect coverage. A probabilistic measure that reflects the defect detection efficiency is used to order the test set so as to enhance the probability of an early detection of faulty chips. The test sequence is further optimized during the testing process by dynamically adjusting the initial test order to adapt to the local defect pattern fluctuations in the lot of chips under test. Experimental results show that the proposed technique delivers significant test time reductions while attaining higher test quality compared to previous adaptive test methodologies. Mingjing Chen, Alex Orailoglu |
ICCD | 2 |
| 2007 | Core-Based Testing of Multiprocessor System-on-Chips Utilizing Hierarchical Functional BusesabstractAn integrated test scheduling methodology for multiprocessor system-on-chips (SOC) utilizing the functional buses for test data delivery is described. The proposed methodology handles both flat bus single processor SOC and hierarchical bus multiprocessor SOC. It is based on a resource graph manipulation and a packet-based packet set scheduling methodology. The resource graph is decomposed into a set of test configuration graphs, which are then used to determine the optimum test configurations and test delivery schedule under a given power constraint. In order to validate the effectiveness of the proposed methodology, a number of experiments are run on several modified benchmark circuits. The results clearly underscore the advantages of the proposed methodology. Fawnizu Azmadi Hussin, Tomokazu Yoneda, Alex Orailoglu, Hideo Fujiwara |
ASP-DAC | 3 |
| 2007 | Improving Circuit Robustness with Cost-Effective Soft-Error-Tolerant Sequential ElementsabstractSoft errors induced by alpha particles and cosmic radiation have become a highly challenging problem in the design of UDSM or nanoscale circuits, making the incorporation of circuit hardening techniques essential. In this paper, a design technique for soft-error-tolerant sequential elements is presented to improve circuit robustness. The proposed technique exploits time and space redundancy using an elaborate flip-flop structure, and provides complete soft error immunity for both the transient faults generated in the combinatorial logic and the particle strikes inside the flip- flops. The proposed technique is developed to be compatible with current digital design technology, thus having minimal impact on design flow and hardware cost. Simulation results confirm the effectiveness of the proposed technique. Mingjing Chen, Alex Orailoglu |
ATS | 2 |
| 2007 | Light-weight synchronization for inter-processor communication acceleration on embedded MPSoCsabstractThe advances in semiconductor technologies have placed MPSoCscenter stage as a standard architecture for embedded applications of ever increasing complexity. Efficient utilization of the ample hardware resources requires applications to be decomposed into fine-grained threads, engendering in turn a large amount of interprocessor communications. While fine-grained on-chip interconnects can reduce the data transfer overhead, the traditional synchronization mechanisms, such as spin locks and barriers, still cause significant contention in polling shared variables. To overcome this issue, in this paper we propose a light-weight distributed synchronization mechanism which statically encodes the semantically correct order of accesses to each shared variable. A sharp reduction in the number of code bits is attained through a reference coloring algorithm, which furthermore enables an implementation within negligible hardware overhead. This light-weight synchronization mechanism allows dependent threads to frequently exchange data during execution, in turn enabling the exploration of fine-grained parallelism for applications with complex dependences. Chengmo Yang, Alex Orailoglu |
CASES | 2 |
| 2007 | Interactive presentation: Logic level fault tolerance approaches targeting nanoelectronics PLAsabstractA regular structure and capability to implement arbitrary logic functions in a two-level logic form have placed crossbar-based programmable logic arrays (PLAs) as promising implementation architectures in the emerging nanoelectronics environment. Yet reliability constitutes an important concern in the nanoelectronics environment, necessitating a thorough investigation and its effective augmentation for crossbar-based PLAs. We investigate in this paper fault masking for crossbar-based nanoelectronics PLAs. Missing nanoelectronics devices at the crosspoints have been observed as a major source of faults in nanoelectronics crossbars. Based on this observation, we present a class of fault masking approaches exploiting logic tautology in two-level PLAs. The proposed approaches enhance the reliability of nanoelectronics PLAs significantly at low hardware cost Wenjing Rao, Alex Orailoglu, Ramesh Karri |
DATE | 2 |
| 2007 | Fault Tolerant Approaches to Nanoelectronic Programmable Logic ArraysabstractProgrammable logic arrays (PLA), which can implement arbitrary logic functions in a two-level logic form, are promising as platforms for nanoelectronic logic due to their highly regular structure compatible with the nano crossbar architectures. Reliability is an important challenge as far as nanoelectronic devices are concerned. Consequently, it is necessary to focus on the fault tolerance aspects of nanoelectronic PLAs to ensure their viability as a foundation for nanoelectronic systems. In this paper, we investigate two types of fault tolerance techniques for nanoelectronic device based PLAs, focusing at the online faults occurring at the cross-points of nano devices. We develop a scheme to precisely locate the faults online, as this is a crucial step for efficient online reconfiguration based fault tolerance schemes. We also propose a tautology based fault masking scheme. We demonstrate that these two types of fault tolerance schemes developed for nano PLAs significantly improve at low hardware cost the reliability of the high fault occurrence nanoelectronic environment. Wenjing Rao, Alex Orailoglu, Ramesh Karri |
DSN | 2 |
| 2007 | Power efficient register file update approach for embedded processorsabstractIn this paper we present an approach for a low power register file in the domain of embedded processors. The suggested approach obtains power savings through tackling the unnecessary writes to register files for short live registers. Writes to register files are essentially redundant when an instruction manages to forward its results to all of its dependents through forwarding hardware. As the percentage of registers that exhibit short liveness is shown to be significant, tackling unnecessary writes contributes to delivering appreciable power savings. In this work we show that tackling the unnecessary writes could be attained efficiently through a register based encoding scheme. The suggested encoding scheme exploits application-specific information and renames all or most of the short live registers to a small subset of the registers that are prespecified during the hardware design. The renaming process is performed at the compiler level. Power savings can be obtained through precluding the set of prespecified registers from writing to the register file. We suggest in this paper efficient algorithms for the purpose of renaming, one algorithm to perform the renaming in the cases of no register pressure and another one for the cases of register pressure. In the cases of register pressure, some of the prespecified registers may need to be turned into normal registers, a process that is managed through the use of reprogrammable hardware support. Although the cases of register pressure could impact power savings, the detailed analysis we outline shows that the size of the prespecified registers subset is typically small which makes register pressure an infrequent event. Experimental analysis on numerical and DSP codes indicates appreciable improvements in power savings. Raid Ayoub, Alex Orailoglu |
ICCD | 2 |
| 2007 | Circuit-level mismatch modelling and yield optimization for CMOS analog circuitsabstractA methodology for constructing circuit-level mismatch models and performing yield optimization is presented for CMOS analog circuits. The methodology combines statistical techniques with direct investigation of circuit behavior, and achieves model simplification and computational efficiency while ensuring sufficient accuracy. The circuit-level mismatch model can be used in performance characterization and yield estimation, both important in providing information for circuit reliability analysis. The proposed yield optimization technique consists of constructing and refining a yield model over the designable parameters, and ensures fast convergence to the global optimal design. The experimental results on two representative circuits confirm the efficiency and effectiveness of the proposed method. Mingjing Chen, Alex Orailoglu |
ICCD | 2 |
| 2007 | Towards Nanoelectronics Processor Architectures
Wenjing Rao, Alex Orailoglu, Ramesh Karri |
J. Electron. Test. | 2 |
| 2007 | On the identification of modular test requirements for low cost hierarchical test path construction
Yiorgos Makris, Alex Orailoglu |
Integr. | 2 |
| 2006 | Power-efficient instruction delivery through trace reuseabstractAs power dissipation inexorably becomes the major bottleneck in system integration and reliability, the front-end instruction delivery path in a traditional out-of-order superscalar processor needs to deliver high application performance in an energy-effective manner. This challenge can be addressed by efficiently reusing the work of fetch and decode performed during preceding loop iterations and resident mostly within the processor itself. As a large percentage of the instructions currently under fetch have previously dispatched copies resident in the Reorder Buffer (ROB), in this paper we develop a mechanism to utilize the ROB as a storage location for previously decoded instructions. Thus instructions can be fed directly from the ROB into the rename and issue stages, enabling the gating off of the fetch and decode logic for large periods of time so as to deliver significant power savings. Power and performance criticality of the ROB requires an efficient reuse identification mechanism; we outline such a cost-efficient Reuse Identification Unit (RIU) which enables effective identification of the matches between the ROB entries and the instructions currently under fetch. Simulation results on both multimedia and SPEC 2000 benchmarks confirm that incorporating the proposed technique on traditional out-of-order superscalar processors results in not only a sight improvement in performance, but also significant savings in the overall system power dissipation, achieved within a limited hardware budget. Chengmo Yang, Alex Orailoglu |
PACT | 2 |
| 2006 | Power efficient branch prediction through early identification of branch addressesabstractEver increasing performance requirements have elevated deeply pipelined architectures to a standard even in the embedded processor domain, requiring the incorporation of dynamic branch prediction subsystems to hide the execution latency of control-altering instructions. In this paper a low power early branch identification technique which enables the design of extremely power-efficient branch predictors and BTBs is proposed. Through static extraction of program information regarding the distance to subsequent branches, this technique enables the calculation of the next branch address as soon as the direction of the current branch has been predicted. Early identification of branch addresses enables a complete elimination of the power hungry BTB lookups normally occurring at every execution cycle, as well as a just-in-time wake-up mechanism when accessing "hibernating" entries in complex predictors, switched to power-saving mode to reduce leakage power dissipation. A cost-efficient Branch Identification Unit (BIU) to calculate branch addresses is presented and analyzed in terms of power and timing characteristics. The effectiveness of the proposed BTB access policy and predictor wake-up mechanism is also confirmed by the simulation results of the SPECint 2000 and Media-bench benchmarks. Chengmo Yang, Alex Orailoglu |
CASES | 2 |
| 2006 | Topology aware mapping of logic functions onto nanowire-based crossbar architecturesabstractHighly regular, nanodevice based architectures have been proposed to replace pure CMOS based architectures in the emerging post CMOS era. Since bottom-up self-assembly is used to build these architectures, regular nanowire crossbars are emerging as a promising candidate. While these regular structures resemble CMOS programmable logic arrays (PLAs), PLA logic synthesis methodologies fail to solve the associated problems since the length and connectivity constraints imposed by individual nanowires in these crossbars translate into challenges hitherto not considered. These strict topological constraints should be considered while mapping Boolean functions onto nanowire crossbars during logic synthesis. We develop a mathematical model for this problem, an algorithm to solve it and three heuristics to improve the algorithm runtime. Wenjing Rao, Alex Orailoglu, Ramesh Karri |
DAC | 2 |
| 2006 | Fault Identification in Reconfigurable Carry Lookahead Adders Targeting Nanoelectronic FabricsabstractOnline repair through reconfiguration is a particularly advantageous approach in the nanoelectronic environment since reconfigurability is naturally supported by the devices. However, precise identification of faulty locations is of critical importance for fine-grain repairs. A CLA is mainly composed of: (1) carry generation blocks and (2) g,p signal generation blocks. In this paper we propose two schemes for fault identification in these two parts correspondingly. For carry generation blocks, an inherently redundant computation path is exploited to identify the faulty block with high precision. As a time redundancy approach, recomputation with rotated operands (RERO) has been utilized in online fault detection for CLA’s [13]. For g,p generation blocks, we exploit the RERO scheme to achieve precise fault identification. A comprehensive analysis is provided for the aliasing in the proposed fault identification approach. It is shown that both the amount of repair hardware overhead and the fault coverage loss for the proposed scheme are very low. Overall, the proposed scheme can perform fast and precise identification of faults in the CLA components with low area overhead, thus facilitating the development of powerful and efficient fault tolerance schemes through online repair for nanoelectronic systems Wenjing Rao, Alex Orailoglu, Ramesh Karri |
ETS | 2 |
| 2006 | Power-Constrained SOC Test Schedules through Utilization of Functional BusesabstractIn this paper, we are proposing a core-based test methodology that utilizes the functional bus for test stimuli and response transportation. An efficient algorithm for the generation of a complete test schedule that efficiently utilizes the functional bus under a power constraint is described. The test schedule is composed of a set of test vector delivery sequences in small chunks, denoted as packets. The utilization of small packet sizes optimizes the functional bus utilization. The experimental results show that the methodology is highly effective compared to previous approaches that do not use the functional bus. The strong results of the proposed approach are particularly highlighted when small bus widths are considered, an important consideration in current SOC designs where increasingly larger bus widths pose routing and reliability challenges. Fawnizu Azmadi Hussin, Tomokazu Yoneda, Alex Orailoglu, Hideo Fujiwara |
ICCD | 3 |
| 2006 | Decision Tree Based Mismatch Diagnosis in Analog CircuitsabstractMismatch is a critical consideration in analog circuit design. Knowledge of mismatch locations and an understanding of their impact on circuit performance are crucial for design optimization and process improvement. We present a circuit level mismatch diagnosis methodology in this paper. The functional parameters with abnormal values are measured as manifestations of mismatch, from which reverse tracing is employed to determine the mismatch source. The methodology is implemented on a representative benchmark and its efficiency confirmed by simulation results. Mingjing Chen, Hosam Haggag, Alex Orailoglu |
VTS | 3 |
| 2006 | Nanofabric Topologies and Reconfiguration Algorithms to Support Dynamically Adaptive Fault ToleranceabstractEmerging nanoelectronics are expected to have very high manufacture-time defect rates and operation-time fault rates. Traditional N-modular redundancy (NMR) exploits the large device densities offered by these nanoelectronics to tolerate these high fault rates by allocating redundant resources according to the worst case fault rates. However, this approach is inflexible when the fault rates are time varying. In this paper, we propose a dynamically adaptive NMR approach by developing: (i) a genre of nanofabric topologies that supports sharing of redundancies in the NMR approach so as to adapt to the time varying fault rates and (ii) reconfiguration algorithms for these topologies to deal with fault tolerance loss caused by manufacturing defects and operation-time online faults, respectively. Simulation results verify that the ability to construct reliable systems, possibly the paramount consideration in constructing working applications in nanoelectronics, is significantly improved with the proposed flexible NMR architecture and the reconfiguration algorithms. Wenjing Rao, Alex Orailoglu, Ramesh Karri |
VTS | 2 |
| 2005 | A unified transformational approach for reductions in fault vulnerability, power, and crosstalk noise & delay on processor busesabstractIn this paper we propose a coding scheme for general-purpose applications that can reduce power dissipation, crosstalk noise and crosstalk delay on the bus lines while simultaneously detecting errors at run time. The reduction in power dissipation can be achieved through reducing the bus switching activity. Not only is the switching activity in individual lines reduced but so is the coupling activity across the adjacent lines, the major contributor to the overall power dissipation in deep submicron technology. Detailed analysis of crosstalk noise and delay shows that eliminating certain patterns of transitions and reducing the infeasible ones in terms of crosstalk noise and power dissipation is a feasible strategy for alleviating these problems. We propose an encoding technique consisting of the use of predefined patterns of transitions, one for each possible combination of input data, to generate the codewords. The restriction to the predefined patterns of transitions enables fast encoding and low hardware overhead. This work presents an extensive analysis of the consequent reduction in crosstalk and power. SPICE derived experimental results show a reduction in worst case crosstalk delay and noise, ranging up to 24% and 10% respectively. Extensive experimental results for various applications show significant reduction in power dissipation ranging up to 44% for switching activity on the bus lines and up to 25% for coupling activity. The results also show a drastic reduction ranging up to 98% in the number of patterns that are most likely to produce crosstalk errors. Raid Ayoub, Alex Orailoglu |
ASP-DAC | 2 |
| 2005 | Fault tolerant nanoelectronic processor architecturesabstractIn this paper we propose a fault-tolerant processor architecture and an associated fault-tolerant computation model capable of fault tolerance in the nanoelectronic environment that is characterized by high and time varying fault rates. The proposed fault tolerant processor architecture not only guarantees the correctness of computation but also is flexible in that it dynamically trades-off computation resources and performance. The core of the architecture is a decentralized instruction control unit called the voter that achieves both fault tolerance and the maximum parallel execution of instructions by exploiting the abundant computational resources provided by nanotechnologies. Although the result of each instruction needs to be confirmed by executing it on multiple computation units, multiple unconfirmed instructions can proceed as speculative branches. The voter implements a hardware-frugal computation unit allocation algorithm to organize the redundant computations and to dynamically control the growth of speculative branches. Wenjing Rao, Alex Orailoglu, Ramesh Karri |
ASP-DAC | 2 |
| 2005 | Forward discrete probability propagation method for device performance characterization under process variationsabstractProcess variations are becoming influential at the device level in deep sub-micron and sub-wavelength design regimes, whereas they used to be a few generations away only influential at circuit level. Process variations cause device performance parameters, such as current or output resistance, to acquire a probability distribution. Estimation of these distributions has been accomplished using Monte Carlo techniques so far. The large number of samples needed by Monte Carlo methods adversely affects the possibility of integrating probabilistic device performance at the circuit level due to run-time inefficiency. In this paper, we introduce a novel technique called Forward Discrete Probability Propagation (FDPP). This method discretizes the probability distributions and effectively propagates these probabilities across a device formula hierarchy, such as the one present in the SPICE3v3 model. Consequently, probability distributions for process parameters are propagated to the device level. It is shown in the paper that with far fewer number of samples, comparable accuracy to a Monte Carlo method is achieved. Rasit Onur Topaloglu, Alex Orailoglu |
ASP-DAC | 2 |
| 2005 | Fault tolerant quantum cellular array (QCA) design using Triple Modular Redundancy with shifted operandsabstractDue to their extremely small feature sizes and ultra low power consumption, Quantum-dot Cellular Automata (QCA) technology is projected to be a promising nanotechnology. However, in nanotechnologies, manufacture time defect levels and operational time fault rates are expected to be quite high. Straightforward Triple Modular Redundancy (TMR) based fault tolerance is inappropriate for QCA nanotechnology since wire delays dominate the logic delays and faults in wires dominate the faults in a QCA based design. Furthermore, long wires are necessary in TMR based designs. In this paper we show that fault-tolerance can be obtained by using TMR with Shifted Operands (TMRSO). TMRSO uses shorter wires of QCA cells and exploits the self-latching property of clocked QCA arrays to provide the same level of fault tolerance capability as straightforward TMR while being significantly faster and smaller. This technique can be applied to a variety of operations; we have validated TMRSO on adders. Implementation results obtained using QCADesigner [6] show that an 8-bit adder using TMRSO has more than 50% area reduction and more than 100% throughput improvement when compared to a TMR implementation. Tongquan Wei, Kaijie Wu 0001, Ramesh Karri, Alex Orailoglu |
ASP-DAC | 4 |
| 2005 | Energy-effcient physically tagged caches for embedded processors with virtual memoryabstractIn this paper we present a low-power tag organization for physically tagged caches in embedded processors with virtual memory support. An exceedingly small subset of tag bits is identified for each application hot-spot so that only these tag bits are used for cache access with no performance sacrifice as they provide complete address resolution. The minimal subset of physical tag bits, i.e. the compressed tag, is dynamically updated following the changes in the physical address space of the application. Special support from the operating system (OS) is introduced in order to maintain the compressed tag during program execution. The compressed tag is updated by the OS to match the current set of physical memory pages allocated to the application. We have proposed efficient algorithms that are incorporated within the memory allocator and the dynamic linker in order to achieve dynamic update of the compressed tags in the cases where the mapping between virtual and physical addresses is modified; such cases include memory allocation/deallocation and swapping physical pages on the secondary memory storage. The only hardware support needed within the l/D-caches is the support for disabling bitlines of the tag arrays. An extensive set of experimental results demonstrates the efficacy of the proposed approach. Peter Petrov, Daniel Tracy, Alex Orailoglu |
DAC | 3 |
| 2005 | A DFT approach for diagnosis and process variation-aware structural test of thermometer coded current steering DACsabstractA design for test (DFT) hardware is proposed to increase the controllability of a thermometer coded current steering digital to analog converter. A procedure is introduced to reduce the diagnosis and structural test time from quadratic to linear using the proposed DFT hardware. To evaluate the applicability of the proposed technique, principal component analysis is used to create virtual process variations to simulate in lieu of semiconductor fabrication data. An architecture specific soft fault model is suggested for the diagnosis problem. Random errors according to the fault model are introduced in the virtual test environment on top of the process variations and it is shown that diagnosis of a fault is possible with high accuracy with the proposed method. The same technique employing principal component analysis is furthermore used to provide process variation-aware reference test comparison values for a structural test of the DAC. The structural test provides a mechanism to test for even unmodeled manufacturing faults. The process variation-aware test values help detect defects even under process variations. The proposed DFT hardware and method are low cost and quite suitable for a built-in self diagnosis and test implementation. Rasit Onur Topaloglu, Alex Orailoglu |
DAC | 2 |
| 2005 | Architectural-Level Fault Tolerant Computation in Nanoelectronic ProcessorsabstractNanoelectronic devices are expected to have extremely high and variable fault rates; thus future processor architectures based on these unreliable devices need to be built with fault tolerance embedded so as to satisfy the fundamental requirement of computational correctness. In this paper an architectural-level computation model is proposed for fault tolerant computations in nanoelectronic processors. The proposed scheme is capable of guaranteeing the correctness of each instruction through exploitation of both hardware and time redundancy, even under high and variable fault rates. Each instruction is confirmed by multiple computation instances. Through a speculative execution based on unconfirmed results, the proposed scheme eliminates the severe performance deterioration typically caused by time redundancy approaches on data dependent instructions. To avoid the exponential growth of resource allocation introduced by the hardware redundancy approaches on the speculations, a hardware allocation framework is developed in the proposed scheme to control the growth of hardware resources while preserving the low latency achieved through the speculative executions. We set up an experimental framework to validate the effectiveness of the proposed scheme as well as to investigate multiple tradeoff points within the proposed approach. Experimental data further confirm that the proposed approach achieves the goal of providing fault tolerance in the pipelined nanoelectronic processors, while at the same time providing high system performance and efficient utilization of hardware resources. Wenjing Rao, Alex Orailoglu, Ramesh Karri |
ICCD | 2 |
| 2005 | Efficient RT-Level Fault Diagnosis
Ozgur Sinanoglu, Alex Orailoglu |
J. Comput. Sci. Technol. | 2 |
| 2005 | The Construction of Optimal Deterministic Partitionings in Scan-Based BIST Fault Diagnosis: Mathematical Foundations and Cost-Effective ImplementationsabstractPartitioning techniques enable identification of fault-embedding scan cells in scan-based BIST. We introduce deterministic partitioning techniques capable of resolving the location of the fault-embedding scan cells. We outline a complete mathematical analysis that identifies the class of deterministic partitioning structures and complement this rigorous mathematical analysis with an exposition of the appropriate cost-effective implementation techniques. We validate the superiority of the deterministic techniques both in an average-case sense by conducting simulation experiments and in a worst-case sense through a thorough mathematical analysis. Ismet Bayraktaroglu, Alex Orailoglu |
IEEE Trans. Computers | 2 |
| 2005 | A reprogrammable customization framework for efficient branch resolution in embedded processorsabstractWe present a customization framework for embedded processors which employs the utilization of application-specific information, thus specializing the processor's microarchitecture to the application needs. The increased processor utilization leads to a low-cost system implementation with no sacrifice in performance requirements and to reduced custom hardware in a typical SOC. We illustrate these ideas through the branch resolution problem, known to impose severe performance degradation on control-dominated embedded applications. A customization approach for early branch resolution and subsequent folding is presented. The application-specific information is captured by the microarchitecture through a low-cost reprogrammable hardware, thus attaining the twin benefits of processor standardization and application-specific customization. Experimental results show that for a representative set of control-dominated applications a reduction in the range of 3--22% in processor cycles can be achieved, thus extending the scope of low-cost embedded processors in complex codesigns for control intensive systems. Peter Petrov, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2005 | Test power reductions through computationally efficient, decoupled scan chain modificationsabstractSOC test time minimization hinges on the attainment of core test parallelism; yet test power constraints hamper this parallelism as excessive power dissipation may damage the SOC being tested. We propose a test power reduction methodology for SOC cores through scan chain modification. By inserting logic gates between scan cells, a given set of test vectors & captured responses is transformed into a new set of inserted stimuli & observed responses that yield fewer scan chain transitions. In identifying the best possible scan chain modification, we pursue a decoupled strategy wherein test data are decomposed into blocks, which are optimized for power in a mutually independent manner. The decoupled handling of test data blocks not only ensures significantly high levels of overall power reduction but it furthermore delivers computational efficiency at the same time. The proposed methodology is applicable to both fully, and partially specified test data; test data analysis in the latter case is performed on the basis of stimuli-directed controllability measures which we introduce. To explore the tradeoff between the test power reduction attained by the proposed methodology & the computational cost, we carry out an analysis that establishes the relationship between block granularity & the number of scan chain modifications. Such an analysis enables the utilization of the proposed methodology in a computationally efficient manner, while delivering solutions that comply with the stringent area & layout constraints in SOC as well. Ozgur Sinanoglu, Alex Orailoglu |
IEEE Trans. Reliab. | 2 |
| 2004 | Efficient RT-level fault diagnosis methodology
Ozgur Sinanoglu, Alex Orailoglu |
ASP-DAC | 2 |
| 2004 | On mismatch in the deep sub-micron era - from physics to circuits
Rasit Onur Topaloglu, Alex Orailoglu |
ASP-DAC | 2 |
| 2004 | CircularScan: A Scan Architecture for Test Cost ReductionabstractScan-based designs are widely used to decrease the complexity of the test generation process; nonetheless, they increase test time and volume. A new scan architecture is proposed to reduce test time and volume while retaining the original scan input count. The proposed architecture allows the use of the captured response as a template for the next pattern with only the necessary bits of the captured response being updated while observing the full captured response. The theoretical and experimental analysis promises a substantial reduction in test cost for large circuits. Baris Arslan, Alex Orailoglu |
DATE | 2 |
| 2004 | Scan Power Minimization through Stimulus and Response TransformationsabstractScan-based cores impose considerable test power challenges due to excessive switching activity during shift cycles. The consequent test power constraints force SOC designers to sacrifice parallelism among core tests, as exceeding power thresholds may damage the chip being tested. Reduction of test power for SOC cores can thus increase the number of cores that can be tested in parallel, improving significantly SOC test application time. In this paper, we propose a scan chain modification technique that inserts logic gates on the scan path. The consequent beneficial test data transformations are utilized to reduce the scan chain transitions during shift cycles and hence test power. We introduce a matrix band algebra that models the impact of logic gate insertion between scan cells on the test stimulus and response transformations realized. As we have successfully modeled the response transformations as well, the methodology we propose is capable of truly minimizing the overall test power. The test vectors and responses are analyzed in an intertwined manner, identifying the best possible scan chain modification, which is realized at minimal area cost. Experimental results justify the efficacy of the proposed methodology as well. Ozgur Sinanoglu, Alex Orailoglu |
DATE | 2 |
| 2004 | Pipelined test of SOC cores through test data transformationsabstractAttaining parallelism among core tests is of crucial importance to the reduction of SOC test costs. In this paper, we propose an SOC test methodology that enhances SOC testapplication throughput with no increase in test pin requirements. In the proposed methodology,the test vector of a core is formed in its scan chain by transforming the response of the preceding core; logic gates inserted between the core scan cells transform the response ofthe preceding core into the core test vector. The consequent core tests can be thought of as being pipelined, thus reducing the time spent for the delivery of the test vectors into the scan cells of the cores being tested in parallel, and hence increasing the throughput of SOC test application. The proposed algorithmic framework identifies the cost-effective hardware that maps the responses of the preceding core onto a maximal number of core test vectors through the utilization of effiient test vector and scan cell reordering heuristics; the impact of these techniques is modeled, enabling their utilization along with the aforementioned transformation techniques. We furthermore investigate various scan chain configuration techniques to enhance the pipeline efficiency, thus minimizing the pipeline period and theSOC test time. The efficacy of the proposed methodology translates into enhanced parallelism in testing SOC cores. Ozgur Sinanoglu, Alex Orailoglu |
ETS | 2 |
| 2004 | Design space exploration for aggressive test cost reduction in CircularScan architecturesabstractScan-based designs effectively reduce test generation complexity and thus deliver improved fault coverage. Nevertheless, the traditional scan architectures suffer from increased test time and test data volume. The CircularScan architecture (Arslan and Orailoglu) provides a flexible environment for test cost reduction. The new scan design enables the use of the captured response of the previously applied test pattern as a template. The subsequent pattern is loaded by efficiently performing the necessary changes on the template through the functionality provided by the new architecture, conceptually exploiting the inherent low specified bit density of the test patterns. We explore the space of possible design alternatives built on the CircularScan architecture; the design alternatives are presented with accompanying test application methods. The experimental results indicate a substantial test cost reduction, reaching 90% levels. The proposed scheme is not only easily scalable but also promises further reductions in test cost when applied to large state of the art ICs. Baris Arslan, Alex Orailoglu |
ICCAD | 2 |
| 2004 | Frugal linear network-based test decompression for drastic test cost reductionsabstractIn This work we investigate an effective approach to construct a linear decompression network in the multiple scan chain architecture. A minimal pin architecture, complemented by negligible hardware overhead, is constructed by mathematically analysing test data relationships, delivering in turn drastic test reductions. The proposed network drives a large number of internal scan chains with a short input vector, thus allowing significant reductions in both test time and test volume. The proposed method constructs an inverter-interconnect based network by exploring the pairwise linear dependencies of the internal scan chain vectors, resulting in a very low cost network that is nonetheless capable of outperforming much costlier compression schemes. We propose an iterative algorithm to construct the network from an initial set of test cubes. The experimental data shows significant reductions in test time and test volume with no loss of fault coverage. Wenjing Rao, Alex Orailoglu, George Su |
ICCAD | 2 |
| 2004 | Extending the Applicability of Parallel-Serial Scan DesignsabstractAlthough scan-based designs are widely used in order to reduce the complexity of test generation, test application time and test data volume are substantially increased. We propose two different methodologies for test cost reduction in scan-based designs. The first methodology improves on the Illinois scan architecture, aiming at reducing the high test cost of the test vectors that necessitate the serial test application mode. The second methodology employs on-chip serial transformations to generate an input stimulus that can be applied efficiently. The transformation-based methodology utilizes the proposed scan design to obtain the minimal cost input stimulus. The experimental results indicate that a substantial test cost reduction, reaching 90% levels, can be obtained. Baris Arslan, Ozgur Sinanoglu, Alex Orailoglu |
ICCD | 3 |
| 2004 | End-to-End Testability Analysis and DfT Insertion for Mixed-Signal PathsabstractIncreasing system complexity and test cost demands new system-level solutions for mixed-signal systems. In this paper, we present a testability analysis and DfT insertion methodology for end-to-end mixed-signal paths. Based on behavioral models and path analysis, testability problems in the path are determined and classified in terms of their bottleneck. Possible solutions to each problem are identified. The DfT insertion problem is then formulated as a min-cost set cover problem to achieve the most cost-efficient solution. In experimental results where test point insertion is used as the DfT approach, nearly 50% reduction in the overall DfT overhead is achieved. Sule Ozev, Alex Orailoglu |
ICCD | 2 |
| 2004 | Test Cost Reduction Through A Reconfigurable Scan ArchitectureabstractScan-based designs are widely used to keep test generation complexity within practical limits; nevertheless, scan-based design substantially increases test application time and test data volume. A novel scan-based design is proposed to reduce the test cost. The new scan-design exploits the low specified bit density of the test sets. The circular structure of the proposed architecture enables the use of the captured response of the previously applied pattern as a template for the subsequent pattern while allowing the full observation of the captured response. The functionality provided by the new architecture is utilized to update the template quickly to obtain the next pattern. The experimental results show a substantial reduction in test cost, reaching 90% levels. Baris Arslan, Alex Orailoglu |
ITC | 2 |
| 2004 | Fault Tolerant Arithmetic with Applications in Nanotechnology based SystemsabstractSeveral emerging nanotechnologies have been displaying the negative differential resistance (NDR) characteristic, which makes them naturally support multi-valued logic with a large number of logic states. Such multi-valued logic with a large number of logic states can support a native digit-level redundant number system and hence a native digit-level carry save arithmetic. We present a new approach to linear block code based fault-tolerant arithmetic in NDR nanotechnologies. Specifically, we show how linear block codes can be used for error checking and error correction in carry save arithmetic operations. The proposed approach significantly improves timing and fault-tolerance of arithmetic operations in the highly unreliable nanoelectronic environment. Since digit-level information redundancy via linear block codes is widely used for fault tolerant communications and storage systems, the proposed scheme also unifies the fault tolerance approaches across arithmetic, interconnection and storage subsystems. Wenjing Rao, Alex Orailoglu, Ramesh Karri |
ITC | 2 |
| 2004 | Autonomous Yet Deterministic Test of SOC CoresabstractIncreased core test parallelism translates into reduced SOC test application time; yet the availability of a limited number of tester channels hampers this parallelism. Furthermore, the test vectors to be delivered into core scan chains need to be stored in the tester memory, imposing considerable costs on SOC tests. We propose an SOC test methodology delivering all the benefits of core self-test, while ensuring fault coverage levels identical to those attained in deterministic test. In the proposed methodology, a single LFSR broadcasts pseudo-random patterns to each core; the LFSR patterns are transformed into the actual test vectors of a core while they are being shifted into the core scan chain. The transformation is realized through the logic gates inserted between the core scan cells. The efficacy and the cost-effectiveness of the proposed methodology reflects into significantly reduced test costs. Ozgur Sinanoglu, Alex Orailoglu |
ITC | 2 |
| 2004 | Searching for Global Test Costs Optimization in Core-Based Systems
Érika F. Cota, Luigi Carro, Marcelo Lubaszewski, Alex Orailoglu |
J. Electron. Test. | 4 |
| 2004 | Fast and energy-frugal deterministic test through efficient compression and compaction techniques
Ozgur Sinanoglu, Alex Orailoglu |
J. Syst. Archit. | 2 |
| 2004 | Tag compression for low power in dynamically customizable embedded processorsabstractWe present a methodology for power reduction by instruction/data cache-tag compression for low-power embedded processors. By statically analyzing the code/data memory layouts for the application hot spots, a variety of proposed schemes for effective tag-size reduction can be employed for power minimization in instruction and data caches. The schemes rely on significantly reducing the number of tag bits stored in the tag arrays for cache-conflict identification, thus considerably decreasing the number of active bitlines, sense amps, and comparator cells. We present a set of tag compression techniques and evaluate each of them separately in terms of efficiency and required hardware support. A detailed very large scale integrated implementation has been performed and a number of experimental results on a set of embedded applications is reported for each technique. Energy dissipation decreases of up to 95% can be observed for the tag arrays, implying significant energy reductions in the range of 50% when amortized across the overall cache subsystem. Peter Petrov, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Enhancing reliability of RTL controller-datapath circuits via Invariant-based concurrent testabstractWe present a low-cost concurrent test methodology for enhancing the reliability of RTL controller-datapath circuits, based on the notion of path invariance. The fundamental observation supporting the proposed methodology is that the inherent transparency behavior of RTL components, typically utilized for hierarchical off-line test, renders rich sources of invariance within a circuit. Furthermore, additional sources of invariance are obtained by examining the algorithmic interaction between the controller, and the datapath of the circuit. A judicious selection & combination of modular transparency functions, based on the algorithm implemented by the controller-datapath pair, yields a powerful set of invariant paths in a design. Compliance to the invariant behavior is checked whenever the latter is activated. Thus, such paths enable a simple, yet very efficient concurrent test capability, achieving fault security in excess of 90% while keeping the hardware overhead below 40% on complicated, difficult-to-test, sequential benchmark circuits. By exploiting fine-grained design invariance, the proposed methodology enhances circuit reliability, and contributes a low-cost concurrent test direction, applicable to general RTL circuits. Yiorgos Makris, Ismet Bayraktaroglu, Alex Orailoglu |
IEEE Trans. Reliab. | 3 |
| 2004 | Design of concurrent test Hardware for Linear analog circuits with constrained hardware overheadabstractConcurrent detection of failures in analog circuits is becoming increasingly more important as safety-critical systems become more widespread. A methodology for automatic design of concurrent failure detection circuitry for linear analog systems is discussed in this paper. The desired hardware bound is specified as a constraint; the methodology aims at providing coverage in terms of all the circuit components while minimizing the loading overhead by reducing the number of internal circuit nodes that need to be tapped. Parameter tolerances are incorporated through either statistical or mathematical analysis to determine the threshold for failure alarm. Sule Ozev, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Low-power instruction bus encoding for embedded processorsabstractAbstract—This paper presents a low-power encoding framework for embedded processor instruction buses. The encoder is capable of adjusting its encoding not only to suit applications but furthermore to suit different aspects of particular program execution. It achieves this by exploiting application-specific knowledge regarding program hot-spots, and thus identifies efficient instruction transformations so as to minimize the bit transitions on the instruction bus lines. Not only is the switching activity on the individual bus lines considered but so is the coupling activity across adjacent bus lines, a foremost contributor to the total power dissipation in the case of nanometer technologies. Low-power codes are utilized in a reprogrammable application specific manner. The restriction to two well-selected classes of simply computable, functional transformations delivers significant storage benefits and ease of reprogrammability, in the process obtaining significant power savings. The microarchitectural support enables reprogrammability of the encoding transformations in order to track code particularities effectively. Such reprogrammability is achieved by utilizing small tables that store relevant application information. The few transformations that result in optimal power reductions for each application hot-spot are selected by utilizing short indices stored into a table, which is accessed only once at the beginning of the transformed bit sequence. Extensive experimental results show significant power reductions ranging up to 80 % for switching activity on bus lines and up to 70 % when bus coupling effects are also considered. I. Peter Petrov, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Extracting Precise Diagnosis of Bridging Faults from Stuck-at Fault InformationabstractAlthough the stuck-at fault model is the standard fault model. the frequently occurring faults in some technologies arc unintentional shorts, denoted as bridging faults. We outline a method that utilizes the information from the stuck-at fault model to accurately diagnose the bridging faults that affect two lines. The proposed method exploits the observation that the bridging fault response matches the stuck-at fault responses on the shorted lines for the failing test vectors and generates a candidate list that accounts for all failures. A further reduction in the size of the candidate set is achieved by extracting information from the test vectors that do not fail. The proposed method uses no layout information whatsoever. Nonetheless, the experimental results indicate that the bridging faults can be accurately diagnosed delivering a reduction in the sizes of the ambiguity sets and full capture of the offending bridging fault. Baris Arslan, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2003 | Test Data Manipulation Techniques for Energy-Frugal, Rapid Scan TestabstractScan-based testing methodologies remedy the testability problem of sequential circuits; yet they suffer from prolonged test time and excessive test power due to numerous shift operations. The significant correlation among test stimuli along with the high density of unspecified bits in test data enables the utilization of the existing test stimulus in the scan chain as the seed for the generation of the subsequent test stimulus, thus reducing both test time and test data volume. The proposed scan-based test scheme accesses only a subset of scan cells for loading the subsequent test stimulus while freezing the remaining scan cells with the preceding test stimulus, thus decreasing scan chain transitions during shift operations. The proposed scan architecture is coupled with test data manipulation techniques which include test stimuli ordering and partitioning algorithms, boosting test time reductions. The experimental results confirm the significant reductions in test application time, test data volume and test power achieved by the proposed scan-based testing methodology. Ozgur Sinanoglu, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2003 | Test application time and volume compression through seed overlappingabstractWe propose in this paper an extension on the Scan Chain Concealment technique to further reduce test time and volume requirement. The proposed methodology stems from the architecture of the existing SCC scheme, while it attempts to overlap consecutive test vector seeds, thus providing increased flexibility in exploiting effectively the large volume of don't-care bits in test vectors. We also introduce modified ATPG algorithms upon the previous SCC scheme and explore various implementation strategies. Experimental data exhibit significant reductions on test time and volume over all current test compression techniques. Wenjing Rao, Ismet Bayraktaroglu, Alex Orailoglu |
DAC | 3 |
| 2003 | Power Efficiency through Application-Specific Instruction Memory Transformations
Peter Petrov, Alex Orailoglu |
DATE | 2 |
| 2003 | Virtual Compression through Test Vector Stitching for Scan Based Designs
Wenjing Rao, Alex Orailoglu |
DATE | 2 |
| 2003 | Customizable Embedded Processor ArchitecturesabstractIn this paper, we present a framework for dynamic application customization for high-performance and low-power embedded processors. The proposed architecture is capable of utilizing application information to boost the performance and lower the power consumption of the most important microarchitectural components such as instruction/data caches and the memory subsystem. We present a design framework, including CAD support infrastructure and reprogrammable hardware support, for a dynamically customizable microarchitecture. We outline the underlying algorithms for compile-time extraction of the utilized application properties and we present the architectural principles of the hardware support. Extensive experimental results confirm the efficacy of this novel embedded processor architecture. Peter Petrov, Alex Orailoglu |
DSD | 2 |
| 2003 | Low-power Branch Target Buffer for Application-Specific Embedded ProcessorsabstractIn this paper we present a methodology for a low-power branch identification mechanism, which enables the design of extremely power efficient branch predictors for embedded processors. The proposed technique utilizes application-specific information regarding the control-flow structure of the program major loops. Such information is used to completely eliminate the power hungry branch target buffer (BTB) lookups which normally occur at every execution cycle. Exact application knowledge regarding the control-flow structure of the program obviates the power expensive BTB operations, thus enabling the utilization of contemporary branch predictors in high-end, yet power-sensitive embedded processors. The utilization of exact application knowledge results not only in the complete elimination of the power hungry BTB structure but also in a perfect branch and target address identification. Cost-efficient and programmable hardware architecture for capturing the control-flow structure of the program is presented thereafter. The hardware complexity of the proposed architecture is carefully analyzed in terms of power, performance and area overhead. The proposed technique delivers power reductions in excess of 90% for a set of embedded benchmarks. Peter Petrov, Alex Orailoglu |
DSD | 2 |
| 2003 | Hierarchical Constraint Conscious RT-level Test GenerationabstractThe increasing complexity of ICs necessitates the use of test generation methodologies at higher levels of abstraction. We propose a computationally efficient RT-level test generation methodology that utilizes a divide and conquer approach. The hierarchical constraints for the module under test are identified through the proposed justification and propagation analysis. These constraints are then taken into account during the local test vector generation for the module under test, enabling the identification of the local test vectors that are guaranteed to be effective not only at the module-level but also at the system-level as well. High quality test sets are thus generated by the proposed methodology in a computationally efficient manner. Experimental results verify the performance boosts attained by the proposed methodology as well. Ozgur Sinanoglu, Alex Orailoglu |
DSD | 2 |
| 2003 | Compiler-Based Register Name Adjustment for Low-Power Embedded Processors
Peter Petrov, Alex Orailoglu |
ICCAD | 2 |
| 2003 | Partial Core Encryption for Performance-Efficient Test of SOCs
Ozgur Sinanoglu, Alex Orailoglu |
ICCAD | 2 |
| 2003 | Virtual Page Tag Reduction for Low-power TLBsabstractWe present a methodology for a power-optimized, software-controlled translation lookaside buffer (TLB) organization. A highly reduced number of virtual page number (VPN) bits sufficient to perform physical address translation is efficiently identified and used when performing TLB lookups, delivering significant power reductions. Information regarding the virtual address space of the program code and data provided by the compiler is augmented with information regarding the dynamically linked libraries and data allocated run-time by the loader, the dynamic linker, and the memory manager. The hardware support needed is constrained to disabling bitlines of the tag arrays associated to the 1-TLB and the D-TLB. Algorithms for identifying the reduced VPNs for power optimized TLB operations together with the required OS support are presented. Peter Petrov, Alex Orailoglu |
ICCD | 2 |
| 2003 | Aggressive Test Power Reduction Through Test Stimuli TransformationabstractExcessive switching activity during shift cycles in scan-based cores imposes considerable test power challenges. To ensure rapid and reliable test of SOCs, we propose a scan chain modification methodology that transforms the stimuli to be inserted to the scan chain through logic gate insertion between scan cells, reducing scan chain transitions. We introduce a novel matrix band algebra to formulate the impact of scan chain modifications on test stimuli transformations. Based on this analysis, we develop algorithms for transforming a set of test vectors into power-optimal test stimuli through cost-effective scan chain modifications. Experimental results show that scan-in power reductions exceeding 90% for test vectors and 99.5% for test cubes can be attained by the proposed methodology. Ozgur Sinanoglu, Alex Orailoglu |
ICCD | 2 |
| 2003 | Modeling Scan Chain Modifications For Scan-in Test Power Minimization
Ozgur Sinanoglu, Alex Orailoglu |
ITC | 2 |
| 2003 | Decompression Hardware Determination for Test Volume and Time Reduction through Unified Test Pattern Compaction and CompressionabstractA methodology for the determination of decompression hardware that guarantees complete fault coverage for a unified compaction/compression scheme is proposed. Test cube information is utilized for the determination of a near optimal decompression hardware. The proposed scheme attains simultaneously high compression levels and reduced pattern counts through a linear decompression hardware. Significant test volume and test application time reductions are delivered through the scheme we propose while a highly cost effective hardware implementation is retained. Ismet Bayraktaroglu, Alex Orailoglu |
VTS | 2 |
| 2003 | Statistical Tolerance Analysis for Assured Analog Test Coverage
Sule Ozev, Alex Orailoglu |
J. Electron. Test. | 2 |
| 2003 | Reducing Average and Peak Test Power Through Scan Chain Modification
Ozgur Sinanoglu, Ismet Bayraktaroglu, Alex Orailoglu |
J. Electron. Test. | 3 |
| 2003 | Concurrent Application of Compaction and Compression for Test Time and Data Volume Reduction in Scan DesignsabstractA test pattern compression scheme for test data volume and application time reduction is proposed. While compression reduces test data volume, the increased number of internal scan chains due to an on-chip, fixed-rate decompressor reduces test application time proportionately. Through on-chip decompression, both the number of virtual scan chains visible to the ATE and the functionality of the ATE are retained intact. Complete fault coverage is guaranteed by constructing the decompression hardware deterministically through analysis of the test pattern set. Ismet Bayraktaroglu, Alex Orailoglu |
IEEE Trans. Computers | 2 |
| 2002 | Test Requirement Analysis for Low Cost Hierarchical Test Path ConstructionabstractWe propose a methodology that examines design modules and identifies appropriate vector justification and response propagation requirements for reducing the cost of hierarchical test path construction. Test requirements are defined as a set of fine-grained input and output bit clusters and pertinent symbolic values. They are independent of actual test sets and are adjusted to the inherent module connectivity and regularity. As a result, they combine the generality required for fast hierarchical test path construction with the precision necessary for minimizing the incurred cost, thus fostering cost-effective hierarchical test. Yiorgos Makris, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2002 | Gate Level Fault Diagnosis in Scan-Based BISTabstractA gate level, automated fault diagnosis scheme is proposed for scan-based BIST designs. The proposed scheme utilizes both fault capturing scan chain information and failing test vector information and enables location identification of single stuck-at faults to a neighborhood of a few gates through set operations on small pass/fail dictionaries. The proposed scheme is applicable to multiple stuck-at faults and bridging faults as well. The practical applicability of the suggested ideas is confirmed through numerous experimental runs on all three fault models. Ismet Bayraktaroglu, Alex Orailoglu |
DATE | 2 |
| 2002 | Test Planning and Design Space Exploration in a Core-Based EnvironmentabstractThis paper proposes a comprehensive model for test planning in a core-based environment. The main contribution of this work is the use of several types of TAMs and the consideration of different optimization factors (area, ping and test time) during the global TAM and test schedule definition. This expansion of concerns makes possible an efficient yet fine-grained search in the huge design space of a reuse-based environment. Experimental results clearly show the variety of trade-offs that can be explored using the proposed model, and its effectiveness on optimizing the system test design. Érika F. Cota, Luigi Carro, Marcelo Lubaszewski, Alex Orailoglu |
DATE | 4 |
| 2002 | Power Efficient Embedded Processor Ip's through Application-Specific Tag Compression in Data CachesabstractIn this paper, we present a methodology for power minimization by data cache tag compression. The set of tags being accessed by the major application loops is analyzed statically during compile time and an efficient and optimal compression scheme is proposed Only a very limited number of tag bits are stored in the tag array for cache conflict identification, thus achieving a significant reduction in the number of active bitlines, sense amps, and comparator cells. The underlying hardware support for dynamically compressing the tags consists of a highly cost and power efficient programmable encoder which lies outside the cache access path, thus not affecting the processor cycle time. A detailed VLSI implementation has been performed and a number of experimental results on a set of embedded applications and numerical kernels is reported Energy dissipation decreases of up to 95% can be observed for the tag arrays, while significant energy reductions in the range of 10%-50% are observed when amortized across the overall cache subsystem. Peter Petrov, Alex Orailoglu |
DATE | 2 |
| 2002 | Reducing Test Application Time Through Test Data Mutation EncodingabstractIn this paper we propose a new compression algorithm geared to reduce the time needed to test scan-based designs. Our scheme compresses the test vector set by encoding the bits that need to be flipped in the current test data slice in order to obtain the mutated subsequent test data slice. Exploitation of the overlap in the encoded data by effective traversal search algorithms results in drastic overall compression. The technique we propose can be utilized as not only a stand-alone technique but also can be utilized on test data already compressed, extracting even further compression. The performance of the algorithm is mathematically analyzed and its merits experimentally confirmed on the larger examples of the ISCAS '89 benchmark circuits. Sherief Reda, Alex Orailoglu |
DATE | 2 |
| 2002 | A novel scan architecture for power-efficient, rapid testabstractScan-based testing methodologies remedy the testability problem of sequential circuits; yet they suffer from prolonged test time and excessive test power due to numerous shift operations. The high density of the unspecified bits in test data enables the utilization of the test response data captured in the scan chain for the generation of the subsequent test stimulus, thus reducing both test time and test data volume. The proposed scan-based test scheme accesses only a subset of scan cells for loading the subsequent test stimulus while freezing the remaining scan cells with the response data captured, thus decreasing the scan chain transitions during shift operations. The experimental results confirm the significant reductions in test application time, test data volume and test power achieved by the proposed scan-based testing methodology. Ozgur Sinanoglu, Alex Orailoglu |
ICCAD | 2 |
| 2002 | Fault Dictionary Size Reduction through Test Response SuperpositionabstractThe exceedingly large size of fault dictionaries constitutes a fundamental obstacle to their usage. We outline a new method to reduce significantly, the size of fault dictionaries. The proposed method partitions the test set and a combined signature is stored for each partition. The new approach aims to provide high diagnostic resolution with a small number of combined signatures. The experimental results show a considerable decrease in the storage requirement of fault dictionaries. Baris Arslan, Alex Orailoglu |
ICCD | 2 |
| 2002 | Cost-Effective Concurrent Test Hardware Design for Linear Analog CircuitsabstractConcurrent detection of failures in analog circuits is becoming increasingly more important as safety-critical systems become more widespread. A methodology for the automatic design of concurrent failure detection circuitry for linear analog systems is discussed in this paper In contrast to previous approaches, the methodology aims at providing coverage in terms of all the circuit components while minimizing the loading overhead by reducing the number of internal circuit nodes that need to be tapped Parameter tolerances are incorporated through either statistical or mathematical analysis to determine the threshold for failure alarm. Experimental results confirm that full coverage can be attained while keeping the hardware overhead within a pre-specified budget. Sule Ozev, Alex Orailoglu |
ICCD | 2 |
| 2002 | Scan Power Reduction Through Test Data Transition Frequency AnalysisabstractSignificant reductions in test application times can be achieved through parallelizing core tests; however, simultaneous test of various cores may result in exceeding power thresholds, endangering the SoC being tested. Test power dissipation is exceedingly high in scan-based environments wherein scan chain transitions during the shift of test data further reflect into significant levels of circuit switching unnecessarily. Scan chain modification helps mitigate this problem as it enables the reduction of transitions in the test stimuli to be inserted to the modified scan chain and in the response to be collected through the scan-out pin. The proposed modifications in the scan chain consist of inverter insertion and scan cell reordering, leading to significant power reductions with neither area nor performance penalty whatsoever A computationally efficient algorithm is presented to identify the optimal scan chain modification based on the transition frequency analysis of the test data. Experimental results confirm the considerable reductions in scan chain transitions. The consequent reduced power dissipation possible under the proposed scheme enables rapid, reliable testing of SoCs. Ozgur Sinanoglu, Ismet Bayraktaroglu, Alex Orailoglu |
ITC | 3 |
| 2002 | Boosting the Accuracy of Analog Test Coverage Computation through Statistical Tolerance AnalysisabstractIncreasing numbers of analog components in today's systems necessitate system level test composition methods that utilize onchip capabilities rather than solely relying on costly DFT approaches. We outline a tolerance analysis methodology for test signal propagation to be utilized in hierarchical test generation for analog circuits. A detailed justification of this proposed novel tolerance analysis methodology is undertaken by comparing our results with detailed SPICE Monte-Carlo simulation data on several combinations of analog modules. The results of our experiments confirm the high accuracy and efficiency of the proposed tolerance analysis methodology. Sule Ozev, Alex Orailoglu |
VTS | 2 |
| 2002 | Test Power Reduction through Minimization of Scan Chain TransitionsabstractParallel test application helps reduce the otherwise considerable test times in SOCs; yet its applicability is limited by average and peak power considerations. The typical test vector loading techniques result infrequent transitions in the scan chain, which in turn reflect into significant levels of circuit switching unnecessarily. Judicious utilization of logic in the scan chain can help reduce transitions while loading the test vector needed. No performance degradation ensues as scan chain modifications have no impact on functional execution. A computationally efficient scheme is proposed to identify, the location and type of the logic to be inserted. The experimental results confirm the significant reductions in test power possible under the proposed scheme. Ozgur Sinanoglu, Ismet Bayraktaroglu, Alex Orailoglu |
VTS | 3 |
| 2002 | Fast Hierarchical Test Path Construction for Circuits with DFT-Free Controller-Datapath Interface
Yiorgos Makris, Jamison Collins, Alex Orailoglu |
J. Electron. Test. | 3 |
| 2002 | Microarchitectural synthesis of performance-constrained, low-power VLSI designsabstractNew portable signal-processing applications such as mobile telephony, wireless computing, and personal digital assistants place stringent power consumption limits on their constituent components. Substantial power savings can be realized if 5 V designs are translated to use the new lower supply voltage standards. This conversion, however, is not achieved easily: a design originally targeted for implementation in a 5 V technology will typically require significant rework to meet timing and throughput requirements at the lower operating voltage. In this paper we describe a high-level synthesis system which assists the designer in performing this task, minimizing the need for manual redesign. Techniques employed in this work include pipelining and a new approach to module selection that minimizes power consumption subject to timing constraints. Using these and other high-level synthesis techniques to target designs to 3.3 V libraries, we show that it is possible to reduce power consumption by as much as 56% as compared to the original 5 V implementation, while meeting specified minimum throughput and maximum latency constraints. Laurence Goodby, Alex Orailoglu, Paul M. Chau |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2001 | Faults in Processor Control Subsystems: Testing Correctness and Performance Faults in the Data Prefetching UnitabstractThe processor control subsystems have for a long time been recognized as a bottleneck in the process of achieving complete fault coverage through various functional test propagation approaches. The difficult-to-test corner cases are further accentuated in fault-resilient control subsystems as no functional effect is incurred as a result of the fault, even though performance suffers. We investigate the construction of software programs, capable of providing full fault coverage at minimal hardware cost, for one such fault resilient subsystem in processor architecture: the data prefetching unit. Experimental results confirm the efficacy of the proposed method. Sobeeh Almukhaizim, Peter Petrov, Alex Orailoglu |
Asian Test Symposium | 3 |
| 2001 | Selecting a PRPG: Randomness, Primitiveness, or Sheer Luck?abstractThe ability of randomness to constitute a quality measure in PRPG selection for logic BIST is investigated. Extensive correlation analyses performed on a rich set of pattern generators and benchmark circuits indicate that higher randomness is no guarantee of improved fault coverage in LFSM-based pseudo-random pattern generators. Further evaluation of fault coverage data indicates that the performance of PRPGs is dependent on both circuit particularities and the number of test patterns employed. Ismet Bayraktaroglu, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2001 | Compaction Schemes with Minimum Test Application TimeabstractTesting embedded cores in a System-On-a-Chip (SoC) necessitates the use of a test access mechanism, which provides for transportation of the test data between the chip and the core I/Os. To relax the requirements on the test access mechanism at the core output side, we outline a space and time compaction scheme which minimizes test application time and required test bandwidth at the same time. We formulate the constraints on a mathematical basis for no aliasing compaction circuitry. The proposed compaction scheme is applicable to both combinational and sequential circuits. The experimental results illustrate that not only test application time is minimized but furthermore the associated area overhead is low as well. Ozgur Sinanoglu, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2001 | Test Volume and Application Time Reduction Through Scan Chain ConcealmentabstractA test pattern compression scheme is proposed in order to reduce test data volume and application time. The number of scan chains that can be supported by an ATE is significantly increased by utilizing an on-chip decompressor. The functionality of the ATE is kept intact by moving the decompression task to the circuit under test. While the number of virtual scan chains visible to the ATE is kept small, the number of internal scan chains driven by the decompressed pattern sequence can be sinificantly increased. Ismet Bayraktaroglu, Alex Orailoglu |
DAC | 2 |
| 2001 | Speeding Up Control-Dominated Applications through Microarchitectural Customizations in Embedded ProcessorsabstractWe present a methodology for microarchitectural customization of embedded processors by exploiting application information, thus attaining the twin benefits of processor standardization and application-specific customization. Such powerful techniques enable increased application fragments to be placed on the processor, with no sacrifice in system requirements, thus reducing the custom hardware and the concomitant area requirements in SOCs. We illustrate these ideas through the branch resolution problem, known to impose severe performance degradation on control-dominated embedded applications. A low-cost late customizable hardware that uses application information to fold out a set of frequently executed branches is described. Experimental results show that for a representative set of control dominated applications a reduction in the range of 7%-22% in processor cycles can be achieved, thus extending the scope of low-cost embedded processors in complex co-designs for control intensive systems. Peter Petrov, Alex Orailoglu |
DAC | 2 |
| 2001 | Diagnosis for scan-based BIST: reaching deep into the signaturesabstractFor partitioning-based diagnosis in a scan-based BIST environment, an exact analysis scheme, capable of identifying all scan cells that receive incorrect data, is proposed. In contrast to previously suggested approaches, the scheme we propose identifies all failing scan cells with no ambiguity whatsoever. Not only do we resolve failing scan cells unambiguously, but we do so at the earliest possible instance through reexamination of already computed signatures. Intensive utilization of this highly precise diagnostic state information leads to prognostic information regarding the usefulness of running upcoming tests which in turn leads to reductions in diagnosis time in excess of 30% compared to previous approaches. Ismet Bayraktaroglu, Alex Orailoglu |
DATE | 2 |
| 2001 | Testability implications in low-cost integrated radio transceivers: a Bluetooth case studyabstractAs the use of wireless communications in daily life increases, attaining low-cost solutions becomes increasingly important due to shrinking profit margins. Cost optimization that solely targets at minimization of the cost of system architecture may result in suboptimal, highly untestable, solutions. Test design and design for testability need to be incorporated into the system design flow to achieve viable solutions. This paper presents an analysis of test requirements, implications and test cost for low-cost Bluetooth systems. Testability problems are identified and possible solutions along with avenues to reduce the test cost by utilizing lower-cost testers are discussed. Christian Olgaard, Sule Ozev, Alex Orailoglu |
ITC | 3 |
| 2001 | Space and time compaction schemes for embedded coresabstractTesting embedded cores in a system-on-a-chip necessitates the use of a test access mechanism, which provides for transportation of the test data between the chip and the core I/Os. We outline an aliasing-free space and time compaction scheme, for both combinational and sequential cores, which minimizes the required test bandwidth and reduces the bandwidth consumption of the test access mechanism at the core output side. The experimental results show that the test bandwidth gain is achieved with no appreciable increase in test application time. Ozgur Sinanoglu, Alex Orailoglu |
ITC | 2 |
| 2001 | Efficient Transparency Extraction and Utilization in Hierarchical TestabstractWe introduce a methodology for identifying transparency behavior appropriate for hierarchical test, based on the theoretical principles of transparency composition. Unlike high level approaches that identify limited, coarse transparency behavior, the proposed methodology is capable of extracting a wide class of fine grained transparency functions for arbitrary sub-word bit clusters. The functions in this class can furthermore be rapidly extracted on the fly and efficiently utilized for hierarchical test translation, thus alleviating the exponential extraction time and storage space requirements of exhaustive approaches. The twin benefits of rapid, automated extraction coupled with the expansion of utilizable transparency scope deliver reduced DFT while enabling cost-effective hierarchical test of high quality. Yiorgos Makris, Alex Orailoglu |
VTS | 3 |
| 2001 | RT-level Fault Simulation Based on Symbolic PropagationabstractThe rapid rise in size and complexity of VLSI circuits has stimulated a need to handle fault simulation at higher levels of abstraction. We outline an RT-level fault simulation technique that utilizes symbolic data to group fault effects. Experimental results show that the proposed methodology provides superior speed-ups and accurate fault coverages. Ozgur Sinanoglu, Alex Orailoglu |
VTS | 2 |
| 2001 | Concurrent test for digital linear systemsabstractInvariant-based concurrent test schemes can provide economical solutions to the problem of concurrent error detection. An invariant-based concurrent error-detection scheme for linear digital systems is proposed. The cost of concurrent error-detection hardware is appreciably reduced due to utilization of a time-extended invariant, which extends the error-checking computation over time and, thus, reduces hardware requirements. Error-detection capabilities of the scheme proposed in this work are analyzed and conditions on the implementation for achieving complete fault coverage are outlined. Implementations fulfilling such conditions have been shown through experiments to provide 100% concurrent fault detection. Ismet Bayraktaroglu, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Performance and power effectiveness in embedded processors customizable partitioned cachesabstractThis paper explores an application-specific customization technique for the data cache, one of the foremost area/power consuming and performance determining microarchitectural features of modern embedded processors. The automated methodology for. customizing the processor microarchitecture that we propose results in increased performance, reduced power consumption and improved determinism of critical system parts while the fixed design ensures processor standardization. The resulting improvements help to enlarge the significant role of embedded processors in modern hardware-software codesign techniques by leading to increased processor utilization and reduced hardware cost. A novel methodology for static analysis and a microarchitecturally field-reprogrammable implementation of a customizable cache controller that implements a partitioned cache structure is proposed. Partitioning the load/store instructions eliminates cache interference; hence, precise knowledge about the hit/miss behavior of the references within each partition becomes available, resulting in significant reduction in tag reads and comparisons. Moreover, eliminating cache interference naturally leads to a significant reduction in the miss rate. The paper presents an algorithm for defining cache partitions, hardware support for customizable cache partitions, and a set of experimental results. The experimental results indicate significant improvements in both power consumption and miss rate. Peter Petrov, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2000 | Accumulation-based concurrent fault detection for linear digital state variable systemsabstractAn algorithmic fault detection scheme for linear digital state variable systems is proposed. The proposed scheme eliminates the necessity of observing the internal states of the system for concurrent fault detection by utilizing an accumulation-based approach. Observation merely of the inputs and the outputs results in significantly reduced area overhead and no performance penalty. Experimental results verify that 100% concurrent fault detection is attainable for linear digital state variable systems. Ismet Bayraktaroglu, Alex Orailoglu |
Asian Test Symposium | 2 |
| 2000 | Fast hierarchical test path construction for DFT-free controller-datapath circuitsabstractWe discuss a hierarchical test generation method for DFT-free controller-datapath pairs. A transparency based scheme is devised for the datapath, wherein locally generated vectors are translated into global design test. The controller is examined through influence tables, used to generate valid control state sequences for testing each module through hierarchical test paths. Fault coverage levels and vector counts thus attained match closely, those of traditional test generation methodologies, while sharply reducing the corresponding computational cost. Yiorgos Makris, Jamison Collins, Alex Orailoglu |
Asian Test Symposium | 3 |
| 2000 | Improved fault diagnosis in scan-based BIST via superpositionabstractAn improved approach for diagnosis of scan-based BIST designs is proposed. The enhancement in diagnosis is achieved by utilizing the superposition principle. Scan cells are partitioned pseudorandomly for observation and the ones provably fault free are removed from the potentially faulty list. Diagnostic resolution is improved by a novel application of the superposition principle, resulting in significant reductions in diagnosis time. 1 Ismet Bayraktaroglu, Alex Orailoglu |
DAC | 2 |
| 2000 | Test Quality and Fault Risk in Digital Filter Datapath BISTabstractAn objective of DSP testing should be to ensure that any errors due to missed faults are infrequent compared to a circuit's intrinsic errors, such as overflow. A method is proposed for quantifying test quality for digital filters by measuring the risk associated with any untested faults. Techniques for finding upper bounds on fault activation rates under worst-case operating conditions are described. These techniques enable test designers to objectively discriminate significant missed faults from near-redundant faults, which are unlikely to be activated in normal operation of the device. This complements fault coverage as a measure of test quality, providing a means of locating high-risk missed faults even in very high coverage test regimes. Laurence Goodby, Alex Orailoglu |
DATE | 2 |
| 2000 | Test Synthesis for Mixed-Signal SOC PathsabstractHigher levels of integration, the need for test re-use, and the mixed-signal nature of today's SOC's necessitate hierarchical test generation and system level test composition to meet stringent market requirements. In this paper a novel methodology for testing analog and digital components in a signal path is discussed. Consequent testability analysis can be utilized to reduce DFT requirements, while test translation provides highly effective low cost test. The proposed approach seamlessly propagates test information across the analog/digital divide. Experimental results substantiate the effectiveness of the proposed mixed-signal test synthesis methodology. Sule Ozev, Ismet Bayraktaroglu, Alex Orailoglu |
DATE | 3 |
| 2000 | Cost effective digital filter design for concurrent testabstractInvariant-based concurrent test schemes can provide economical solutions to the problem of concurrent testing of digital filters. Design methodologies for digital filters ensuring concurrent testability are outlined. Experimental results confirm through fault simulation 100% fault coverage within area cost comparable to that of DfT for off-line test and with error detection latency well below human response times. Ismet Bayraktaroglu, Alex Orailoglu |
ICASSP | 2 |
| 2000 | Unifying methodologies for high fault coverage concurrent and off-line test of digital filtersabstractA low-cost on-line test scheme for digital filters that additionally provides an off-line BIST solution is proposed. The scheme utilizes an invariant of the digital filter in order to detect on-line possible circuit malfunctions. The on-line checking hardware is shared with off-line BIST. The analysis performed indicates that exact 100% fault secureness is attained when the digital filter is designed according to design criteria that we identify in the paper. Furthermore, fault simulations show near 100% fault coverage for off-line BIST. Ismet Bayraktaroglu, Alex Orailoglu |
ISCAS | 2 |
| 2000 | Transparency-based hierarchical test generation for modular RTL designsabstractWe discuss a novel hierarchical test generation methodology for RTL designs, based on the concept of modular transparency. We introduce the channel notion, a powerful mechanism that captures modular transparency in terms of bijection functions defined on variable bitwidth signal entities. Through a recursive search algorithm, transparency channels are further combined into reachability paths suitable for translating local test vectors for each module into global design test. A divide and conquer hierarchical test generation methodology is described, resulting in significant test generation time speed-up and comparable fault coverage and vector count to complete circuit gate-level ATPG. Yiorgos Makris, Jamison Collins, Alex Orailoglu, Praveen Vishakantaiah |
ISCAS | 3 |
| 2000 | Deterministic partitioning techniques for fault diagnosis in scan-based BISTabstractA deterministic partitioning technique for fault diagnosis in scan-based BIST is proposed. Properties of high quality partitions for improved fault diagnosis times are identified and low cost hardware implementations of high quality deterministic partitions are outlined. The superiority of the partitions generated by the proposed approach is confirmed through mathematical analysis. Theoretical analyses, worst case bounds, and experimental simulation data all confirm the superiority of the proposed deterministic approaches. Ismet Bayraktaroglu, Alex Orailoglu |
ITC | 2 |
| 2000 | Invariance-Based On-Line Test for RTL Controller-Datapath CircuitsabstractWe present a low-cost on-line test methodology for RTL controller-datapath pairs, based on the notion of path invariance. The fundamental observation supporting the proposed methodology is that the transparency behavior inherent in RTL components renders rich sources of invariance in a design. Furthermore, the algorithmic controller-datapath interaction provides additional sources of invariance. A judicious selection and combination of modular transparency, based on the algorithm implemented by the controller-datapath pair, yields a powerful set of invariant paths. Such paths enable a simple, yet very efficient on-line test capability, achieving fault security in excess of 90% while keeping the hardware overhead below 40% on complicated, difficult to test, benchmarks. Yiorgos Makris, Ismet Bayraktaroglu, Alex Orailoglu |
VTS | 3 |
| 2000 | Test Selection Based on High Level Fault Simulation for Mixed-Signal SystemsabstractMixed-signal design and test tools are failing to keep pace with the increasing necessity for design exploration in the early stages. We outline a methodology and toolset to enable test selection at the early design stages by providing a high level fault simulator and associated block-level modeling and traversal capabilities. Experimental results show that the outlined methodology provides superior fault simulation speed-ups while helping to minimize the test time for a mixed-signal receiver system. Sule Ozev, Alex Orailoglu |
VTS | 2 |
| 2000 | On-line test for fault-secure fault identificationabstractIn an increasing number of applications, reliability is essential. On-line resistance to permanent faults is a difficult and important aspect of providing reliability. Particularly vexing is the problem of fault identification. Current methods are either domain specific or expensive. We have developed a fault-secure methodology for permanent fault identification through algorithmic duplication without necessitating complete functional unit replication. Fault identification is achieved through a unique binding methodology during high-level synthesis based on an extension of parity-like error correction equations in the domain of functional units. The result is an automated chip-level approach with extremely low area and cost overhead. Samuel Norman Hamilton, Alex Orailoglu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Self Recovering Controller and Datapath CodesignabstractAs society has become more reliant on electronics, the need for fault tolerant ICs has increased. This has resulted in significant research into both fault tolerant controller design, and mechanisms for datapath fault tolerance insertion. By treating these two issues separately, previous work has failed to address compatibility issues, as well as efficient codesign methodologies. In this paper, we present a unified approach to detecting control and datapath faults through the datapath, along with a method for fault identification and reconfiguration. By detecting control faults in the datapath, we avoid the area and performance overhead of detecting control faults through duplication or error checking codes. The result is a complete design methodology for self recovering architectures capable of far more efficient solutions than previous approaches. Samuel Norman Hamilton, Alex Orailoglu, Andre Hertwig |
DATE | 2 |
| 1999 | Channel-Based Behavioral Test Synthesis for Improved Module ReachabilityabstractWe introduce a novel behavioral test synthesis methodology that attempts to increase module reachability, driven by powerful global design path analysis. Based on the notion of transparency channels, test justification and propagation bottlenecks are revealed for each module in the design. Subsequently the proposed behavioral test synthesis scheme eliminates during scheduling, allocation and binding, as many reachability bottlenecks, as possible. Furthermore, it identifies the control states and provides the templates required for translating each module's test into global design rest. We demonstrate our scheme on a representative example, unveiling the potential of path analysis based techniques to accurately identify and eliminate module reachability bottlenecks, thus guiding behavioral rest synthesis. Yiorgos Makris, Alex Orailoglu |
DATE | 2 |
| 1999 | Low-Cost On-Line Test for Digital FiltersabstractA low-cost on-line test scheme for digital filters is proposed. The scheme uses an invariant of the digital filter, the frequency response at specific points, in order to detect possible malfunctioning of the circuit. The analysis performed indicates that 100% fault secureness is possible, if certain design constraints are followed. Ismet Bayraktaroglu, Alex Orailoglu |
VTS | 2 |
| 1999 | Redundancy and testability in digital filter datapathsabstractTest issues in application-specific digital filter datapaths are investigated. It is found that such designs can contain hundreds of redundant faults, making it difficult to accurately determine fault coverage. Since these redundant faults tend to appear in the same general location as test-resistant faults, the presence of many redundant faults can hide significant untested faults despite high overall test coverage. Classes of redundant faults that arise in digital filter datapaths are described, and we propose a suite of techniques for identifying and eliminating the most common redundancies based on arithmetic optimization. The approach is suitable as a front-end to more accurate fault simulation, or can be used in the design process to eliminate redundant logic. The approach is validated as a tool for developing very high-coverage built-in self-test circuits, showing that 100% fault coverage can be achieved in 24k-gate filters with as little as 1% area overhead. When used as a datapath optimization technique, the average area reduction over 15 designs was 8.9%, compared with moderately optimized designs. As a front-end to fault simulation, the approach yielded a 97.9% average reduction in the number of undetected faults across the 15 designs. Laurence Goodby, Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | An Examination of PRPG Selection Approaches for Large, Industrial DesignsabstractA study for selecting effective pseudo-random pattern generators (PRPGs) for large industrial designs is undertaken. For the PRPG selection process, guidance from a set of ISCAS benchmark circuits is initially sought. The set of benchmark circuits used is shown to be ineffective in providing material guidance. The alternative of PRPG selection through actual design experimentation is examined and its weaknesses identified and outlined. A brief DFT analysis is outlined to indicate the importance of early PRPG selection. Ismet Bayraktaroglu, K. Udawatta, Alex Orailoglu |
Asian Test Symposium | 3 |
| 1998 | Concurrent Error Recovery with Near-Zero Latency in Synthesized ASICsabstractThe importance of fault tolerant design has been steadily increasing as reliance on error free electronics continues to rise in critical military, medical, and automated transportation applications. While rollback and checkpointing techniques facilitate area efficient fault tolerant designs, they are inapplicable to a large class of time-critical applications. We have developed a novel synthesis methodology that avoids rollback, and provides both zero reduction in throughput and near-zero error latency. In addition, our design techniques reduce power requirements associated with traditional approaches to fault tolerance. Samuel Norman Hamilton, Alex Orailoglu |
DATE | 2 |
| 1998 | DFT guidance through RTL test justification and propagation analysisabstractWe introduce a formal mechanism for capturing test justification and propagation related behavior of blocks. Based on the identified test translation behavior, an RTL testability analysis methodology for hierarchical designs is derived. An algorithm for pinpointing the local-to-global test translation controllability and observability bottlenecks is presented. The analysis results are validated through an ATPG-based experimental flow and the applicability of the scheme for addressing test challenges in large designs by guiding DFT decisions is discussed. Yiorgos Makris, Alex Orailoglu |
ITC | 2 |
| 1998 | RTL Test Justification and Propagation Analysis for Modular Designs
Yiorgos Makris, Alex Orailoglu |
J. Electron. Test. | 2 |
| 1998 | On-Line Fault Resilience Through Gracefully Degradable ASICs
Alex Orailoglu |
J. Electron. Test. | 1 |
| 1997 | Frequency-Domain Compatibility in Digital Filter BISTabstractWe examine frequency-domain issues in the design and selectionof on-chip test generators for built-in self-test (BIST) of high-performancedigital filters. Test-generator/circuit compatibility isidentified as a significant factor in testing large filters. A fault-injectionexperiment is used to show that when an incompatible testgenerator is used, high fault coverage (over 99%) does not guaranteethat all serious faults will be detected. The frequency-domaincharacteristics of some basic test generation schemes are examined,and guidelines for test generator selection are proposed. Analyticaltechniques for identifying frequency-related testability problemsare discussed, and several test generation schemes are evaluated byfault simulating them against lowpass, bandpass, and highpass filters.A mixed test generation scheme is shown to reduce the numberof untested faults by a factor of two to three over a standard linear-feedbackshift-register (LFSR) based test scheme, at little added cost. Laurence Goodby, Alex Orailoglu |
DAC | 2 |
| 1997 | Microarchitectural synthesis for rapid BIST testingabstractThe impact of testability on design cost necessitates its consideration during the earliest stages of synthesis. Built-in self test (BIST) is an accepted testing approach, but its application to many designs is limited by the long test application time required to achieve high fault coverage. This work addresses the problem of BIST test time for high fault coverage by targeting test concurrency during high-level and structural synthesis. High-level synthesis generates RTL circuits which guarantee concurrent controllability and observability of all hardware components from test registers. Structural synthesis for testability completes the microarchitectural definition by specifying the test registers in the circuit, and defining a BIST test plan for the circuit. Remaining test conflicts are avoided without reducing test throughput by using partial-intrusion BIST to enable test data to be pipelined through nontest registers. The use of pipelined BIST testing in conjunction with high-level synthesis for test conflict minimization enables reduced test time through high test concurrency. Partial-intrusion BIST reduces the number of test registers by using nontest registers to propagate test data. Both behavioral and structural synthesis are directed by testability metrics which measure the test concurrency of a design. The use of these metrics in an integrated behavioral synthesis system pioneers the inclusion of test concurrency issues in microarchitectural synthesis. Experimental results using this synthesis system, SYNCBIST, show that designs generated by this approach are testable with few patterns, and using few test registers. Alex Orailoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | Pseudorandom-Pattern Test Resistance in High-Performance DSP DatapathsabstractThe testability of basic DSP datapath structures using pseudorandom built-in self-test techniques is examined.The addition of variance mismatched signals is identified as a testing problem, and the associated fault detection probabilities are derived in terms of signal probability distributions.A method of calculating these distributions is described, and it is shown how these distributions can be used to predict testing problems that arise from the correlation properties of test sequences generated using linear-feedback shift registers.Finally, it is shown empirically that variance matching using associativity transformations can reduce the number of untested faults by a factor of eight over variance mismatched designs. Laurence Goodby, Alex Orailoglu |
DAC | 2 |
| 1996 | Variance mismatch: identifying random-test resistance in DSP datapathsabstractPseudorandom built-in self-test (BIST) is an attractive means of testing DSP datapath structures as long as the delay and area overhead can be kept low. One obstacle to low-overhead BIST is random-pattern test-resistant datapath structures. We examine the causes of this resistance in some typical DSP datapaths, and introduce variance mismatch as an analytical tool for identifying these test problems. By understanding the mechanisms behind random-pattern test resistance, it is possible to make design decisions that are compatible with pseudorandom BIST, resulting in reduced test length and higher fault coverage. Variance matching theory is applied to the design of two large FIR filters, resulting in an 88% reduction in the number of missed faults for a fixed test length, and two orders-of-magnitude reduction in test length for a specific fault coverage target. Laurence Goodby, Alex Orailoglu |
ICASSP | 2 |
| 1996 | Microarchitectural synthesis of gracefully degradable, dynamically reconfigurable ASICsabstractIn this paper, we propose a novel fault-tolerance scheme, band reconfiguration, to handle multiple permanent faults in functional units of general ASIC designs. An associated high-level synthesis procedure that automatically generates such fault-tolerant systems is also presented. The proposed scheme permits multiple levels of graceful degradation. During each reconfiguration the system instantly reconfigures itself through operation rescheduling and hardware rebinding. The design objectives are optimization of resource utilization rate under each configuration, and reduction of hardware and performance overheads. The proposed high-level synthesis approach enables fast and area-effective implementations of gracefully degradable ASICs. Alex Orailoglu |
ICCD | 1 |
| 1996 | Can Defect-Tolerant Chips Better Meet the Quality Challenge?
R. L. Campbell, P. Kuekes, David Y. Lepejian, Wojciech Maly, Michael Nicolaidis, Alex Orailoglu |
VTS | 6 |
| 1996 | Automatic Synthesis of Self-Recovering VLSI SystemsabstractWe describe an integrated system for synthesizing self-recovering microarchitectures called /spl Sscr//spl Yscr//spl Nscr//spl Cscr//spl Escr//spl Rscr//spl Escr/ in the /spl Sscr//spl Yscr//spl Nscr//spl Cscr//spl Escr//spl Rscr//spl Escr/ model for self-recovery, transient faults are detected using duplication and comparison, while recovery from transient faults is accomplished via checkpointing and rollback. /spl Sscr//spl Yscr//spl Nscr//spl Cscr//spl Escr//spl Rscr//spl Escr/ initially inserts checkpoints subject to designer specified recovery time constraints. Subsequently, /spl Sscr//spl Yscr//spl Nscr//spl Cscr//spl Escr//spl Rscr//spl Escr/ incorporates detection constraints by ensuring that two copies of the computation are executed on disjoint hardware. Towards ameliorating the dedicated hardware required for the original and duplicate computations, /spl Sscr//spl Yscr//spl Nscr//spl Cscr//spl Escr//spl Rscr//spl Escr/ imposes intercopy hardware disjointness at a sub-computation level instead of at the overall computation level. The overhead is further moderated by restructuring the pliable input representation of the computation. /spl Sscr//spl Yscr//spl Nscr//spl Cscr//spl Escr//spl Rscr//spl Escr/ has successfully derived numerous self-recovering microarchitectures. Towards validating the methodology for designing fault-tolerant VLSI ICs, we carried out a physical design of a self-recovering 16-point FIR filter. Alex Orailoglu, Ramesh Karri |
IEEE Trans. Computers | 1 |
| 1996 | Time-constrained scheduling during high-level synthesis of fault-secure VLSI digital signal processorsabstractAdvances in VLSI technology are making it feasible to pack millions of transistors on a single chip. A consequent increase in the number of on-chip faults as well as the growing importance of quality-metrics such as reliability and fault-tolerance are making on-chip fault-tolerance mandatory. On-chip realization of a computation is fault-secure if an observable error in the computation is detected. Components used in life-critical systems should be secured against all faults. While fault-security can be realized by duplicating the computation on disjoint hardware and voting on the result(s), such straightforward strategies entail appreciable hardware overhead. This paper presents computer-aided behavioral synthesis of fault-secure microarchitectures which require less than proportional increase in hardware. The strategy selects intermediate computations for additional voting. The resulting class of fault-secure microarchitectures supplants the enormous hardware requirements of naive fault-secure strategies with enhanced hardware utilization afforded by securing the intermediate computations. Experimental results show that fault-security can be implemented at a less than proportional increase in hardware overhead. Ramesh Karri, Alex Orailoglu |
IEEE Trans. Reliab. | 2 |
| 1995 | Towards 100% Testable FIR Digital FiltersabstractTestability problems that arise in the design of fixed-coefficient finite impulse response (FIR) filters are examined. A class of redundant faults that naturally derive from the structure and behavior of these filters are examined, and design-for-test (DFT) techniques based on scaling theory are used to eliminate the redundancies. Eliminating these redundancies makes it possible for built-in self-test (BIST) approaches to reach 100% coverage, and automatic test-pattern generation (ATPG) based approaches can benefit by more than an order of magnitude reduction in test generation time. A case study provides a demonstration of the approach. Laurence Goodby, Alex Orailoglu |
ITC | 2 |
| 1995 | Testability metrics for synthesis of self-testable designs and effective test plansabstractWe propose a set of unified metrics for self-testability which are portable across different phases of synthesis. Furthermore, applicability of the proposed test metrics is verified through extensive experiments on benchmark designs. Mahsa Vahidi, Alex Orailoglu |
VTS | 2 |
| 1994 | Microarchitectural Synthesis of VLSI Designs with High Test ConcurrencyabstractThe testability of a VLSI design is strongly aected by its register-transfer level (RTL) structure.Since the high-level synthesis process determines the RTL structure, it is necessary to consider testability during high-level synthesis.A synthesis system composed of scheduling and binding components minimizes the number of hardware sharing con icts between tests in the test schedule.Novel test con ict estimates are used to direct the synthesis process.The test con ict estimation is based on examination of the in terconnect structure of the partial design state during synthesis.Test con ict estimates enable our synthesis system to select design options which increase test concurrency, thereby decreasing test time.Experimen tal results show that designs generated by this approach are testable in a highly concurrent manner. Ian G. Harris, Alex Orailoglu |
DAC | 2 |
| 1994 | Area-Efficient Fault Detection During Self-Recovering Microarchitecture SynthesisabstractWe will present the area-efficient fault-detection synthesis component of SYNCERE, an integrated system for synthesizing area-efficient self-recovering microarchitectures. In the SYNCERE model for self-recovery, transient fault detection is based on duplication and comparison, while recovery from transient faults is accomplished via checkpointing and rollback. SYNCERE minimizes the overhead of duplication using two complementary area-optimization techniques. Whereas imposing inter-copy hardware disjointness at a sub-computation level instead of at the overall computation level ameliorates the dedicated hardware required for the original and duplicate computations, restructuring the pliable input representation of the duplicate computation further moderates the overall hardware. Ramesh Karri, Alex Orailoglu |
DAC | 2 |
| 1994 | Simulated annealing based yield enhancement of layoutsabstractThis paper presents DEFT, a system for synthesizing defect-tolerant layouts, that in-grains tolerance to fabrication induced defects. This is accomplished by dispersing nets with large overlaps into nonadjacent tracks. DEFT also affords trade-offs between area (measured as the number of tracks) and yield of the resulting layout. The defect-tolerant layouts synthesized by DEFT have been consistently superior to those generated by other layout synthesis systems.> Ramesh Karri, Alex Orailoglu |
Great Lakes Symposium on VLSI | 2 |
| 1994 | Microarchitectural Synthesis of Performance-Constrained, Low-Power VLSI DesignsabstractNew portable signal-processing applications such as mobile telephony, wireless computing, and personal digital assistants place stringent power consumption limits on their constituent components. Substantial power servings can be realized if 5 V designs are translated to use the new lower supply voltage standards. This conversion, however is not achieved easily: a design originally targeted for implementation in a 5 V technology will typically require significant re-work to meet timing and throughput requirements at the lower operating voltage. We describe a high-level synthesis system which assists the designer in performing this task, minimizing the need for manual re-design. Techniques employed in this work include pipelining and a new approach to module selection which minimizes power consumption subject to timing constraints. Using these and other high-level synthesis techniques to target designs to 3.3 V libraries, we show that it is possible to reduce power consumption by as much as 56% as compared to the original 5 V implementation, while meeting specified minimum throughput and maximum latency constraints.> Laurence Goodby, Alex Orailoglu, Paul M. Chau |
ICCD | 2 |
| 1994 | SYNCBIST: SYNthesis for Concurrent Built-In-Self-TestabilityabstractWe present a system which synthesizes, from a behavioral description, an RTL circuit which is testable with a high degree of test concurrency. The system produces a datapath containing test registers, and a BIST test plan for the testing of the chip. All design decisions are made using an estimate of test conflicts, which is based on an analysis of the reachability of each component port from I/O pins and test registers. Chip testing according to the partial-intrusion BIST methodology is assumed. Empirical results show the effect of test conflicts on the test application time, and highlight the benefit of using the proposed synthesis approach for test conflict reduction.> Ian G. Harris, Alex Orailoglu |
ICCD | 2 |
| 1994 | Integrating Binding Constraints in the Synthesis of Area-Efficient Self-Recovering MicroarchitecturesabstractFault-tolerance increases hardware reliability of a VLSI design but it also has the disadvantage of increasing chip area. This area overhead can however be minimized by consideration of fault-tolerance during the high-level synthesis stage of design. We introduce an intertwined scheduling and binding algorithm for self-recovering fault-tolerant designs. Two copies in a self-recovering design have to be performed by disjoint sets of hardware to maintain fault-tolerance under a single fault assumption. The proposed algorithm is the first that satisfies the hardware disjointness condition which is necessary for fault-tolerant design. Area efficient self-recovering microarchitectures are achieved through novel binding techniques which influence scheduling decisions. Various synthesis benchmarks are used to illustrate the effectiveness of this approach.> Karin Högstedt, Alex Orailoglu |
ICCD | 2 |
| 1994 | Synthesis of fault-tolerant and real-time microarchitectures
Alex Orailoglu, Ramesh Karri |
J. Syst. Softw. | 1 |
| 1994 | Coactive scheduling and checkpoint determination during high level synthesis of self-recovering microarchitecturesabstractThe growing trend towards VLSI implementation of crucial tasks in critical applications has increased both the demand for and the scope of fault-tolerant VLSI systems. In this paper, we present a self-recovering microarchitecture synthesis system. In a self-recovering microarchitecture, intermediate results are compared at regular intervals, and if correct saved in registers (checkpointing). On the other hand, on detecting a fault, the self-recovering microarchitecture rolls back to a previous checkpoint and retries. The proposed synthesis system comprises of a heuristic and an optimal subsystem. The heuristic synthesis subsystem has two components. Whereas the checkpoint insertion algorithm identifies good checkpoints by successively eliminating clock cycle boundaries that either have a high checkpoint overhead or violate the retry period constraint, the novel edge-based schedule, assigns edges to clock cycle boundaries, in addition to scheduling nodes to clock cycles. Also, checkpoint insertion and edge-based scheduling are intertwined using a flexible synthesis methodology. We additionally show an Integer Linear Programming model for the self-recovering microarchitecture synthesis problem. The resulting ILP formulation can minimize either the number of voters or the overall hardware, subject to constraints on the number of clock cycles the retry period, and the number of checkpoints.> Alex Orailoglu, Ramesh Karri |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1993 | High-Level Synthesis of Fault-Secure MicroarchitecturesabstractAdvances in VLSI technology are making it feasible to pack millions of transistors on a single chip.A consequent increase in the number of on-chip faults as well as the growing import of quality metrics such as reliability and fault-tolerance are necessitating on-chip fault-tolerance.On-chip realization of a computation is fault-secure if no fault in the computation goes undetected.In this paper, we present high-level synthesis of fattlt-secure microarchitectures which require less than proportional increase in hardware.The proposed strategy selects intermediate computations for additional voting.The resulting class of fanltsecure microarchitectures supplants the enormous hardware requirements of naive fault-secure strategies with enhanced hardware utilization afforded bysecuring the intermediate computations.1 Introduction Shrinking device dimensions and low operating voltages have rendered VLSI systems susceptible to faults [12], thereby mandating on-chip fault-tolerance.Nevertheless, design of fault-tolerant ICS is not only complex but entails both an area overhead, and a performance penalty.Fault-tolerance becomes manageable at higher levels of design abstraction while preserving the numerous design options in terms of area us performance trade-offs.Fault-tolerance in general refers to a collection of techniques to mask/detect/recover-from/diagnose faults.Depending on the target environment a particular technique becomes appropriate.For example, systems used in Iife-criticaf applications cannot tolerate any faulty results.Consequently, it is crucial to detect alf faults in a timely fashion.Faultsecurity is a technique that can detect all faults in a sys- tem and it can be done online.A computation on a set of processors is fault-secure if no faultl in the computation (generated by a faulty processor) goes undetected.A microarchitecture can be secured against faults by straightforward duplication and voting.Since the hardware overhead of such a naive fault-securing strategy is enormous, we propose synthesis of alternate, low cost, fault-secure Ramesh Karri, Alex Orailoglu |
DAC | 2 |
| 1993 | Test Path Generation and Test Scheduling for Self-Testable DesignsabstractThe high cost of chip testing makes testability an important aspect of any chip design. Early inclusion of test considerations into the design process can reduce test time and area overhead. We have developed an algorithm which defines built-in self-testing (BIST) tests for an RTL datapath. Datapath registers are chosen to be upgraded to testable registers to execute the define tests. Test time of an individual test is reduced by considering pattern randomness and error masking transparency properties of modules involved in the test. Parallelism between different tests is increased by reducing the number of conflicts between tests. Intertwining the creation of tests with the selection of testable registers in the datapath guides the choice of testable registers to a minimal area solution. The use of these metrics provides our system with an accurate estimate of test time, and therefore facilitates the definition of tests which reduce test application time. Experimental results show that our algorithm defines tests which require low test time.> Alex Orailoglu, Ian G. Harris |
ICCD | 1 |
| 1993 | Intertwined Scheduling, Module Selection and Allocation in Time-and-Area
Ian G. Harris, Alex Orailoglu |
ISCAS | 2 |
| 1992 | Transformation-Based High-Level Synthesis of Fault-Tolerant ASICs
Ramesh Karri, Alex Orailoglu |
DAC | 2 |
| 1992 | High-Level Synthesis of Self-Recovering MicroArchitecturesabstractA methodology for the computer aided synthesis of microarchitectures that can recover from transient faults is presented. The synthesis is formulated as a two-step procedure of checkpoint insertion followed by duplication. The checkpoint insertion technique minimizes the voting overhead subject to input constraints on maximum allowable recovery time and the maximum number of retries. Additionally checkpoint insertion is interspersed with the scheduling decisions of a novel edge-based scheduler. Self-recovering microarchitectures which perform optimally but require less than proportional increase in hardware are generated by exploiting cost minimizing transformations.> Alex Orailoglu, Ramesh Karri |
ICCD | 1 |
| 1991 | ALPS: An Algorithm for Pipeline Data Path SynthesisabstractArticle Free Access Share on ALPS: an algorithm for pipeline data path synthesis Authors: Ramesh Karri Department of Computer Science and Engineering, University of California, San Diego, La Jolla, CA Department of Computer Science and Engineering, University of California, San Diego, La Jolla, CAView Profile , Alex Orailoğlu Department of Computer Science and Engineering, University of California, San Diego, La Jolla, CA Department of Computer Science and Engineering, University of California, San Diego, La Jolla, CAView Profile Authors Info & Claims MICRO 24: Proceedings of the 24th annual international symposium on MicroarchitectureSeptember 1991 Pages 124–132https://doi.org/10.1145/123465.123490Published:01 September 1991Publication History 2citation268DownloadsMetricsTotal Citations2Total Downloads268Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Ramesh Karri, Alex Orailoglu |
MICRO | 2 |
| 1986 | Flow graph representation
Alex Orailoglu, Daniel Gajski |
DAC | 1 |