EDBT 2026 Demo / reviewers in the wild / expert
Ankur Srivastava 0001
dblp:83/1695
· DBLP profile ↗
134ranked-venue papers
8as first author
22since 2021 · last 2026
0000-0002-5445-904XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 127 · 8 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 3 since 2021Security and privacy · 2Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TroLL: Exploiting Structural Similarities Between Logic Locking and Hardware TrojansabstractLogic locking and hardware Trojans are two fields in hardware security that have been mostly developed independently from each other. In this paper, we identify the relationship between these two fields. We find that a common structure that exists in many logic locking techniques has desirable properties of hardware Trojans (HWT). We then construct a novel type of HWT, called Trojans based on Logic Locking (TroLL), in a way that can evade state-of-the-art ATPG-based HWT detection techniques. In an effort to detect TroLL, we propose customization of existing state-of-the-art ATPG-based HWT detection approaches as well as adapting the SAT-based attacks on logic locking to HWT detection. In our experiments, we use random sampling as reference. It is shown that the customized ATPG-based approaches are the best performing but only offer limited improvement over random sampling. Moreover, their efficacy also diminishes as TroLL’s triggers become longer (i. e. have more bits specified). We thereby highlight the need to find a scalable HWT detection approach for TroLL. Yuntao Liu 0001, Aruna Jayasena, Prabhat Mishra 0001, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | SCONE: A Logic Locking Technique Utilizing SMT Solver and Circuit Encoding Scheme for Efficient Hardware IP ProtectionabstractMultiple intellectual property (IP) protections have emerged to defeat security threats in integrated circuit (IC) supply chain. Among these, logic locking is regarded as a promising IP protection for its security. A state-of-the-art work uses stripped-functionality logic locking (SFLL) technique with protected input patterns (PIPs) satisfying the distance of at least 2 (Dist2) property, or D2PIPs, for ensuring resilience against both input-output (I/O)-based and structural attacks. However, this approach has research challenges in scalability, flexibility, and security, as stated and discussed in our paper. Our paper solves these challenges by (i) utilizing a satisfiability modulo theories (SMT) solver and (ii) developing a secure circuit encoding scheme. SCONE, our secure logic locking technique, combines the two methods and meets all three challenges simultaneously. Our results show that SCONE improves scalability $\mathbf{3 5 0} \times$ on the IBEX processor (16 K gates) and remains resilient against five I/O or structural attacks. Index Terms-logic locking, encoding scheme, SMT solver. Zhaokun Han, Daniel Xing, Kostas Amberiadis, Ankur Srivastava 0001, Jeyavijayan Rajendran |
DAC | 4 |
| 2025 | SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic ReasoningabstractOptimizing Register Transfer Level (RTL) code is crucial for improving the efficiency and performance of digital circuits in the early stages of synthesis. Manual rewriting, guided by synthesis feedback, can yield high-quality results but is time-consuming and error-prone. Most existing compiler-based approaches have difficulty handling complex design constraints. Large Language Model (LLM)-based methods have emerged as a promising alternative to address these challenges. However, LLM-based approaches often face difficulties in ensuring alignment between the generated code and the provided prompts. This paper introduces SymRTLO, a neuron-symbolic framework that integrates LLMs with symbolic reasoning for the efficient and effective optimization of RTL code. Our method incorporates a retrieval-augmented system of optimization rules and Abstract Syntax Tree (AST)-based templates, enabling LLM-based rewriting that maintains syntactic correctness while minimizing undesired circuit behaviors. A symbolic module is proposed for analyzing and optimizing finite state machine (FSM) logic, allowing fine-grained state merging and partial specification handling beyond the scope of pattern-based compilers. Furthermore, a fast verification pipeline, combining formal equivalence checks with test-driven validation, further reduces the complexity of verification. Experiments on the RTL-Rewriter benchmark with Synopsys Design Compiler and Yosys show that SymRTLO improves power, performance, and area (PPA) by up to 43.9%, 62.5%, and 51.1%, respectively, compared to the state-of-the-art methods. We will release the code as open source upon the paper's acceptance. Wanghao Ye, Ping Guo 0007, Yexiao He, Bowei Tian, Shwai He, Guoheng Sun, Zheyu Shen, Ankur Srivastava 0001, Qingfu Zhang 0001, Gang Qu 0001, Ang Li 0005 |
NeurIPS | 11 |
| 2024 | A High Level Approach to Co-Designing 3D ICsabstract3D ICs promise increased logic density and reduced routing congestion over conventional monolithic 2D ICs. High level synthesis (HLS) tools promise reduced design complexity by approaching the design from a higher abstraction level and allow for more optimization flexibility. We propose improving timing closure of 3D ICs by co-designing the architecture and physical design by integrating HLS and 3D IC macro placement into the same holistic loop. On average our method is able to reduce estimated total negative slack (TNS) by 62% and 92% when compared to a traditional binding and placement technique for 2D and 3D ICs respectively. Daniel Xing, Ankur Srivastava 0001 |
DAC | 2 |
| 2024 | HQ-DTM: A Hierarchical Q-learning Algorithm for Dynamic Thermal Management of Multi-core ProcessorsabstractIn this paper, we propose a hierarchical Q-learning algorithm for dynamic thermal management (HQ-DTM) of multi-core processors. The proposed technique aims to maximize the performance of the processor subject to temperature constraints, while minimizing temperature measurement overhead. In order to achieve this, a Q-learning-based temperature measurement module is deployed alongside a Q-learning-based working mode selection module. The temperature measurement module determines the granularity of thermal model to be used under different workloads to ensure that temperature measurement overhead remains minimal while satisfying temperature constraints. The temperature measured is provided as input to the working mode selection module. It uses dynamic voltage and frequency scaling as control to maximize the throughput of each core while maintaining temperature and temporal temperature gradient (TTG) constraints. Our proposed technique has been evaluated for the SPLASH2 benchmarks. The results shows clear evidence that HQ-DTM mitigates the trade-off between processor performance and maintaining optimum thermal behavior. Abir Ahsan Akib, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2024 | 3D-Aware Low Power High-level Resource Binding and Co-DesignabstractWhile IC interconnect switching power contributes to overall dynamic power, reductions in interconnect power are only readily tackled during physical design, after architectural decisions have already been made. Since a chip's interconnect structure depends on the overall architecture's module connectivity, minimizing interconnect switching power requires the design flexibility that architectural decisions, such as operation binding, enable. We propose a co-design flow that integrates an interconnect-aware power optimal HLS binding methodology together with a dataflow and power aware physical placement tool to reduce overall switching power consumed not just within modules but also at the interconnects that connect them. We test our proposed co-design flow on functions extracted from MediaBench and show that, averaged over all tested benchmarks, our method reduces overall switching power by 31% and 28% for 2D and 3D IC floorplans respectively when compared to a conventional binding and timing-aware design process. Daniel Xing, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2024 | Removal of SAT-Hard Instances in Logic Obfuscation Through Inference of FunctionalityabstractLogic obfuscation is a prominent approach to protect intellectual property within integrated circuits during fabrication. Many attacks on logic locking have been proposed, particularly in the Boolean satifiability (SAT) attack family, leading to the development of stronger obfuscation techniques. Some obfuscation techniques, including Full-Lock and InterLock, resist SAT attacks by inserting SAT-hard instances into the design, making the SAT attack infeasible. In this work, we observe that this class of obfuscation leaves most of the original design topology visible to an attacker, who can reverse-engineer the original design given the functionality of the SAT-hard instance. We show that an attacker can expose the SAT-hard instance functionality of Full-Lock or InterLock with a polynomial number of queries of its inputs and outputs. We then develop a mathematical framework showing how the functionality can be inferred using only a black-box oracle, as is commonly used in attacks in the literature. Using this framework, we develop a novel attack that allows a SAT-capable attacker to efficiently unlock designs obfuscated with Full-Lock. Our attack recovers the intellectual property from these obfuscation techniques that were previously thought secure. We empirically demonstrate the potency of our novel sensitization attack against benchmark circuits obfuscated with Full-Lock. Isaac McDaniel, Michael Zuzak, Ankur Srivastava 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Security Evaluation of State Space Obfuscation of Hardware IP through a Red Team-Blue Team PracticeabstractDue to the inclination towards a fab-less model of integrated circuit (IC) manufacturing, several untrusted entities get white-box access to the proprietary intellectual property (IP) blocks from diverse vendors. To this end, the untrusted entities pose security-breach threats in the form of piracy, cloning, and reverse-engineering, sometimes threatening national security. Hardware obfuscation is a prominent countermeasure against such issues. Obfuscation allows for preventing the usage of the IP blocks without authorization from the IP owners. Due to finite state machine (FSM) transformation-based hardware obfuscation, the design’s FSM gets transformed to make it difficult for an attacker to reverse-engineer the design. A secret key needs to be applied to make the FSM functional, thus preventing the usage of the IP for unintended purposes. Although several hardware obfuscation techniques have been proposed, due to the inability to analyze the techniques from the attackers’ standpoint, numerous vulnerabilities inherent to the obfuscation methods go undetected unless a true adversary discovers them. In this article, we present a collaborative approach between two entities—one acting as an attacker or red team and another as a defender or blue team , the first systematic approach to replicate the real attacker-defender scenario in the hardware security domain, which in return strengthens the FSM transformation-based obfuscation technique. The blue team transforms the underlying FSM of a gate-level netlist using state space obfuscation. The red team plays the role of an adversary or evaluator and tries to unlock the design by extracting the unlocking key or recovering the obfuscation circuitries. As the key outcome of this red team–blue team effort, a robust state space obfuscation methodology is evolved showing security promises. Md. Moshiur Rahman 0001, Jim Geist, Daniel Xing, Yuntao Liu 0001, Ankur Srivastava 0001, Travis Meade, Yier Jin, Swarup Bhunia |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2023 | TimingCamouflage+ DecamouflagedabstractIn today's world, sending a chip design to a third party foundry for fabrication poses a serious threat to one's intellectual property. To keep designs safe from adversaries, design obfuscation techniques have been developed to protect the IP details of the design. This paper explains how the previously considered secure algorithm, TimingCamouflage+, can be thwarted and the original circuit can be recovered [15]. By removing wave-pipelining false paths, the TimingCamouflage+ algorithm is reduced to the insecure TimingCamouflage algorithm [16]. Since the TimingCamouflage algorithm is vulnerable to the TimingSAT attack, this reduction proves that TimingCamouflage+ is also vulnerable to TimingSAT and not a secure camouflaging technique [7]. This paper describes how wave-pipelining paths can be removed, and this method of handling false paths is tested on various benchmarks and shown to be both functionally correct and feasible in complexity. Priya Mittu, Yuntao Liu 0001, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Low Power Logic Obfuscation Through System Level Clock GatingabstractLogic locking methods such as Stripped Functionality Logic Locking (SFLL) tend to yield high overheads. SFLL only corrupts a small part of the input space by design in order to maintain good SAT resilience and in doing so selects high frequency inputs to corrupt (protect) and therefore increases locking's impact on system level error. This implies that much of the time stripped modules are doing unnecessary work while the restore units are correcting the computations. We propose taking advantage of this fact to selectively clock gate the modules when protected inputs are being processed. Under the highest possible level of attack resilience, this alone can yield up to 24.5 % dynamic power savings when protected inputs are applied to synthesized MediaBench benchmarks. We also propose a system-level design approach that utilizes the data-flow graph to also gate operations that fully depend on other gated operations. In conjunction with modifying operation binding, this increases power savings to 32.9 % under the same strict security constraints. Daniel Xing, Yuntao Liu 0001, Ankur Srivastava 0001 |
ISLPED | 3 |
| 2023 | Introduction to the Special Issue on CAD for Security: Pre-silicon Security Sign-off Solutions Through Design CycleabstractThis introduction welcomes all readers to this ACM JETC special issue on CAD for Security: Pre-silicon Security Sign-off Solutions Through Design Cycle. The articles published in this special issue reflect how computer-aided design (CAD) tools are developed to expand the notion of automated security verification throughout the system-on-chip (SoC) design cycle. This special issue aims to demonstrate how the semiconductor industry must look for security-oriented metrics and evaluation as part of automatic CAD solution development to aid analysis, identifying, root-causing, and mitigating SoC security problems. Throughout this introductory note, we first represent the need for such a security-oriented sign-off solution for the ASIC design flow, then it is followed by providing an overview of the articles published in this special issue and how they address such requirements. Farimah Farahmandi, Ankur Srivastava 0001, Giorgio Di Natale, Mark Tehranipoor |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2023 | Security-Aware Resource Binding to Enhance Logic ObfuscationabstractLogic obfuscation mitigates the unauthorized use of design IP by untrusted partners during integrated circuit (IC) fabrication. To do so, these techniques produce gate-level errors that derail typical applications run on the IC. Recent research has derived a link between the error rate and the Boolean satisfiability (SAT) attack resilience of logic obfuscation. As a result, it has been shown to be difficult for obfuscation to inject sufficient gate-level error to derail application-level function while maintaining resilience to SAT-style attacks. In this work, we explore use of architectural knowledge during the resource binding phase of high-level synthesis to automate the design of locked architectures capable of high-corruption and SAT resilience simultaneously. To do so, we bifurcate logic obfuscation schemes into two families based on their error profile: distributed error locking and critical minterm locking. We then develop security-focused binding/locking algorithms for each locking family and use them to bind/lock 11 MediaBench benchmarks. For distributed error locking, our proposed security-aware binding algorithms designed locked circuits capable of corrupting a typical application for 52% more wrong keys than a circuit bound with conventional algorithms. For critical minterm locking, our proposed security-aware binding algorithms designed locked circuits capable of corrupting a typical application for 100% of wrong keys while also exhibiting$26\times $more application errors than a circuit bound with conventional algorithms. Regardless of locking family, our security-aware algorithms improved corruption without degrading SAT resilience or incurring sizable design overheads to do so. Obfuscation applied post-binding could not achieve high-corruption and SAT resilience simultaneously in these benchmarks. Michael Zuzak, Yuntao Liu 0001, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Hardware IP Protection against Confidentiality Attacks and Evolving Role of CAD ToolabstractWith growing use of hardware intellectual property (IP) based integrated circuits (IC) design and increasing reliance on a globalized supply chain, the threats to confidentiality of hardware IPs have emerged as major security concerns to the IP producers and owners. These threats are diverse, including reverse engineering (RE), piracy, cloning, and extraction of design secrets, and span different phases of electronics life cycle. The academic research community and the semiconductor industry have made significant efforts over the past decade on developing effective methodologies and CAD tools targeted to protect hardware IPs against these threats. These solutions include watermarking, logic locking, obfuscation, camouflaging, split manufacturing, and hardware redaction. This paper focuses on key topics on confidentiality of hardware IPs encompassing the major threats, protection approaches, security analysis, and metrics. It discusses the strengths and limitations of the major solutions in protecting hardware IPs against the confidentiality attacks, and future directions to address the limitations in the modern supply chain ecosystem. Swarup Bhunia, Amitabh Das, Saverio Fazzari, Vivian Kammler, David Kehlet, Jeyavijayan Rajendran, Ankur Srivastava 0001 |
ICCAD | 7 |
| 2022 | A Combined Logical and Physical Attack on Logic ObfuscationabstractLogic obfuscation protects integrated circuits from an untrusted foundry attacker during manufacturing. To counter obfuscation, a number of logical (e.g. Boolean satisfiability) and physical (e.g. electro-optical probing) attacks have been proposed. By definition, these attacks use only a subset of the information leaked by a circuit to unlock it. Countermeasures often exploit the resulting blind-spots to thwart these attacks, limiting their scalability and generalizability. To overcome this, we propose a combined logical and physical attack against obfuscation called the CLAP attack. The CLAP attack leverages both the logical and physical properties of a locked circuit to prune the keyspace in a unified and theoretically-rigorous fashion, resulting in a more versatile and potent attack. To formulate the physical portion of the CLAP attack, we derive a logical formulation that provably identifies input sequences capable of sensitizing logically expressive regions in a circuit. We prove that electro-optically probing these regions infers portions of the key. For the logical portion of the attack, we integrate the physical attack results into a Boolean satisfiability attack to find the correct key. We evaluate the CLAP attack by launching it against four obfuscation schemes in benchmark circuits. The physical portion of the attack fully specified 60.6% of key bits and partially specified another 10.3%. The logical portion of the attack found the correct key in the physical-attack-limited keyspace in under 30 minutes. Thus, the CLAP attack unlocked each circuit despite obfuscation. Michael Zuzak, Yuntao Liu 0001, Isaac McDaniel, Ankur Srivastava 0001 |
ICCAD | 4 |
| 2022 | A Black-Box Sensitization Attack on SAT-Hard Instances in Logic ObfuscationabstractLogic obfuscation is a prominent approach to protect intellectual property within integrated circuits during fabrication. In response to logic obfuscation, the Boolean satisfiability attack was developed and demonstrated to unlock a great deal of existing obfuscation configurations. This drove the development of new SAT-resistant obfuscation countermeasures. Some of these, including Full-Lock and InterLock, resist SAT attacks by inserting SAT-hard instances, rapidly scaling the runtime of each SAT attack iteration. In this work, we demonstrate that while such countermeasures resist SAT-style attack strategies, an attacker with access to the inputs and outputs of the SAT-hard instance Full-Lock has inserted into an oracle circuit can infer the design’s intended functionality in linear time, thereby unlocking the circuit. We also observe that this class of obfuscation leaves most of the original design topology intact and show how this enables an attacker to sensitize the SAT-hard instance within a black-box oracle and make inferences about the instance’s input-output relationship from the oracle’s primary inputs and outputs. We develop a novel attack which uses this leakage to allow an attacker to efficiently unlock designs obfuscated with Full-Lock without the special assumption of access to the SAT-hard instance’s inputs and outputs. This recovers the intellectual property and renders these obfuscation techniques insecure. We empirically demonstrate the potency of our novel sensitization attack against benchmark circuits obfuscated with SAT-hard instances. Our proposed attack was able to unlock all 6 benchmark circuits containing 384-bit keys and 3 out of 4 benchmarks with a 960-bit key within 48 hours. In comparison, the conventional SAT attack was only able to unlock 3 of 6 benchmarks with 384 key bits and none of the 4 benchmarks with 960 key bits in the same 48 hour timeout period. Isaac McDaniel, Michael Zuzak, Ankur Srivastava 0001 |
ICCD | 3 |
| 2022 | Evaluating the Security of Logic-Locked Probabilistic CircuitsabstractLogic locking is a design-for-security scheme to thwart attacks by an untrusted foundry. Prior work exposed the vulnerability of logic-locked circuits using Boolean satisfiability (SAT). While these attacks are effective against deterministic circuits, they cannot unlock probabilistic/approximate designs, which have become increasingly popular. In this work, we expand SAT-style attacks to locked circuits with a probabilistic behavior. We proposeStatSAT, an attack incorporating statistical techniques into the SAT attack to unlock probabilistic designs. We then propose a countermeasure, called high error rate keys (HERKs), to thwart StatSAT and other attacks on probabilistic circuits. HERKs leverage high error wires, caused by the probabilistic behavior, to hide the correct key under stochastic noise. Michael Zuzak, Ankit Mondal, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Invited: Independent Verification and Validation of Security-Aware EDA Tools and IPabstractSecure silicon requires a seamless integration of new tools, new IP, and design flows to help designers protect integrated circuits from increasingly sophisticated attacks. Independent Validation and Verification (IV&V) of this integrated technology is important to ensure that the tools actually deliver on their security claims when used by independent parties (i.e., people who were not involved in designing the tools). This work discusses the principles and approaches for IV&V of such a complex design environment, including validation of the security strength of the various hardware security techniques, such as combinational and sequential logic locking, Trojan Detection, side-channel mitigation, and blockchain-based asset management. The main challenge in running an IV&V effort is to ensure that the process provides rigorous, methodical and provable evaluation of the claims of not only the component tools and IP, but whether such an integrated environment can produce security-hardened designs by a non-security expert. CCS Concepts • Hardware $\rightarrow$ Very large scale integration design; Methodologies for EDA; • Security and privacy $\rightarrow$ Security in hardware. Benjamin Tan 0001, Siddharth Garg, Ramesh Karri, Yuntao Liu 0001, Michael Zuzak, Abhisek Chakraborty, Ankur Srivastava 0001, Omid Aramoon, Qian Xu 0022, Gang Qu 0001, Adam A. Porter, Jeno Szep, Warren Savage |
DAC | 7 |
| 2021 | A Resource Binding Approach to Logic ObfuscationabstractLogic locking has been proposed to counter security threats during IC fabrication. Such an approach restricts unauthorized use by injecting sufficient module level error to derail application level IC functionality. However, recent research has identified a trade-off between the error rate of logic locking and its resilience to a Boolean satisfiablity (SAT) attack. As a result, logic locking often cannot inject sufficient error to impact an IC while maintaining SAT resilience. In this work, we propose using architectural context available during resource binding to co-design architectures and locking configurations capable of high corruption and SAT resilience simultaneously. To do so, we propose 2 security-focused binding/locking algorithms and apply them to bind/lock 11 MediaBench benchmarks. The resulting circuits showed a 26x and 99x increase in the application errors of a fixed locking configuration while maintaining SAT resilience and incurring minimal overhead compared to other binding schemes. Locking applied post-binding could not achieve a high application error rate and SAT resilience simultaneously. Michael Zuzak, Yuntao Liu 0001, Ankur Srivastava 0001 |
DAC | 3 |
| 2021 | Robust and Attack Resilient Logic Locking with a High Application-Level ImpactabstractLogic locking is a hardware security technique aimed at protecting intellectual property against security threats in the IC supply chain, especially those posed by untrusted fabrication facilities. Such techniques incorporate additional locking circuitry within an integrated circuit (IC) that induces incorrect digital functionality when an incorrect verification key is provided by a user. The amount of error induced by an incorrect key is known as the effectiveness of the locking technique. A family of attacks known as “SAT attacks” provide a strong mathematical formulation to find the correct key of locked circuits. To achieve high SAT resilience (i.e., complexity of SAT attacks), many conventional logic locking schemes fail to inject sufficient error into the circuit when the key is incorrect. For example, in the case of SARLock and Anti-SAT, there are usually very few (or only one) input minterms that cause any error at the circuit output. The state-of-the-art s tripped functionality logic locking (SFLL) technique provides a wide spectrum of configurations that introduced a tradeoff between SAT resilience and effectiveness. In this work, we prove that such a tradeoff is universal among all logic locking techniques. To attain high effectiveness of locking without compromising SAT resilience, we propose a novel logic locking scheme, called Strong Anti-SAT (SAS). In addition to SAT attacks, removal-based attacks are another popular kind of attack formulation against logic locking where the attacker tries to identify and remove the locking structure. Based on SAS, we also propose Robust SAS (RSAS) that is resilient to removal attacks and maintains the same SAT resilience and effectiveness as SAS. SAS and RSAS have the following significant improvements over existing techniques. (1) We prove that the SAT resilience of SAS and RSAS against SAT attack is not compromised by increase in effectiveness . (2) In contrast to prior work that focused solely on the circuit-level locking impact, we integrate SAS-locked modules into an 80386 processor and show that SAS has a high application-level impact. (3) Our experiments show that SAS and RSAS exhibit better SAT resilience than SFLL and their effectiveness is similar to SFLL. Yuntao Liu 0001, Michael Zuzak, Yang Xie 0001, Abhishek Chakraborty 0001, Ankur Srivastava 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2021 | Evaluating the Security of Delay-Locked CircuitsabstractIn order to enhance the security of logic obfuscation schemes, delay locking has been proposed in combination with traditional functional logic locking approaches. A circuit obfuscated using this approach preserves the original functionality only when both correct functional and delay keys are provided. In this article, we develop a novel SAT formulation-based attack approach called TimingSAT to deobfuscate the functionalities of such delay-locked designs. The proposed technique models the timing characteristics of various types of gates present in a design as Boolean functions to build a timing profile embedded SAT formulation in terms of targeted key inputs. TimingSAT attack works in two stages. In the first stage, the functional key is found using the conventional SAT attack approach, and in the second stage, the delay key is determined using the aforementioned timing profile embedded SAT formulation of the circuit. In both stages of the attack, wrong keys are iteratively eliminated till a key belonging to the correct equivalence class is obtained. We perform experiments to demonstrate the effectiveness of our proposed TimingSAT attack to break delay-locked benchmarks within a few hours. Subsequently, we propose a countermeasure called stripped-functionality delay locking (SFDL) which not only thwarts TimingSAT attack but also resists all known attacks against logic obfuscation. SFDL combines the concept of delay locking with a stripped-functionality-based logic locking approach to realize an effective IP security solution for hardware designs. Unlike existing logic locking schemes, SFDL simultaneously achieves strong SAT attack resiliency as well as significantly high output corruptibility. Abhishek Chakraborty 0001, Yuntao Liu 0001, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Trace Logic Locking: Improving the Parametric Space of Logic LockingabstractTo protect against an untrusted foundry, logic locking must 1) inject sufficient error to ensure critical application failures for any wrong key (error severity) and 2) resist any attack against it (attack resilient). We begin our work by deriving a fundamental tradeoff between these two goals which exists underlying all logic locking, regardless of construction. This relationship forces integrated circuit (IC) designers to sacrifice the error severity of logic locking to increase its attack resilience and vice versa. We proceed by exploring the consequences of this tradeoff through architectural simulations of ICs incorporating locking sweeping over the derived parametric space. We find that the efficacy of logic locking is severely limited by this tradeoff. In response, we propose trace logic locking (TLL), a novel enhancement of module level logic locking which enables existing art to secure arbitrary length sequences of input minterms, referred to as traces. Doing so injects an additional degree of freedom into the parametric space of locking, enabling locking techniques to overcome the limitations of our derived tradeoff. We both theoretically and empirically prove this by using TLL to enhance cutting edge locking. In ten large benchmarks, we show that TLL-enhanced logic locking provides exponentially stronger attack resilience than conventional locking with only modest additional overhead. Finally, we demonstrate the efficacy of TLL in a processor IC using architectural simulations. Despite prior art being unable to secure this IC, we find that TLL concurrently achieves strong error severity and attack resilience. Michael Zuzak, Yuntao Liu 0001, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Ising-FPGA: A Spintronics-based Reconfigurable Ising Model SolverabstractThe Ising model has been explored as a framework for modeling NP-hard problems, with several diverse systems proposed to solve it. The Magnetic Tunnel Junction– (MTJ) based Magnetic RAM is capable of replacing CMOS in memory chips. In this article, we propose the use of MTJs for representing the units of an Ising model and leveraging its intrinsic physics for finding the ground state of the system through annealing. We design the structure of a basic MTJ-based Ising cell capable of performing the functions essential to an Ising solver. The hardware overhead of the Ising model is analyzed, and a technique to use the basic Ising cell for scaling to large problems is described. We then go on to propose Ising-FPGA, a parallel and reconfigurable architecture that can be used to map a large class of NP-hard problems, and show how a standard Place and Route tool can be utilized to program the Ising-FPGA. The effects of this hardware platform on our proposed design are characterized and methods to overcome these effects are prescribed. We discuss how three representative NP-hard problems can be mapped to the Ising model. Further, we suggest ways to simplify these problems to reduce the use of hardware and analyze the impact of these simplifications on the quality of solutions. Simulation results show the effectiveness of MTJs as Ising units by producing solutions close/comparable to the optimum and demonstrate that our design methodology holds the capability to account for the effects of the hardware. Ankit Mondal, Ankur Srivastava 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2020 | Hardware-Assisted Intellectual Property Protection of Deep Learning ModelsabstractThe protection of intellectual property (IP) rights of well-trained deep learning (DL) models has become a matter of major concern, especially with the growing trend of deployment of Machine Learning as a Service (MLaaS). In this work, we demonstrate the utilization of a hardware root-of-trust to safeguard the IPs of such DL models which potential attackers have access to. We propose an obfuscation framework called Hardware Protected Neural Network (HPNN) in which a deep neural network is trained as a function of a secret key and then, the obfuscated DL model is hosted on a public model sharing platform. This framework ensures that only an authorized end-user who possesses a trustworthy hardware device (with the secret key embedded on-chip) is able to run intended DL applications using the published model. Extensive experimental evaluations show that any unauthorized usage of such obfuscated DL models result in significant accuracy drops ranging from 73.22 to 80.17% across different neural network architectures and benchmark datasets. In addition, we also demonstrate the robustness of proposed HPNN framework against a model fine-tuning type of attack. Abhishek Chakraborty 0001, Ankit Mondal, Ankur Srivastava 0001 |
DAC | 3 |
| 2020 | StatSAT: A Boolean Satisfiability based Attack on Logic-Locked Probabilistic CircuitsabstractThe outsourcing of chip designs for fabrication has raised concerns regarding the protection of Intellectual Property (IP) from an untrustworthy foundry. Logic locking is a design-for-security technique that has the potential to thwart attacks from such an adversary. On the other hand, the notions of approximate and probabilistic computing have been popularized due to their low energy consumption characteristics and their potential application in error-tolerant frameworks. Prior work has looked into and exposed the vulnerability of logic-locked circuits using concepts of Boolean Satisfiability (SAT), but mostly from the perspective of deterministic designs. Despite existing attack frameworks not being directly applicable, we show in this work that circuits exhibiting probabilistic behavior also face the same threat. We propose StatSAT, an attack methodology incorporating statistical techniques into the existing SAT attack, that can overcome the hurdles imposed by the probabilistic behavior. Our attack results show that the adversary is capable of unlocking the circuit to an extent good for all practical purposes. Ankit Mondal, Michael Zuzak, Ankur Srivastava 0001 |
DAC | 3 |
| 2020 | Energy-efficient Design of MTJ-based Neural Networks with Stochastic ComputingabstractHardware implementations of Artificial Neural Networks (ANNs) using conventional binary arithmetic units are computationally expensive, energy-intensive, and have large area overheads. Stochastic Computing (SC) is an emerging paradigm that replaces these conventional units with simple logic circuits and is particularly suitable for fault-tolerant applications. We propose an energy-efficient use of Magnetic Tunnel Junctions (MTJs), a spintronic device that exhibits probabilistic switching behavior, as Stochastic Number Generators (SNGs), which forms the basis of our NN implementation in the SC domain. Further, the error resilience of target applications of NNs allows approximating the synaptic weights in our MTJ-based NN implementation, in ways brought about by properties of the MTJ-SNG, to achieve energy-efficiency. An algorithm is designed that, given an error tolerance, can perform such approximations in a single-layer NN in an optimal way owing to the convexity of the problem formulation. We then use this algorithm and develop a heuristic approach for approximating multi-layer NNs. Classification problems were evaluated on the optimized NNs and results showed substantial savings in energy for little loss in accuracy. Ankit Mondal, Ankur Srivastava 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | Keynote: A Disquisition on Logic LockingabstractThe fabless business model has given rise to many security threats, including piracy of intellectual property (IP), overproduction, counterfeiting, reverse engineering (RE), and hardware Trojans (HT). Such threats severely undermine the benefits of the fabless model. Among the countermeasures developed to thwart piracy and RE attacks, logic locking has emerged as a promising and versatile solution that is being adopted by both academia and industry. The idea behind logic locking is to lock the design using a “keying” mechanism; only the rightful owner has control over the locked design. Therefore, the design remains nonfunctional without the knowledge of the key. In this article, we survey the evolution of logic locking over the last decade. We introduce various “cat-and-mouse” games involved in logic locking along with its novel applications-including, processor pipelines, graphics processing units (GPUs), and analog circuits. We aim this article to be a primer for researchers interested in developing new logic-locking techniques and employing logic locking in different application domains. Abhishek Chakraborty 0001, Nithyashankari Gummidipoondi Jayasankaran, Yuntao Liu 0001, Jeyavijayan Rajendran, Ozgur Sinanoglu, Ankur Srivastava 0001, Yang Xie 0001, Muhammad Yasin, Michael Zuzak |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | In Situ Stochastic Training of MTJ Crossbars With Machine Learning AlgorithmsabstractOwing to high device density, scalability, and non-volatility, magnetic tunnel junction (MTJ)-based crossbars have garnered significant interest for implementing the weights of neural networks (NNs). The existence of only two stable states in MTJs implies a high overhead of obtaining optimal binary weights in software. This article illustrates that the inherent parallelism in the crossbar structure makes it highly appropriate for in situ training, wherein the network is taught directly on the hardware. It leads to significantly smaller training overhead as the training time is independent of the size of the network, while also circumventing the effects of alternate current paths in the crossbar and accounting for manufacturing variations in the device. We show how the stochastic switching characteristics of MTJs can be leveraged to perform probabilistic weight updates using the gradient descent algorithm. We describe how the update operations can be performed on crossbars implementing NNs and restricted Boltzmann machines, and perform simulations on them to demonstrate the effectiveness of our techniques. The results reveal that stochastically trained MTJ-crossbar feed-forward and deep belief nets achieve a classification accuracy nearly the same as that of real-valued weight networks trained in software and exhibit immunity to device variations. Ankit Mondal, Ankur Srivastava 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2019 | Anti-SAT: Mitigating SAT Attack on Logic LockingabstractLogic locking is a technique that is proposed to protect outsourced IC designs from piracy and counterfeiting by untrusted foundries. A locked IC preserves the correct functionality only when a correct key is provided. Recently, the security of logic locking is threatened by a new attack called SAT attack, which can decipher the correct key of most logic locking techniques within a few hours even for a reasonably large key-size. This attack iteratively solves SAT formulas which progressively eliminate the incorrect keys till the circuit is unlocked. In this paper, we present a circuit block (referred to as Anti-SAT block) to enhance the security of existing logic locking techniques against the SAT attack. We show using a mathematical proof that the number of SAT attack iterations to reveal the correct key in a circuit comprising an Anti-SAT block is an exponential function of the key-size thereby making the SAT attack computationally infeasible. Besides, we address the vulnerability of the Anti-SAT block to various removal attacks and investigate obfuscation techniques to prevent these removal attacks. More importantly, we provide a proof showing that these obfuscation techniques for making Anti-SAT un-removable would not weaken the Anti-SAT block's resistance to SAT attack. Through our experiments, we illustrate the effectiveness of our approach to securing modern chips fabricated in untrusted foundries. Yang Xie 0001, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Enhanced Phase-Driven Q-Learning-Based DRM for Multicore ProcessorsabstractIn this paper, we propose a new dynamic reliability management technique for multicore processors using phase-driven Q-learning-based method. Our technique considers a wide range of long-term reliability issues and maximizes the throughput of the processor subject to the reliability constraint. We employ ON/OFF switching actions and dynamic voltage and frequency scaling as control knobs (i.e., working modes) to tune the state of cores of the processor. In order to achieve this, our technique detects program phases and adaptively determines the optimal working modes for each phase using the Q-learning-based method. By integrating the phase detection into the Q-learning-based management, our technique can provide efficient management for the programs with highly diverse phases. We also propose three additional modules to improve the management efficiency of our technique. In order to evaluate our technique, we use it to manage a 3-D CPU with high-diver programs. Several failure mechanisms are considered in this case study. Our proposed technique is compared with two existing Q-learning-based techniques. The experimental results demonstrate that when the number of phases is smaller than the number of working modes, our technique can achieve more than 1.36× improvement in performance with 60% memory space savings. Zhiyuan Yang 0001, Caleb Serafy, Tiantao Lu, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Reducing Timing Side-Channel Information Leakage Using 3D IntegrationabstractRecently, following the work pioneered by Kocher [1] , using cache behavior as a timing side-channel to leak critical system information has received lots of attentions because of its easy-to-implement nature and amazingly good results. Recent attacks have been demonstrated to successfully leak the full key from many commonly used encryption algorithms including RSA, AES, etc. These attacks pose great threats to applications that depend on these encryption methods such as banking systems, military systems, etc. To mitigate the increasing threat, numerous countermeasures, mostly software patches, have been proposed. Hardware mitigations, however, have been less pursued. 3D integration, which stacks multiple dies vertically, offers shorter wire-length and improves system performance. It may be used to offset the performance overhead these countermeasures incur. In this paper, we investigate several possible ways in which the availability of 3D integration can be exploited to mitigate timing side-channel attacks while still obtaining superior performance over a baseline 2D system. Simulation results using Gem5 simulator show that our techniques can significantly reduce timing information leakage while still achieving 20.43 percent performance gain over a 2D baseline system. Chongxi Bao, Ankur Srivastava 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2018 | GPU obfuscation: attack and defense strategiesabstractConventional attacks against existing logic obfuscation techniques rely on the presence of an activated hardware for analysis. In reality, obtaining such activated chips may not always be practical, especially if the on-chip test structures are disabled. In this paper, we develop an iterative SAT formulation based attack strategy for deobfuscating many-core GPU hardware without any requirement of an activated chip. Our experiments on a real testbed using NVIDIA's SASSIFI framework reveal that more than 95% of the application runs on such an approximately unlocked GPU result in correct outcomes with 95% confidence-level and 5% confidence-interval. To counter the proposed attack, we develop a Cache Locking countermeasure which significantly degrades the performance of GPGPU applications for a wrong cache-key. Abhishek Chakraborty 0001, Yang Xie 0001, Ankur Srivastava 0001 |
DAC | 3 |
| 2018 | TimingSAT: timing profile embedded SAT attackabstractIn order to enhance the security of logic obfuscation schemes, delay based logic locking has been proposed in combination with traditional functional logic locking approaches in recent literature. A circuit obfuscated using the aforementioned approach preserves the correct functionality only when both correct functional and delay keys are provided. In this paper, we develop a novel SAT formulation based approach called TimingSAT to deobfuscte the functionalities of such delay locked designs within a reasonable amount of time. The proposed technique models the timing characteristics of various types of gates present in the design as Boolean functions to build timing profile embedded SAT formulations in terms of targeted key inputs. TimingSAT attack works in two stages: In the first stage the functional keys are found using traditional SAT attack approach and in the second stage the delay keys are deciphered utilizing the timing profile embedded SAT formulation of the circuit. In both stages of the attack, wrong keys are iteratively eliminated till a key belonging to the correct equivalence class is obtained. The experimental results highlight the effectiveness of the proposed TimingSAT attack to break delay logic locked benchmarks within few hours. Abhishek Chakraborty 0001, Yuntao Liu 0001, Ankur Srivastava 0001 |
ICCAD | 3 |
| 2018 | In-situ Stochastic Training of MTJ Crossbar based Neural NetworksabstractOwing to high device density, scalability and non-volatility, Magnetic Tunnel Junction-based crossbars have garnered significant interest for implementing the weights of an artificial neural network. The existence of only two stable states in MTJs implies a high overhead of obtaining optimal binary weights in software. We illustrate that the inherent parallelism in the crossbar structure makes it highly appropriate for in-situ training, wherein the network is taught directly on the hardware. It leads to significantly smaller training overhead as the training time is independent of the size of the network, while also circumventing the effects of alternate current paths in the crossbar and accounting for manufacturing variations in the device. We show how the stochastic switching characteristics of MTJs can be leveraged to perform probabilistic weight updates using the gradient descent algorithm. We describe how the update operations can be performed on crossbars both with and without access transistors and perform simulations on them to demonstrate the effectiveness of our techniques. The results reveal that stochastically trained MTJ-crossbar NNs achieve a classification accuracy nearly same as that of real-valued-weight networks trained in software and exhibit immunity to device variations. Ankit Mondal, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2018 | Value-driven Synthesis for Neural Network ASICsabstractIn order to enable low power and high performance evaluation of neural network (NN) applications, we investigate new design methodologies for synthesizing neural network ASICs (NN-ASICs). An NN-ASIC takes a trained NN and implements a chip with customized optimization. Knowing the NN topology and weights allows us to develop unique optimization schemes which are not available to regular ASICs. In this work, we investigate two types of value-driven optimized multipliers which exploit the knowledge of synaptic weights and we develop an algorithm to synthesize the multiplication of trained NNs using these special multipliers instead of general ones. The proposed method is evaluated using several Deep Neural Networks. Experimental results demonstrate that compared to traditional NNPs, our proposed NN-ASICs can achieve up to 6.5x and 55x improvement in performance and energy efficiency (i.e. inverse of Energy-Delay-Product), respectively. Zhiyuan Yang 0001, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2018 | A Combined Optimization-Theoretic and Side- Channel Approach for Attacking Strong Physical Unclonable FunctionsabstractThe promise of strong physical unclonable functions (PUF) is to utilize the manufacturing variations of circuit elements to produce an independent and unpredictable response to any input challenge vector. Attacks on PUFs that predict the responses to input challenge vectors offer an interesting research problem. An attacking approach based on the optimization theory and side-channel information is proposed where we estimate the manufacturing variations of the circuit elements and predict the PUF's responses to challenge vectors whose actual responses are not known. We apply this attacking approach on some popular PUF designs, including the Arbiter PUFs, the Memristor Crossbar PUFs, and the XOR Arbiter PUFs. Simulations show a substantial reduction in attack complexity compared with previously proposed machine-learning (ML)-based attacks: we achieve an average reduction of 66% in attack time compared with the ML approach. Despite some overhead, our approach is also applicable when the PUF responses are noisy. Yuntao Liu 0001, Yang Xie 0001, Chongxi Bao, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Delay Locking: Security Enhancement of Logic Locking against IC Counterfeiting and OverproductionabstractLogic locking is a technique that has been proposed to thwart IC counterfeiting and overproduction by untrusted foundry. Recently, the security of logic locking is threatened by a new attack called SAT attack, which can effectively decipher the correct key of most logic locking techniques. In this paper, we propose a new technique called delay locking to enhance the security of existing logic locking techniques. For delay locking, the key into a locked circuit not only determines its functionality, but also its timing profile. A functionality-correct but timing-incorrect key will result in timing violations and thus making the circuit malfunction. Yang Xie 0001, Ankur Srivastava 0001 |
DAC | 2 |
| 2017 | Phase-driven Learning-based Dynamic Reliability Management For Multi-core ProcessorsabstractIn this paper, we propose a phase-driven Q-learning based dynamic reliability management (DRM) technique for multi-core processors to solve DRM problems of maximizing the processor performance subject to a large class of reliability constraints by turning ON/OFF cores and dynamic voltage frequency scaling. Our technique utilizes the existing methods to detect program phases (i.e. [17]) and learns (rather than obtaining at the off-line stage) the optimal configuration of the multi-core processor for each phase. Our technique outperforms the existing learning-based DRM methods in managing programs with highly diverse phases. Our proposed technique is evaluated by solving a DRM problem in 3D CPUs of maximizing processor performance subject to the electromigration induced power delivery network reliability constraint. Compared to the latest Q-learning based DRM technique [11], our method can achieve more than 1.3x improvement in performance with 77% memory savings. Zhiyuan Yang 0001, Caleb Serafy, Tiantao Lu, Ankur Srivastava 0001 |
DAC | 4 |
| 2017 | Design Space Modeling and Simulation for Physically Constrained 3D CPUsabstractDesign space exploration (DSE) is becoming increasingly complex and 3D integration compounds the problem by imposing a more complex design space. Moreover, 3D design is a new frontier of CPU design, expected to rely more heavily on statistical modeling than designer intuition. Past work has used regression modeling with random sampling of the design space. In this paper we propose a directed simulation technique where intermediate predictions more efficiently direct simulation resources. Our results show over 98% model accuracy while simulating less than 5% of the design space and reducing simulation time more than 4.5x compared to random sampling. Caleb Serafy, Zhiyuan Yang 0001, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Template Attack Based Deobfuscation of Integrated CircuitsabstractLogic encryption algorithms have gained wide popularity to safeguard Integrated Circuits (ICs) from being pirated or counterfeited in untrusted third-party foundries. However, an untrusted foundry can reverse engineer the netlist and gain important insight regarding the design of a chip. In this paper, we demonstrate how an adversary can monitor side-channel information of an activated chip and analyze the corresponding reverse engineered netlist to successfully deobfuscate the functionality of the circuit. In particular, our proposed attack is based on a Template Analysis (TA) approach which deciphers the key inputs of a locked netlist by exploiting power side-channel traces of the activated chip. The proposed methodology utilizes the fact that various key-gates of a netlist (locked using standard logic encryption algorithms) are located at different logic depths, which in turn enables a side-channel adversary to unlock the circuit functionality level-by-level following an iterative approach. The experimental results confirm that netlists locked using Random Logic Encryption, Strong Logic Encryption, and state-of-the-art point-function schemes can all be broken with a limited number of power side-channel traces by utilizing our proposed TA attack. Abhishek Chakraborty 0001, Yang Xie 0001, Ankur Srivastava 0001 |
ICCD | 3 |
| 2017 | Neural TrojansabstractWhile neural networks demonstrate stronger capabilities in pattern recognition nowadays, they are also becoming larger and deeper. As a result, the effort needed to train a network also increases dramatically. In many cases, it is more practical to use a neural network intellectual property (IP) that an IP vendor has already trained. As we do not know about the training process, there can be security threats in the neural IP: the IP vendor (attacker) may embed hidden malicious functionality, i.e neural Trojans, into the neural IP. We show that this is an effective attack and provide three mitigation techniques: input anomaly detection, re-training, and input preprocessing. All the techniques are proven effective. The input anomaly detection approach is able to detect 99.8% of Trojan triggers although with 12.2% false positive. The re-training approach is able to prevent 94.1% of Trojan triggers from triggering the Trojan although it requires that the neural IP be reconfigurable. In the input preprocessing approach, 90.2% of Trojan triggers are rendered ineffective and no assumption about the neural IP is needed. Yuntao Liu 0001, Yang Xie 0001, Ankur Srivastava 0001 |
ICCD | 3 |
| 2017 | Introducing TFUE: The trusted foundry and untrusted employee model in IC supply chain securityabstractIn contrast to other studies in IC supply chain security where foundries are classified as either untrusted or trusted, a more realistic threat model is that the foundries are legally and economically obliged to perform trustworthy service, and it is the individual employees that introduce security risks. We call the above as the trusted foundry and untrusted employee (TFUE) model. Based on this model, we investigate new opportunities of establishing trustworthy operations in foundries made possible by double patterning lithography (DPL). DPL is used to setup two independent mask development lines which do not need to share any information. Under this setup, we consider the attack model where the untrusted employee(s) may try to insert Trojans into the circuit. As a countermeasure, we customize DPL to decompose the layout into two sub-layouts in such a way that each sub-layout individually expose minimum information to the untrusted employee. Yuntao Liu 0001, Chongxi Bao, Yang Xie 0001, Ankur Srivastava 0001 |
ISCAS | 4 |
| 2017 | Power optimizations in MTJ-based Neural Networks through Stochastic ComputingabstractArtificial Neural Networks (ANNs) have found widespread applications in tasks such as pattern recognition and image classification. However, hardware implementations of ANNs using conventional binary arithmetic units are computationally expensive, energy-intensive and have large area overheads. Stochastic Computing (SC) is an emerging paradigm which replaces these conventional units with simple logic circuits and is particularly suitable for fault-tolerant applications. Spintronic devices, such as Magnetic Tunnel Junctions (MTJs), are capable of replacing CMOS in memory and logic circuits. In this work, we propose an energy-efficient use of MTJs, which exhibit probabilistic switching behavior, as Stochastic Number Generators (SNGs), which forms the basis of our NN implementation in the SC domain. Further, error resilient target applications of NNs allow us to introduce Approximate Computing, a framework wherein accuracy of computations is traded-off for substantial reductions in power consumption. We propose approximating the synaptic weights in our MTJ-based NN implementation, in ways brought about by properties of our MTJ-SNG, to achieve energy-efficiency. We design an algorithm that can perform such approximations within a given error tolerance in a single-layer NN in an optimal way owing to the convexity of the problem formulation. We then use this algorithm and develop a heuristic approach for approximating multi-layer NNs. To give a perspective of the effectiveness of our approach, a 43% reduction in power consumption was obtained with less than 1% accuracy loss on a standard classification problem, with 26% being brought about by the proposed algorithm. Ankit Mondal, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2017 | TSV-Based 3-D ICs: Design Methods and ToolsabstractVertically integrated circuits (3-D ICs) may revitalize Moore's law scaling which has slowed down in recent years. 3-D stacking is an emerging technology that stacks multiple dies vertically to achieve higher transistor density independent of device scaling. They provide high-density vertical interconnects, which can reduce interconnect power and delay. Moreover, 3-D ICs can integrate disparate circuit technologies into a single chip, thereby unlocking new system-on-chip architectures that do not exist in 2-D technology. While 3-D integration could bring new architectural opportunities and significant performance enhancement, new thermal, power delivery, signal integrity and reliability challenges emerge as power consumption grows, and device density increases. Moreover, the significant expansion of CPU design space in 3-D requires new architectural models and methodologies for design space exploration (DSE). New design tools and methods are required to address these 3-D-specific challenges. This keynote paper focuses on the state of the art, ongoing advances and future challenges of 3-D IC design tools and methods. The primary focus of this paper is TSV-based 3-D ICs, although we also discuss recent advances in monolithic 3-D ICs. The objective of this paper is to provide a unified perspective on the fundamental opportunities and challenges posed by 3-D ICs especially from the context of design tools and methods. We also discuss the methodology of co-design to address more complicated and interdependent design problems in 3-D IC, and conclude with a discussion of the remaining challenges and open problems that must be overcome to make 3-D IC technology commercially viable. Tiantao Lu, Caleb Serafy, Zhiyuan Yang 0001, Sandeep Kumar Samal, Sung Kyu Lim, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Low-Power Clock Tree Synthesis for 3D-ICsabstractWe propose efficient algorithms to construct a low-power clock tree for through-silicon-via (TSV)-based 3D-ICs. We use shutdown gates to save clock trees’ dynamic power, which selectively turn off certain clock tree branches to avoid unnecessary clock activities when the modules in these tree branches are inactive. While this clock gating technique has been extensively studied in 2D circuits, its application in 3D-ICs is unclear. In 3D-ICs, a shutdown gate is connected to a control signal unit through control TSVs, which may cause placement conflicts with existing clock TSVs in the layout due to TSV’s large physical dimension. We develop a two-phase clock tree synthesis design flow for 3D-ICs: (1) 3D abstract clock tree generation based on K-means clustering and (2) clock tree embedding with simultaneous shutdown gates’ insertion based on simulated annealing (SA) and a force-directed TSV placer. Experimental results indicate that (1) the K-means clustering heuristic significantly reduces the clock power by clustering modules with similar switching behavior and close proximity, and (2) the SA algorithm effectively inserts the shutdown gates to a 3D clock tree, while considering control TSV’s placement. Compared with previous 3D clock tree synthesis techniques, our K-means clustering-based approach achieves larger reduction in clock tree power consumption while ensuring zero clock skew. Tiantao Lu, Ankur Srivastava 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | Mitigating SAT Attack on Logic Locking
Yang Xie 0001, Ankur Srivastava 0001 |
CHES | 2 |
| 2016 | ECO Based Placement and Routing Framework for 3D FPGAs with Micro-fluidic CoolingabstractIntegrated micro-fluidic (MF) cooling is a promising technique to solve the thermal problems in 3D FPGAs [1] (As shown in Figure 1). However, this cooling method has some nonideal properties such as non-uniform heat removal capacity along the flow direction. Existing 3D FPGA placement and routing (P&R) tools are unaware of micro-fluidic cooling, thus leading to large on-chip temperature variation which is harmful to the reliability of 3D FPGAs. In this paper we demonstrate that we can incorporate micro-fluidic cooling considerations in existing 3D FPGA P&R tools simply with a cooling-aware Engineering Change Order (ECO) based placement framework. Taking the placement result of an existing P&R tool, the framework modifies the node positions to improve the on-chip temperature uniformity accounting for fluidic cooling structures. Hence we do not need to invest in a stand alone fluidic cooling aware 3D FPGA CAD framework. Zhiyuan Yang 0001, Caleb Serafy, Ankur Srivastava 0001 |
FCCM | 3 |
| 2016 | Physical Design of 3D FPGAs Embedded with Micro-channel-based Fluidic CoolingabstractThrough Silicon Via (TSV) based 3D integration technology is a promising technology to increase the performance of FPGAs by achieving shorter global wire-length and higher logic density. However, 3D FPGAs also suffer from severe thermal problems due to the increase in power density and thermal resistance. Moreover, past work has shown that leakage power can account for 40\% of the total power at current technology nodes and leakage power increases non-linearly with temperature. This intensifies the thermal problem in 3D FPGAs and more aggressive cooling methods such as micro-channel based fluidic cooling are required to fully exploit their benefits. The interaction between micro-channel heat sink design and the performance of a 3D FPGA is very complicated and a comprehensive approach is required to identify the optimal design of 3D FPGAs subject to thermo-electrical constraints. In this work, we propose an analysis framework for 3D FPGAs embedded with micro-channel-based fluidic cooling to study the impact of channel density on cooling and performance. According to our simulation results, we provide guidelines for designing 3D FPGAs embedded with micro-channel cooling and identify the optimal design for each benchmark. Compared to naive 3D FPGA designs which use fixed thermal heat sink, the optimal design identified using our framework can improve the operating frequency and energy efficiency by up to 80.3% and 124.0%. Zhiyuan Yang 0001, Ankur Srivastava 0001 |
FPGA | 2 |
| 2016 | An optimization-theoretic approach for attacking physical unclonable functionsabstractPhysical unclonable functions (PUFs) utilize manufacturing variations of circuit elements to produce unpredictable response to any challenge vector. The attack on PUF aims to predict the PUF response to all challenge vectors while only a small number of challenge-response pairs (CRPs) are known. The target PUFs in this paper include the Arbiter PUF (ArbPUF) and the Memristor Crossbar PUF (MXbarPUF). The manufacturing variations of the circuit elements in the targeted PUF can be characterized by a weight vector. An optimization-theoretic attack on the target PUFs is proposed. The feasible space for a PUF's weight vector is described by a convex polytope confined by the known CRPs. The centroid of the polytope is chosen as the estimate of the actual weight vector, while new CRPs are adaptively added into the original set of known CRPs. The linear behavior of both ArbPUF and MXbarPUF is proven which ensures that the feasible space for their weight vectors is convex. Simulation shows that our approach needs 71.4% fewer known CRPs and 86.5% less time than the state-of-the-art machine learning based approach. Yuntao Liu 0001, Yang Xie 0001, Chongxi Bao, Ankur Srivastava 0001 |
ICCAD | 4 |
| 2016 | Voltage Noise Induced DRAM Soft Error Reduction Technique for 3D-CPUsabstractThree-dimensional integration enables stacking DRAM on top of CPU, providing high bandwidth and short latency. However, non-uniform voltage fluctuation and local thermal hotspot in CPU layers are coupled into DRAM layers, causing a non-uniform bit-cell leakage (thereby bit flip) distribution. We propose a performance-power-resilience simulation framework to capture DRAM soft error in 3D multi-core CPU systems. A dynamic resilience management (DRM) scheme is investigated, which adaptively tunes CPU's operating points to adjust DRAM's voltage noise and thermal condition during runtime. The DRM uses dynamic frequency scaling to achieve a resilience borrow-in strategy, which effectively enhances DRAM's resilience without sacrificing performance. Tiantao Lu, Caleb Serafy, Zhiyuan Yang 0001, Ankur Srivastava 0001 |
ISLPED | 4 |
| 2016 | On Reverse Engineering-Based Hardware Trojan DetectionabstractDue to design and fabrication outsourcing to foundries, the problem of malicious modifications to integrated circuits (ICs), also known as hardware Trojans (HTs), has attracted attention in academia as well as industry. To reduce the risks associated with Trojans, researchers have proposed different approaches to detect them. Among these approaches, test-time detection approaches have drawn the greatest attention. Many test-time approaches assume the existence of a Trojan-free (TF) chip/model also known as “golden model.” Prior works suggest using reverse engineering (RE) to identify such TF ICs for the golden model. However, they did not state how to do this efficiently. In fact, RE is a very costly process which consumes lots of time and intensive manual effort. It is also very error prone. In this paper, we propose an innovative and robust RE scheme to identify the TF ICs. We reformulate the Trojan-detection problem as clustering problem. We then adapt a widely used machine learning method, ${K}$ -means clustering, to solve our problem. Simulation results using state-of-the-art tools on several publicly available circuits show that the proposed approach can detect HTs with high accuracy rate. A comparison of this approach with our previously proposed approach [1] is also conducted. Both the limitations and application scenarios of the two methods are discussed in detail. Chongxi Bao, Domenic Forte, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Unlocking the True Potential of 3-D CPUs With Microfluidic Coolingabstract3-D integration is a promising technology to sustain transistor density scaling in the future, as well as facilitating new architectural designs that were not possible with traditional integration techniques. However, 3-D integration comes with some serious challenges, chief among them heat removal. A promising technology for thermal issues is microfluidic (MF) cooling. In this paper, we perform a design space analysis study on 3-D CPUs. We show that aggressive cooling solutions such as MF cooling are necessary to unlock the true potential of 3-D ICs. Without such cooling the thermal feasibility region of the design space is significantly reduced. We observe that interactions between thermal, electrical, and physical aspects of 3-D CPUs with MF cooling are substantial, and must be cooptimized during our analysis to correctly identify optimal design points. We simulate a spectrum of 3-D CPU architectures which offer vast improvements to performance, but are energy inefficient and thermally infeasible with air cooling. Furthermore, we show a 2.30× (1.59×) improvement in performance (energy efficiency) when MF cooling and floorplan cooptimization are added to our design space analysis simulation flow. Caleb Serafy, Avram Bar-Cohen, Ankur Srivastava 0001, Donald Yeung |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Electromigration-aware Clock Tree Synthesis for TSV-based 3D-ICsabstractIn 3D-IC technology, electromigration (EM) degradation has become severe due to the high thermal-mechanical stress induced by the Through-Silicon-Vias (TSVs). However, little has been done on designing an EM-robust clock tree for 3D-ICs. In this paper, we propose a systematic EM-aware clock tree synthesis design flow, to enhance the 3D clock tree's EM reliability, with little interference to clock tree's performance metrics such as total wire length and clock skew. We develop a simple TSV's EM objective function based on multi-physics of the mass transportation equation, and validate it against the finite element method (FEM) simulation. Then we use this objective function to formulate a heuristic, based on integer linear programming (ILP), which places the clock TSVs such that the clock tree's EM reliability is maximized. Results show that the our heuristic is able to increase 3D clock's EM lifetime by more than 3.63x with little wire length overhead, while maintaining zero clock skew. Tiantao Lu, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | 3D Integration: New opportunities in defense against cache-timing side-channel attacksabstractRecently, following the work pioneered by Kocher [1], using cache behavior as a timing side-channel to leak critical system information has received lots of attentions because of its easy-to-implement nature and amazingly good results. Recent attacks have been demonstrated to successfully leak the full key from many commonly used encryption algorithms including RSA, AES, etc. These attacks pose great threats to applications that depend on these encryption methods such as banking systems, military systems, etc. To mitigate the increasing threat, numerous countermeasures, mostly software patches, have been proposed. Hardware mitigations, however, have been less pursued. In this paper, we show that emerging 3D integration technology offers new opportunities in defense against these attacks. We propose two cache design mechanisms that can make the attacker's job harder, even impossible. Experimental results show that using our cache design, the side-channel leakage is significantly reduced while still achieving performance gains over a conventional 2D system. Chongxi Bao, Ankur Srivastava 0001 |
ICCD | 2 |
| 2015 | Temperature Tracking: Toward Robust Run-Time Detection of Hardware TrojansabstractThe hardware Trojan threat has motivated development of Trojan detection schemes at all stages of the integrated circuit (IC) lifecycle. While the majority of existing schemes focus on ICs at test-time, there are many unique advantages offered by post-deployment/run-time Trojan detection. However, run-time approaches have been underutilized with prior work highlighting the challenges of implementing them with limited hardware resources. In this paper, we propose three innovative low-overhead approaches for run-time Trojan detection which exploit the thermal sensors already available in many modern systems to detect deviations in power/thermal profiles caused by Trojan activation. The first one is a local sensor-based approach that uses information from thermal sensors together with hypothesis testing to make a decision. The second one is a global approach that exploits correlation between sensors and maintains track of the ICs thermal profile using a Kalman filter (KF). The third approach incorporates leakage power into the system dynamic model and apply extended KF (EKF) to track ICs thermal profile. Simulation results using state-of-the-art tools on ten publicly available Trojan benchmarks verify that all three proposed approaches can detect active Trojans quickly and with few false positives. Among three approaches, EKF is flawless in terms of the ten benchmarks tested but would require the most overhead. Chongxi Bao, Domenic Forte, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | TSV Replacement and Shield Insertion for TSV-TSV Coupling Reduction in 3-D Global PlacementabstractThrough silicon via (TSV) cross coupling can seriously degrade circuit performance of 3-D ICs if it is not considered during design. In this paper, we propose two algorithms which combine coupling-aware TSV placement with shield insertion to yield better results than either technique alone. We first introduce an algorithm for TSV placement assuming a fixed standard cell placement. The result of this algorithm is a 13% reduction in worst case coupling across all TSV pairs and a 90% reduction in the total number of TSV pairs violating an imposed coupling threshold. We then introduce a second algorithm that perturbs a given standard cell and TSV placement to improve coupling. This second algorithm yields a 17% reduction in worst case coupling and removes all coupling violations. Both algorithms cause wirelength (WL) to increase no more than 5%. Our algorithms offer a large improvement to TSV-TSV coupling at the expense of only a meager degradation of total WL, and in many designs and applications this trade-off is well justified. Caleb Serafy, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Modeling and Layout Optimization for Tapered TSVsabstractThrough-silicon-via (TSV) offers vertical connections for 3-D ICs. Due to its large dimensions and nonideal etching process, TSVs layout needs to be carefully optimized to balance peak current density and delay for digital circuit. This brief investigates the TSVs tapering effect (which is an inevitable byproduct of deep reactive Ion etching-based manufacturing) and its impact on the TSVs electrical properties. We show that the current crowding effect is more severe in realistic tapered TSVs than ideal cylindrical TSVs. We propose a nonuniform current density model for tapered TSVs, which achieves considerable accuracy and speedup in estimating the current density distribution, when compared with the existing models developed for cylindrical TSVs. We apply our model to perform a detailed study on: 1) impact of TSVs tapering on peak current density and 2) wire sizing problem to minimize TSV-involved path delay under second-order delay model while keeping the peak current density within tolerable levels. A new dynamic programming-based heuristic is proposed to find the optimal wire configuration, which reduces both peak current density and delay, thereby improving the reliability and performance. Tiantao Lu, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Gated low-power clock tree synthesis for 3D-ICsabstractIn this paper, we minimize 3D clock power using shutdown gates to selectively turn off unnecessary clock activities. In 3D-IC, shutdown signals require large-sized Through-Silicon-Vias(TSVs), so we propose a simulated annealing(SA) based algorithm along with a force-directed TSV placer to decide the selection of shutdown gates and the locations of TSVs under layout whitespace constraint. Furthermore, we recognize optimal power saving is achieved when the clock tree itself is designed simultaneously with the shutdown network. Experimental results show that our heuristic decreases the total clock power by more than 20% with less than 1.5% wirelength overhead while ensuring zero clock skew. Tiantao Lu, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2014 | Unlocking the true potential of 3D CPUs with micro-fluidic coolingabstractAs technology scaling is coming to an end, 3D integration is a promising technology to continue transistor density scaling in the future and facilitate new architectural designs. However heat removal is a serious chalenge in 3D ICs. A promising solution is micro-fluidic (MF) cooling. In this paper we argue that aggressive cooling methods are necessary to unlock the true potential of 3D ICs. We simulate a spectrum of 3D CPU architectures which offer vast improvements to performance, but are inefficient and thermally infeasible with air cooling alone. Our results show that integrating micro-fluidic cooling can increase average performance by 2.62x and energy efficiency by 1.78x by unlocking new architectural configurations. Caleb Serafy, Ankur Srivastava 0001, Donald Yeung |
ISLPED | 2 |
| 2014 | Coupling-aware force driven placement of TSVs and shields in 3D-IC layoutsabstractIn 3D ICs, TSV cross coupling can seriously degrade circuit performance if it is not sufficiently considered in a design. Cross coupling is heavily dependent on how TSVs are placed, and should be considered during the floorplanning of the chip. In this work we propose a coupling-aware TSV placement algorithm that attempts to reduce both wirelength and TSV cross coupling. TSV shielding is another method for coupling mitigation, and the proposed algorithm combines coupling-aware TSV placement with shield insertion to yield better results than either technique alone. With regard to the most heavily coupled TSV pair in a design, our results show that applying both techniques simultaneously produces a 12.3% improvement in maximum S-parameter compared to traditional TSV placement which optimizes wirelength only. Using coupling-aware TSV placement or shield insertion alone produces a 4.2% and 4.8% improvement respectively. The improvement offered by using both techniques simultaneously is actually more than the sum of the improvement offered by using each technique on its own. This implies that the two techniques are not independent of one another, and that when used simultaneously each technique increases the effectiveness of the other, giving strong motivation for using them both simultaneously. Furthermore, the percent increase in wirelength due to using these two techniques is an order of magnitude less than the percent improvement to coupling, justifying the tradeoff made by our algorithm. Caleb Serafy, Ankur Srivastava 0001 |
ISPD | 2 |
| 2014 | A geometric approach to chip-scale TSV shield placement for the reduction of TSV coupling in 3D-ICs
Caleb Serafy, Bing Shi 0001, Ankur Srivastava 0001 |
Integr. | 3 |
| 2014 | Optimized Micro-Channel Design for Stacked 3-D-ICsabstractThe three dimensional circuit (3-D-IC) achieves high performance by stacking several layers of active electronic components vertically. Despite its impact on performance improvement, 3-D-IC also brings great challenges to chip thermal management due to its high heat density. Microchannel-based liquid cooling shows great potential in removing the high density heat inside 3-D circuits. The current microchannel heat sink designs spread the entire surface to be cooled with microchannels. This approach, though it provides sufficient cooling, consumes significant amount of extra cooling power. In this paper, we investigate the design of non-uniformly distributed microchannel cooling systems which provide sufficient cooling with less cooling power. The experiments show that, compared with the conventional design which spreads microchannels all over the chip, our non-uniform microchannel design achieves up to 80% cooling power savings. Bing Shi 0001, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Statistical Framework for Designing On-Chip Thermal Sensing Infrastructure in Nanoscale SystemsabstractThermal/power issues have become increasingly important with more and more transistors being placed on a single chip. Many dynamic thermal/power management techniques have been proposed to address such issues but they all depend heavily on accurate knowledge of the chip's thermal state during runtime. In this paper, we describe a unified statistical framework for designing an on-chip thermal sensing infrastructure that can be used to track the chip's thermal state at runtime. Specifically, we address the following problems in this statistical framework: 1) sensor placement; 2) sensor data compression; 3) sensor data fusion; and 4) overall interplay. Our methods exploit the correlations between temperatures in different parts of the chip to drive sensor placement, data compression, and data fusion in both noiseless and noisy sensor cases. Our framework is also capable of choosing the appropriate degree of compression for each sensor while accounting for their local space constraints during deployment. The experimental results show that the root-mean-square error of the thermal estimates produced by our sensing infrastructure is on average 35% better than an equivalent system that uses a range-based placement scheme and a uniform compression scheme. It took our methods at most about 9 s to decide the overall solution for placement, compression, and data fusion at the design stage. This demonstrates the effectiveness and applicability of our unified statistical design methodology. Yufu Zhang, Bing Shi 0001, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Geometric approach to chip-scale TSV shield placement for the reduction of TSV coupling in 3D-ICsabstractIn 3D ICs, interlayer communication is achieved using through-silicon-vias (TSVs), which can suffer from cross coupling if placed naïvely. In this paper, cross coupling between TSVs is modeled, and a chip-scale TSV coupling mitigation scheme is presented using TSV shielding. A geometric coupling model is developed which is simple enough to quickly estimate the pairwise coupling between TSVs, unlike circuit models of coupling that have been proposed in previous works. Our geometric model's ability to make fast accurate estimations of chip-scale cross coupling make it a good model to use for shield placement optimization. Caleb Serafy, Bing Shi 0001, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | Thermal stress aware 3D-IC statistical static timing analysisabstractIt is widely known that fabrication and thermal variations influence circuit delay. In three dimensional circuits (3D-ICs), due to the incorporation of through-silicon-vias (TSVs), thermal stress also becomes an increasing contributor to gate delay. As a result, thermal variations cause not only direct impact to the circuit parameters, they also cause 3D-IC stress variations resulting in an additional source of variability which needs to be accounted for in statistical static timing analysis (SSTA). In this paper, we study the impact (both direct and indirect - through thermal stress) of thermal variations on gate delay, and propose Monte Carlo (MC) simulation based fabrication, temperature and thermal stress variations aware SSTA methodology that accounts for this complex effect. We show that thermal stress variations cause extra -4~10% delay changes compared with the SSTA that only accounts for fabrication and direct thermal impacts, indicating that thermal stress should be considered together with direct temperature effect. We then develop a more efficient canonical thermal stress aware SSTA method that can achieve good accuracy and 843x speedup compared with MC based method. Bing Shi 0001, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Co-optimization of TSV assignment and micro-channel placement for 3D-ICsabstractThe three dimensional circuit (3D-IC) brings forth new challenges to physical design such as allocation and management of through-silicon-vias (TSVs). Meanwhile, the thermal issues in 3D-IC becomes significant necessitating the use of active cooling schemes such as micro-channel liquid coolings. Both TSVs and micro-channels go through the interlayer regions of 3D-IC resulting in potential resource conflict. This paper investigates the co-optimization of TSV assignment to interlayer nets and micro-channel allocation such that both wirelength and micro-channel cooling energy are co-optimized. We propose a multi-commodity flow based formulation to solve the co-optimization. The experimental results show that, our approach achieves 51% cooling power savings or 6.08% wire length reduction compared with the approaches that assign TSVs and allocate micro-channels separately. Bing Shi 0001, Caleb Serafy, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | Temperature tracking: an innovative run-time approach for hardware Trojan detectionabstractThe hardware Trojan threat has motivated development of Trojan detection schemes at all stages of the integrated circuit (IC) lifecycle. While the majority of existing schemes focus on ICs at test-time, there are many unique advantages offered by post-deployment/run-time Trojan detection. However, run-time approaches have been underutilized with prior work highlighting the challenges of implementing them with limited hardware resources. In this paper, we propose innovative low-overhead approaches for run-time Trojan detection which exploit the thermal sensors already available in many modern systems to detect deviations in power/thermal profiles caused by Trojan activation. Simulation results using state-of-the-art tools on publicly available Trojan benchmarks verify that our approaches can detect active Trojans quickly and with few false positives. Domenic Forte, Chongxi Bao, Ankur Srivastava 0001 |
ICCAD | 3 |
| 2013 | Improving the Quality of Delay-Based PUFs via Optical Proximity CorrectionabstractSilicon physically unclonable functions (PUFs) are circuits that exploit modern manufacturing variations to generate unique signatures for chip authentication and cryptographic key generation. Existing research has focused on improving PUF quality at architectural or design levels, but has ignored opportunities available during fabrication, which is the source of systematic and random variation in (ICs)/PUFs. For typical ICs (where security is not a concern), optical proximity correction (OPC) is used to suppress both these types of variations. However, several prior works have shown that only systematic variations negatively impact PUF quality and random variations are beneficial for PUFs. In this paper, we propose two PUF-aware OPC cost functions: 1) P-OPC generates a PUF lithography mask that increases all variations in PUF circuitry (the opposite of state-of-the-art OPC), and 2) SVC-OPC generates mask patterns that reduce the systematic variation found in PUFs for better quality. Simulation results for ring oscillator (RO) PUFs show that the proposed techniques can improve PUF signature quality compared to current state-of-the-art OPC. Domenic Forte, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Energy- and Thermal-Aware Video Coding via Encoder/Decoder Workload BalancingabstractVideo coding and compression are essential components of multimedia services but are known to be computationally intensive and energy demanding. Traditional video coding paradigms, predictive and distributed video coding (PVC and DVC), result in excessive computation at either the encoder (PVC) or decoder (DVC). Several recent papers have proposed a hybrid PVC/DVC codec which shares the video coding workload between encoder and decoder. In this article, we propose a controller for such hybrid coders that considers energy and temperature to dynamically split the coding workload of a system comprised of one encoder and one decoder. We also present two heuristic algorithms for determining safe operating temperatures in the controller solution: (1) stable state thermal modeling algorithm, which focuses on long term temperatures, and (2) transient thermal modeling algorithm, which is better for short-term thermal behavior. Results show that the proposed algorithms result in more balanced energy utilization, improve overall system lifetime, and reduce operating temperatures when compared to strictly PVC and DVC systems. Domenic Forte, Ankur Srivastava 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Resource-aware architectures for adaptive particle filter based visual target trackingabstractThere are a growing number of visual tracking applications now being envisioned for mobile devices. However, since computer vision algorithms such as particle filtering have large computational demands, they can result in high energy consumption and temperatures in mobile devices. Conventional approaches for distributed target tracking with a camera node and a receiver node are either sender-based (SB) or receiver-based (RB). The SB approach uses little energy and bandwidth, but requires a sender with large computational resources. The RB approach fits applications where computational resources are completely unavailable to the sender, but requires very large energy and bandwidth. In this article, we propose three architectures for distributed particle filtering that (i) reduce particle filtering workload and (ii) allow for dynamic migration of workload between nodes participating in tracking. We also discuss an adaptive particle filtering extension that adapts particle filter computational complexity and can be applied to both the conventional and proposed architectures for improved energy efficiency. Results show that the proposed solutions require low additional overhead, improve on tracking system lifetime, balance node temperatures, maintain track of the desired target, and are more effective than conventional approaches in many scenarios. Domenic Forte, Ankur Srivastava 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2013 | Thermal-aware sensor scheduling for distributed estimationabstractA sensor network is a distributed system where sensor nodes autonomously collect local data and collaborate to solve global problems. Recent work has shown that sensor functionality varies with node temperature. Extreme temperatures can decrease node/network lifetime by leading to premature hardware failure and reducing battery capacity. Furthermore, high temperatures can increase sensor measurement noise and disrupt communication between overheated sensor nodes, thereby interfering with their ability to contribute valuable information to collaborative tasks. In the past, sensor networks only consisted of low-end devices with limited power, computational capabilities, and available bandwidth. Such devices would only experience high temperatures in harsh environments. However, sensor networks are now envisioned for applications that require higher-end devices, such as smart cameras, smart phones, and laptops. The power dissipated by such devices is much larger than low-end sensors and can create thermal emergencies in sensor hardware even in calm environments. In this article, we present unique management opportunities for distributed estimation tasks in sensor networks consisting of high-end devices prone to thermal issues. We attempt to balance both thermal- and performance-related constraints by examining trade-offs between sensor sampling rate, number of sensors, node temperature, and state estimation error. Initially, we devise a scheduling algorithm which can achieve a desired real-time performance constraint while maintaining a thermal limit on temperature assuming identical nodes in the network. Then, we extend the concept to a network consisting of heterogeneous sensor nodes. Analytical results and simulation experiments are done for state estimation with a Kalman filter for simplicity, but our main contributions should easily extend to any form of estimation with measurable error. Results show that our policies can successfully balance the trade-offs between thermal- and performance-related constraints. Note that our analyses, schemes, and results are less applicable to low-end sensors whose operation does not cause high node temperature. This work is most suited for high-performance sensors and upper-tier sensors which experience greater workloads. Domenic Forte, Ankur Srivastava 0001 |
ACM Trans. Sens. Networks | 2 |
| 2013 | Dynamic Thermal Management Under Soft Thermal ConstraintsabstractIn this paper, we investigate dynamic thermal management (DTM) policies under soft thermal constraint that allow the thermal constraint to be violated occasionally for boosting system performance. First, we investigate soft-constraint DTM using lumped radio control (RC) thermal models. We develop analytical expressions for the optimal core frequency policies that maximize overall performance under soft thermal constraint for both single-core and homogeneous multicore processors. We then generalize the problem to heterogeneous multicore processor and use a more accurate distributed RC thermal model to account for the spatial thermal variation. The generalized problem also takes into account the impact of increased temperature on transistor delay and leakage power. The problem is solved by convex optimization. Experimental results indicate that for a two-core processor, a mere 10 °C increase in the core temperature for 100 s results in about 30% performance gain. Bing Shi 0001, Yufu Zhang, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | On improving the uniqueness of silicon-based physically unclonable functions via optical proximity correctionabstractPhysically Unclonable Functions (PUFs) are effective for security applications because they generate unique signatures that are resistant to cloning attempts as well as physical tampering. A silicon PUF is a special circuit embedded in an IC that relies on random fabrication process variations to produce a unique signature for its native IC. While current research directions have focused on improving PUF quality at the architectural level, little work has explicitly targeted their fundamental source of randomness, the fabrication process. During IC fabrication, Optical Proximity Correction (OPC) is typically used to suppress manufacturing variations. In this paper, we recognize that this is actually counterintuitive for PUFs. We provide a novel framework which enables OPC to increase the effects of manufacturing variations within PUF circuitry and produce more randomness in PUFs for greater uniqueness and reliability. The proposed OPC techniques are validated using a population of 100 ring oscillator PUFs. Results show that our schemes provide over five times larger variation in ring oscillator delay, improve PUF uniqueness by 5%, and improve PUF reliability by as much as 70% when compared to conventional OPC. Domenic Forte, Ankur Srivastava 0001 |
DAC | 2 |
| 2012 | TSV-constrained micro-channel infrastructure design for cooling stacked 3D-ICsabstractMicro-channel based liquid cooling has significant capability of removing high density heat in 3D-ICs. The conventional micro-channel structures investigated for cooling 3D-ICs use straight channels. However, the presence of TSVs which form obstacles to the micro-channels prevents distribution of straight micro-channels. In this paper, we investigate the methodology of designing TSV-constrained micro-channel infrastructure. Specifically, we decide the locations and geometry of micro-channels with bended structure so that the cooling effectiveness is maximized. Our micro-channel structure could achieve up to 87% pumping power savings compared with the structure using straight micro-channels. Bing Shi 0001, Ankur Srivastava 0001 |
ISPD | 2 |
| 2012 | Accelerating Gate Sizing Using Graphics Processing UnitsabstractIn this paper, we investigate the gate sizing problem and develop techniques for improving the runtime by effectively exploiting the graphics processing unit (GPU) resources. Theoretically, we investigate a randomized cutting plane-based convex optimization technique which is highly parallelizable and can effectively exploit the single instruction multiple data structure imposed by GPUs. In order to further improve the runtime, we also develop GPU-oriented implementation guidelines that exploit the specific structure that convex gate sizing formulations impose. We implemented our method on NVIDIA Tesla 10 GPU and obtain 21× to 400× speedup compared to the MOSEK optimization tool implemented on conventional CPU. The quality of solution of our method is very close to that achieved by MOSEK, since both are optimal. Bing Shi 0001, Yufu Zhang, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Non-uniform micro-channel design for stacked 3D-ICsabstractMicro-channel cooling shows great potential in removing high density heat in 3D circuits. The current micro-channel heat sink designs spread the entire surface to be cooled with micro-channels. This approach, though might provide sufficient cooling, requires quite high pumping power. In this paper, we investigate the non-uniform allocation of micro-channels to provide sufficient cooling with less pumping power. Specifically, we decide the count, location and pumping pressure drop/flow rate of micro-channels such that acceptable cooling is achieved at minimum pumping power. Thermal wake effect and runtime pressure drop/flow rate control are also considered. The experiments showed that, compared with the conventional design which spreads micro-channels all over the chip, our non-uniform microchannel design achieves 55--60% pumping power saving. Bing Shi 0001, Ankur Srivastava 0001 |
DAC | 2 |
| 2011 | Adaptable architectures for distributed visual target trackingabstractThere are a growing number of visual tracking applications for mobile devices. However, the computer vision algorithms which process real-time video to track moving targets are demanding. Since a single mobile device possesses limited computational capabilities, energy, etc. to fully support target tracking, some works have investigated architectures which migrate a portion of tracking duties to another device at the cost of transmission bandwidth and energy. In this paper, we investigate the resource utilization in such architectures and present an adaptable architecture which balances tracking workload among the participating devices based on current resource availability (energy, temperature, bandwidth). Results show that the proposed solution requires low additional overhead, can improve on tracking system lifetime by reducing energy consumption, and is more effective in maintaining safe operating temperatures within participants as compared to previously investigated architecture Domenic Forte, Ankur Srivastava 0001 |
ICCD | 2 |
| 2011 | Energy-aware and quality-scalable data placement and retrieval for disks in video server environmentsabstractAs the popularity of video streaming over the Internet grows, energy consumption in video server environments which store and retrieve video data increases as well. Previous work has shown that video quality delivered to clients can be scaled in order to serve more concurrent video requests and/or reduce energy consumption of server disks. We propose a data placement strategy for such quality scaling methods which distributes video data within a disk based on its priority/importance. Results show that in doing so the disk can retrieve data with greater efficiency and serve lower quality video to more clients than previously investigated strategies. Domenic Forte, Ankur Srivastava 0001 |
ICCD | 2 |
| 2011 | Accurate Temperature Estimation Using Noisy Thermal Sensors for Gaussian and Non-Gaussian CasesabstractMulticore system-on-chips (SOCs) rely on runtime thermal monitoring using on-chip thermal sensors for dynamic thermal management (DTM). However, on-chip sensors are highly susceptible to noise due to fabrication randomness,VDDfluctuations, etc. This causes discrepancy between the actual temperature and the one observed by thermal sensor. In this paper, we address the problem of estimating the accurate temperature of on-chip thermal sensor when the sensor reading has been corrupted by noise. We present statistical techniques for the following: 1) when the underlying randomness exhibits jointly-Gaussian characteristics we present the optimal solution for temperature estimation; 2) for close to Gaussian cases we give a heuristic based on Moment Matching; 3) when the underlying randomness is non-Gaussian a hypothesis testing framework is used to predict the sensor temperatures. The previous three techniques are investigated in both single sensor and multisensor scenarios, respectively. The latter tries to estimate the actual temperatures for several sensors simultaneously while exploiting the correlations in temperature and circuit parameters among different sensors. The experiments showed that using our estimation schemes the root mean square (RMS) error can be reduce (with very small runtime overhead) by 71.5% as compared to blindly trusting the sensors to be noise-free. Yufu Zhang, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Adaptive and autonomous thermal tracking for high performance computing systemsabstractMany DTM schemes rely heavily on the accurate knowledge of the chip's dynamic thermal state to make optimal performance/temperature trade-off decisions. This information is typically generated using a combination of thermal sensor inputs and various estimation schemes such as Kalman filter. A basic assumption used by such schemes is that the statistical characteristics of the power consumption do not change. This is problematic since such characteristics are heavily application dependent. In this paper, we first present autonomous schemes for detecting the change in the statistical characteristics of power and then propose adaptive schemes for capturing such new statistical parameters dynamically. This could enable accurate temperature estimation during runtime given dynamically changing power statistical states. Our schemes use a combination of hypothesis testing and residual whitening methods and can improve the accuracy by 67% as compared to the traditional non-adaptive schemes. Yufu Zhang, Ankur Srivastava 0001 |
DAC | 2 |
| 2010 | Thermal-Aware Sensor Scheduling for Distributed Estimation
Domenic Forte, Ankur Srivastava 0001 |
DCOSS | 2 |
| 2010 | Energy and thermal-aware video coding via encoder/decoder workload balancingabstractEven with consistent advances in storage and transmission capacity, video coding and compression are essential components of multimedia services. Traditional video coding paradigms result in excessive computation at either the encoder or decoder. However, several recent papers have proposed a hybrid PVC/DVC (Predictive/Distributed Video Coding) codec which shares the video coding workload. In this paper, we propose a controller for such hybrid coders that considers energy and temperature to dynamically split the coding workload of a system comprised of one encoder and one decoder. Results show that the proposed controller results in more balanced energy utilization, improving overall system lifetime and reducing operating temperatures when compared to strictly PVC and DVC systems. Domenic Forte, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2010 | Dynamic thermal management for single and multicore processors under soft thermal constraintsabstractIn this paper, we investigate Dynamic Thermal Management (DTM) policies under soft thermal constraint that allows the thermal constraint to be violated for a user specified period. For single core processor, we develop analytical expression for the optimal frequency policy under the soft constraint such that maximal performance can be extracted. We extend this problem to multi-core processor and provide optimal frequency policy when all cores run at same frequency. We also present LP based approximated formulation that generates frequency policies where each core has separate frequency control and considers leakage power. We use frequency legalization to approximate the frequency into discrete values. Experimental results indicate that 10 degree celsius increase in core temperature for 100sec results in 13% performance gain for single core processor, and 30% performance gain for two-core processor. Without Tmax constraint, the performance improves almost 100% for two-core processor. Bing Shi 0001, Yufu Zhang, Ankur Srivastava 0001 |
ISLPED | 3 |
| 2010 | A statistical framework for designing on-chip thermal sensing infrastructure in nano-scale systemsabstractThermal/power issues have become increasingly important with more and more transistors being put on a single chip. Many dynamic thermal/power management techniques have been proposed to address such issues but they all heavily depend on accurate knowledge of the chip's thermal state during runtime. In this paper we describe a unified statistical framework for designing an on-chip thermal sensing infrastructure which can be used to track the chip's thermal state at runtime. Specifically we address the following problems: (1)sensor placement; (2)sensor data compression; (3)sensor data fusion; (4)overall interplay. Our methods exploit the thermal correlation to generate the overall solution in both the noiseless and noisy sensor settings. Our framework is also capable of choosing the appropriate degree of compression for each sensor while accounting for their local space constraints when doing the sensor deployment. The experimental results showed that our infrastructure can improve the temperature estimation accuracy by 27% (on average) as compared to an equivalent system that uses range-based placement and uniform compression. It took our methods about 6.3 seconds to decide the overall solution for placement, compression and data fusion at design stage. This demonstrates the effectiveness and applicability of our unified statistical design methodology. Yufu Zhang, Bing Shi 0001, Ankur Srivastava 0001 |
ISPD | 3 |
| 2010 | On-chip sensor-driven efficient thermal profile estimation algorithmsabstractThis article addresses the problem of chip-level thermal profile estimation using runtime temperature sensor readings. We address the challenges of: (a) availability of only a few thermal sensors with constrained locations (sensors cannot be placed just anywhere); (b) random chip power density characteristics due to unpredictable workloads and fabrication variability. Firstly we model the random power density as a probability density function. Given such statistical characteristics and the runtime thermal sensor readings, we exploit the correlation in power dissipation among different chip modules to estimate the expected value of temperature at each chip location. Our methods are optimal if the underlying power density has Gaussian nature. We give a heuristic method to estimate the chip-level thermal profile when the underlying randomness is non-Gaussian. An extension of our method has also been proposed to address the dynamic case. Several speedup strategies are carefully investigated to improve the efficiency of the estimation algorithm. Experimental results indicated that, given only a few thermal sensors, our method can generate highly accurate chip-level thermal profile estimates within a few milliseconds. Yufu Zhang, Ankur Srivastava 0001, Mohamed Zahran 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | Accurate temperature estimation using noisy thermal sensorsabstractMulticore SOCs rely on runtime thermal measurements using on-chip sensors for DTM. In this paper we address the problem of estimating the actual temperature of on-chip thermal sensor when the sensor reading has been corrupted by noise. Thermal sensors are prone to noise due to fabrication randomness, VDD fluctuations etc. This causes discrepancy between actual temperature and the one predicted by thermal sensor. Our experiments estimate this variation to be around 30%. In this paper we present a statistical methodology for predicting the actual temperature for a given sensor reading. We present two techniques: single sensor prediction and multi-sensor prediction. The latter tries to estimate the actual temperature for each sensor (of the many on-chip sensors) simultaneously while exploiting the correlations between temperature and noise of different sensors. When the underlying randomness follows a Gaussian characteristic, we present optimal schemes of estimating the expected temperature. We also present heuristic schemes for the case where the Gaussian assumption fails to hold. The experiments showed that using our estimation schemes the RMS error can be reduce as much as 67% as compared to blindly trusting the sensors to be noise free. Yufu Zhang, Ankur Srivastava 0001 |
DAC | 2 |
| 2008 | Chip level thermal profile estimation using on-chip temperature sensorsabstractThis paper addresses the problem of chip level thermal profile estimation using runtime temperature sensor readings. We address the challenges of a) availability of only a few thermal sensors with constrained locations (sensors cannot be placed just anywhere) b) random on-chip power density characteristics due to unpredictable workloads and fabrication variability. Firstly we model the random power density as a probability density function. Given this random characteristic and runtime thermal sensor readings, we exploit the correlation between power dissipation of different chip modules to estimate the expected value of temperature at each chip location. Our methods are optimal if the underlying power density has Gaussian nature. We also present a heuristic to generate the chip level thermal profile estimates when the underlying randomness is non-Gaussian. Experimental results indicate that our method generates highly accurate thermal profile estimates of the entire chip at runtime using only a few thermal sensors. Yufu Zhang, Ankur Srivastava 0001, Mohamed Zahran 0001 |
ICCD | 2 |
| 2008 | Variability-Driven Formulation for Simultaneous Gate Sizing and Postsilicon Tunability AllocationabstractProcess variations cause design performance to become unpredictable in deep submicrometer technologies. Several statistical techniques (timing analysis, gate sizing, and buffer insertion) have been proposed to counter these variations during the optimization phase of the design flow to get a better timing yield. Another interesting approach to improve the timing yield is postsilicon-tunable (PST) clock tree. In this paper, we propose such an integrated framework that performs simultaneous statistical gate sizing in the presence of PST clock-tree buffers for minimizing binning yield loss (YL) and tunability costs by determining the ranges of delay tuning to be provided at each buffer. The simultaneous gate sizing and PST-buffer range determination problem is proved to be a convex-stochastic programming formulation under longest path-delay constraints and, hence, solved optimally. We further extend the formulation into a heuristic to additionally consider shortest path-delay constraints. We make experimental comparisons using nominal gate sizing followed by PST-buffer management using the work of Tsai as a base case. We take the solution obtained from this approach and perform the following: 1) sensitivity-based statistical gate sizing while retaining the PST clock tree and 2) simultaneous gate sizing and PST-buffer range determination as proposed in this paper. On an average, the base-case approach gave 23% timing YL, the sensitivity approach gave 15% YL, whereas our proposed algorithm gave only 4% YL. Vishal Khandelwal, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Algorithmic and Architectural Optimizations for Computationally Efficient Particle FilteringabstractIn this paper, we analyze the computational challenges in implementing particle filtering, especially to video sequences. Particle filtering is a technique used for filtering nonlinear dynamical systems driven by non-Gaussian noise processes. It has found widespread applications in detection, navigation, and tracking problems. Although, in general, particle filtering methods yield improved results, it is difficult to achieve real time performance. In this paper, we analyze the computational drawbacks of traditional particle filtering algorithms, and present a method for implementing the particle filter using the Independent Metropolis Hastings sampler, that is highly amenable to pipelined implementations and parallelization. We analyze the implementations of the proposed algorithm, and, in particular, concentrate on implementations that have minimum processing times. It is shown that the design parameters for the fastest implementation can be chosen by solving a set of convex programs. The proposed computational methodology was verified using a cluster of PCs for the application of visual tracking. We demonstrate a linear speed-up of the algorithm using the methodology proposed in the paper. Aswin C. Sankaranarayanan, Ankur Srivastava 0001, Rama Chellappa |
IEEE Trans. Image Process. | 2 |
| 2008 | Variability Driven Gate Sizing for Binning Yield OptimizationabstractHigh performance applications are highly affected by process variations due to considerable spread in their expected frequencies after fabrication. Typically ldquobinningrdquo is applied to those chips that are not meeting their performance requirement after fabrication. Using binning, such failing chips are sold at a loss (e.g., proportional to the degree that they are failing their performance requirement). This paper discusses a gate-sizing algorithm to minimize ldquoyield-lossrdquo associated with binning. We propose a binning yield-loss function as a suitable objective to be minimized. We show this objective is convex with respect to the size variables and consequently can be optimally and efficiently solved. These contributions are yet made without making any specific assumptions about the sources of variability or how they are modeled. We show computation of the binning yield-loss can be done via any desired statistical static timing analysis (SSTA) tool. The proposed technique is compared with a recently proposed sensitivity-based statistical sizer, a deterministic sizer with worst-case variability estimate, and a deterministic sizer with relaxed area constraint. We show consistent improvement compared to the sensitivity-based approach in quality of solution (final binning yield-loss value) as well as huge run-time gain. Moreover, we show that a deterministic sizer with a relaxed area constraint will also result in reasonably good binning yield-loss values for the extra area overhead. Azadeh Davoodi, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Monte-Carlo driven stochastic optimization framework for handling fabrication variabilityabstractIncreasing effects of fabrication variability have inspired a growing interest in statistical techniques for design optimization. In this work, we propose a Monte-Carlo driven stochastic optimization framework that does not rely on the distribution of the varying parameters (unlike most other existing techniques). Stochastic techniques like Successive Sample Mean Optimization (SSMO) and Stochastic Decomposition present a strong framework for solving linear programming formulations in which the parameters behave as random variables. We consider Binning-Yield Loss (BYL) as the optimization objective and show that we can get a provably optimal solution under a convex BYL function. We apply this framework for the MTCMOS sizing problem [21] using SSMO and Stochastic Decomposition techniques. The experimental results show that the solution obtained from stochastic decomposition based framework had 0% yield-loss, while the deterministic solution [21] had a 48% yield-loss. Vishal Khandelwal, Ankur Srivastava 0001 |
ICCAD | 2 |
| 2007 | Statistical timing analysis using Kernel smoothingabstractWe have developed a new statistical timing analysis approach that does not impose any assumptions on the nature of manufacturing variability and takes into account an arbitrary model of spatial correlation as well as all types of functional correlations (e.g. reconvergence-based correlations). The starting point for statistical timing analysis is small scale Monte Carlo (MC) simulation. In order to speed-up the MC simulation process we use stratified balanced sampling and postprocessing of the simulation data using non-parametric kernel estimation. The MC simulation and the statistical analysis procedure are interleaved with the calculation of the critical paths. In order to speed up simulation, we identify and simulate only gates relevant for calculation of the clock cycle time. The application of statistical techniques enable not only accurate statistical timing analysis, but also stability and scalability analysis. The approach is evaluated using MCNC benchmarks and yields more than six orders of magnitude speed improvement compared with the standard MC simulation. Jennifer Wong-Ma, Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001, Miodrag Potkonjak |
ICCD | 4 |
| 2007 | Variability-driven formulation for simultaneous gate sizing and post-silicon tunability allocationabstractProcess variations cause design performance to become unpredictable in deep sub-micron technologies. Several statistical techniques (timing analysis, gate-sizing) have been proposed to counter these variations during design optimization. Another interesting approach to improve timing yield is post-silicon tunable (PST) clock-tree. In this work, we propose an integrated framework that performs simultaneous statistical gate-sizing in presence of PST clock-tree buffers for minimizing binning-yield loss (BYL) and tunability costs by determining the ranges of tuning to be provided at each buffer. The simultaneous gate-sizing and PST bu er range deter- mination problem is proved to be a convex stochastic programming formulation under longest path delay constraints and hence solved optimally. We further extend the formulation into a heuristic to additionally consider shortest path delay constraints. We make experimental comparisons using nominal gate sizing followed by PST bu er management using [12] as a base-case. We take the solution obtained from this approach and perform 1) Sensitivity-based statistical gate-sizing while retaining the PST clock tree 2) Simultaneous gate sizing and PST buffer range determination as proposed in this work. On an average, the BYL obtained from our approach is 98% lower than the base-case ([12]) and 95% lower than the sensitivity-based algorithm. On an average the base-case approach [12] gave 22% timing yield loss (YL), the sensitivity approach gave 19% YL, where as our proposed algorithm gave only 3% YL. The total PST tuning buffer range that is allocated through the proposed algorithm is comparable to that obtained from [12]. The proposed algorithm had a 2.2x runtime speedup compared to the sensitivity-based algorithm. Vishal Khandelwal, Ankur Srivastava 0001 |
ISPD | 2 |
| 2007 | Active mode leakage reduction using fine-grained forward body biasing strategy
Vishal Khandelwal, Ankur Srivastava 0001 |
Integr. | 2 |
| 2007 | Leakage Control Through Fine-Grained Placement and Sizing of Sleep TransistorsabstractMultithreshold CMOS (MTCMOS) technology has become a popular technique for standby power reduction. Sleep transistor insertion in circuits is an effective application of MTCMOS technology for reducing leakage power. In this paper, we present a fine-grained approach where each gate in the circuit is provided with an independent sleep transistor. Key advantages of this approach include better circuit slack utilization and improvements in ground-bounce-related signal integrity (which is a major disadvantage in clustering-based approaches). To this end, we propose an optimal polynomial-time fine-grained sleep transistor sizing algorithm. We also prove the selective sleep transistor placement problem as NP-complete and propose an effective heuristic. Finally, in order to reduce the sleep transistor area penalty, we propose a placement-area-constrained sleep transistor sizing formulation. Our experiments show that, on average, the sleep transistor placement and optimal sizing algorithms gave 50.9% and 46.5% savings in leakage power as compared with the conventional fixed-delay penalty algorithms for 5% and 7% circuit slowdown, respectively. Moreover, the postplacement area penalty was less than 5%, which is comparable to clustering schemes. Vishal Khandelwal, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | A Quadratic Modeling-Based Framework for Accurate Statistical Timing Analysis Considering CorrelationsabstractThe impact of parameter variations on timing due to process variations has become significant in recent years. In this paper, we present a statistical timing analysis (STA) framework with quadratic gate delay models that also captures spatial correlations. Our technique does not make any assumption about the distribution of the parameter variations, gate delays, and arrival times. We propose a Taylor-series expansion-based quadratic representation of gate delays and arrival times which are able to effectively capture the nonlinear dependencies that arise due to increasing parameter variations. In order to reduce the computational complexity introduced due to quadratic modeling during STA, we also propose an efficient linear modeling driven quadratic STA scheme. We ran two sets of experiments assuming the global parameters to have uniform and Gaussian distributions, respectively. On an average, the quadratic STA scheme had 20.5times speedup in runtime as compared to Monte Carlo simulations with an rms error of 0.00135 units between the two timing cummulative density functions (CDFs). The linear modeling driven quadratic STA scheme had 51.5times speedup in runtime as compared to Monte Carlo simulations with an rms error of 0.0015 units between the two CDFs. Our proposed technique is generic and can be applied to arbitrary variations in the underlying parameters under any spatial correlation model Vishal Khandelwal, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Variability driven gate sizing for binning yield optimizationabstractProcess variations result in a considerable spread in the frequency of the fabricated chips. In high performance applications, those chips that fail to meet the nominal frequency after fabrication are either discarded or sold at a loss which is typically proportional to the degree of timing violation. The latter is called binning. In this paper we present a gate sizing-based algorithm that optimally minimizes the binning yield-loss. Specifically we make the following contributions: 1) prove the binning yield function to be convex, 2) the proof does not make any assumptions about the sources of variability, their distributions (Gaussian/Non-Gaussian) or correlation, 3) by using Kelley's cutting-plane method for convex programs, we integrate our strategy with statistical timing analysis tools (STA), without making any assumptions about how STA is done, 4) if the objective is to optimize the traditional yield (and not binning yield) our approach can still optimize the same to a very large extent. Comparison of our approach with sensitivity-based approaches under fabrication variability shows an improvement of on average 72% in the binning yield-loss with an area overhead of an average 6%, while achieving a 2.69 times speedup under a stringent timing constraint. Moreover we show that a worstcase deterministic approach fails to generate a solution for certain delay constraints. We also show that optimizing the binning yield-loss minimizes the traditional yield-loss (although it is not a direct objective) with a 61% improvement from a sensitivity-based approach. Azadeh Davoodi, Ankur Srivastava 0001 |
DAC | 2 |
| 2006 | Probabilistic evaluation of solutions in variability-driven optimizationabstractVLSI design optimization requires evaluation of different solutions, to compare superiority of one over the other. Typically, a solution is superior if it has a better associated timing and cost. In the presence of fabrication variability, the timing and cost of a solution become random variables with spatial and functional correlations. Therefore the evaluation of solutions shall be performed probabilistically to determine the probability that a solution has better cost and timing. In this paper we propose and evaluate three methods for fast and accurate probabilistic comparison of solutions: 1) regular Monte Carlo simulation (as a basis of comparison), 2) joint-pdf approximation using moment matching, and 3) bound-based Conditional Monte Carlo simulation.We integrated these methods in a variability-driven leakage optimization framework using dual threshold voltages. Experimental results show that joint-pdf based approximation is very fast, however it results in sub-optimal solutions due to lower accuracy. Conditional Monte Carlo method is on average 25 times faster than regular Monte Carlo, but slower than approximating joint-pdf. It also results in additional improvement in expected leakage, when compared to joint-pdf method. Monte Carlo simulation is extremely slow and inapplicable to an optimization framework. Deterministic approaches that are based on worst-case estimates had the highest expected leakage. Azadeh Davoodi, Ankur Srivastava 0001 |
ISPD | 2 |
| 2006 | Probabilistic Evaluation of Solutions in Variability-Driven OptimizationabstractVery large-scale integration design optimization requires comparison of different solutions to evaluate superiority of one over the other. Typically, a solution is superior if it has a better associated timing and cost. In the presence of fabrication variability, the timing and cost of a solution become random variables with spatial and functional correlations. Therefore, the evaluation of solutions shall be performed probabilistically to determine the probability that a solution has better cost and timing. In this paper, the authors propose/evaluate three methods for fast and accurate computation of this probability: 1) regular Monte Carlo (MC) simulation (as a basis of comparison); 2) joint probability density function (jpdf) approximation using moment matching; and 3) bound-based conditional-MC simulation. They integrated these methods in a variability-driven leakage optimization framework using dual threshold voltages. Their results show that jpdf approximation is efficient; however, it results in suboptimal solutions due to lower accuracy approximating jpdf Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | A statistical methodology for wire-length predictionabstractIn this paper, the classic wire-length estimation problem is addressed and a new statistical wire-length estimation approach that captures the probability distribution function of net lengths after placement and before routing is proposed. These types of models are highly instrumental in formalizing a complete and consistent probabilistic approach to design automation and design closure where, along with optimizing the pertinent cost function, the associated prediction error is also considered. The wire-length prediction model was developed using a combination of parametric and nonparametric statistical techniques. The model predicts not only the length of the net using input parameters extracted from the floorplan of a design, but also probability distributions that a net with given characteristics after placement will have a particular length. The model is validated using the learn-and-test and resubstitution techniques. The model can be used for a variety of purposes, including the generation of a large number of statistically sound, and therefore realistic, instances of designs. The net models were applied to the probabilistic buffer-insertion problem and substantial improvement was obtained in net delay after routing (~ 20%) when compared to a traditional bounding box (BBOX)-based buffer-insertion strategy Jennifer Wong-Ma, Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001, Miodrag Potkonjak |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | Effective techniques for the generalized low-power binding problemabstractThis article proposes two very fast graph theoretic heuristics for the low power binding problem given fixed number of resources and multiple architectures for the resources. First, the generalized low power binding problem is formulated as an Integer Linear Programming (ILP) problem that happens to be an NP-complete task to solve. Then two polynomial-time heuristics are proposed that provide a speedup of up to 13.7 with an extremely low penalty for power when compared to the optimal ILP solution for our selected benchmarks. Azadeh Davoodi, Ankur Srivastava 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2005 | Simultaneous floorplanning and resource binding: a probabilistic approachabstractIn this work we present a probabilistic approach to simultaneous floorplanning and resource binding for low power. Traditional approaches iteratively perform floorplanning and resource binding while using crude deterministic wire-length estimates like bounding box (since we do not have routing information for inter module inter-connect). Non-availability of accurate wire-length results in suboptimal design and failure of timing closure. In this work we model the wire-lengths as probability distributions and propose a novel probabilistic optimization methodology. Experimental results using state of the art commercial and academic tools were conducted. The novelty in this work is in the higher chance of ending with a feasible design that is synthesizable without losing in overall power (interconnect + module + register). Experimental results show that on-average the number of unsynthesized modules after routing for Mediabench benchmarks were 2 in the conventional case, while on average our probabilistic approach had all modules synthesized after routing. Azadeh Davoodi, Ankur Srivastava 0001 |
ASP-DAC | 2 |
| 2005 | Wake-up protocols for controlling current surges in MTCMOS-based technologyabstractThis paper proposes strategies to control the wake-up noise for circuits implemented in MTCMOS technology. In MTCMOS circuits, during the switchings between the active and standby modes, sudden surges in current happens due to floating voltages at the nodes. These surges might violate the reliability of the circuit. In this paper we address the above problem by developing wake-up strategies to control these current surges as the circuit is getting turned on. Through gradually turning on a circuit a smaller current will be drawn from the power-grid network. A novel partitioning technique is proposed for MTCMOS circuits under a given constraint of maximum drawn-current from the power-grid network. Two approaches are proposed in this paper; the optimal ILP-based formulation and a polynomial-time heuristic. Experimental results show that up to 90.7% improvement in peak drawn-current is obtained with a maximum of 4 clock cycles time to turn on the circuit. Also result show the effectiveness of the heuristic in terms of the quality of solution and a run-time of up to 6600 times faster than the ILP approach for larger circuits. Azadeh Davoodi, Ankur Srivastava 0001 |
ASP-DAC | 2 |
| 2005 | A general framework for accurate statistical timing analysis considering correlationsabstractThe impact of parameter variations on timing due to process and environmental variations has become significant in recent years. With each new technology node this variability is becoming more prominent. In this work, we present a general Statistical Timing Analysis (STA) framework that captures spatial correlations between gate delays. Our technique does not make any assumption about the distributions of the parameter variations, gate delay and arrival times. We propose a Taylor-series expansion based polynomial representation of gate delays and arrival times which is able to effectively capture the non-linear dependencies that arise due to increasing parameter variations. In order to reduce the computational complexity introduced due to polynomial modeling during STA, we propose an efficient linear-modeling driven polynomial STA scheme. On an average the degree-2 polynomial scheme had a 7.3x speedup as compared to Monte Carlo with 0.049 units of rms error w.r.t Monte Carlo. Our technique is generic and can be applied to arbitrary variations in the underlying parameters. Vishal Khandelwal, Ankur Srivastava 0001 |
DAC | 2 |
| 2005 | VLSI CAD tool protection by birthmarking design solutionsabstractMany techniques have been proposed in the past for the protection of VLSI design IPs (intellectual property). CAD tools and algorithms are intensively used in all phases of modern VLSI designs; however, little has been done to protect them. Basically, given a problem Ρ and a solution Σ, we want to be able to determine whether Σ is obtained by a particular tool or algorithm.We propose two techniques that intentionally leave some trace or birthmark, which refers to certain easy detectable properties, in the design solutions to facilitate CAD tool tracing and protection. The pre-processing technique provides the ideal protection at the cost of losing control of solution's quality. The post-processing technique balances the level of protection and design quality.We conduct a case study on how to protect a timing-driven gate duplication algorithm. Experimental results on a large set of MCNC benchmarks confirm that the pre-processing technique results in a significant reduction (about 48%) of the optimization power of the tool, while the post-processing technique has almost no penalty (less than 2%) on the tool's performance. Gang Qu 0001, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Variability-Driven Buffer Insertion Considering CorrelationsabstractIn this work, we investigate the buffer insertion problem under process variations. Sub 100-nm fabrication process causes significant variations on many design parameters. We propose a probabilistic buffer insertion method assuming variations on both interconnect and buffer parameters and consider their correlations due to common sources of variation. Our proposed method is compatible with the more accurate DSM wire-delay model, as well as the Elmore delay model. In addition, a probabilistic pruning criterion is proposed to evaluate potential solutions, while considering their correlations. Experimental results demonstrate that considering correlations using the more accurate DSM delay model results in meeting the timing constraint with an average probability of 0.63. However probabilistic buffer insertion ignoring correlations and deterministic methods, meet the timing constraint with an average probability of 0.25 and 0.19 respectively. Azadeh Davoodi, Ankur Srivastava 0001 |
ICCD | 2 |
| 2005 | Algorithmic and Architectural Design Methodology for Particle Filters in HardwareabstractIn this paper, we present algorithmic and architectural methodology for building particle filters in hardware. Particle filtering is a new paradigm for filtering in presence of nonGaussian nonlinear state evolution and observation models. This technique has found wide-spread application in tracking, navigation, detection problems especially in a sensing environment. So far most particle filtering implementations are not lucrative for real time problems due to excessive computational complexity involved. In this paper, we re-derive the particle filtering theory to make it more amenable to simplified VLSI implementations. Furthermore, we present and analyze pipelined architectural methodology for designing these computational blocks. Finally, we present an application using the bearing only tracking problem and evaluate the proposed architecture and algorithmic methodology. Aswin C. Sankaranarayanan, Rama Chellappa, Ankur Srivastava 0001 |
ICCD | 3 |
| 2005 | Probabilistic dual-Vth leakage optimization under variabilityabstractIn this paper we address the problem of growing leakage variability through effective dual-threshold voltage assignment. We propose a probabilistic dynamic programming-based method to assign dual-threshold voltages such that the overall expected leakage is minimized under a given probability of violating the timing constraint (timing yield). The key characteristics of our strategy are two pruning criteria that stochastically identify pareto-optimal solutions and prune the sub-optimal ones. Compared to other variability-driven dual-threshold voltage assignment schemes, the main advantages of our approach are 1) considering correlations due to common sources of variation, 2) providing controllable runtime, which in one of the proposed strategies is comparable to the deterministic algorithm, and 3) performing optimization based on all the signal paths simultaneously, as opposed to one path at a time. Experimental results indicate that the proposed probabilistic scheme is significantly better than a comparable deterministic dual-threshold voltage assignment, both in terms of expected leakage and the probability of violating the timing constraint Azadeh Davoodi, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2005 | On effective slack management in postscheduling phaseabstractIn this paper, we propose techniques for effective slack management in high-level synthesis. Our design methodology improves the usability of slack. This manifests itself in the form of relaxed latency constraints on resources. Relaxed latency constraints could be exploited to generate designs with better power, area, routability, and other measures. The slack-management engine has two key components: delay budgeting and resource binding. We propose a left edge traversal-based algorithm for delay budgeting. For resource binding, we developed an algorithm that applies a locally optimal binding procedure at each clock step. In order to demonstrate the effectiveness of our strategy, we built an experimental flow that integrated SUIF, Synopsys Design Compiler, Cadence Silicon Ensemble, and our own optimization tools. Experiments with the MediaBench suite shows that our methodology could generate designs with better quality than designs and faster design closure when compared with designs generated without slack management. Ankur Srivastava 0001, Seda Ogrenci Memik, Bo-Kyung Choi, Majid Sarrafzadeh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Voltage scheduling under unpredictabilities: a risk management paradigmabstractThis article addresses the problem of voltage scheduling in unpredictable situations. The voltage scheduling problem assigns voltages to operations such that the power is minimized under a clock delay constraint. In the presence of unpredictabilities, meeting the clock latency constraint cannot be guaranteed. This article proposes a novel risk management based technique to solve this problem. Here, the risk management paradigm assigns a quantified value to the amount of risk the designer is willing to take on the clock cycle constraint. The algorithm then assigns voltages in order to meet the expected value of clock cycle constraint while keeping the maximum delay within the specified “risk” and minimizing the power. The proposed algorithm is based on dynamic programming and is optimal for trees. Experimental results show that the traditional voltage scheduling approach is incapable of handling unpredictabilities. Our approach is capable of generating an effective tradeoff between power and “risk”: the more the risk, the less the power. The results show that a small increase in design risk positively affects the power dissipation. Azadeh Davoodi, Ankur Srivastava 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2005 | Power-driven simultaneous resource binding and floorplanning: a probabilistic approachabstractFloorplanning information is integrated during resource binding for better modeling of the interconnect effects on timing and power. Although this integration improves the estimation of the interconnect effects, nonavailability of exact net-lengths can result in suboptimal solutions, because global routing is not yet performed. In this work we propose a probabilistic approach to integrate floorplanning and resource binding by modeling the distribution of the net-lengths from a given floorplan. The advantage of this approach is that a probabilistic technique can better capture the inaccuracy associated with net-length estimation, and consequently, the inaccuracy in estimation of net-delay and net-power. The result is higher chance of successful synthesis, and therefore faster timing closure. Additionally, due to better management of uncertainty, it has a better overall post-synthesis power. These results are illustrated in our experiments that were conducted using state of the art commercial and academic tools. Azadeh Davoodi, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Simultaneous Vt selection and assignment for leakage optimizationabstractThis paper presents a novel approach for leakage optimization through simultaneous V/sub t/ selection and assignment. V/sub t/ selection implies deciding the right value for V/sub t/ and assignment implies deciding which gates should be assigned a particular threshold voltage. We also include the effect of variability in threshold voltage on delay and leakage due to fabrication process variations in our formulations and present a scheme that lets the designer control the leakage and delay variability in his design. The proposed algorithm is a general mathematical formulation that has been shown to trivially extend to multiple threshold voltages. Vishal Khandelwal, Azadeh Davoodi, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | High level techniques for power-grid noise immunityabstractPower-grid networks are very important aspects of large scale integrated systems. In the modern deep sub-micron era these networks are prone to many sources of noise hence making the voltage supply uctuate. This Vdd-Ground noise can have detrimental effect on design quality. This paper presents a unique strategy of achieving noise immunity through voltage scheduling in Data Flow Graphs (DFGs). A dynamic programming based approach is applied to obtain noise immunity by imposing a grid on the voltage axis. We also present a unique way of including resource binding information into the algorithm. Experimental results indicated that considerable amount of Vdd-noise immunity is achieved for the selected benchmarks. Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2004 | Variability inspired implementation selection problemabstractGiven a directed acyclic graph and different possible implementations for each node, the implementation selection problem (ISP) selects the appropriate implementation for each node such that a given global design objective is optimized, ISP is a generic formulation that is explicitly or implicitly solved in several design automation problems like leakage optimization using dual V/sub th/, gate sizing, etc. An implementation of a node results in an associated delay and perhaps cost for the node. In the presence of different sources of uncertainty and fabrication variability, fixed estimates of delays and costs of a node are extremely erroneous. We investigate a probabilistic approach to solve ISP by considering probability density functions for delays and costs of a node. We propose a dynamic-programming based approach in a probabilistic sense and introduce effective pruning criteria when dealing with probability distributions for identifying co-optimal solution at each stage. A case study of leakage optimization using dual V/sub th/ is presented where we show the effectiveness of a probabilistic approach considering V/sub th/ variability over a traditional deterministic one. Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001 |
ICCAD | 3 |
| 2004 | Efficient statistical timing analysis through error budgetingabstractWe propose a technique for optimizing the runtime in statistical timing analysis. Given a global acceptable error budget at the primary output which signifies the difference in the area of the accurate and approximate timing CDFs, we propose a formulation of budgeting this global error across all nodes in the circuit. This node error budget is used to simplify the computation of arrival time CDFs at each node using approximations. This simplification reduces the runtime of statistical timing analysis. We investigate two ways of exploiting this node error budget, firstly through piecewise linear approximation (see ibid., A. Devgan and C. Kashyap, 2003) and secondly though hierarchical quadratic approximation. Experimental results on ISCAS/MCNC benchmarks show that our approach is at most 3 times faster than accurate statistical timing analysis and had a very small error. We also found quadratic piecewise approximation to be more accurate than linear approximation but at lesser gains in runtime. Vishal Khandelwal, Azadeh Davoodi, Ankur Srivastava 0001 |
ICCAD | 3 |
| 2004 | Leakage control through fine-grained placement and sizing of sleep transistorsabstractLeakage power is increasingly gaining importance with technology scaling. Multi-threshold CMOS (MTCMOS) technology has become a popular technique for standby power reduction. Sleep transistor insertion in circuits is an effective application of MTCMOS technology for reducing leakage power. In This work we present a fine grained approach where each gate in the circuit is provided an independent sleep transistor. Key advantages of this approach include better circuit slack utilization and improvements in signal integrity (which is a major disadvantage in clustering based approaches). To this end, we propose an optimal polynomial time fine grained sleep transistor sizing algorithm. We also prove the selective sleep transistor placement problem as NP-complete and propose an effective heuristic. Finally, in order to reduce the sleep transistor area penalty (which might get high since clustering is not performed), we propose a placement area constrained sleep transistor sizing formulation. Our experiments show that on an average the sleep transistor placement and optimal sizing algorithm gave 69.7% and 59.0% savings in leakage power as compared to the conventional fixed delay penalty algorithms for 5 and 7% circuit slowdown respectively. Moreover the post placement area penalty was less than 5% which is comparable to clustering schemes according to Mohab Anis et al. (2003). Vishal Khandelwal, Ankur Srivastava 0001 |
ICCAD | 2 |
| 2004 | Wire-length prediction using statistical techniquesabstractWe address the classic wire-length estimation problem and propose a new statistical wire-length estimation approach that captures the probability distribution function of net lengths after placement and before routing. The wire-length prediction model was developed using a combination of parametric and non-parametric statistical techniques. The model predicts not only the length of the net using input parameters extracted from the floorplan of a design, but also probability distributions that a net with given characteristics obtained after placement will have a particular length. The model is validated using both learn-and-test and resubstitution techniques. The model can be used for a variety of purposes, including the generation of a large number of statistically sound and therefore realistic instances of designs. We applied the net models to the probabilistic buffer insertion problem and obtained substantial improvement in net delay after routing. Jennifer Wong-Ma, Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001, Miodrag Potkonjak |
ICCAD | 4 |
| 2004 | Active mode leakage reduction using fine-grained forward body biasing strategyabstractLeakage power minimization has become an important issue with technology scaling. Variable threshold voltage schemes have become popular for standby power reduction. In this work we look at another emerging aspect of this potent problem which is leakage power reduction in active mode of operation. In gate level circuits, a large number of gates are not switching in active mode at any given point in time but nevertheless are consuming leakage power. We propose a fine-grained Forward Body Biasing (FBB) Scheme for active mode leakage power reduction in gate level circuits without any delay penalty. Our results show that our optimal polynomial time FBB allocation scheme results in 70.2% reduction in leakage currents. We also present a novel placement-driven FBB allocation algorithm that effectively reduces the area penalty using the post-placement area slack and results in 39.7%, 64.7% and 67.1% reduction in leakage currents for 0%, 4% and 8% area slack respectively. Vishal Khandelwal, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2004 | Empirical models for net-length probability distribution and applicationsabstractIn this paper, we propose a novel, empirical, and parameterizable model for estimating the probability distribution of wire length for each net in a placed netlist. The model is simple and fast to compute. We did extensive experimentation with state-of-the-art commercial (Cadence) and academic (Parquet and Labyrinth) tools and validated our model. Our distribution model was around three times more accurate than assuming half-perimeter bounding box as the fixed net-length estimate. Since the model is parameterizable it can be easily tailored for different routing tools and benchmarks. This model would be very useful in defining a full fledged probabilistic design automation methodology in which various design metrics are optimized from a probabilistic point of view. We also discuss the application of our model in a novel probabilistic approach to the buffer insertion problem. Azadeh Davoodi, Vishal Khandelwal, Ankur Srivastava 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | Timing driven gate duplicationabstractIn the past few years, gate duplication has been studied as a strategy for cutset minimization in partitioning problems. This paper addresses the problem of delay optimization by gate duplication. We present an algorithm to solve the gate duplication problem. It traverses the network from primary outputs(PO) to primary inputs(PI) in topologically sorted order evaluating tuples at the input pins of gates. The tuple's first component corresponds to the input pin required time if that gate is not duplicated. The second component corresponds to the input pin required time if that gate were duplicated. After tuple evaluation the algorithm traverses the network from PI to PO in topologically sorted order, deciding the gates to be duplicated. The last and final traversal is again from PO to PI, in which the gates are physically duplicated. Our algorithm uses the dynamic programming structure. We report delay improvements over other optimization methodologies. Gate duplication, along with other optimization strategies, can be used for meeting the stringent delay constraints in today's ultra complex designs. Ankur Srivastava 0001, Ryan Kastner, Chunhong Chen, Majid Sarrafzadeh |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2003 | A Probabilistic Approach to Buffer Insertion
Vishal Khandelwal, Azadeh Davoodi, Akash Nanavati, Ankur Srivastava 0001 |
ICCAD | 4 |
| 2003 | Achieving Design Closure Through Delay Relaxation Parameter
Ankur Srivastava 0001, Seda Ogrenci Memik, Bo-Kyung Choi, Majid Sarrafzadeh |
ICCAD | 1 |
| 2003 | Effective graph theoretic techniques for the generalized low power binding problemabstractThis paper proposes two very fast graph theoretic heuristics for the low power binding problem given fixed number of resources and multiple architectures for the resources. First the generalized low power binding problem is formulated as an Integer Linear Programming(ILP) problem which happens to be an NP-complete task to solve. Then two polynomial-time heuristics are proposed that provide a speedup of up to 13.7 with an extremely low penalty for power when compared to the optimal ILP solution for our selected benchmarks. Azadeh Davoodi, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2003 | Voltage scheduling under unpredictabilities: a risk management paradigmabstractThis paper addresses the problem of voltage scheduling in unpredictable situations. The voltage scheduling problem assigns voltages to operations such that the power is minimized under a clock cycle constraint. In presence of unpredictabilities meeting the clock constraint cannot be guaranteed. This paper proposes a novel risk management based technique to solve this problem. The risk management paradigm assigns a quantified value to the amount of risk the designer is willing to take on the clock cycle constraint. The algorithm then assigns voltages in order to meet the expected value of clock cycle constraint while keeping the maximum delay within the specified "risk" and minimizing the power. Azadeh Davoodi, Ankur Srivastava 0001 |
ISLPED | 2 |
| 2003 | Simultaneous Vt selection and assignment for leakage optimizationabstractThis paper presents a novel approach for leakage optimization through simultanous Vt selection and assignment. Vt selection implies deciding the right value for $V_t$ and assignment implies deciding which gates should be assigned which thresh-hold value. The proposed algorithm is a general mathematical formulation that can be trivially extended to multiple thresh-hold voltages (more than two). Traditional leakage optimization strategies either assume the prespecification of thresh-hold values or are good only for two thresh-holds. The presented formulation is based on linear programming approach under the piecewise linear approximation of delay/leakage vs thresh-hold curves. The algortihm was incorporated in SIS. Experimental results indicate that on some benchmarks having more that two thresh-holds was beneficial for leakage. Ankur Srivastava 0001 |
ISLPED | 1 |
| 2002 | Predictability: definition, ananlysis and optimizationabstractPredictability is the quantified from of accuracy. We propose a predictability driven design methodology. The novelty lies in defining and using the idea of predictability. In order to illustrate the basic concepts we focus on the low power binding problem. The binding problem for low power was solved in [3], [5], but in the presence of in-accuracies, their claims of optimality are imprecise. Our experiments show that these inaccuracies could be as high as 33%. Our methodology could improve this unpredictability to as low as 11% with minimal power penalty (7% on average). Ankur Srivastava 0001, Majid Sarrafzadeh |
ICCAD | 1 |
| 2002 | Early evaluation techniques for low power bindingabstractThis paper presents effective metrics to evaluate the power dissipation of scheduled data flow graphs (DFGs). This enables early evaluation of schedules without performing the computationally expensive resource-binding step. Our metrics correlate heavily (as high as 0.95 and > 0.75 for most test cases) with power dissipation values obtained after resource binding and rescheduling for power optimization steps. An experimental flow that integrates path-based scheduling, power optimal binding and power driven iterative rescheduling stages is constructed. The flow integrates commercial tools; like Synopsys, VSS and academic compilers like SUIF in a common optimization framework. Experimental results on DFGs from MediaBench suit also demonstrate the fact that metric evaluation is on average 42.6 times faster than performing optimal binding and iterative power improvement. Hence metric based evaluation enables fast design exploration at early stages. Eren Kursun, Ankur Srivastava 0001, Seda Ogrenci Memik, Majid Sarrafzadeh |
ISLPED | 2 |
| 2002 | Budget Management with Applications
Chunhong Chen, Elaheh Bozorgzadeh, Ankur Srivastava 0001, Majid Sarrafzadeh |
Algorithmica | 3 |
| 2001 | Timing driven gate duplication in technology independent phaseabstractWe propose a timing driven gate duplication algorithm for the technology independent phase. Our algorithm is a generalization of the gate duplication strategy suggested in [1]. Our technique gets a more global view by duplicating multiple gates at a time. We compare the minimum circuit delay obtained by SIS [2] with the delay obtained by using our gate duplication. Results show that up to 11% improvement in delay can be obtained. Our algorithm does not have an adverse effect on the overall synthesis time, indicating that gate duplication is an efficient strategy for timing optimization. Ankur Srivastava 0001, Chunhong Chen, Majid Sarrafzadeh |
ASP-DAC | 1 |
| 2001 | Layout aware retimingabstractArticle Share on Layout aware retiming Authors: A. Ranjan Monterey Design Systems, Sunnyvale, CA Monterey Design Systems, Sunnyvale, CAView Profile , A. Srivastava Computer Science Department, Univeristy of California, Los Angeles, CA Computer Science Department, Univeristy of California, Los Angeles, CAView Profile , V. Karnam Advanced Micro Devices, Sunnyvale, CA Advanced Micro Devices, Sunnyvale, CAView Profile , M. Sarrafzadeh Computer Science Department, Univeristy of California, Los Angeles, CA Computer Science Department, Univeristy of California, Los Angeles, CAView Profile Authors Info & Claims GLSVLSI '01: Proceedings of the 11th Great Lakes symposium on VLSIMarch 2001 Pages 25–30https://doi.org/10.1145/368122.368153Online:01 March 2001Publication History 6citation186DownloadsMetricsTotal Citations6Total Downloads186Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ankur Srivastava 0001, V. Karnam, Majid Sarrafzadeh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2001 | Design and analysis of physical design algorithmsabstractWe will review a few key algorithmic and analysis concepts with application to physical design problems. We argue that design and detailed analysis of algorithms is of fundamental importance in developing better physical design tools and to cope with the complexity of present-day designs. Majid Sarrafzadeh, Elaheh Bozorgzadeh, Ryan Kastner, Ankur Srivastava 0001 |
ISPD | 4 |
| 2001 | Activity-driven clock designabstractIn this paper, we investigate reducing the power consumption of a synchronous digital system by minimizing the total power consumed by the clock signals. We construct activity-driven clock trees wherein sections of the clock tree are turned off by gating the clock signals. Since gating the clock signal implies that additional control signals and gates are needed, there exists a tradeoff between the amount of clock tree gating and the total power consumption of the clock tree. We exploit similarities in the switching activity of the clocked modules to reduce the number of clock gates. Assuming a given switching activity of the modules, we propose three novel activity-driven problems: a clock tree construction problem, a clock gate insertion problem, and a zero-skew clock gate insertion problem. The objective of these problems is to minimize the system's power consumption by constructing an activity-driven clock tree. We propose an approximation algorithm based on recursive matching to solve the clock tree construction problem. We also propose an exact algorithm employing the dynamic programming paradigm to solve the gate insertion problems. Finally, we present experimental results that verify the effectiveness of our approach. This paper is a step in understanding how high-level decisions (e.g., behavioral design) can affect a low-level design (e.g., clock design). Amir H. Farrahi, Chunhong Chen, Ankur Srivastava 0001, Gustavo E. Téllez, Majid Sarrafzadeh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | On the complexity of gate duplicationabstractIn this paper, we show that both the global and local gate duplication problems for delay optimization are NP-complete under certain delay models. Ankur Srivastava 0001, Ryan Kastner, Majid Sarrafzadeh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | On gate level power optimization using dual-supply voltagesabstractIn this paper, we present an approach for applying two supply voltages to optimize power in CMOS digital circuits under the timing constraints. Given a technology-mapped network, we first analyze the power/delay model and the timing slack distribution in the network. Then a new strategy is developed for timing-constrained optimization issues by making full use of stacks. Based on this strategy, the power reduction is translated into the polynomial-time-solvable maximal-weighted-independent-set problem on transitive graphs. Since different supply voltages used in the circuit lead to totally different power consumption, we propose a fast heuristic approach to predict the optimum dual-supply voltages by looking at the lower bound of power consumption in the given circuit. To deal with the possible power penalty due to the level converters at the interface of different supply voltages, we use a "constrained F-M" algorithm to minimize the number of level converters. We have implemented our approach under an SIS environment. Experiment shows that the resulting lower bound of power is tight for most circuits and that the predicted "optimum" supply voltages are exactly or very close to the best choice of actual ones. The total power saving of up to 26% (average of about 20%) is achieved without degrading the circuit performance, compared to the average power improvement of about 7% by the gate sizing technique based on a standard cell library. Our technique provides the power-delay tradeoff by specifying different timing constraints in circuits for power optimization. Chunhong Chen, Ankur Srivastava 0001, Majid Sarrafzadeh |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Timing Driven Gate Duplication: Complexity Issues and AlgorithmsabstractDelay optimization is a fundamental goal in logic synthesis. This paper presents gate duplication as a strategy for performance optimization. Many timing optimization strategies have been proposed over the past few years. In the past few years, the research community has looked at gate duplication extensively as a method of reducing the cut-set of partitions. Strategies of logic duplication for cut-set minimization have been suggested. The strength of gate duplication as a cut-set minimizing strategy has been demonstrated. However applicability of this strategy in reducing the circuit delay has not been studied in detail. In this paper we prove the problem of partitioning a set of fanouts between a gate and it's replica (both gates have the same fanins) such that the required time constraint at the input pin is met, to be NP-Complete. Hence even the local optimization by gate duplication problem (formally defined later) is also NP-Complete. We then present an algorithm for gate duplication which is based on the dynamic programming approach. Since the problem of partitioning a set of fanouts between a node and it's replica is NP-Complete, we use a heuristic for making this decision which is optimal under specific conditions. We report delay improvements as high as 8% over highly optimized results generated by SIS. The rest of this paper is organized as follows. Section 2 deals with the delay model and provides basic definitions. Section 3 reviews the complexity of the global gate duplication problem. Section 4 presents the proof of NP-Completeness for the local gate duplication problem. Section 5 describes a heuristic for gate duplication in detail, followed by results in Section 6. This is followed by some observations and conclusion in Section 7. Ankur Srivastava 0001, Ryan Kastner, Majid Sarrafzadeh |
ICCAD | 1 |