EDBT 2026 Demo / reviewers in the wild / expert
Wayne P. Burleson
dblp:b/WaynePBurleson
· DBLP profile ↗
109ranked-venue papers
11as first author
4since 2021 · last 2025
0000-0002-7068-4351ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 91 · 8 first-author · 4 since 2021Security and privacy · 10Software engineering, systems software and programming languages · 9 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5Computer networks · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Integrated Security Mechanisms for Weight Protection in Memristive Crossbar ArraysabstractMemristive crossbar arrays enable in-memory computing by performing parallel analog computations directly within memory, making them well-suited for machine learning, neural networks, and neuromorphic systems. However, despite their advantages, non-volatile memristors are vulnerable to security threats (such as adversarial extraction of stored weights when the hardware is compromised. Protecting these weights is essential since they represent valuable intellectual property resulting from lengthy and costly training processes using large, often proprietary, datasets. As a solution we propose two security mechanisms: Keyed Permutor and Watermark Protection Columns; where both safeguard critical weights and establish verifiable ownership (even in cases of data leakage). Our approach integrates efficiently with existing memristive crossbar architectures without significant design modifications. Simulations across$45\ \text{nm}, 22\ \text{nm}$, and 7nm CMOS nodes, using a realistic interconnect model and a large RF dataset, show that both mechanisms offer robust protection with under 10% overhead in area, delay and power. We also present initial experiments employing the widely known MNIST dataset; further highlighting the feasibility of securing memristive in-memory computing systems with minimal performance trade-offs. Muhammad Faheemur Rahman, Wayne P. Burleson |
ASAP | 2 |
| 2025 | Gradient Approximation of Approximate Multipliers for High-Accuracy Deep Neural Network RetrainingabstractApproximate multipliers (AppMults) are widely employed in deep neural network (DNN) accelerators to reduce the area, delay, and power consumption. However, the inaccuracies of AppMults degrade DNN accuracy, necessitating a retraining process to recover accuracy. A critical step in retraining is computing the gradient of the AppMult, i.e., the partial derivative of the approximate product with respect to each input operand. Conventional methods approximate this gradient using that of the accurate multiplier (AccMult), often leading to suboptimal retraining results, especially for AppMults with relatively large errors. To address this issue, we propose a difference-based gradient approximation of AppMults to improve retraining accuracy. Experimental results show that compared to the state-of-the-art methods, our method improves the DNN accuracy after retraining by 4.10% and 2.93% on average for the VGG and ResNet models, respectively. Moreover, after retraining a ResNet18 model using a 7-bit AppMult, the final DNN accuracy does not degrade compared to the quantized model using the 7-bit AccMult, while the power consumption is reduced by 51%. Chang Meng, Wayne P. Burleson, Weikang Qian, Giovanni De Micheli |
DATE | 2 |
| 2024 | RareLS: Rarity-Reducing Logic Synthesis for Mitigating Hardware Trojan ThreatsabstractHardware Trojan (HT) poses a critical security threat to integrated circuits, which can change circuit functionality or leak sensitive data. HTs are typically activated under low-probability conditions by exploiting "rare signals" in logic circuits. In this paper, we propose RareLS, rarity-reducing logic synthesis for mitigating HT threats. Specifically, RareLS reduces the number of rare signals through rarity-oriented technology-independent optimization and technology mapping. Experimental results show that RareLS reduces rare signals by 63.4% on average, with a small overhead of 4.0% in area, 1.9% in delay, and 6.2% in power. Moreover, RareLS complicates HT insertion for attackers by reducing HT trigger logic by 92.94%, and aids defenders in detecting HTs by shortening the test length by more than 80.83%. Chang Meng, Mingfei Yu, Wayne P. Burleson, Giovanni De Micheli |
ICCAD | 4 |
| 2023 | Voltage Sensor Implementations for Remote Power Attacks on FPGAsabstractThis article presents a study of two types of on-chip FPGA voltage sensors based on ring oscillators (ROs) and time-to-digital converter (TDCs), respectively. It has previously been shown that these sensors are often used to extract side-channel information from FPGAs without physical access. The performance of the sensors is evaluated in the presence of circuits that deliberately waste power, resulting in localized voltage drops. The effects of FPGA power supply features and sensor sensitivity in detecting voltage drops in an FPGA power distribution network (PDN) are evaluated for Xilinx Artix-7, Zynq 7000, and Zynq UltraScale+ FPGAs. We show that both sensor types are able to detect supply voltage drops, and that their measurements are consistent with each other. Our findings show that TDC-based sensors are more sensitive and can detect voltage drops that are shorter in duration, while RO sensors are easier to implement because calibration is not required. Furthermore, we present a new time-interleaved TDC design that sweeps the sensor phase. The new sensor generates data that can reconstruct voltage transients on the order of tens of picoseconds. Shayan Moini, Aleksa Deric, Xiang Li 0158, George Provelengios, Wayne P. Burleson, Russell Tessier, Daniel E. Holcomb |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2019 | Combining Clock and Voltage Noise Countermeasures Against Power Side-Channel AnalysisabstractThe power side-channel continues to be a major vulnerability in many critical systems. Numerous countermeasures have been proposed since its discovery as a serious vulnerability, including both hardware and software implementations. Each countermeasure has its own drawback, with some of the more popular countermeasures incurring 3-4x overhead in area and power. While this is acceptable for some devices, other lightweight devices can not tolerate this amount of overhead. In addition, most countermeasures are quite invasive to the design process, requiring modification to the design and additional validation. This work explores two relatively non-invasive countermeasures, both the detailed implementations and the interactions: 1) clock randomization, and 2) supply voltage noise, by real collecting power traces from a state-of-the-art FPGA board and using the Correlation Power Analysis (CPA) attack. Counter-attacks and their impact on the countermeasure efficacy are also explored. The key result is that the combined effects of the two countermeasures is greater than the impact of either countermeasure when used independently. Jacqueline Lagasse, Christopher Bartoli, Wayne P. Burleson |
ASAP | 3 |
| 2019 | Guest Editorial Special Section on Security Challenges and Solutions With Emerging Computing TechnologiesabstractMultiple emerging computing technologies based on, e.g., graphene, spintronics, resistive RAM, quantum computing, and others are being developed to enhance the capabilities of logic devices and circuits. The rapid growth in these technologies is synchronized with the decline of Moore’s law, thus promises to herald the era of Beyond CMOS technologies with a significant improvement in energy efficiency, reliability, performance, and manufacturability. These devices enable very different computing paradigms, e.g., neuromorphic computing, non-Boolean computing, and in-memory computing, thus making these platforms an interesting playground for circuit and application-developers alike. Anupam Chattopadhyay, Swaroop Ghosh, Wayne P. Burleson, Debdeep Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Enhancing Power, Performance, and Energy Efficiency in Chip Multiprocessors Exploiting Inverse Thermal DependenceabstractTechnology scaling of complementary metal-oxide-semiconductor has resulted in new thermal behavior where increase in operating temperature results in reduced circuit propagation delay. This paper exploits this inverse thermal dependence (ITD) for power, performance, and temperature optimization in single-core and multicore processor architectures for various thermally hot and cold applications. Since ITD increases the maximum achievable operating frequency of a processor at high temperatures, it is used to reduce the execution time of applications. Dynamic thermal management (DTM) techniques, such as activity migration (AM), dynamic voltage frequency scaling (DVFS), and throttling, are modified to leverage ITD to either enhance the performance or the energy efficiency. While recent work observed the ITD effect for 45- and 32-nm technologies, in this paper, we explore future technologies through predictive SPICE models for 20-, 14-, 10-, and 7-nm technologies. The results show that the ITD-aware techniques reduce the execution time, energy-delay product (EDP) and energy-delay-square product by up to 28%, 33%, and 48%, respectively. Moreover, the ITD-aware DVFS yields the lowest execution time while resulting in the most uniform thermal profile and thus enhanced reliability. Overall, the ITD-aware techniques reduce the execution time and EDP, both in combination with DTM techniques or stand-alone, especially at lower than nominal operating voltages. Katayoun Neshatpour, Wayne P. Burleson, Amin Khajeh, Houman Homayoun |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | An asynchronous NoC router in a 14nm FinFET library: Comparison to an industrial synchronous counterpartabstractAn asynchronous high-performance low-power 5-port network-on-chip (NoC) router is introduced. The proposed router integrates low-latency input buffers using a circular FIFO design, and a novel end-to-end credit-based virtual channel (VC) flow control for a replicated switch architecture. This asynchronous router is then compared to an AMD synchronous router, in a realistic advanced 14nm FinFET library. This is the first such comparison, to the best of our knowledge, using a real synchronous router baseline already fabricated in several commercial products. Initial post-synthesis pre-layout experiments show dominating results for the asynchronous router, when compared to the synchronous router. In particular, 55% less area and 28% latency improvement are observed for the asynchronous implementation. Also, 88% and 58% savings in idle and active power, respectively, are obtained. Weiwei Jiang 0002, Davide Bertozzi, Gabriele Miorandi, Steven M. Nowick, Wayne P. Burleson, Greg Sadowski |
DATE | 5 |
| 2016 | Invited - Who is the major threat to tomorrow's security?: you, the hardware designerabstractMore and more security attacks today are perpetrated by exploiting the hardware: memory errors can be exploited to take over systems, side-channel attacks leak secrets to the outside worlds, weak random number generators render cryptography ineffective, etc. At the same time, many of the tenets of efficient design are in tension with guaranteeing security. For instance, classic secure hardware does not allow to optimize common execution patterns, share resources, or provide deep introspection. Wayne P. Burleson, Onur Mutlu, Mohit Tiwari |
DAC | 1 |
| 2016 | Persistent Clocks for Batteryless Sensing DevicesabstractSensing platforms are becoming batteryless to enable the vision of the Internet of Things, where trillions of devices collect data, interact with each other, and interact with people. However, these batteryless sensing platforms—that rely purely on energy harvesting—are rarely able to maintain a sense of time after a power failure. This makes working with sensor data that is time sensitive especially difficult. We propose two novel, zero-power timekeepers that use remanence decay to measure the time elapsed between power failures. Our approaches compute the elapsed time from the amount of decay of a capacitive device, either on-chip Static Random-Access Memory (SRAM) or a dedicated capacitor. This enables hourglass-like timers that give intermittently powered sensing devices a persistent sense of time. Our evaluation shows that applications using either timekeeper can keep time accurately through power failures as long as 45s with low overhead. Josiah D. Hester, Nicole Tobias, Amir Rahmati, Lanny Sitanayah, Daniel E. Holcomb, Kevin Fu, Wayne P. Burleson, Jacob Sorber |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2015 | Reinforcement Learning for Thermal-aware Many-core Task AllocationabstractTo maintain reliable operation, task allocation for many-core processors must consider the heat interaction of processor cores and network-on-chip routers in performing task assignment. Our approach employs reinforcement learning, machine learning algorithm that performs task allocation based on current core and router temperatures and a prediction of which assignment will minimize maximum temperature in the future. The algorithm updates prediction models after each allocation based on feedback regarding the accuracy of previous predictions. Our new algorithm is verified via detailed many-core simulation which includes on-chip routing. Our results show that the proposed technique is fast (scheduling performed in <1 ms) and can efficiently reduce peak temperature by up to 8°C in a 49-core processor (4.3°C on average) versus a competing task allocation approach for a series of SPLASH-2 benchmarks. Shiting (Justin) Lu, Russell Tessier, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Revisiting Dynamic Thermal Management Exploiting Inverse Thermal DependenceabstractAs CMOS technology scales down towards nanometer regime and the supply voltage approaches the threshold voltage, increase in operating temperature results in increased circuit current, which in turn reduces circuit propagation delay. This paper exploits this new phenomenon, known as inverse thermal dependence (ITD) for power, performance, and temperature optimization in processor architecture. ITD changes the maximum achievable operating frequency of the processor at high temperatures. Dynamic thermal management techniques such as activity migration, dynamic voltage frequency scaling, and throttling are revisited in this paper, with a focus on the effect of ITD. Results are obtained using the predictive technology models of 7nm, 10nm 14nm and 20nm technology nodes and with extensive architectural and circuit simulations. The results show that based on the design goals, various design corners should be re-investigated for power, performance and energy-efficiency optimization. Architectural simulations for a multi-core processor and across standard benchmarks show that utilizing ITD-aware schemes for thermal management improves the performance of the processor in terms of speed and energy-delay-product by 8.55% and 4.4%, respectively. Katayoun Neshatpour, Houman Homayoun, Amin Khajeh, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | Virtual Proofs of Reality and their Physical ImplementationabstractWe discuss the question of how physical statements can be proven over digital communication channels between two parties (a "prover" and a "verifier") residing in two separate local systems. Examples include: (i) "a certain object in the prover's system has temperature X°C", (ii) "two certain objects in the prover's system are positioned at distance X", or (iii) "a certain object in the prover's system has been irreversibly altered or destroyed". As illustrated by these examples, our treatment goes beyond classical security sensors in considering more general physical statements. Another distinctive aspect is the underlying security model: We neither assume secret keys in the prover's system, nor do we suppose classical sensor hardware in his system which is tamper-resistant and trusted by the verifier. Without an established name, we call this new type of security protocol a "virtual proof of reality" or simply a "virtual proof" (VP). In order to illustrate our novel concept, we give example VPs based on temperature sensitive integrated circuits, disordered optical scattering media, and quantum systems. The corresponding protocols prove the temperature, relative position, or destruction/modification of certain physical objects in the prover's system to the verifier. These objects (so-called "witness objects") are prepared by the verifier and handed over to the prover prior to the VP. Furthermore, we verify the practical validity of our method for all our optical and circuit-based VPs in detailed proof-of-concept experiments. Our work touches upon, and partly extends, several established concepts in cryptography and security, including physical unclonable functions, quantum cryptography, interactive proof systems, and, most recently, physical zero-knowledge proofs. We also discuss potential advancements of our method, for example "public virtual proofs" that function without exchanging witness objects between the verifier and the prover. Ulrich Rührmair, J. L. Martinez-Hurtado, Xiaolin Xu 0001, Christian Kraeh, Christian Hilgers, Dima Kononchuk, Jonathan J. Finley, Wayne P. Burleson |
IEEE Symposium on Security and Privacy | 8 |
| 2015 | Reliable Physical Unclonable Functions Using Data Retention Voltage of SRAM CellsabstractPhysical unclonable functions (PUFs) are circuits that produce outputs determined by random physical variations from fabrication. The PUF studied in this paper utilizes the variation sensitivity of static random access memory (SRAM) data retention voltage (DRV), the minimum voltage at which each cell can retain state. Prior work shows that DRV can uniquely identify circuit instances with 28% greater success than SRAM power-up states that are used in PUFs [1]. However, DRV is highly sensitive to temperature, and until now this makes it unreliable and unsuitable for use in a PUF. In this paper, we enable DRV PUFs by proposing a DRV-based hash function that is insensitive to temperature. The new hash function, denoted DRV-based hashing (DH), is reliable across temperatures because it utilizes the temperature-insensitive ordering of DRVs across cells, instead of using the DRVs in absolute terms. To evaluate the security and performance of the DRV PUF, we use DRV measurements from commercially available SRAM chips, and use data from a novel DRV prediction algorithm. The prediction algorithm uses machine learning for fast and accurate simulation-free estimation of any cell's DRV, and the prediction error in comparison to circuit simulation has a standard deviation of 0.35 mV. We demonstrate the DRV PUF using two applications-secret key generation and identification. In secret key generation, we introduce a new circuit-level reliability knob as an alternative to error correcting codes. In the identification application, our approach is compared to prior work and shown to result in a smaller false-positive identification rate for any desired true-positive identification rate. Xiaolin Xu 0001, Amir Rahmati, Daniel E. Holcomb, Kevin Fu, Wayne P. Burleson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2014 | Efficient Power and Timing Side Channels for Physical Unclonable Functions
Ulrich Rührmair, Xiaolin Xu 0001, Jan Sölter, Ahmed Mahmoud, Mehrdad Majzoobi, Farinaz Koushanfar, Wayne P. Burleson |
CHES | 7 |
| 2014 | Modeling and Experimental Demonstration of Accelerated Self-Healing TechniquesabstractIn this paper we postulate that future electronics systems will use sleep time as an active recovery period essential for their overall performance. Our hypothesis is that by explicitly controlling the ratio of sleep vs. active and sleep conditions (e.g. higher temperatures, negative voltages), we can deeply rejuvenate electronic systems periodically to improve their metrics. We perform a series of stress and recovery experiments using commercial FPGAs to demonstrate several cases where we bring stressed chips back to within 90% of their original margin by actively rejuvenating for only 1/4 of the stress time. We validate our experiments against extracted models and present potential applications to multi-core systems. Xinfei Guo, Wayne P. Burleson, Mircea R. Stan |
DAC | 2 |
| 2014 | Special session: How secure are PUFs really? On the reach and limits of recent PUF attacksabstractJust over a decade ago, Physical Unclonable Functions (PUFs) have been introduced as a new cryptographic and security primitive in a number of seminal publications. Due to their assumed security and cost advantages, they have attracted substantial attention both from the security industry and the academic community, and are also gaining ground in commercial applications. Nevertheless, a number of recent works have presented successful attacks on PUF core properties, such as their digital and physical unclonability. How strong and relevant are these attacks, and how secure are PUFs really? This question is addressed in a dedicated hot topic session at DATE 2014. This paper provides a short and easily accessible overview of the session. Ulrich Rührmair, Ulf Schlichtmann, Wayne P. Burleson |
DATE | 3 |
| 2014 | Hybrid side-channel/machine-learning attacks on PUFs: A new threat?abstractMachine Learning (ML) is a well-studied strategy in modeling Physical Unclonable Functions (PUFs) but reaches its limits while applied on instances of high complexity. To address this issue, side-channel attacks have recently been combined with modeling techniques to make attacks more efficient [25][26]. In this work, we present an overview and survey of these so-called “hybrid modeling and side-channel attacks” on PUFs, as well as of classical side channel techniques for PUFs. A taxonomy is proposed based on the characteristics of different side-channel attacks. The practical reach of some published side-channel attacks is discussed. Both challenges and opportunities for PUF attackers are introduced. Countermeasures against some certain side-channel attacks are also analyzed. To better understand the side-channel attacks on PUFs, three different methodologies of implementing side-channel attacks are compared. At the end of this paper, we bring forward some open problems for this research area. Xiaolin Xu 0001, Wayne P. Burleson |
DATE | 2 |
| 2014 | Parametric Trojans for Fault-Injection Attacks on Cryptographic HardwareabstractWe propose two extremely stealthy hardware Trojans that facilitate fault-injection attacks in cryptographic blocks. The Trojans are carefully inserted to modify the electrical characteristics of predetermined transistors in a circuit by altering parameters such as doping concentration and do pant area. These Trojans are activated with very low probability under the presence of a slightly reduced supply voltage (0.001 for 20% Vdd reduction). We demonstrate the effectiveness of the Trojans by utilizing them to inject faults into an ASIC implementation of the recently introduced lightweight cipher PRINCE. Full circuit-level simulation followed by differential cryptanalysis demonstrate that the secret key can be reconstructed after around 5 fault-injections. Raghavan Kumar, Philipp Jovanovic, Wayne P. Burleson, Ilia Polian |
FDTC | 3 |
| 2014 | Modeling and analysis of Phase Change Materials for efficient thermal managementabstractDirect placement of Phase Change Materials (PCMs) on the chip has been recently explored as a passive temperature management solution. PCMs provide the ability to store large amounts of heat at a close-to-constant temperature during the phase change (solid to liquid and vice versa). This latent heat capacity can be used to provide higher performance while reducing hot spots. Detailed modeling of the phase change behavior is essential for the design and evaluation of systems with PCM. This paper proposes an accurate phase change model that is integrated into the commonly used thermal simulation tool, HotSpot. It also provides validation of the proposed model by carrying out computational fluid dynamics (CFD) simulations using COMSOL Multiphysics®. This paper also explores the impact of PCM properties on the thermal profile of a processor, and demonstrates that PCM material choices can affect peak temperatures by up to 20.1°C. Experimental results show that dynamic policy decisions change dramatically when using the proposed detailed phase change model, as prior simpler PCM models can substantially over/under-estimate temperature and PCM melting duration. The proposed model helps design more effective dynamic management policies and enables realistic evaluation of systems with PCM. Fulya Kaplan, Charlie De Vivero, Samuel Howes, Manish Arora, Houman Homayoun, Wayne P. Burleson, Dean M. Tullsen, Ayse K. Coskun |
ICCD | 6 |
| 2014 | Hybrid modeling attacks on current-based PUFsabstractPhysically Unclonable Functions have emerged as a possible candidate to replace traditional cryptography. However, majority of the strong PUFs are vulnerable to modeling attacks. In this work, we take a closer look at the possible attacks on one of the strong PUF architectures known as Current-based PUFs, which exploit irregularities in transistor currents to generate unique signatures. We demonstrate that the fault-injection attacks when coupled with a machine learning (ML) algorithm can considerably push the limits of prediction accuracies. Based on simulations, we observed that the stand-alone ML algorithms suffer from error prone CRPs especially for higher length PUFs. In such scenarios, hybrid attacks exploiting the unreliable responses pushed the prediction accuracies up to 99% for higher length Current-based PUF circuits. Raghavan Kumar, Wayne P. Burleson |
ICCD | 2 |
| 2014 | Keynote talk I: Security and privacy in implantable medical devices: An ongoing concernabstractIn the last 5 years, there has been a considerable increase in research and awareness of security and privacy issues in implantable medical devices. This has been driven by the rapid introduction and deployment of an increasing diversity of medical devices. Although wearable and clinical devices are also of interest, the implantable nature of many recent devices significantly increases risks as well as engineering challenges. This talk will use several case studies to review the multi-disciplinary process in which vulnerabilities are analyzed, disclosed, and repaired. Business incentives and regulations play an important role in the overall process. Technical challenges include energy harvesting, side-channel protection, lightweight crypto, personalization, wireless communications, security/safety tradeoffs, and human factors. Several open research areas will be discussed. Wayne P. Burleson |
MEMOCODE | 1 |
| 2014 | Dynamic synchronizer flip-flop performance in FinFET technologiesabstractThe use of fine-grain Dynamic Voltage and Frequency Scaling (DVFS) has increased the number of distinct clock domains on a given Network-on-Chip (NoC). This necessitates robust synchronizers to prevent clock domain communication failures, even as FinFET devices have begun to replace planar devices. This paper presents simulation results and comparisons between dynamic (requiring reset) and non-dynamic synchronizer flip-flops implemented in predictive models for both planar technologies and future FinFET technologies. Results demonstrate that synchronizers built with FinFET devices 1) exhibit a tau value which continues to scale with fan-out of four delay and 2) can be improved with forward biasing, but 3) are more sensitive to temperature. Dynamic flip-flops settled metastability fastest when using standard technology voltages, but previously couldn't be used in non-dynamic systems. For this reason, a new synchronizer design is also presented which exploits the benefits of dynamic flip-flops without the need for a dedicated reset signal. Mark Buckler, Arpan Vaidya, Wayne P. Burleson |
NOCS | 4 |
| 2014 | Dynamic On-Chip Thermal Sensor Calibration Using Performance CountersabstractNumerous sensors are currently deployed in modern processors to collect thermal information for fine-grained dynamic thermal management (DTM). Due to process variation and silicon aging, on-chip thermal sensors require periodic calibration before use in DTM. However, the calibration cost for thermal sensors can be prohibitively high as the number of on-chip sensors increases. In this paper, a model which is suitable for online calculation is employed to estimate the temperatures of multiple sensor locations on the silicon die. The estimated sensor and actual sensor thermal profile show a very high similarity with correlation coefficient${\sim}{\rm 0.9}$for most tested benchmarks. Our calibration approach combines potentially inaccurate temperature values obtained from two sources: temperature readings from thermal sensors and temperature estimations using system performance counters. A data fusion strategy based on Bayesian inference, which combines information from these two sources, is demonstrated along with a temperature estimation approach using performance counters. The average absolute error of the corrected sensor temperature readings is${<}{1.5}^{\circ}{\rm C}$and the standard deviation of error is less than${<}{\rm 0.5}^{\circ}{\rm C}$for tested benchmarks. Shiting (Justin) Lu, Russell Tessier, Wayne P. Burleson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | Stealthy Dopant-Level Hardware Trojans
Georg T. Becker, Francesco Regazzoni 0001, Christof Paar, Wayne P. Burleson |
CHES | 4 |
| 2013 | Balancing security and utility in medical devices?abstractImplantable Medical Devices (IMDs) are being embedded increasingly often in patients' bodies to monitor and help treat medical conditions. To facilitate monitoring and control, IMDs are often equipped with wireless interfaces. While convenient, wireless connectivity raises the risk of malicious access to an IMD that can potentially infringe patients' privacy and even endanger their lives. Masoud Rostami, Wayne P. Burleson, Farinaz Koushanfar, Ari Juels |
DAC | 2 |
| 2013 | Run-time probabilistic detection of miscalibrated thermal sensors in many-core systemsabstractMany-core architectures use large numbers of small temperature sensors to detect thermal gradients and guide thermal management schemes. In this paper a technique to identify thermal sensors which are operating outside a required accuracy is described. Unlike previous on-chip temperature estimation approaches, our algorithms are optimized to run on-line while thermal management decisions are being made. The accuracy of a sensor is determined by comparing its readings to expected values from a probability distribution function determined from surrounding sensors. Experiments show that a sensor operating outside a desired accuracy can be identified with a detection rate of over 90% and an average false alarm rate of < 6%, with a confidence level of 90%. The run time of our method is shown to be around 3× lower than a recently-published temperature estimation method, enhancing its suitability for run-time implementation. Shiting (Justin) Lu, Wayne P. Burleson, Russell Tessier |
DATE | 3 |
| 2013 | Low-power Networks-on-Chip: Progress and remaining challengesabstractAfter a long period of academic and industrial research, networks-on-chips (NoCs) are starting to be incorporated into commercial multi-processor designs. NoCs have proven themselves to scale better than bus-based designs and they are here to stay. It is still important to note, however, that even well-designed NoCs consume a large portion of a given system's power budget. This brief paper and accompanying presentation discuss what options are available to designers who need to reduce NoC power consumption, their benefits, and their limitations. Techniques discussed here include general NoC system design as well as disruptive interconnect mediums and their associated strategies. Mark Buckler, Wayne P. Burleson, Greg Sadowski |
ISLPED | 2 |
| 2013 | Litho-aware and low power design of a secure current-based physically unclonable functionabstractPhysically Unclonable Functions (PUFs) are lightweight cryptographic primitives for generating unique signatures from complex manufacturing variations. In this work, we present a current-based PUF designed using a generalized lithographic simulation framework for improving inter-die and inter-wafer uniqueness. The sensitivity of the circuit to manufacturing variations is enhanced by placing the gate structures at pitches closer to forbidden zone, where the sensitivity of Critical Dimension (CD) to the pitch variations is very high. Simulation results show that the litho-aware current based PUF has improved inter- and intra-distance over the conventional current-based PUF. The litho-aware PUF consumes about 0.034 pico joules of energy per response bit, which is substantially better than delay-based PUF implementations. Raghavan Kumar, Wayne P. Burleson |
ISLPED | 2 |
| 2013 | Efficient E-Cash in Practice: NFC-Based Payments for Public Transportation Systems
Gesine Hinterwälder, Christian T. Zenger, Foteini Baldimtsi, Anna Lysyanskaya, Christof Paar, Wayne P. Burleson |
Privacy Enhancing Technologies | 6 |
| 2013 | PUF Modeling Attacks on Simulated and Silicon DataabstractWe discuss numerical modeling attacks on several proposed strong physical unclonable functions (PUFs). Given a set of challenge-response pairs (CRPs) of a Strong PUF, the goal of our attacks is to construct a computer algorithm which behaves indistinguishably from the original PUF on almost all CRPs. If successful, this algorithm can subsequently impersonate the Strong PUF, and can be cloned and distributed arbitrarily. It breaks the security of any applications that rest on the Strong PUF's unpredictability and physical unclonability. Our method is less relevant for other PUF types such as Weak PUFs. The Strong PUFs that we could attack successfully include standard Arbiter PUFs of essentially arbitrary sizes, and XOR Arbiter PUFs, Lightweight Secure PUFs, and Feed-Forward Arbiter PUFs up to certain sizes and complexities. We also investigate the hardness of certain Ring Oscillator PUF architectures in typical Strong PUF applications. Our attacks are based upon various machine learning techniques, including a specially tailored variant of logistic regression and evolution strategies. Our results are mostly obtained on CRPs from numerical simulations that use established digital models of the respective PUFs. For a subset of the considered PUFs-namely standard Arbiter PUFs and XOR Arbiter PUFs-we also lead proofs of concept on silicon data from both FPGAs and ASICs. Over four million silicon CRPs are used in this process. The performance on silicon CRPs is very close to simulated CRPs, confirming a conjecture from earlier versions of this work. Our findings lead to new design requirements for secure electrical Strong PUFs, and will be useful to PUF designers and attackers alike. Ulrich Rührmair, Jan Sölter, Frank Sehnke, Xiaolin Xu 0001, Ahmed Mahmoud, Vera Stoyanova, Gideon Dror, Jürgen Schmidhuber, Wayne P. Burleson, Srini Devadas |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2013 | Timing Uncertainty in 3-D Clock Trees Due to Process Variations and Power Supply NoiseabstractClock distribution networks are affected by different sources of variations. The resulting clock uncertainty significantly affects the frequency of a circuit. To support this analysis, a statistical model of skitter, which consists of clock skew and jitter, for 3-D clock trees is introduced. The effect of skitter on both the setup and hold time slacks is modeled. The variation of skitter is shown to be underestimated up to 36% if process variations and dynamic power supply noise are considered separately, which highlights the importance of this unified treatment. Potential scenarios of supply noise in 3-D integrated circuits (ICs) are investigated. 3-D circuits generated from industrial benchmarks are simulated to show the skitter under these scenarios. The mean and standard deviation of skitter can vary up to 60% and 51%, respectively, due to the different amplitudes and phases of supply noise. The tradeoff between skitter and the power consumed by clock trees is also shown. A set of guidelines are presented to decrease skitter in 3-D ICs. By applying these guidelines to industrial benchmarks, simulations show a decrease in the mean skitter up to 31%. Hu Xu 0002, Vasilis F. Pavlidis, Xifan Tang, Wayne P. Burleson, Giovanni De Micheli |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Design challenges for secure implantable medical devicesabstractImplantable medical devices, or IMDs, are increasingly being used to improve patients' medical outcomes. Designers of IMDs already balance safety, reliability, complexity, power consumption, and cost. However, recent research has demonstrated that designers should also consider security and data privacy to protect patients from acts of theft or malice, especially as medical technology becomes increasingly connected to other systems via wireless communications or the Internet. This survey paper summarizes recent work on IMD security. It discusses sound security principles to follow and common security pitfalls to avoid. As trends in power efficiency, sensing, wireless systems and bio-interfaces make possible new and improved IMDs, they also underscore the importance of understanding and addressing security and privacy concerns in an increasingly connected world. Wayne P. Burleson, Shane S. Clark, Benjamin Ransford, Kevin Fu |
DAC | 1 |
| 2012 | Distributed sensor data processing for many-coresabstractFuture many-core systems will rely heavily on a wide variety of sensors which provide run-time information about on-chip environment and workload. In this paper, a new dedicated infrastructure for distributed sensor processing for many-core systems is described. This infrastructure includes a sparse array of dedicated processors which evaluate on-chip sensor data and a two-level hierarchical network-on-chip (NoC) which allows for efficient sensor data collection. This design is evaluated using benchmark driven simulations for a three-dimensional (3D) stack, necessitating inter-layer sensor data communication. The experimental results for up to 1024 cores indicate that for typical sensor data collection rates, one sensor data processor (SDP) per 64 cores is optimal for sensor data latency. The use of a two-level NoC is shown to provide an average of 65% sensor data latency improvement versus a flat sensor data NoC structure for a 256-core system. Russell Tessier, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 3 |
| 2012 | Collaborative calibration of on-chip thermal sensors using performance countersabstractThermal sensors are currently deployed in processors to collect thermal information for dynamic thermal management (DTM). The calibration cost for thermal sensors can be prohibitively high as the number of on-chip sensors increases. We propose an on-line multi-sensor calibration method which combines potentially inaccurate temperature values obtained from two sources: temperature readings from thermal sensors and temperature estimations using system performance counters. A data fusion strategy based on Bayesian inference, which combines information from these two sources, is demonstrated along with a temperature estimation approach using performance counters. The approaches are verified via simulation for an AMD Athlon 64 processor with 24 on-chip temperature sensors scaled to a 45nm technology node. Our results show that the standard deviation of temperature sensor measurement errors can be reduced from 3 ~ 4 °C to ≤ 1 °C using the proposed method. Additionally, our MATLAB implementation shows that the new approach runs at least 67x faster than competing approaches based on Kalman filtering making it highly appropriate for run-time use. Shiting (Justin) Lu, Russell Tessier, Wayne P. Burleson |
ICCAD | 3 |
| 2012 | TARDIS: Time and Remanence Decay in SRAM to Implement Secure Protocols on Embedded Devices without Clocks
Amir Rahmati, Mastooreh Salajegheh, Daniel E. Holcomb, Jacob Sorber, Wayne P. Burleson, Kevin Fu |
USENIX Security Symposium | 5 |
| 2012 | An architecture-independent instruction shuffler to protect against side-channel attacksabstractEmbedded cryptographic systems, such as smart cards, require secure implementations that are robust to a variety of low-level attacks. Side-Channel Attacks (SCA) exploit the information such as power consumption, electromagnetic radiation and acoustic leaking through the device to uncover the secret information. Attackers can mount successful attacks with very modest resources in a short time period. Therefore, many methods have been proposed to increase the security against SCA. Randomizing the execution order of the instructions that are independent, i.e., random shuffling , is one of the most popular among them. Implementing instruction shuffling in software is either implementation specific or has a significant performance or code size overhead. To overcome these problems, we propose in this work a generic custom hardware unit to implement random instruction shuffling as an extension to existing processors. The unit operates between the CPU and the instruction cache (or memory, if no cache exists), without any modification to these components. Both true and pseudo random number generators are used to dynamically and locally provide the shuffling sequence. The unit is mainly designed for in-order processors, since the embedded devices subject to these kind of attacks use simple in-order processors. More advanced processors (e.g., superscalar, VLIW or EPIC processors) are already more resistant to these attacks because of their built-in ILP and wide word size. Our experiments on two different soft in-order processor cores, i.e., OpenRISC and MicroBlaze, implemented on FPGA show that the proposed unit could increase the security drastically with very modest resource overhead. With around 2% area, 1.5% power and no performance overhead, the shuffler increases the effort to mount a successful power analysis attack on AES software implementation over 360 times. Ali Galip Bayrak, Nikola Velickovic, Paolo Ienne, Wayne P. Burleson |
ACM Trans. Archit. Code Optim. | 4 |
| 2012 | Detecting Software Theft in Embedded Systems: A Side-Channel ApproachabstractSource code plagiarism has become a serious problem for the industry. Although there exist many software solutions for comparing source codes, they are often not practical in the embedded environment. Today's microcontrollers have frequently implemented a memory read protection that prevents a verifier from reading out the necessary source code. In this paper, we present three verification methods to detect software plagiarism in embedded software without knowing the implemented source code. All three approaches make use of side-channel information that is obtained during the execution of the suspicious code. The first method is passive, i.e., no previous modification of the original code is required. It determines the Hamming weights of the executed instructions of the suspicious device and uses string matching algorithms for comparisons with a reference implementation. In contrast, the second method inserts additional code fragments as a watermark that can be identified in the power consumption of the executed source code. As a third method, we present how this watermark can be extended by using a signature that serves as a proof-of-ownership. We show that particularly the last two approaches are very robust against code-transformation attacks. Georg T. Becker, Daehyun Strobel, Christof Paar, Wayne P. Burleson |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2012 | Design and Validation of Arbiter-Based PUFs for Sub-45-nm Low-Power Security ApplicationsabstractHarnessing unique physical properties of integrated circuits to enhance hardware security and IP protection has been extensively explored in recent years. Physical unclonable functions (PUFs) can sense inherent manufacturing variations as chip identifications. To enable the integration of PUFs into low-power and security applications, we study the impacts of process technology and supply voltage scaling on arbiter-based PUF circuit design. A Monte Carlo-based statistical analysis has demonstrated that advanced technologies and reduced supply voltage can improve the PUF uniqueness due to increased delay sensitivity. A linear regression approach has been leveraged to generate PUF delay profile by factoring in device, supply voltage and temperature variations. An accurate SVM-based software modeling analysis is used to verify the PUF additive delay behavior. Finally, postsilicon validation on arbiter-based PUF test chips in 45 nm SOICMOS technology has been correlated to simulation results and the inconsistency has been discussed. The test chips can resist the basic support vector machine attack due to the dynamic circuit effects and the limitation of our delay model. Lang Lin, Sudheendra Srivathsa, Dilip Kumar Krishnappa, Prasad Shabadi, Wayne P. Burleson |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2012 | Compact Expressions for Supply Noise Induced Period Jitter of Global Binary Clock TreesabstractPeriod jitter plays a critical role in global clock distribution design, because it directly impacts the time available for logic operations between sequential elements. Moreover, time-varying supply noise injected in global clock drivers can worsen the timing margin of critical paths by modulating period jitter. In the planning stage of global clock distribution for a high-end microprocessor, it is very critical to differentiate and understand the impacts of different design parameters on period jitter. However, it is hard to achieve due to complex relationship among different independent/dependent design parameters: supply noise amplitude, supply noise frequency, clock driver size, physical structures of interconnects, number of clock stages, temperature, process corners, etc. Jinwook Jang, Olivier Franza, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Hardware security in VLSIabstractThis special session addresses the increasingly critical area of hardware security in VLSI. As computing becomes ubiquitous, most systems need to implement various layers of security, all of which rely on the security of the underlying hardware. At the same time, persistent trends in VLSI are producing challenges and opportunities in the design of security mechanisms. Increased levels of integration, performance and power efficiency, allow the implementation of strong security protocols, however challenges arise due to the emergence of side-channel attacks and the ability to insert hardware Trojans that are difficult to detect. But on the other side, imperfections in the manufacturing process and on-chip noise can be leveraged to implement security primitives such as unique IDs and random numbers. The lessons learned from Hardware Security in advanced CMOS also have more general application. Statistical design and variation- and noise-aware methodologies are all current hot topics in VLSI that share themes with hardware security. Wayne P. Burleson, Yusuf Leblebici |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Robust signaling techniques for through silicon via bundlesabstractIn high performance 3D ICs with increasing trend for multi-core and NoC architectures, signaling techniques play a crucial role in determining the overall performance of the system. In this work, we explored single ended and differential signaling techniques for Through Silicon Via (TSV) bundles and analyzed their behavior in the presence of power supply noise. We obtained maximum data rate and energy/bit values for each of the signaling techniques and identified the dominant factors that determine these values. Simple analysis is carried out to understand the impact of fault tolerant scheme on the performance of the signaling technique. Frequency dependent RLGC parasitics of 3x3 and 4x4 TSV bundles are extracted using Ansoft Q3D Extractor. NCSU 45nm PDK and HSPICE simulation tool are used. For robustness to supply noise analysis, noise amplitudes of 2.5%, 5%, 7.5% and 10% of supply voltage and a noise frequency of 200MHz is considered. It is observed that Inter Symbol Interference (ISI), power supply noise and fault tolerant architecture play crucial role in determining the robust and high performance signaling technique for TSV bundles. Krishna C. Chillara, Jinwook Jang, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 3 |
| 2011 | A 45.6μ2 13.4μw 7.1v/v resolution sub-threshold based digital process-sensing circuit in 45nm CMOSabstractProcess optimization and yield enhancement techniques rely on physical sensors to provide them with feedback to perform accurate process-characterization across the die. The increased susceptibility of circuit performance characteristics to parameter variations in deep sub-micron technologies motivates the need for low-overhead (area/power) and high-sensitivity process monitoring circuits which can be used by compensation schemes to tune the process and meet frequency targets or power-budgets. To this end, we have proposed and implemented in 45nm CMOS-SOI, a novel, digital, on-chip process sensing circuit with a high sensitivity of 7.1V per volt variation in NMOS Vth, low power dissipation of 13.4¼W and a compact layout occupying 45.6¼2. The design is robust towards temperature variations and supply induced noise incurring an accuracy loss of 5mV (worst-case) and 4mV (1-A) respectively of the NMOS threshold-voltage. Basab Datta, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | A high sensitivity and process tolerant digital thermal sensing scheme for 3-D IcsabstractThermal sensing is a pressing need in stacked 3-D chips with limited number of vertical heat conduits. In 3-D systems with active temperature control, the controller is reliant on sensors placed on individual planes to provide the necessary thermal feedback. To this end we propose a delay-line based thermal-sensor which provides a high temperature-sensitivity, high level of process-robustness and is amenable to the TSV-based communication paradigm used in 3D systems. The temperature-sensitive piece mitigates the high process-susceptibility of 3-D circuits through usage of multiple logic-stages composed of long-channel devices and elimination of common-mode noise on the delay-line pair. The thermal information is conveyed to the controller in the form of a signal-frequency and hence is insensitive to path-mismatch in TSV wires. A high post-digitization temperature sensitivity of 0.82%/°C was achieved. The 1-σ accuracy loss due to process-variations and supply-noise was limited to 0.78°C and 1.06°C respectively indicating a high level of process-tolerance. Basab Datta, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | An arbiter based on-chip droop detector systemabstractThis paper presents a new design of supply droop detector system with a 10-20mV resolution at a very fast sampling rate (3GHz) in IBM 45nm SOI process. The detector system samples noisy supply voltage and stable reference voltages, and these sampled voltages are converted into different propagation delay values by using current starved inverters. Simple arbiters are used to convert these propagation delay differences into digitally reconstructed supply droop voltage. In addition, the detector has only 1 clock latency; a digital value of supply droop voltage is available after one clock cycle. This new droop detector system can greatly assist an adaptive frequency scaling system with a digital control and short latency. Jinwook Jang, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Implementing hardware Trojans: Experiences from a hardware Trojan challengeabstractHardware Trojans have become a growing concern in the design of secure integrated circuits. In this work, we present a set of novel hardware Trojans aimed at evading detection methods, designed as part of the CSAW Embedded System Challenge 2010. We introduced and implemented unique Trojans based on side-channel analysis that leak the secret key in the reference encryption algorithm. These side-channel-based Trojans do not impact the functionality of the design to minimize the possibility of detection. We have demonstrated the statistical analysis approach to attack such Trojans. Besides, we introduced Trojans that modify either the functional behavior or the electrical characteristics of the reference design. Novel techniques such as a Trojan draining the battery of a device do not have an immediate impact and hence avoid detection, but affect the long term reliability of the system. Georg T. Becker, Ashwin Lakshminarasimhan, Lang Lin, Sudheendra Srivathsa, Vikram B. Suresh, Wayne P. Burleson |
ICCD | 6 |
| 2011 | A Dedicated Monitoring Infrastructure for Multicore ProcessorsabstractOn-chip monitoring of environmental information, such as temperature, voltage, and error data, is becoming increasingly important. To address this need, a low-overhead architectural approach to monitor data collection and use in multicore systems is described. A key aspect of our standalone monitoring subsystem is a low-complexity, on-chip network designed to transport monitor data with multiple priority levels. Collected monitor information is evaluated by a dedicated processor. Experimental results using architectural and interconnect simulators show that the new low-overhead subsystem facilitates employment of thermal and delay-aware dynamic voltage and frequency scaling. In contrast to using existing on-chip interconnect resources to communicate monitor data, the new subsystem provides necessary bandwidth for monitor data traffic without impacting application data traffic. Synthesis results show that our dedicated monitoring approach consumes about 0.2% of multicore area and power resources for an 8-core system based on AMD Athlon 64 processor cores. Sailaja Madduri, Ramakrishna Vadlamani, Wayne P. Burleson, Russell Tessier |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Multicore soft error rate stabilization using adaptive dual modular redundancyabstractThe use of dynamic voltage and frequency scaling (DVFS) in contemporary multicores provides significant protection from unpredictable thermal events. A side effect of DVFS can be an increased processor exposure to soft errors. To address this issue, a flexible fault prevention mechanism has been developed to selectively enable a small amount of per-core dual modular redundancy (DMR) in response to increased vulnerability, as measured by the processor architectural vulnerability factor (AVF). Our new algorithm for DMR deployment aims to provide a stable effective soft error rate (SER) by using DMR in response to DVFS caused by thermal events. The algorithm is implemented in real-time on the multicore using a dedicated monitor network-on-chip and controller which evaluates thermal information and multicore performance statistics. Experiments with a multicore simulator using standard benchmarks show an average 6% improvement in overall power consumption and a stable SER by using selective DMR versus continuous DMR deployment. Ramakrishna Vadlamani, Wayne P. Burleson, Russell Tessier |
DATE | 3 |
| 2010 | Analysis and mitigation of NBTI-impact on PVT variability in repeated global interconnect performanceabstractRepeated interconnects remain the design choice for high-performance global communication due to their clearly defined performance metrics and smooth amalgamation to the VLSI CAD flow. Simultaneous variability in process-voltage-temperature (PVT) causes interconnect performance to fluctuate from its nominal value and hence it's an essential analysis needed to decide on timing margins for global wires. Negative Bias Temperature Instability (NBTI) has emerged as the most compelling device reliability concern in deep sub-micron technologies. In this paper, we thoroughly investigate the effects on NBTI-stress on PVT-induced variability in repeated global interconnect performance. The delay-spread due to PVT can deviate by 5-12% in the presence of NBTI stress. We propose 2 schemes to mitigate NBTI-stress effects on interconnect-delay and re-align the delay-distribution with its nominal state. The 1st involves upsizing all repeaters uniformly. A 5% size increment is found to be sufficient to counter the worst-case realistic deviation in delay-spread. The 2nd is a dynamic solution which involves using an NBTI detector circuit to monitor signal activity and assess device-status over a portion of the wire and increasing the drive strength of tunable buffers in the event of NBTI-stress. Basab Datta, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Circuit-level NBTI macro-models for collaborative reliability monitoringabstractThe increasing significance of Negative Bias Temperature Instability (NBTI) induced device-reliability degradation presents a compelling reason to perform efficient circuit-level reliability tracking. We propose a novel collaborative monitoring frame-work to track circuit level performance degradation caused specifically by NBTI. We use heterogeneous on-chip sensors to measure environmental and stress parameters and a macro-model to map the device-level degradation information into circuit-level reliability estimates. The macro-model is built using curve-fitted data and provides a practical upper bound of the path-delay-degradation to expect under a given set of dynamic parameters which includes operating conditions, process and stress parameters. Through usage of on-chip sensing resources we minimize the need for extensive circuit-specific analyses and also, the pessimism caused by assuming worst-case operating corners. We validate our approach on ISCAS-85 benchmarks and observe excellent correlation (>0.99) between worst-case SPICE observed and model-predicted path-delay degradation. Basab Datta, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Thermal-aware voltage droop compensation for multi-core architecturesabstractAs the rated performance of microprocessors increases, voltage droop emergencies become a significant problem. In this paper, two new techniques to combat voltage droop emergencies are explored. First, a direct connection between temperature and processor clock frequency modulation during voltage droops is established. In general, a higher temperature leads to a lower voltage droop with the same processor activity. Thus, processor frequencies can be reduced less at high temperature in an effort to prevent voltage emergencies. Through experimentation, the benefits of temperature-flexible frequency scaling are explored. Second, processor signatures consisting of performance statistics are used to identify when voltage droop compensation is needed in a multicore environment. The use of an independent on-chip interconnect network allows for the sharing of signatures across cores at run time. Signature sharing in combination with frequency throttling is shown to provide an improvement in average run-time performance in a number of cases for an eight-core multiprocessor. Basab Datta, Wayne P. Burleson, Russell Tessier |
ACM Great Lakes Symposium on VLSI | 3 |
| 2010 | Low-power sub-threshold design of secure physical unclonable functionsabstractThe unique and unpredictable nature of silicon enables the use of physical unclonable functions (PUFs) for chip identification and authentication. Since the function of PUFs depends on minute uncontrollable process variations, a low supply voltage can benefit PUFs by providing high sensitivity to variations and low power consumption as well. Motivated by this, we explore the feasibility of sub-threshold arbiter PUFs in 45nm CMOS technology. By modeling process variations and interconnect imbalance effects at the post-layout design level, we optimize the PUF supply voltage for the minimum power-delay product and investigate the trade-offs on PUF uniqueness and reliability. Moreover, we demonstrate that such a design optimization does not compromise the security of PUFs regarding modeling attacks and side-channel analysis attacks. Our final 64-stage sub-threshold PUF design only needs 418 gates and consumes 0.047 pJ energy per cycle, which is very promising for low-power wireless sensing and security applications. Lang Lin, Daniel E. Holcomb, Dilip Kumar Krishnappa, Prasad Shabadi, Wayne P. Burleson |
ISLPED | 5 |
| 2009 | Trojan Side-Channels: Lightweight Hardware Trojans through Side-Channel Engineering
Lang Lin, Markus Kasper, Tim Güneysu, Christof Paar, Wayne P. Burleson |
CHES | 5 |
| 2009 | Analysis and mitigation of process variation impacts on Power-Attack ToleranceabstractEmbedded cryptosystems show increased vulnerabilities to implementation attacks such as power analysis. CMOS technology trends are causing increased process variations which impact the data-dependent power of deep submicron cryptosystem designs. In this paper, we use Monte Carlo methods in SPICE circuit simulations to analyze the statistical properties of the data-dependent power with predictive 45nm CMOS device and ITRS process variation models. In addition to the "measurement to disclosure" (MTD) used in [3], we define a lower level metric, Power-Attack Tolerance (PAT), to model both dynamic power and leakage power data-dependence. We show that the PAT of a typical cryptographic component implementation using CMOS standard-cells can significantly deteriorate due to process variations, thus increasing the component's vulnerability to power attacks. Power-attack-resistant logic styles (e.g. SABL [9]) have been developed which increase PAT by an order of magnitude by balancing power consumption at the gate level with considerable overhead. However in the presence of process variations, the degradation probability of MTD is 57%. To mitigate this problem, we demonstrate a transistor sizing optimization method that can reduce such negative impacts to only 18% with minimal power and area overhead. Lang Lin, Wayne P. Burleson |
DAC | 2 |
| 2009 | A monitor interconnect and support subsystem for multicore processorsabstractIn many current SoCs, the architectural interface to on-chip monitors is ad hoc and inefficient. In this paper, a new architectural approach which advocates the use of a separate low-overhead subsystem for monitors is described. A key aspect of this approach is an on-chip interconnect specifically designed for monitor data with different priority levels. The efficiency of our monitor interconnect is assessed for a multicore system using both an interconnect and a system-level simulator. Collected monitor information is used by a dedicated processor to control the frequency and voltage of individual multicore processors. Experimental results show that the new low-overhead subsystem facilitates employment of thermal and delay-aware dynamic voltage and frequency scaling. Sailaja Madduri, Ramakrishna Vadlamani, Wayne P. Burleson, Russell Tessier |
DATE | 3 |
| 2009 | Low-power, process-variation tolerant on-chip thermal monitoring using track and hold based thermal sensorsabstractHigh die temperatures adversely impact CMOS device operation degrading both performance and reliability of integrated circuits. Dynamic thermal management (DTM) schemes rely on physical sensors to provide them with feedback to ensure an accurate and closed-loop throttling mechanism. Power-density trends in current-generation, high-performance processors motivate the need for multiple, low-power and highly accurate thermal monitoring circuits. Studies performed on power-thermal-characteristics of current processors suggest that better sensor-measurement accuracy translates into power-savings and better system-performance. To this end, we propose a novel, track and hold-based on-chip thermal sensor which provides an accuracy of <1ºC and has a low-power consumption of <25 W. Our design operates using the nominal VDD while at the same time takes advantage of the vastly-increased thermal sensitivity in sub-threshold mode. We also propose two schemes to improve the tolerance of thermal sensors to process-induced variations. The 1st scheme makes use of a leakage-current-sensor to identify the process-corner, selects the appropriate corresponding calibration-table and reduces the sensor-measurement error by 45-70%. The 2nd scheme provides a statistical approach towards mitigating the non-idealities by averaging response from multiple sensor-copies and reduces the measurement error by 65-85%. Basab Datta, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2009 | MOLES: Malicious off-chip leakage enabled by side-channelsabstractEconomic incentives have driven the semiconductor industry to separate design from fabrication in recent years. This trend leads to potential vulnerabilities from untrusted circuit foundries to covertly implant malicious hardware Trojans into a genuine design. Hardware Trojans provide back doors for on-chip manipulation, or leak secret information off-chip once the compromised IC is deployed in the field. This paper explores the design space of hardware Trojans and proposes a novel technique, "Malicious Off-chip Leakage Enabled by Side-channels" (MOLES), which employs power side-channels to convey secret information off-chip. An experimental MOLES circuit is designed with fewer than 50 gates and is embedded into an Advanced Encryption Standard (AES) cryptographic circuit in a predictive 45nm CMOS technology model. Engineered by a spread-spectrum technique, the MOLES technique is capable of leaking multi-bit information below the noise power level of the host IC to evade evaluators' detections. In addition, a generalized methodology for a class of MOLES circuits and design verification by statistical correlation analysis are presented. The goal of this work is to demonstrate the potential threats of MOLES on embedded system security. Nevertheless, MOLES could be constructively used for hardware authentication, fingerprinting and IP protection. Lang Lin, Wayne P. Burleson, Christof Paar |
ICCAD | 2 |
| 2009 | Power-Up SRAM State as an Identifying Fingerprint and Source of True Random NumbersabstractIntermittently powered applications create a need for low-cost security and privacy in potentially hostile environments, supported by primitives including identification and random number generation. Our measurements show that power-up of SRAM produces a physical fingerprint. We propose a system of fingerprint extraction and random numbers in SRAM (FERNS) that harvests static identity and randomness from existing volatile CMOS memory without requiring any dedicated circuitry. The identity results from manufacture-time physically random device threshold voltage mismatch, and the random numbers result from runtime physically random noise. We use experimental data from high-performance SRAM chips and the embedded SRAM of the WISP UHF RFID tag to validate the principles behind FERNS. For the SRAM chip, we demonstrate that 8-byte fingerprints can uniquely identify circuits among a population of 5,120 instances and extrapolate that 24-byte fingerprints would uniquely identify all instances ever produced. Using a smaller population, we demonstrate similar identifying ability from the embedded SRAM. In addition to identification, we show that SRAM fingerprints capture noise, enabling true random number generation. We demonstrate that a 512-byte SRAM fingerprint contains sufficient entropy to generate 128-bit true random numbers and that the generated numbers pass the NIST tests for runs, approximate entropy, and block frequency. Daniel E. Holcomb, Wayne P. Burleson, Kevin Fu |
IEEE Trans. Computers | 2 |
| 2008 | Low-power clock distribution in a multilayer core 3d microprocessorabstractClock distribution networks are extremely critical from a performance and power standpoint. They account for about 20-30% of the total power dissipated in current generation microprocessors. Many three-dimensional (3D) schemes propose to reduce interconnect length to improve performance and decrease power consumption. In this paper we propose a clock distribution network for a 3D multilayer core microprocessor. The 3D microprocessor floor plan has a single core folded onto multiple layers. A separate layer for the clock distribution network is proposed in the 3D microprocessor. This arrangement of a 3D chip stack reduces (a) power lost in long interconnects at block level and (b) in the clock distribution. Simulation results indicate a 15-20% power saving for this clock distribution scheme as compared to a 2D structure. A methodology for turning off the global clock grid along with the logic for an entire layer in a 3D stack is also proposed. Simulation results indicate an additional 8-10% savings in power with minimal impact on the critical parameters of the clock grid. Venkatesh Arunachalam, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Collaborative sensing of on-chip wire temperatures using interconnect based ring oscillatorsabstractHigh die temperatures adversely impact CMOS circuit operation degrading performance and reliability of both devices and interconnect. Current thermal scaling trends in multilevel low-k interconnect structures suggest an increasing heat density for the metal layers as a result of which the temperature dependence of logic is matched or often exceeded by that of interconnect. This motivates the need to perform 3-D thermal sensing in deep nanometer designs. We propose a novel sensor design that alleviates the complexities associated with time-to-digital conversion in wire-based thermal sensing. The sensing circuit makes use of wire-segments between individual stages of a ring-oscillator to perform thermal sensing using the oscillator frequency value as the mapping to corresponding wire temperature. Alternatively, the sensor can be tuned to strengthen the thermal sensitivity of the devices over that of interconnects to perform substrate-based sensing. We propose a collaborative scheme to sample the thermal status of the different metal levels and the substrate. The proposed sensor provides a resolution of 1°C while consuming an active power of 65-112µW and its sensitivity to process and supply noise can be minimized through design optimizations. Basab Datta, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Leakage-based differential power analysis (LDPA) on sub-90nm CMOS cryptosystemsabstractSince the vulnerability of cryptosystems to differential power analysis (DPA) was reported in 1999, various power analysis attacks and corresponding countermeasures have been studied. With the scaling down of supply voltage and CMOS technology below 90 nm, leakage power plays an increasing role in the overall power dissipation. Future cryptosystems need to address this trend, though it has not been of concern yet in low- cost cryptosystems such as smartcards and RFED tags which currently use older technologies and low performance transistors. In this paper, we explore the impact of leakage power on conventional DPA and the feasibility of a novel leakage-based DPA (LDPA). We first use SPICE simulations to explore the leakage dependence on input patterns of logic gates implemented in 90 nm, 65 nm, and 45 nm CMOS technologies. Then we simulate a successful LDPA on a subset of a DES cryptosystem with only 120 rounds, in contrast to the 200 rounds reported for a conventional DPA in 180 nm technology. Furthermore, we demonstrate how even a DES implementation using a DPA-resistant logic style can be broken with LDPA in 2000 rounds, compared with the conventional DPA using more than 5000 rounds. Lang Lin, Wayne P. Burleson |
ISCAS | 2 |
| 2008 | Reconfigurable Hardware for High-Security/ High-Performance Embedded Systems: The SAFES PerspectiveabstractEmbedded systems present significant security challenges due to their limited resources and power constraints. This paper focuses on the issues of building secure embedded systems on reconfigurable hardware and proposes a security architecture for embedded systems (SAFES). SAFES leverages the capabilities of reconfigurable hardware to provide efficient and flexible architectural support for security standards and defenses against a range of hardware attacks. The SAFES architecture is based on three main ideas: (1) reconfigurable security primitives; (2) reconfigurable hardware monitors; and (3) a hierarchy of security controllers at the primitive, system and executive level. Results are presented for reconfigurable AES and RC6 security primitives and highlight the value of such an architecture. This paper also emphasizes that reconfigurable hardware is not just a technology for hardware accelerators dedicated to security primitives as has been focused on by most studies but a real solution to provide high-security and high-performance for a system. Guy Gogniat, Tilman Wolf, Wayne P. Burleson, Jean-Philippe Diguet, Lilian Bossuet, Romain Vaslin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | Thermal Impacts on NoC InterconnectsabstractThermal issues are an increasing concern in microelectronics due to increased power density as well as the increasing vulnerability of the system to temperature effects (delay, leakage, reliability). NoCs promise to relieve many of the scaling problems that arise with increasing levels of on-chip system integration. This paper addressed the impacts on NoC interconnect circuits under harsh uniform temperature changes and non-uniform spatial temperature distribution profiles. Temporal and spatial thermal variations were addressed in 65 nm, 45 nm and 32 nm interconnect circuits. Standard repeater insertion and differential current sensing techniques have been implemented. The circuits were analyzed in temperatures as high as 150degC for the temporal variations, with a maximum temperature difference through wire of up to 50degC. High temperature caused more delay and power overhead in smaller technologies, i.e. 45 nm and 32 nm, by as much as 71% at 150degC for a given wirelength of 3 mm in 32 nm. Spatial temperature distribution profile influenced the propagation delay by 14.7% for a maximum thermal gradient of 50degC in the worst case for a 32 nm, 3 mm repeated wire. However, the delay degradation of an alternative differential current sensing (DCS) technique is largely determined by the amplifier temperature. Future work may consider the modeling of self-heating of the interconnect circuits Ibis Benito, Wayne P. Burleson |
NOCS | 3 |
| 2007 | Low power on-chip thermal sensors based on wiresabstractCurrent thermal scaling trends in multilevel low-k interconnect structures suggest an increasing heat density as we move from substrate to higher metal levels. Thus, the deterioration of interconnect performance at extreme temperatures has the capability to offset the degradation in device performance when operating at higher-than-normal temperatures. Existing thermal sensing approaches rely heavily on devices (MOS/diodes). They are optimized for a low area and power overhead but continue to suffer from leakage and self-heating and also, tend to disregard the thermal impact on interconnects. We propose an alternate approach of using interconnects to perform the thermal sensing. With feature-size shrinking, metal layers are closer to the substrate suggesting a strong correlation between interconnect temperature and thermal profile of the underlying substrate. Thus, in addition to quantifying the temperature impact on interconnect signal delay; output of proposed sensors can be used to estimate substrate thermal status as well. The simplistic schemes proposed allow reuse of existing on-chip resources such as drivers and time-digitizers, have a low power requirement and are robust against variations in wire dimensions, non-uniform temperature distribution and supply noise. Basab Datta, Wayne P. Burleson |
VLSI-SoC | 2 |
| 2007 | Current-Sensing and Repeater Hybrid Circuit Technique for On-Chip InterconnectsabstractIn this paper, hybrids based on current-sensing and repeaters are proposed for on-chip interconnects in an effort to overcome the limitations of these techniques. A novel receiver for current-sensing results in static power savings and allows an easier transition from current-sensing to traditional full rail voltage signals. Measurements of hybrids on a 0.18-m CMOS technology show significant gains over repeater insertion in delay across wire lengths. Hybrids can also be used in placement constrained and low-noise scenarios to achieve delay and power benefits. Atul Maheshwari, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Synchro-Tokens: A Deterministic GALS Methodology for Chip-Level Debug and TestabstractThis paper describes a novel deterministic globally-asynchronous locally-synchronous (GALS) methodology called "synchro-tokens". Wrappers around the synchronous blocks keep the system globally asynchronous while ensuring that each transition, although arriving at a nondeterministic time, is sensed by the synchronous block during a deterministic cycle of the local clock. This determinism facilitates debug and test methodologies, such as the use of stored-pattern testers, which are effective only when the system behavior is predictable and repeatable. Applications of synchro-tokens to GALS systems with two or more synchronous blocks and one or more asynchronous data channels are shown. Synchro-tokens supports both pipelined and unpipelined channels and a variety of clock generation methodologies. Novel schematic level designs of the wrapper components in a 180-nm technology are used to compare the performance of several different deterministic GALS design styles. Matthew W. Heath, Wayne P. Burleson, Ian G. Harris |
IEEE Trans. Computers | 2 |
| 2005 | An energy-aware active smart cardabstractDespite recent advances in smart card technology, most modern smart cards continue to rely on card readers for power and clocking, creating a potential security gap. In this paper, we present an energy-aware smart card architecture that operates using an embedded battery and crystal. This low-power VLSI system is continually active and provides enhanced security through periodic internal update when the card is detached from a reader. Our architecture achieves reduced power consumption by deactivating the majority of its circuitry, including an embedded microcontroller, for the vast majority of the card's lifetime. A proof-of-concept prototype implementation of the architecture has been developed including register-transfer-level and gate-level designs which have been synthesized to silicon. To permit extended operation for up to 18 months, critical design logic has been implemented using ultralow-power (adiabatic) circuit techniques. Russell Tessier, David Jasinski, Atul Maheshwari, Aiyappan Natarajan, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2005 | A reconfigurable, power-efficient adaptive Viterbi decoderabstractError-correcting convolutional codes provide a proven mechanism to limit the effects of noise in digital data transmission. Although hardware implementations of decoding algorithms, such as the Viterbi algorithm, have shown good noise tolerance for error-correcting codes, these implementations require an exponential increase in very large scale integration area and power consumption to achieve increased decoding accuracy. To achieve reduced decoder power consumption, we have examined and implemented decoders based on the reduced-complexity adaptive Viterbi algorithm (AVA). Run-time dynamic reconfiguration is performed in response to varying communication channel-noise conditions to match minimized power consumption to required error-correction capabilities. Experimental calculations indicate that the use of dynamic reconfiguration leads to a 69% reduction in decoder power consumption over a nonreconfigurable field-programmable gate array implementation with no loss of decode accuracy. Russell Tessier, Sriram Swaminathan, Ramaswamy Ramaswamy, Dennis Goeckel, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2004 | Synchro-Tokens: Eliminating Nondeterminism to Enable Chip-Level Test of Globally-Asynchronous Locally-Synchronous SoC?sabstractGlobally asynchronous locally synchronous (GALS) clocking applied to a system-on-a-chip (SoC) results in a design in which each core is a synchronous block (SB) of logic with a locally generated clock. Inter-core communication is asynchronous and controlled by wrapper logic around the cores. The nondeterministic synchronization used by most GALS architectures makes chip-level silicon debug and functional test difficult and costly. Deterministic GALS methodologies make dataflow assumptions which are only valid for a very limited set of applications. This paper describes a novel deterministic GALS methodology called "synchro-tokens" whose parameterized wrappers are flexible enough to be useful for a wide range of applications while supporting synchronous debug and test methodologies such as 1149.1 and P1500. The validation of determinism, estimation of area overhead, and analysis of performance impact are detailed. Matthew W. Heath, Wayne P. Burleson, Ian G. Harris |
DATE | 2 |
| 2004 | Mitigating static power in current-sensed interconnectsabstractInterconnects are an increasing concern in recent years, resulting in novel techniques such as current sensing. However these techniques must be designed to tradeoff delay and both dynamic and static power consumption. This paper presents an innovative approach to reduce static power in differential current-sensed interconnects. This system uses a self-timed shut-off system to reduce static currents used to bias the current sense amplifier. Results indicated that the self timed shut-off system reduced static power by 23.4% for a 10mm line in 250nm technology with no overhead in performance. On an average it reduced static power by 9.7% for 4mm-9mm lines over 180nm, 130nm, 100nm and 65nm technologies and 6% from 10mm-15mm line over the same set of technologies as before. Physical design of the system was implemented in 250nm technology along with the implementation of a test circuit, ready to be fabricated. Extensions of this shut-off mechanism may be useful for mitigating leakage power in a variety of interconnect circuits. Vishak Venkatraman, Atul Maheshwari, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 3 |
| 2004 | Dynamically Configurable Security for SRAM FPGA BitstreamsabstractSummary form only given. We propose a solution to improve the security of SRAM FPGAs through bitstream encryption. This proposition is distinct from other works because it uses the latest capabilities of SRAM FPGAs like partial and dynamic reconfiguration. It doesn't need any external battery to store the secret key. It opens a new way of application partitioning according to the security policy. Lilian Bossuet, Guy Gogniat, Wayne P. Burleson |
IPDPS | 3 |
| 2004 | Differential current-sensing for on-chip interconnectsabstractThis paper presents a differential current-sensing technique as an alternative to existing circuit techniques for on-chip interconnects. Using a novel receiver circuit, it is shown that, delay-optimal current-sensing is a faster (20% on an average) option as compared to the delay-optimal repeater insertion technique for single-cycle wires. Delay benefit for current-sensing increases with an increase in wire width. Unlike repeaters, current-sensing does not require placement of buffers along the wire, and hence, eliminates any placement constraints. Inductive effects are negligible in differential current-sensing. Current-sensing also provides a tighter bound on delay with respect to process variations. However, current-sensing has some drawbacks. It is power inefficient due to the presence of static-power dissipation. Current-sensing is essentially a low-swing signaling technique, and hence, it is sensitive to full swing aggressor noise. Atul Maheshwari, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Trading off transient fault tolerance and power consumption in deep submicron (DSM) VLSI circuitsabstractHigh fault tolerance for transient faults and low-power consumption are key objectives in the design of critical embedded systems. Systems like smart cards, PDAs, wearable computers, pacemakers, defibrillators, and other electronic gadgets must not only be designed for fault tolerance but also for ultra-low-power consumption due to limited battery life. In this paper, a highly accurate method of estimating fault tolerance in terms of mean time to failure (MTTF) is presented. The estimation is based on circuit-level simulations (HSPICE) and uses a double exponential current-source fault model. Using counters, it is shown that the transient fault tolerance and power dissipation of low-power circuits are at odds and allow for a power fault-tolerance tradeoff. Architecture and circuit level fault tolerance and low-power techniques are used to demonstrate and quantify this tradeoff. Estimates show that incorporation of these techniques results either in a design with an MTTF of 36 years and power consumption of 102 /spl mu/W or a design with an MTTF of 12 years and power consumption of 20 /spl mu/W. Depending on the criticality of the system and the power budget, certain techniques might be preferred over others, resulting in either a more fault tolerant or a lower power design, at the sacrifice of the alternative objective. Atul Maheshwari, Wayne P. Burleson, Russell Tessier |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Repeater and current-sensing hybrid circuits for on-chip interconnectsabstractDesigning interconnects is becoming an increasingly challenging problem with a few solutions. In this paper hybrid circuit based on the well known delay-optimal repeaters and the recently proposed differential current-sensing is presented. Comparison in terms of delay, power and area is drawn between various versions of the hybrid circuit with delay-optimal repeater insertion and differential current sensing in order to derive at the best possible solution. It is shown that driving 25% of the wire with repeaters and remaining with current-sensing is the best solution from delay standpoint (about 30% faster than delay-optimal repeaters). Not only do hybrid circuits consume less area, they are also a more acceptable solution from placement point of view due to fewer repeaters and a long segment of uninterrupted wire. Static power consumption inherited from differential current-sensing is the biggest drawback of the hybrid circuits. Atul Maheshwari, Wayne P. Burleson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2003 | A hybrid adiabatic content addressable memory for ultra low-power applicationsabstractThis paper presents a hybrid adiabatic content addressable memory (CAM). The CAM uses an adiabatic switching technique to reduce the energy consumption in the match line while keeping the performance for the read/write operation. The adiabatic CAM is suitable for ultra low-power, low performance applications such as smart cards and portable devices. This CAM uses a clocked power supply for the match line while the rest of the circuit is the same as the basic CAM. A novel smart card application which uses the adiabatic CAM is illustrated. The circuit simulations for a 16x16 and 32x32 CAM were done in Hspice using 0.18 μm Berkeley models and the energy dissipation was compared with a basic CAM. The results show three orders of magnitude in energy savings for the 16x16 CAM and one order of magnitude savings for the 32x32 CAM when operated at 2Mhz. The maximum frequency of operation for which there was considerable energy savings was found to be 200 Mhz with a 20% and 45% energy savings for 16x16 and 32x32 CAM respectively. Aiyappan Natarajan, David Jasinski, Wayne P. Burleson, Russell Tessier |
ACM Great Lakes Symposium on VLSI | 3 |
| 2003 | Adaptive system on a chip (ASOC): a backbone for power-aware signal processing coresabstractFor motion estimation (ME) and discrete cosine transform (DCT) of MPEG video encoding, content variation and perceptual tolerance in video signals can be exploited to gracefully trade quality for low power. As a result, power-aware hardware cores have been proposed for these video encoding subsystems. Adaptive system-on-a-chip, aSoC, supports power-aware cores by providing an on-chip communications framework designed to promote scalability and flexibility in system-on-a-chip designs. This paper describes aSoC's ability to dynamically control voltage and frequency scaling through a simple voltage and frequency selection scheme. A small demonstration system is tested and shows up to 90% reduction in core power when the aSoC voltage scaling features are enabled. Andrew Laffely, Russell Tessier, Wayne P. Burleson |
ICIP (3) | 4 |
| 2002 | A dynamically reconfigurable adaptive viterbi decoderabstractThe use of error-correcting codes has proven to be an effective way to overcome data corruption in digital communication channels. Although widely-used, the most popular communications decoding algorithm, the Viterbi algorithm, requires an exponential increase in hardware complexity to achieve greater decode accuracy. In this paper, we describe the analysis and implementation of a reduced-complexity decode approach, the adaptive Viterbi algorithm (AVA). Our AVA design is implemented in reconfigurable hardware to take full advantage of algorithm parallelism and specialization. Run-time dynamic reconfiguration is used in response to changing channel noise conditions to achieve improved decoder performance. Implementation parameters for the decoder have been determined through simulation and the decoder has been implemented on a Xilinx XC4036-based PCI board. An overall decode performance improvement of 7.5X for AVA has been achieved versus algorithm implementation on a Celeron-processor based system. The use of dynamic reconfiguration leads to a 20% performance improvement over a static implementation with no loss of decode accuracy. Sriram Swaminathan, Russell Tessier, Dennis Goeckel, Wayne P. Burleson |
FPGA | 4 |
| 2002 | Boosters for driving long onchip interconnects - design issues, interconnect synthesis, and comparison with repeatersabstract50-62 Ankireddy Nalamalpu, Wayne P. Burleson |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | Dynamically parameterized algorithms and architectures to exploit signal variations for improved performance and reduced powerabstractSignal processing algorithms and architectures can use dynamic reconfiguration to exploit variations in signal statistics with the objectives of improved performance and reduced power consumption. Parameters provide a simple and formal way to characterize incremental changes to a computation and its computing mechanism. This paper examines five parameterized computations which are typically implemented in hardware for a wireless multimedia terminal: (1) motion estimation, (2) discrete cosine transform, (3) Lempel-Ziv lossless compression, (4) 3D graphics light rendering and (5) Viterbi decoding. Each computation is examined for the capability of dynamically adapting the algorithm and architecture parameters to variations in their respective input signals. Dynamically reconfigurable low-power implementations of each computation are currently underway. Wayne P. Burleson, Russell Tessier, Dennis Goeckel, Sriram Swaminathan, Prashant Jain, Jeongseon Euh, Subramanian Venkatraman, Vidhya Thyagarajan |
ICASSP | 1 |
| 2001 | Boosters for driving long on-chip interconnects: design issues, interconnect synthesis and comparison with repeatersabstractTrends in CMOS technology and VLSI architectures are causing interconnect to play an increasing role in overall performance, power consumption and design effort. Traditionally, repeaters are used for driving long on-chip interconnects, however recent studies indicate that repeaters are using increasing area, power, and design resources as well as having an inherent limit in how much they can improve performance [4,10,14]. This paper presents a new circuit called a {\bf booster} which compares favorably with repeaters in terms of area, performance, power and placement sensitivity. Boosters also have the advantage of being bidirectional and providing a low impedance termination to improve signal integrity. Driver edge rates are slower and peak power is drastically reduced compared to repeaters, thus improving signal integrity and mitigating inductive effects. Boosters are shown to be more than 20\% faster for driving a variety of interconnect loads over conventional repeaters in 0.16 $\mu$m CMOS technology. Boosters are typically inserted three times less frequently than repeaters for optimal performance, resulting in fewer boosters for driving the same interconnect lengths thereby saving on area, power and placement effort. Ankireddy Nalamalpu, Wayne P. Burleson |
ISPD | 2 |
| 2000 | Low power digital design in FPGAs (poster abstract): a study of pipeline architectures implemented in a FPGA using a low supply voltage to reduce power consumptionabstractNo abstract available. Andrés David García García, Jean-Luc Danger, Wayne P. Burleson |
FPGA | 3 |
| 2000 | Low power digital design in FPGAs: a study of pipeline architectures implemented in a FPGA using a low supply voltage to reduce power consumptionabstractSome techniques for low power operation in VLSI using the lowest possible supply voltage coupled with an architectural optimization have shown that we can save power even if we increase silicon area. In this paper we present a strategy to reduce power consumption in FPGAs based on pipeline architectures working with a low supply voltage. Andrés David García García, Wayne P. Burleson, Jean-Luc Danger |
ISCAS | 2 |
| 2000 | Repeater insertion in deep sub-micron CMOS: ramp-based analytical model and placement sensitivity analysisabstractRepeaters are now widely used to increase the performance of long on-chip interconnections in CMOS VLSI. In this paper, we take an updated look at repeater insertion in state-of-the-art CMOS, using a new more detailed model. In spite of the more complex model, we present closed form expressions for the delay and the optimal repeater spacing and sizing. Our model is based on the alpha-power law to account for the short-channel effects and resistive loads that arise in deep sub-micron technologies. Unlike previous work, we model the repeater input as a ramp and accurately model both linear and saturation regions of operation for estimating the propagation delay. Our analytical repeater model is applied for estimating the performance of driving various repeated RC loads and exhibits a maximum error of only 5% when compared with SPICE in a 0.13 /spl mu/m CMOS technology. In practice, it is not always feasible to insert the repeaters at the exact optimal locations along an interconnect. We present a placement sensitivity analysis to quantify the effect of the sub-optimal repeater placement on performance. Closed form expressions are derived to re-size the repeaters to compensate for the sub-optimal placement. Ankireddy Nalamalpu, Wayne P. Burleson |
ISCAS | 2 |
| 1999 | Configuration Cloning: Exploiting Regularity in Dynamic DSP ArchitecturesabstractArticle Free Access Share on Configuration cloning: exploiting regularity in dynamic DSP architectures Authors: S. R. Park Dept. of ECE, University of Massachusetts, Amherst, MA Dept. of ECE, University of Massachusetts, Amherst, MAView Profile , W. Burleson Dept. of ECE, University of Massachusetts, Amherst, MA Dept. of ECE, University of Massachusetts, Amherst, MAView Profile Authors Info & Claims FPGA '99: Proceedings of the 1999 ACM/SIGDA seventh international symposium on Field programmable gate arraysFebruary 1999 Pages 81–89https://doi.org/10.1145/296399.296433Published:01 February 1999Publication History 7citation322DownloadsMetricsTotal Citations7Total Downloads322Last 12 Months7Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF S. R. Park, Wayne P. Burleson |
FPGA | 2 |
| 1999 | The spring scheduling coprocessor: a scheduling acceleratorabstractThe spring scheduling coprocessor is a novel very large scale integration (VLSI) accelerator for multiprocessor real-time systems. The coprocessor can be used for static as well as online scheduling. Many different policies and their combinations can be used (e.g., earliest deadline first, highest value first, or resource-oriented policies such as earliest available time first). In this paper, we describe a coprocessor architecture, a CMOS implementation, an implementation of the host/coprocessor interface and a study of the overall performance improvement. We show that the current VLSI chip speeds up the main portion of the scheduling operation by over three orders of magnitude. We also present an overall system improvement analysis by accounting for the operating system overheads and identify the next set of bottlenecks to improve. The scheduling coprocessor includes several novel VLSI features. It is implemented as a parallel architecture for scheduling that is parameterized for different numbers of tasks, numbers of resources, and internal wordlengths. The architecture was implemented using a single-phase clocking style in several novel ways. The 328 000 transistor custom 2-/spl mu/m VLSI accelerator running with a 100-MHz clock, combined with careful hardware/software co-design results in a considerable performance improvement, thus removing a major bottleneck in real-time systems. Wayne P. Burleson, Jason Ko, Douglas Niehaus, Krithi Ramamritham, John A. Stankovic, Gary Wallace, Charles C. Weems |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | Reconfiguration for power saving in real-time motion estimationabstractMotion estimation presents a class of algorithms well-suited to reconfigurable hardware due to their variable computational load, highly structured array architectures, robust reduced complexity algorithms, and a motivation for low power implementations in portable video products. Motion estimation is the most computationally demanding part of video compression algorithms and hence usually requires hardware support for real-time implementation. However, dedicated hardware usually requires that the algorithm and most of its parameters be hardwired. Reconfigurable hardware based on FPGAs allows the parallelism of hardware implementations with the flexibility of software. The statistics of motion vectors can be monitored on a frame by frame basis to choose appropriate algorithm and hardware configurations. Unlike some proposed applications of dynamic reconfiguration, this rate can easily be supported by existing FPGA technology. Another novel aspect of this work is that we use power savings as a motivation for the reconfiguration. Although FPGAs are not a very power efficient technology; careful design of array architectures can allow power to be saved by avoiding unnecessary computation by adjusting the search area according to the changing characteristics of an input video signal. Another more general result is that further power saving can be achieved by utilizing free FPGA resources as local memory to avoid power-hungry off-chip communication. Practical implementation issues using Xilinx 6200 series FPGAs are also discussed. S. R. Park, Wayne P. Burleson |
ICASSP | 2 |
| 1998 | Wave-pipelining: a tutorial and research surveyabstractWave-pipelining is a method of high-performance circuit design which implements pipelining in logic without the use of intermediate latches or registers. The combination of high-performance integrated circuit (IC) technologies, pipelined architectures, and sophisticated computer-aided design (CAD) tools has converted wave-pipelining from a theoretical oddity into a realistic, although challenging, VLSI design method. This paper presents a tutorial of the principles of wave-pipelining and a survey of wave-pipelined VLSI chips and CAD tools for the synthesis and analysis of wave-pipelined circuits. Wayne P. Burleson, Maciej J. Ciesielski, Fabian Klass |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1998 | Efficient VLSI for Lempel-Ziv compression in wireless data communication networksabstractWe present a parallel algorithm, architecture, and implementation for efficient Lempel-Ziv (LZ)-based data compression. The parallel algorithm exhibits a scalable, parameterized, and regular structure and is well suited for VLSI array implementation. Based on our parallel algorithm and systematic design methodologies, two semisystolic array architectures have been developed which are low power and area efficient. The first architecture trades off the compression speed for the area and has a low run-time overhead for multichannel compression. The second architecture achieves a high compression rate (one data symbol per clock) at the expense of the area due to a large clock load and global wiring. Compared to a recent state-of-the-art parallel architecture, our first array structure requires significantly less chip area (/spl sime/330 k versus /spl sime/36 k transistors) and more than an order of magnitude less power (/spl ap/1.0 W versus /spl ap/70 mW) while still providing the compression speed required for most data communication applications. Hence, data compression can be adopted in portable data communication as well as wireless local area networks. The second architecture has at least three times less area and power while providing the same constant compression rate. To demonstrate the correctness of our design, a prototype module for the first architecture has been implemented using 1.2 /spl mu/ complementary metal-oxide-semiconductor (CMOS) technology. The compression module contains 32 simple and identical processors, has an average compression rate of 12.5 million bytes/s, and consumes 18.34 mW without the dictionary (/spl ap/70 mW with a 4.1k SRAM for the dictionary) while operating at a 100 MHz clock rate (simulated). Bongjin Jung, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | Performance optimization of wireless local area networks through VLSI data compression
Bongjin Jung, Wayne P. Burleson |
Wirel. Networks | 2 |
| 1997 | An FPGA-based data acquisition system for a 95 GHz W-band radarabstractWe describe a 95 GHz radar for an unmanned aerial vehicle (UAV). The radar measures vertical profiles of the reflectivity and Doppler velocity of clouds, which are then telemetered to the ground for storage. Telemetry bandwidth requires that substantial real-time data processing be done on the UAV in a low-power (less than 100 watts) and small size (less than 1 cubic foot) system. A prototype was developed in less than a year, thus a flexible programmable technology was required. Although typical remote sensing radars use DSP chips, it was determined that our power, size, performance and design-time requirements were best met using FPGA technology. Our system is based on the Giga-Ops Spectrum system which uses Xilinx FPGAs on a novel modular PCI board. Unlike numerous recent FPGA-based signal processors, this presents a new class of applications and embedded system requirements. Reconfigurable capabilities are currently being explored to support radar algorithms which can adapt to a changing environment. Michael Petronino, Ray Bambha, James R. Carswell, Wayne P. Burleson |
ICASSP | 4 |
| 1997 | VLSI array algorithms and architectures for RSA modular multiplicationabstractWe present two novel iterative algorithms and their array structures for integer modular multiplication. The algorithms are designed for Rivest-Shamir-Adelman (RSA) cryptography and are based on the familiar iterative Horner's rule, but use precalculated complements of the modulus. The problem of deciding which multiples of the modulus to subtract in intermediate iteration stages has been simplified using simple look-up of precalculated complement numbers, thus allowing a finer-grain pipeline. Both algorithms use a carry save adder scheme with module reduction performed on each intermediate partial product which results in an output in carry-save format. Regularity and local connections make both algorithms suitable for high-performance array implementation in FPGA's or deep submicron VLSI. The processing nodes consist of just one or two full adders and a simple multiplexor. The stored complement numbers need to be precalculated only when the modulus is changed, thus not affecting the performance of the main computation. In both cases, there exists a bit-level systolic schedule, which means the array can be fully pipelined for high performance and can also easily be mapped to linear arrays for various space/time tradeoffs. Yongjin Jeong, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | Low-power encodings for global communication in CMOS VLSIabstractTechnology trends and especially portable applications are adding a third dimension (power) to the previously two-dimensional (speed, area) VLSI design space. A large portion of power dissipation in high performance CMOS VLSI is due to the inherent difficulties in global communication at high rates and we propose several approaches to address the problem. These techniques can be generalized at different levels in the design process. Global communication typically involves driving large capacitive loads which inherently require significant power. However, by carefully choosing the data representation, or encoding, of these signals, the average and peak power dissipation can be minimized. Redundancy can be added in space (number of bus lines), time (number of cycles) and voltage (number of distinct amplitude levels). The proposed codes can be used on a class of terminated off-chip board-level buses with level signaling, or on tristate on-chip buses with level or transition signaling. Mircea R. Stan, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1996 | Two dimensional codes for low powerabstractCoding was previously proposed for reducing power consumption in CMOS. The original formulations use extra redundancy in space (number of bus lines) for reducing the bus transition activity (and consequently the dynamic power and simultaneous switching noise). This paper proposes several new coding techniques for low power. First it looks at codes in which redundancy in time is used for reduced bus activity. Two-dimensional codes with redundancy in both time and space can then be developed for extra power reduction. Interestingly, these two-dimensional codes can be unrolled in either space or time in order to obtain new one-dimensional codes in the other dimension. More powerful codes using Run-Length Limited (RLL), phase-modulation and amplitude-modulation techniques are finally proposed. Mircea R. Stan, Wayne P. Burleson |
ISLPED | 2 |
| 1995 | Equivalence Checking of Datapaths Based on Canonical Arithmetic ExpressionsabstractAbstract| Numerous formal veri cation systems have been proposed and developed for Finite Sate Machine based control units (notably SMV [19] as well as others).However, most research on the equivalence checking of datapaths is still con ned to the bit-level.F ormal veri cation of arithmetic expressions and synthesized datapaths, especially considering nite word-length computation, has not been addressed.Thus formal veri cation techniques have been prohibited from more extensive applications in numerical and Digital Signal Processing.In this paper a formal system, called Conditional Term Rewriting on Attribute Syntax Trees (ConTRAST) is developed and demonstrated for verifying the equivalence between two dierently synthesized datapaths.This result arises from a sophisticated integration of attribute grammars, which provide expressive data structures for syntactic and semantic information about designed datapaths, and term rewriting systems, which transform functionally equivalent datapaths into the same canonical form.The equivalence relation is de ned as a congruence closure in the rewriting system, which can be generated from arbitrary axioms, such as associativity, commutativity, etc. in a certain algebraic system.Furthermore, the eect of nite word-lengths and their associated arithmetic precision are also considered in the de nition of equivalence classes.As a particular application of ConTRAST, a formal veri cation system is designed to check equivalence under precision constraints.The results of initial DSP synthesis experiments are displayed, where two dierently implemented IIR lters in direct II and cascaded architectures are automatically compared under given precision constraints. Wayne P. Burleson |
DAC | 2 |
| 1995 | Coding a terminated bus for low powerabstractCoding was proposed as a general method of decreasing power dissipation for the I/O. Lower power dissipation can be obtained by using extra bus liner for coding the data. This paper presents an application of the general theory of limited-weight codes for a class of parallel terminated buses with pull-up terminators (e.g. Rambus). Power dissipation on such a bus-line is larger for a logical 1 and it follows that patterns with few 1s should be chosen. A perfect k/2-limited weight code equivalent to the previously proposed Bus-Invert method and a novel non-perfect 3-limited weight code are described. Both codes can be algorithmically generated and practical issues related to their implementation on the Rambus are discussed. Mircea R. Stan, Wayne P. Burleson |
Great Lakes Symposium on VLSI | 2 |
| 1995 | High-Level Estimation of High-Performance Architectures for Reed-Solomon DecodingabstractReed-Solomon (RS) codes are a standard error-detection and correction approach in diverse communication and computer systems. The major obstacle to wider application has been a computationally complex decoding structure not native to conventional languages and microprocessors due to finite field arithmetic. Based on our previous work on finite field arithmetic, this paper shows a detailed computational structure of RS decoding algorithms, which are expressed by parallel hardware implementations using an array synthesis method. The purpose of this paper is to explore detailed computational structures of the RS decoder as a purely parallel implementation and make high level estimations on area and performance without actually building hardware. Depending on design criteria, we can then apply various mapping strategies to obtain a set of implementations which satisfy a broad range of performance, cost, power and reliability requirements. Yongjin Jeong, Wayne P. Burleson |
ISCAS | 2 |
| 1995 | Bus-invert coding for low-power I/OabstractTechnology trends and especially portable applications drive the quest for low-power VLSI design. Solutions that involve algorithmic, structural or physical transformations are sought. The focus is on developing low-power circuits without affecting too much the performance (area, latency, period). For CMOS circuits most power is dissipated as dynamic power for charging and discharging node capacitances. This is why many promising results in low-power design are obtained by minimizing the number of transitions inside the CMOS circuit. While it is generally accepted that because of the large capacitances involved much of the power dissipated by an IC is at the I/O little has been specifically done for decreasing the I/O power dissipation. We propose the bus-invert method of coding the I/O which lowers the bus activity and thus decreases the I/O peak power dissipation by 50% and the I/O average power dissipation by up to 25%. The method is general but applies best for dealing with buses. This is fortunate because buses are indeed most likely to have very large capacitances associated with them and consequently dissipate a lot of power.> Mircea R. Stan, Wayne P. Burleson |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1994 | Distributed control synthesis for data-dependent iterative algorithmsabstractData-dependent control flow changes are typically implemented in complex general-purpose controllers. However, in medium to fine-grained iterative algorithms found in DSP and arithmetic, it is desirable for both cost and performance reasons to develop simplified and distributed control structures throughout the array architectures. We present a transformation technique to systematically convert an iterative algorithm whose loop bounds are data-dependent to an equivalent data-independent regular algorithm. We use the worst-case situation to define the structure of a given computation. Then, depending upon applications, various optimizations are performed to minimize cost overhead associated with data-dependent characteristics. To demonstrate the feasibility of our methodologies, we have designed and simulated parallel algorithms for two important applications using the proposed transformation. Based on the resulting algorithms, we have developed VLSI array architectures using conventional methods of regular array synthesis (S.Y. Kung, 1988). Compared with previously reported structures, we have achieved substantial improvements in terms of both performance and array costs.> Bongjin Jung, Yongjin Jeong, Wayne P. Burleson |
ASAP | 3 |
| 1994 | Forum: Wave-pipelining: Is it Practical?abstractWave-pipelining has recently drawn considerable interest from both the academic and industrial communities as a method for high-speed pipelining without the use of intermediate latching. Although high-performance IC technologies, pipelined RISC and DSP architectures, and sophisticated CAD tools have enabled and driven several recent demonstration circuits, significant challenges remain before wave-pipelining will be used extensively in commercial chips. In this forum, we discuss whether these challenges are surmountable, and if so, what are the techniques that will be needed both now and with future VLSI technologies.> Wayne P. Burleson, Leonard W. Cotten, Fabian Klass, Maciej J. Ciesielski |
ISCAS | 1 |
| 1994 | A VLSI Systolic Array Architecture for Lempel-Ziv-Based Data CompressionabstractWe present a parallel algorithm, architecture, and implementation for Lempel-Ziv-based data compression. The parallel algorithm exhibits a regular structure and is well suited for parallel VLSI array implementation. Based on our parallel algorithm, a word-parallel systolic array has been developed using systematic design methodologies. Compared to a recent systolic architecture, our array structure is substantially faster, with latency of N/2+M compared to 2N+M where M is the maximum allowable length of symbols to be encoded at each encoding step and N is the length of symbols in an encoding buffer which have already been encoded. Furthermore the architecture consumes significantly less area and has a faster clock rate.> Bongjin Jung, Wayne P. Burleson |
ISCAS | 2 |
| 1993 | VLSI array synthesis for polynomial GCD computationabstractPolynomial GCD (greatest common divisor) finding is an important problem in algebraic computation, especially in decoding error correcting codes. The authors show a new systolic array structure for the polynomial GCD problem using a systematic array synthesis technique. The VLSI implementation of the array structure is area-efficient and achieves maximum throughput with pipelining. The dependency graph (DG) of the Euclid GCD algorithm is drawn using iterated polynomial division. The resulting DG is data-dependent and variable-sized. The authors consider the worst-case implementation to make the DG data-dependent and fixed-size, where data-dependences are hidden inside by introducing four different working modes in each DG node. This novel approach requires just a few additional multiplexors and can be generalized for other data-dependent and variable-sized computation. The authors then map the DG to a one-dimensional systolic array using a linear mapping. The new array structure has m/sub 0/ + n/sub 0/ + 1 processing elements, where m/sub 0/ and n/sub 0/ are degrees of two polynomials. It can find a GCD of any two polynomials of total degree less than or equal to m/sub 0/ + n/sub 0/. The block pipeline period is one, which means that it can start a new GCD computation immediately in the next cycle. Unlike the array of Brent and Kung, a pre-processing step for extracting a common factor X/sup i/ is not necessary and the size of the processing element (PE) does not depend on m/sub 0/ and n/sub 0/. The authors extend this new array structure to the extended polynomial GCD algorithm, which is closely related to the decoding of BCH and Reed-Solomon codes. To verify the structure, they have used the VERILOG simulator, and implemented a 2 /spl mu/ CMOS test chip.> Yongjin Jeong, Wayne P. Burleson |
ASAP | 2 |
| 1993 | Node merging: A transformation on bit-level dependence graphs for efficient VLSI array designabstractThe authors present a transformation technique, called node merging, on bit-level dependence graphs to systematically explore tradeoffs between area and various system performances, such as clock period, pipelining period, block pipelining period, computation time, and dynamic power dissipation to obtain optimal VLSI array processors for bit-level regular algorithms. By merging several DG nodes into one node, multi-bit level array processors can be designed using formal regular array synthesis methods thereby significantly reducing the number of pipelining registers required compared to bit-pipelined array processors. In general, delay paths within a node in bit-level dependence graphs are unbalanced, and the clock period is determined by the critical path delay. By merging nodes along noncritical paths, the authors improve computation time as well as VLSI area with a relatively small increase in clock period and pipelining period. They also expand the complexity of node functions enough to apply a meaningful logic optimization or performance enhancement using well-known logic synthesis tools such as SIS for even further improvement. Since the transformation results in a new DG, it can be easily combined with conventional VLSI array synthesis techniques for efficient bit-level array processor design. Therefore, the method provides an efficient way to explore a significantly broader design space in VLSI array processor design.> Bongjin Jung, Wayne P. Burleson |
ASAP | 2 |
| 1993 | Formal descriptions, semantics and verification of VLSI array processorsabstractThe authors present two description languages for specifying algorithms and describing implementations in ARREST - a diagram environment for VLSI array architectures. They then define their operational semantics based on feedback products of sequential machines, which provide a precise and intuitive meaning of concurrency and composition. Next the authors develop a formal specification language based on modal functions for describing the behavior of abstract product machines, which formally interprets both of the languages defined before and provides the basis for formal verification in ARREST. Finally, they classify the correctness and equivalence properties to be verified, define eight important types of equivalence, and discuss their relations with the stuttering-equivalence in the semantics.> Wayne P. Burleson |
ASAP | 2 |
| 1993 | The Spring Scheduling Co-Processor: A Scheduling AcceleratorabstractWe present a novel co-processor for multiprocessor scheduling in the Spring real-time operating system. Since most dynamic scheduling problems are NP-complete, we use a heuristic algorithm which uses a smart searching scheme to find a feasible schedule for a set of specified tasks and hard deadlines. A parallel VLSI architecture for scheduling is developed that can be scaled for different numbers of tasks, numbers of resources, internal wordlengths, and future IC technologies. The scheduling architecture is implemented in a 0.8/spl mu/ CMOS technology and uses an advanced clocking scheme to allow further scaling to future technologies. With an internal clock rate of 100 MHz, a speed increase of two orders of magnitude is expected for scheduling tasks, thus removing a major bottleneck in real-time systems.> Wayne P. Burleson, Jason Ko, Douglas Niehaus, Krithi Ramamritham, John A. Stankovic, Gary Wallace, Charles C. Weems |
ICCD | 1 |
| 1993 | Rank-order Filtering Algorithms: A Comparison of VLSI Implementations
J. David Narkiewicz, Wayne P. Burleson |
ISCAS | 2 |
| 1993 | The Spring Scheduling Co-Processor: Design, Use, and PerformanceabstractWe present a novel VLSI co-processor for real-time multiprocessor scheduling. The co-processor can be used for sophisticated static scheduling as well as for online scheduling using many different algorithms such as earliest deadline first, highest value first, or the Spring scheduling algorithm. When such an algorithm is used online it is important to assess the performance impact of the interface of the co-processor to the host system, in this case, the Spring kernel. We focus on the interface and its implications for overall scheduling performance. We show that the current VLSI chip speeds up the main portion of the scheduling operation by over three orders of magnitude and speeds up the overall scheduling operation 30 fold. The parallel VLSI architecture for scheduling is briefly presented. This architecture can be scaled for different numbers of tasks, resources, and internal word lengths. The implementation uses an advanced clocking scheme to allow further scaling using future IC technologies.> Douglas Niehaus, Krithi Ramamritham, John A. Stankovic, Gary Wallace, Charles C. Weems, Wayne P. Burleson, Jason Ko |
RTSS | 6 |
| 1992 | ARREST: an interactive graphic analysis tool for VLSI arraysabstractThe authors present a graphical CAD tool, Array Estimator (ARREST), for VLSI array architectures. In real VLSI arrays, piece-wise regular computations are spread across space and time and occur at a fine-grain, which can make visualization quite difficult. Consequently, a graphical interface environment is desirable to enhance the design, verification, and analysis of VLSI arrays by providing feedback at all levels of the design process. ARREST reads a high level description of structured VLSI algorithms in terms of affine recurrence equations (AREs) and permits a broad range of transformations on the algorithm. The system does not target a fully automated design process, instead it provides a designer with a means to systematically explore various array architectures and evaluate design trade-offs between VLSI cost and performance. To allow a human designer better insight into the design process, ARREST uses the Xt/MOTIF window system for graphics and interfaces to the Cadence VERILOG simulator.> Wayne P. Burleson, Bongjin Jung |
ASAP | 1 |
| 1991 | The partitioning problem on VLSI arrays: I/O and local memory complexityabstractThe space/time costs of implementing the nonlocal communication of partitioned algorithms are studied. In addition, the author looks at the often-neglected issue of I/O complexity for partitioned algorithms. He illustrates the design method and tradeoffs with an example in DSP, the fixed-point multiply-accumulator (MAC).> Wayne P. Burleson |
ICASSP | 1 |
| 1991 | A Simulator for General Purpose Optical ArraysabstractArchitectures based on optical arrays offer the promise of massive parallelism and three-dimensional computing. The design of an integrated general purpose optical or electrooptical machine is a formidable task. The simulation of these technologies using electronic computers allows a number of designs for such machines to be explored. The general purpose simulator for electrooptical arrays presented demonstrates the practicality and presents the limitations of such technologies. The behavior of bi-level optically active gate and lens materials is simulated as solutions to well-known equations in a low level optical simulation. A second higher level optical simulation is used to simulate optical arrays at the gate level using modified ray tracing techniques. A third level provides a simulation of the electrooptical integration of an example array processor, and allows simulation of machine code for a minimal classical machine to be mapped onto an example array architecture.> Walter B. Marvin, Wayne P. Burleson |
ICCD | 2 |