Russell Tessier

dblp:23/2353 · DBLP profile ↗
← Back
114ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-0591-7566ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 95 · 6 first-author · 14 since 2021Software engineering, systems software and programming languages · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Security and privacy · 4 · 1 since 2021Computer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 Sky to Edge-Cloud: Heterogeneous Computation Offloading for Energy-Efficient Drone Computing
abstract
Offloading computation from edge devices to an edge cloud can save energy and enhance performance, but managing workloads from heterogeneous devices such as CPUs, GPUs, and NPUs remains challenging. This paper presents a prototype heterogeneous offloading system using an FPGA-based edge cloud capable of handling diverse computation models. Using drone computing as a use-case, we demonstrate that applications such as depth estimation, object detection, and Simultaneous Localization and Mapping (SLAM) can efficiently offload CPU, GPU, and DSP tasks to a reconfigurable accelerator. By intelligently mapping heterogeneous kernels to the reconfigurable FPGA edge cloud, our system reduces drone energy by up to 90%, underscoring its heterogeneous kernel offloading capability and real-time performance and energy gains.
Zhehang Zhang, Bharadwaj Madabhushi, Sandip Kundu, Russell Tessier
ACM Great Lakes Symposium on VLSI4
2025 Managing Computation Offloading from Edge Devices to a Reconfigurable Edge Cloud
abstract
Edge cloud computing is becoming increasingly vital to meet the computational demands of billions of interconnected devices, many of which operate under strict power and latency constraints. Modern 5G networks and advanced low-latency communication technologies enable rapid data transfer between edge devices and the cloud, making computational offloading both feasible and efficient. While traditional edge platforms have primarily relied on multi-core microprocessors, the growing architectural diversity of edge devices necessitates new approaches that support heterogeneous edge computing. This paper presents a novel edge computing framework that harnesses the capabilities of FPGAs to address the complexities of heterogeneous edge environments. A dynamically reconfigurable edge cloud enables seamless support for architecturally diverse edge devices. Central to our solution is an intelligent Offload Management System (OMS) that makes real-time decisions about whether to offload tasks or execute them locally, based on resource availability, energy efficiency, and deadline constraints. We validate our approach using an experimental setup featuring quad-core ARM Cortex-A76 and Qualcomm Snapdragon processors found in edge devices paired with an AMD ZCU104 FPGA board as an edge cloud node. Our results demonstrate how multiple edge devices can collaboratively utilize shared cloud resources via intelligent offloading. We specifically assess machine learning workloads processed with deep learning processing unit (DPU)-based models, showcasing the significant potential of a reconfigurable edge cloud in servicing edge devices.
Zhehang Zhang, Bharadwaj Madabhushi, Sandip Kundu, Russell Tessier
ASAP4
2025 Diwall: A Lightweight Host Intrusion Detection System Against Jamming and Packet Injection Attacks
abstract
The rapid growth of Internet of Things (IoT) applications in various sectors has led to a significant increase in the number of IoT devices. This has led to the deployment of numerous IoT protocols to provide greater connectivity. However, this extensive adoption has also left them vulnerable to attack. In particular, attacks targeting wireless communication capabilities represent a significant threat. Such attacks exploit various vulnerabilities in the wireless connectivity unit, compromising its security. To counter this threat, this article proposes a Host Intrusion Detection System (HIDS) for detecting wireless attacks. Its components are customized to support IoT end-devices using low-GHz and sub-GHz data rate protocols. The HIDS deploys a hardware tracer to monitor microarchitecture and network metrics using hardware performance counters (HPCs). It performs monitoring of network and microarchitecture metrics for a 32-bit RISC-V-based wireless connectivity unit. The HIDS uses analysis and classification of monitored data for detecting memory corruption and jamming attacks. We evaluate the effectiveness of the HIDS in detecting packet injection and jamming attacks. Our Field?Programmable Gate Array (FPGA) implementation of HIDS has a logic overhead of about 14.30% and 22.89% of flip flops (FFs) and lookup tables (LUTs), respectively, compared to the CV32E40P baseline on an Arty A7 100T board. The design frequency and code size penalties are less than 1% for a RISC-V processor with a LoRaWAN protocol stack.
Mohamed El Bouazzati, Philippe A. Tanguy, Guy Gogniat, Russell Tessier
ACM Trans. Embed. Comput. Syst.4
2024 Reliability and Security of AI Hardware
abstract
In recent years, Artificial Intelligence (AI) systems have achieved revolutionary capabilities, providing intelligent solutions that surpass human skills in many cases. However, such capabilities come with power-hungry computation workloads. Therefore, the implementation of hardware acceleration becomes as fundamental as the software design to improve energy efficiency, silicon area, and latency of AI systems. Thus, innovative hardware platforms, architectures, and compiler-level approaches have been used to accelerate AI workloads. Crucially, innovative AI acceleration platforms are being adopted in application domains for which dependability must be paramount, such as autonomous driving, healthcare, banking, space exploration, and industry 4.0. Unfortunately, the complexity of both AI software and hardware makes the dependability evaluation and improvement extremely challenging. Studies have been conducted on both the security and reliability of AI systems, such as vulnerability assessments and countermeasures to random faults and analysis for side-channel attacks. This paper describes and discusses various reliability and security threats in AI systems, and presents representative case studies along with corresponding efficient countermeasures.
Dennis Gnad, Martin Gotthard, Jonas Krautter, Angeliki Kritikakou, Vincent Meyers, Paolo Rech, Josie E. Rodriguez Condia, Annachiara Ruospo, Ernesto Sánchez 0001, Fernando Santos 0001, Olivier Sentieys, Mehdi Baradaran Tahoori, Russell Tessier, Marcello Traiola
ETS13
2024 SNOWWI: A Three-Frequency InSAR for Snow Science Applications
abstract
In this paper we describe the development and motivation behind the development of a NASA-sponsored airborne instrument, SNOWWI (Snow Water-equivalent Wide Swath Interferometer) that is being developed for exploring the volume scattering and penetration depth characteristics of the snowpack at three different frequencies (5.4 GHz, C-band; 13.64 GHz known as Ku-low; 17.24 GHz known as Ku-high). The system, as it is being constructed is able to receive co- and cross-polarized (VV and VH) returns in an interferometric configuration. By implementing these components of the radar signature on the same platform, we will be able to explore the relationship between snow depth, density and snow water equivalent on the overall radar signature. This work is being done in conjunction with a strong modeling component being led by the University of Michigan, a ground campaign component supported by Boise State University and the US Army Corps of Engineers Cold Regions Research and Engineering Laboratory (CRREL), and a spaceborne concept development being led by Capella Space.
Paul Siqueira, Marc Closa Tarrés, Max Adam, Eric Sutherland, Joseph Maloyan, Takuya Seaver, Russell Tessier, Leung Tsang, Firoz Kanti Borah, H. P. Marshall, Elias Deeb, Gordon Farquharson
IGARSS7
2024 First Results From a Dual Ku- and C-Band Airborne SAR for Snowpack Measurements
abstract
This article presents the first results of the newly conceived airborne Synthetic Aperture Radar system, SNOWWI.SNOWWI is a dual Ku- and C-Band interferometric and dualpolarized (VV and VH) system operating at 13.64 GHz, 17.24 GHz, and 5.39 GHz. The system aims to deliver snowpack observations to quantify Snow Depth (SD) and Snow Water Equivalent (SWE), which have been included as Targeted Observables in the National Academies’ 2017 Decadal Strategy for Earth Observation from Space. This manuscript includes results from the system’s first deployment in Grand Mesa, CO, in January and March 2024.
Marc Closa Tarrés, Paul Siqueira, Max Adam, Eric Sutherland, Joseph Maloyan, Takuya Seaver, Russell Tessier, Leung Tsang, Firoh Borah, HP Marshall, Elias Deeb, Gordon Farquharson
IGARSS7
2024 On the Malicious Potential of Xilinx's Internal Configuration Access Port (ICAP)
abstract
Field Programmable Gate Arrays (FPGAs) have become increasingly popular in computing platforms. With recent advances in bitstream format reverse engineering, the scientific community has widely explored static FPGA security threats. For example, it is now possible to convert a bitstream to a netlist, revealing design information, and apply modifications to the static bitstream based on this knowledge. However, a systematic study of the influence of the bitstream format understanding in regards to the security aspects of the dynamic configuration process, particularly for Xilinx’s Internal Configuration Access Port (ICAP), is lacking. This article fills this gap by comprehensively analyzing the security implications of ICAP interfaces, which primarily support dynamic partial reconfiguration. We delve into the Xilinx bitstream file format, identify misconceptions in official documentation, and propose novel configuration (attack) primitives based on dynamic reconfiguration, i.e., create/read/update/delete circuits in the FPGA, without requiring pre-definition during the design phase. Our primitives are consolidated in a novel Stealthy Reconfigurable Adaptive Trojan framework to conceal Trojans and evade state-of-the-art netlist reverse engineering methods. As FPGAs become integral to modern cloud computing, this research presents crucial insights on potential security risks, including the possibility of a malicious tenant or provider altering or spying on another tenant’s configuration undetected.
Nils Albartus, Maik Ender, Jan-Niklas Möller, Marc Fyrbiak, Christof Paar, Russell Tessier
ACM Trans. Reconfigurable Technol. Syst.6
2023 A Practical Remote Power Attack on Machine Learning Accelerators in Cloud FPGAs
abstract
The security and performance of FPGA-based accelerators play vital roles in today's cloud services. In addition to supporting convenient access to high-end FPGAs, cloud vendors and third-party developers now provide numerous FPGA accelerators for machine learning models. However, the security of accelerators developed for state-of-the-art Cloud FPGA environments has not been fully explored, since most remote accelerator attacks have been prototyped on local FPGA boards in lab settings, rather than in Cloud FPGA environments. To address existing research gaps, this work analyzes three existing machine learning accelerators developed in Xilinx Vitis to assess the potential threats of power attacks on accelerators in Amazon Web Services (AWS) F1 Cloud FPGA platforms, in a multi-tenant setting. The experiments show that malicious co-tenants in a multi-tenant environment can instantiate voltage sensing circuits as register-transfer level (RTL) kernels within the Vitis design environment to spy on co-tenant modules. A methodology for launching a practical remote power attack on Cloud FPGAs is also presented, which uses an enhanced time-to-digital (TDC) based voltage sensor and auto-triggered mechanism. The TDC is used to capture power signatures, which are then used to identify power consumption spikes and observe activity patterns involving the FPGA shell, DRAM on the FPGA board, or the other co-tenant victim's accelerators. Voltage change patterns related to shell use and accelerators are then used to create an auto-triggered attack that can automatically detect when to capture voltage traces without the need for a hard-wired synchronization signal between victim and attacker. To address the novel threats presented in this work, this paper also discusses defenses that could be leveraged to secure multi-tenant Cloud FPGAs from power-based attacks.
Shanquan Tian, Shayan Moini, Daniel E. Holcomb, Russell Tessier, Jakub Szefer
DATE4
2023 A Lightweight Intrusion Detection System against IoT Memory Corruption Attacks
abstract
Attacks against internet-of-things (IoT) end-devices represent a significant threat since their wireless communication capabilities provide a potential attack entry point. To address this threat, we demonstrate the use of hardware performance counters (HPCs) in a host-based intrusion detection system (HIDS). The counter-based monitors are customized to support IoT end-devices which use low data rate GHz and sub-GHz protocols. Our solution implements a hardware unit that performs data tracing for a 32-bit RISC-V based wireless connectivity unit. The unit can detect ongoing remote attacks in real time. We demonstrate the effectiveness of our system by detecting a packet injection exploit. Our FPGA implementation of HIDS has a logic overhead of about 6% and design frequency penalty of less than 1% for a RISC-V processor.
Mohamed El Bouazzati, Russell Tessier, Philippe A. Tanguy, Guy Gogniat
DDECS2
2023 Fault Recovery from Multi-Tenant FPGA Voltage Attacks
abstract
As multi-tenant FPGA applications continue to scale in size and complexity, their need for resilience against environmental effects and malicious actions continues to grow. To ensure continuously correct computation, faults in the compute fabric must be identified, isolated, and suppressed in the nanosecond to microsecond range. In this paper, we detail a circuit and system-level methodology to detect compute failure conditions due to on-FPGA voltage attacks. Our approach rapidly suppresses incorrect results and regenerates potentially-tainted results before they propagate, allowing time for an attacker to be suppressed. Instrumentation includes voltage sensors to detect error conditions induced by attackers. This analysis is paired with focused remediation approaches involving data buffering, fault suppression, results recalculation, and computation restart. Our approach has been demonstrated using an RSA encryption circuit implemented on a Stratix 10 FPGA. We show that a voltage attack using on-FPGA power wasters can be effectively detected and computation halted in 15 ns, preventing the injection of timing faults. Potentially tainted results are successfully regenerated, allowing for fault-free circuit operation. A full characterization of the latency and resource overheads of fault detection and recovery is provided.
Shayan Moini, Dhruv Kansagara, Daniel E. Holcomb, Russell Tessier
ACM Great Lakes Symposium on VLSI4
2023 A Visionary Look at the Security of Reconfigurable Cloud Computing
abstract
Field-programmable gate arrays (FPGAs) have become critical components in many cloud computing platforms. These devices possess the fine-grained parallelism and specialization needed to accelerate applications ranging from machine learning to networking and signal processing, among many others. Unfortunately, fine-grained programmability also makes FPGAs a security risk. Here, we review the current scope of attacks on cloud FPGAs and their remediation. Many of the FPGA security limitations are enabled by the shared power distribution network in FPGA devices. The simultaneous sharing of FPGAs is a particular concern. Other attacks on the memory, host microprocessor, and input/output channels are also possible. After examining current attacks, we describe trends in cloud architecture and how they are likely to impact possible future attacks. FPGA integration into cloud hypervisors and system software will provide extensive computing opportunities but invite new avenues of attack. We identify a series of system, software, and FPGA architectural changes that will facilitate improved security for cloud FPGAs and the overall systems in which they are located.
Mirjana Stojilovic, Kasper Bonne Rasmussen, Francesco Regazzoni 0001, Mehdi Baradaran Tahoori, Russell Tessier
Proc. IEEE5
2023 Jitter-based Adaptive True Random Number Generation Circuits for FPGAs in the Cloud
abstract
In this article, we present and evaluate a true random number generator (TRNG) design that is compatible with the restrictions imposed by cloud-based Field Programmable Gate Array (FPGA) providers such as Amazon Web Services (AWS) EC2 F1. Because cloud FPGA providers disallow the ring oscillator circuits that conventionally generate TRNG entropy, our design is oscillator-free and uses clock jitter as its entropy source. The clock jitter is harvested with a time-to-digital converter (TDC) and a controllable delay line that is continuously tuned to compensate for process, voltage, and temperature variations. After describing the design, we present and validate a stochastic model that conservatively quantifies its worst-case entropy. We deploy and model the design in the cloud on 60 EC2 F1 FPGA instances to ensure sufficient randomness is captured. TRNG entropy is further validated using NIST test suites, and experiments are performed to understand how the TRNG responds to on-die power attacks that disturb the FPGA supply voltage in the vicinity of the TRNG. After introducing and validating our basic TRNG design, we introduce and validate a new variant that uses four instances of a linkable sampling module to increase the entropy per sample and improve throughput. The new variant improves throughput by 250% at a modest 17% increase in CLB count.
Xiang Li 0158, Peter Stanwicks, George Provelengios, Russell Tessier, Daniel E. Holcomb
ACM Trans. Reconfigurable Technol. Syst.4
2023 Voltage Sensor Implementations for Remote Power Attacks on FPGAs
abstract
This article presents a study of two types of on-chip FPGA voltage sensors based on ring oscillators (ROs) and time-to-digital converter (TDCs), respectively. It has previously been shown that these sensors are often used to extract side-channel information from FPGAs without physical access. The performance of the sensors is evaluated in the presence of circuits that deliberately waste power, resulting in localized voltage drops. The effects of FPGA power supply features and sensor sensitivity in detecting voltage drops in an FPGA power distribution network (PDN) are evaluated for Xilinx Artix-7, Zynq 7000, and Zynq UltraScale+ FPGAs. We show that both sensor types are able to detect supply voltage drops, and that their measurements are consistent with each other. Our findings show that TDC-based sensors are more sensitive and can detect voltage drops that are shorter in duration, while RO sensors are easier to implement because calibration is not required. Furthermore, we present a new time-interleaved TDC design that sweeps the sensor phase. The new sensor generates data that can reconstruct voltage transients on the order of tens of picoseconds.
Shayan Moini, Aleksa Deric, Xiang Li 0158, George Provelengios, Wayne P. Burleson, Russell Tessier, Daniel E. Holcomb
ACM Trans. Reconfigurable Technol. Syst.6
2022 Precise Fault Injection to Enable DFIA for Attacking AES in Remote FPGAs
abstract
Differential Fault Intensity Analysis (DFIA) is a class of biased-fault attacks that aim to recover secret keys from block ciphers such as Advanced Encryption Standard (AES). In DFIA an attacker collects a set of ciphertexts generated while carefully controlling the fault intensity, and then performs an analysis on the results that reveals the secret encryption key. In AES, DFIA requires injecting varied intensity faults during exactly the 9th round of encryption, which could be accomplished using clock or supply voltage glitching, although previous works give scant consideration to shaping the fault within a realistic scenario.In this work, we demonstrate DFIA against an FPGA implementation of AES without assuming arbitrary external control of clock or supply voltage. Instead we use on-chip ring oscillators (ROs) to create a precise and controllable voltage drop in the vicinity of the AES circuit, which causes timing faults to occur. The fault intensity is finely controlled by changing the number of activated ROs, and we explore how to optimize the timing of the RO activation to cause a fault in the 9th round as is required in DFIA. We use this approach to perform DFIA against AES on Xilinx Spartan-7 FPGA, show that it successfully extracts AES key bytes, and discuss its performance.
Xiang Li 0158, Russell Tessier, Daniel E. Holcomb
FCCM2
2022 The Future of FPGA Acceleration in Datacenters and the Cloud
abstract
In this article, we survey existing academic and commercial efforts to provide Field-Programmable Gate Array (FPGA) acceleration in datacenters and the cloud. The goal is a critical review of existing systems and a discussion of their evolution from single workstations with PCI-attached FPGAs in the early days of reconfigurable computing to the integration of FPGA farms in large-scale computing infrastructures. From the lessons learned, we discuss the future of FPGAs in datacenters and the cloud and assess the challenges likely to be encountered along the way. The article explores current architectures and discusses scalability and abstractions supported by operating systems, middleware, and virtualization. Hardware and software security becomes critical when infrastructure is shared among tenants with disparate backgrounds. We review the vulnerabilities of current systems and possible attack scenarios and discuss mitigation strategies, some of which impact FPGA architecture and technology. The viability of these architectures for popular applications is reviewed, with a particular focus on deep learning and scientific computing. This work draws from workshop discussions, panel sessions including the participation of experts in the reconfigurable computing field, and private discussions among these experts. These interactions have harmonized the terminology, taxonomy, and the important topics covered in this manuscript.
Christophe Bobda, Joel Mandebi, Paul Chow, Mohammad Ewais, Naif Tarafdar, Juan Camilo Vega, Kenneth Eguro, Dirk Koch, Suranga Handagala, Miriam Leeser, Martin C. Herbordt, Hafsah Shahzad, H. Peter Hofstee, Burkhard Ringlein, Jakub Szefer, Ahmed Sanaullah, Russell Tessier
ACM Trans. Reconfigurable Technol. Syst.17
2021 Remote Power Side-Channel Attacks on BNN Accelerators in FPGAs
abstract
Multi-tenant FPGAs have recently been proposed, where multiple independent users simultaneously share a remote FPGA. Despite its benefits for cost and utilization, multi-tenancy opens up the possibility of malicious users extracting sensitive information from co-located victim users. To demonstrate the dangers, this paper presents a remote, power-based side-channel attack on a binarized neural network (BNN) accelerator. This work shows how to remotely obtain voltage estimates as the BNN circuit executes, and how the information can be used to recover the inputs to the BNN. The attack is demonstrated with a BNN used to recognize handwriting images from the MNIST dataset. With the use of precise time-to-digital converters (TDCs) for remote voltage estimation, the MNIST inputs can be successfully recovered with a maximum normalized cross-correlation of 75% between the input image and the recovered image.
Shayan Moini, Shanquan Tian, Daniel E. Holcomb, Jakub Szefer, Russell Tessier
DATE5
2021 Remote Power Attacks on the Versatile Tensor Accelerator in Multi-Tenant FPGAs
abstract
Architectural details of machine learning models are crucial pieces of intellectual property in many applications. Revealing the structure or types of layers in a model can result in a leak of confidential or proprietary information. This issue becomes especially concerning when the machine learning models are executed on accelerators in multi-tenant FPGAs where attackers can easily co-locate sensing circuitry next to the victim's machine learning accelerator. To evaluate such threats, we present the first remote power attack that can extract details of machine learning models executed on an off-the-shelf domain-specific instruction set architecture (ISA) based neural network accelerator implemented in an FPGA. By leveraging a time-to-digital converter (TDC), an attacker can deduce the composition of instruction groups executing on the victim accelerator, and recover parameters of General Matrix Multiplication (GEMM) instructions within a group, all without requiring physical access to the FPGA. With this information, an attacker can then reverse-engineer the structure and layers of machine learning models executing on the accelerator, leading to potential theft of proprietary information.
Shanquan Tian, Shayan Moini, Adam Wolnikowski, Daniel E. Holcomb, Russell Tessier, Jakub Szefer
FCCM5
2021 On the Design and Misuse of Microcoded (Embedded) Processors - A Cautionary Note
Nils Albartus, Clemens Nasenberg, Florian Stolz, Marc Fyrbiak, Christof Paar, Russell Tessier
USENIX Security Symposium6
2020 Power Wasting Circuits for Cloud FPGA Attacks
abstract
Recent research has exposed a number of security issues related to the use of FPGAs in cloud computing environments. Circuits that deliberately waste power can be carefully crafted by a malicious cloud FPGA user and deployed to cause denial-of-service and fault injection attacks. The main defense strategy used by FPGA cloud services involves checking user-submitted designs for circuit structures that are known to aggressively consume power. In this work, we evaluate a variety of circuit power wasting techniques that typically are not flagged by design rule checks imposed by FPGA cloud computing vendors. We demonstrate that a multi-stage circuit based on standard logic operations can be exploited to induce delay faults in co-located circuits. The efficiency of five power wasting circuits, including our new design, is evaluated in terms of power consumed per logic resource.
George Provelengios, Daniel E. Holcomb, Russell Tessier
FPL3
2020 NestedNet: A Container-based Prototyping Tool for Hierarchical Software Defined Networks
abstract
Emulators for software-defined networks (SDNs) are important prototyping tools in validating network hardware performance under a broad range of topologies and parameters. Modern SDNs typically contain hierarchical collections of network nodes, each with interconnected compute devices. These devices often have widely varying compute environments making accurate emulation using discrete processes in a virtual machine (VM) difficult. In this paper, we describe NestedNet, a new container-based prototyping environment for hierarchical SDN systems. Each network node is represented as a Docker container. The node internals inside the container are implemented as nested Docker containers interconnected via an Open vSwitch. Unlike previous emulators, the execution of each heterogeneous component in a network node can be accurately performed using native code within the target execution environment. To demonstrate the flexibility of our rapid prototyping system, we emulate a mobile ad hoc network (MANET) topology of twelve interconnected nodes of five components each and evaluate its performance using throughput and latency metrics. Emulated throughput values of up to 32 Gbps per link are achieved.
Xuzhi Zhang, Narendra Prabhu, Russell Tessier
RSP3
2020 CoNFV: A Heterogeneous Platform for Scalable Network Function Virtualization
abstract
Network function virtualization (NFV) is a powerful networking approach that leverages computing resources to perform a time-varying set of network processing functions. Although microprocessors can be used for this purpose, their performance limitations and lack of specialization present implementation challenges. In this article, we describe a new heterogeneous hardware-software NFV platform called CoNFV that provides scalability and programmability while supporting significant hardware-level parallelism and reconfiguration. Our computing platform takes advantage of both field-programmable gate arrays (FPGAs) and microprocessors to implement numerous virtual network functions (VNF) that can be dynamically customized to specific network flow needs. The most distinctive feature of our system is the use of global network state to coordinate NFV operations. Traffic management and hardware reconfiguration functions are performed by a global coordinator that allows for the rapid sharing of network function states and continuous evaluation of network function needs. With the help of state sharing mechanism offered by the coordinator, customer-defined VNF instances can be easily migrated between heterogeneous middleboxes as the network environment changes. A resource allocation and scheduling algorithm dynamically assesses resource deployments as network flows and conditions are updated. We show that our deployment algorithm can successfully reallocate FPGA and microprocessor resources in a fraction of a second in response to changes in network flow capacity and network security threats including intrusion.
Xuzhi Zhang, Xiaozhe Shao, George Provelengios, Naveen Kumar Dumpala, Lixin Gao 0001, Russell Tessier
ACM Trans. Reconfigurable Technol. Syst.6
2020 Power Distribution Attacks in Multitenant FPGAs
abstract
The increased use of field-programmable gate arrays (FPGAs) in the cloud and embedded computing environments has led to a number of potential security risks. The sizable amount of logic resources in these devices makes them amenable to sharing across multiple untrusted tenants. However, the co-location of multiple independent circuits presents the possibility of malicious fault injection into an unsuspecting circuit. In this article, the ability of one tenant's FPGA circuit to inject delay faults into another tenant's application located at points across the FPGA die via deliberate supply voltage modulation is investigated. To illustrate the risks involved, a Rivest-Shamir-Adleman (RSA) encryption key extraction attack is performed by introducing delay faults in hardware via voltage manipulations. This attack does not require modification to the encryption core nor require attack activation synchronized with specific encryption operations. Our work characterizes the magnitude of on-chip voltage changes and fault injections over time in relation to the on-chip location of the malicious circuit once an attack is initiated. Strategies to identify power manipulation using low-cost monitoring circuits that can locate the source of an attack are highlighted.
George Provelengios, Daniel E. Holcomb, Russell Tessier
IEEE Trans. Very Large Scale Integr. Syst.3
2019 Characterization of Long Wire Data Leakage in Deep Submicron FPGAs
abstract
The simultaneous use of FPGAs by multiple tenants has recently been shown to potentially expose sensitive information without the victim's knowledge. For example, neighboring long wires in SRAM-based FPGAs have been shown to allow for clandestine data exfiltration. In this work, we explore distinct characteristics of this signal crosstalk that could be used to enhance or prevent information leakage. First, we develop a mechanism to characterize the crosstalk coupling that exists between neighboring wires at the femtosecond scale. Second, we show that it is possible to reverse engineer channel layouts by determining which pairs of routing resources/links in the channel exhibit coupling to each other even if this information is not provided by the FPGA vendor. To fully characterize these effects, we examine long wire coupling on different types of wires across three devices implemented in different technology nodes from 65 to 20 nm. We experimentally demonstrate that information leakage is apparent for all three FPGA families.
George Provelengios, Chethan Ramesh, Shivukumar B. Patil, Kenneth Eguro, Russell Tessier, Daniel E. Holcomb
FPGA5
2019 Characterizing Power Distribution Attacks in Multi-User FPGA Environments
abstract
Multi-tenant FPGAs that contain circuits from multiple users are emerging as a new usage model in cloud and embedded computing environments. Interactions between untrusting tenant applications in an FPGA can enable new security exposures and the risk of side channel attacks or fault injection. In this work, we investigate the ability for aggressive power consumption of one application to disturb the power network to an extent that causes delay faults in a second application on the same FPGA. In particular, we identify the mechanisms by which the supply voltage is disturbed by the attack, and we characterize the magnitude of the disturbance as a function of time, power consumed by attacker, and position of the victim relative to the attacker. We highlight strategies that can be used to mitigate attacks, including low-cost monitoring circuits that can identify the source of an attack so that the attacker's use of the FPGA can be revoked.
George Provelengios, Daniel E. Holcomb, Russell Tessier
FPL3
2019 Closed-Loop Proportion-Derivative Control of Suppressing Seizures in a Neural Mass Model
abstract
In this work, we present an analytical approach of closed-loop Proportional-Derivative (PD) control to determine the stimulation parameters for suppressing high-amplitude epileptic activity in a neural mass model. Closed-loop PD control to suppress epileptic activity in the Jansen's neural mass model (Jansen's NMM) has been studied. This work shows that the output signal of the Jansen's NMM model without the PD control feedback is high amplitude epileptic seizure activity which turns into low amplitude activity with the intervention feedback of a PD controller. A graphical stability analysis method was employed to determine the stability region of the PD controller in the gain parameter space. Therefore, this approach draws a region of PD controller parameters that is empirically chosen to stabilize epileptic seizure activities in the chosen NMM. Furthermore, this approach allows us to explore the relationship between the model parameters of inducing epileptic activity and the feedback controller parameters to foster a better understanding of the mechanism to suppress epileptic seizure activity by applying closed-loop stimulation (pharmacology stimulation, electrical stimulation or optogenetic stimulation etc.).
Lijuan Xia, Ahmed Soltan, Xuzhi Zhang, Andrew Jackson 0001, Russell Tessier, Patrick Degenaar
ISCAS5
2019 HAL - The Missing Piece of the Puzzle for Hardware Reverse Engineering, Trojan Detection and Insertion
abstract
Hardware manipulations pose a serious threat to numerous systems, ranging from a myriad of smart-X devices to military systems. In many attack scenarios an adversary merely has access to the low-level, potentially obfuscated gate-level netlist. In general, the attacker possesses minimal information and faces the costly and time-consuming task of reverse engineering the design to identify security-critical circuitry, followed by the insertion of a meaningful hardware Trojan. These challenges have been considered only in passing by the research community. The contribution of this work is threefold: First, we present HAL, a comprehensive reverse engineering and manipulation framework for gate-level netlists. HAL allows automating defensive design analysis (e.g., including arbitrary Trojan detection algorithms with minimal effort) as well as offensive reverse engineering and targeted logic insertion. Second, we present a novel static analysis Trojan detection technique ANGEL which considerably reduces the false-positive detection rate of the detection technique FANCI. Furthermore, we demonstrate that ANGEL is capable of automatically detecting Trojans obfuscated with DeTrust. Third, we demonstrate how a malicious party can semi-automatically inject hardware Trojans into third-party designs. We present reverse engineering algorithms to disarm and trick cryptographic self-tests, and subtly leak cryptographic keys without any a priori knowledge of the design's internal workings.
Marc Fyrbiak, Sebastian Wallat, Pawel Swierczynski, Max Hoffmann 0001, Sebastian Hoppach, Matthias Wilhelm 0002, Tobias Weidlich, Russell Tessier, Christof Paar
IEEE Trans. Dependable Secur. Comput.8
2019 Introduction to the Special Section on Security in FPGA-accelerated Cloud and Datacenters
abstract
editorial Free Access Share on Introduction to the Special Section on Security in FPGA-accelerated Cloud and Datacenters Editors: Chistophe Bobda University of Florida Russell Tessier, University of Massachusetts Amherst University of Florida Russell Tessier, University of Massachusetts AmherstView Profile , Ken Eguro Microsoft Research Ryan Kastner, University of California, San Diego Microsoft Research Ryan Kastner, University of California, San DiegoView Profile Authors Info & Claims ACM Transactions on Reconfigurable Technology and SystemsVolume 12Issue 3September 2019 Article No.: 11epp 1–3https://doi.org/10.1145/3352060Published:13 September 2019Publication History 0citation190DownloadsMetricsTotal Citations0Total Downloads190Last 12 Months44Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteView all FormatsPDF
Christophe Bobda, Russell Tessier, Kenneth Eguro, Ryan Kastner
ACM Trans. Reconfigurable Technol. Syst.2
2019 Loop Unrolling for Energy Efficiency in Low-Cost Field-Programmable Gate Arrays
abstract
Field-programmable gate arrays (FPGAs) are used for a wide variety of computations in low-cost embedded systems. Although these systems often have modest performance constraints, their energy consumption must typically be limited. Many FPGA applications employ repetitive loops that cannot be straightforwardly split into parallel computations. Performing a loop sequentially generally requires high-speed clocks that consume considerable clock power and sometimes require clock generation using a phase-locked loop (PLL). Loop unrolling addresses the high-speed clock issue, but its use often leads to significant combinational glitch power. In this work, a computer-aided design (CAD) approach that unrolls loops for designs targeted to low-cost FPGAs is described. Our approach considers latency constraints in an effort to minimize energy consumption for loop-based computation. To reduce glitch power, a glitch-filtering approach is introduced that provides a balance between glitch reduction and design performance. Glitch-filter enable signals are generated and routed to the filters using resources best suited to the target FPGA. Our approach automatically inserts glitch filters and associated control logic into a design prior to processing with FPGA synthesis, place, and route tools. Our energy-saving loop-unrolling approach has been evaluated using five benchmarks often used in low-cost FPGAs. The energy-saving capabilities of the approach have been evaluated for an Intel Cyclone IV and a Xilinx Artix-7 FPGA using board-level power measurement. The use of unrolling and glitch filtering is shown to reduce energy by at least 65% for an Artix-7 device and 50% for a Cyclone IV device while meeting design latency constraints.
Naveen Kumar Dumpala, Shivukumar B. Patil, Daniel E. Holcomb, Russell Tessier
ACM Trans. Reconfigurable Technol. Syst.4
2019 Efficient PUF-Based Key Generation in FPGAs Using Per-Device Configuration
abstract
Reconfigurable systems often require secret keys to encrypt and decrypt data. Applications requiring high security commonly generate keys based on physical unclonable functions (PUFs), circuits that use random manufacturing variations to produce secret keys that are unique to each device. Implementing PUFs on field-programmable gate arrays (FPGAs) is usually difficult, because the designer has limited control over layout, and each PUF system requires a large area overhead to correct errors in the PUF response bits. In this paper, we extend the state of the art for FPGA-based weak PUFs using a novel methodology of per-device configuration and a new PUF variant derived from the popular FPGA-specific Anderson PUF. The PUF is evaluated using Xilinx XC7Z020 programmable systemon-chips from the Virtex-7 family on Zynq ZedBoard platforms. The design we propose has several advantages over existing work including the Anderson PUF on which it is based. Our design is tunable to minimize the response bias and can be implemented using the common SLICEL components on Xilinx FPGAs. Moreover, the proposed PUF design enables an efficient per-device configuration that reduces bit error rate by over 10× at room temperature and improves response stability by over 2× across all temperatures. We demonstrate that the proposed per-device PUF configuration step leads to roughly 2× savings in area resources for PUFs and error correction as used in key generation.
Mohammad A. Usmani, Shahrzad Keshavarz, Eric Matthews, Lesley Shannon, Russell Tessier, Daniel E. Holcomb
IEEE Trans. Very Large Scale Integr. Syst.5
2018 A Bandwidth-Optimized Routing Algorithm for Hybrid FPGA Networks-on-Chip
abstract
In this paper, a heuristic routing algorithm that is tuned for routing traffic in hybrid FPGA NoCs is presented. This multi-iteration routing algorithm requires a limited amount of hardware for prescheduled stream-based routing while allowing for bandwidth-optimized usage of NoC routing resources. By efficiently scheduling data streams, the remaining NoC bandwidth can be used for bursty, packet-switched traffic. We demonstrate our approach using a hybrid NoC and show an average 11% data stream bandwidth improvement for a collection of five benchmark traffic patterns.
Shivukumar B. Patil, Russell Tessier
FCCM3
2018 FPGA Side Channel Attacks without Physical Access
abstract
As FPGA use becomes more diverse, the shared use of these devices becomes a security concern. Multi-tenant FPGAs that contain circuits from multiple independent sources or users will soon be prevalent in cloud and embedded computing environments. The recent discovery of a new attack vector using neighboring long wires in Xilinx SRAM FPGAs presents the possibility of covert information leakage from an unsuspecting user's circuit. The work described in this paper makes two contributions that dramatically extend this finding. First, we rigorously evaluate several Intel SRAM FPGAs and confirm that long wire information leakage is also prevalent in these devices. Second, we present the first successful attack on an unsuspecting circuit in an FPGA using information passively obtained from neighboring long-lines. Information obtained from a single AES S-box input wire combined with analysis of encrypted output is used to rapidly expose an AES key. This attack is performed remotely without modifying the victim circuit, using electromagnetic probes or power measurements, or modifying the FPGA in any way. We show that our approach is effective for three different FPGA devices. Our results demonstrate that the attack can recover encryption keys from AES circuits running at 10MHz, and has the capability to scale to much higher frequencies.
Chethan Ramesh, Shivukumar B. Patil, Siva Nishok Dhanuskodi, George Provelengios, Sébastien Pillement, Daniel E. Holcomb, Russell Tessier
FCCM7
2018 Lynq: A Lightweight Software Layer for Rapid SoC FPGA Prototyping
abstract
Modern FPGAs include a diverse collection of heterogeneous processing elements including microprocessors. However, in many cases, specialized knowledge is required to integrate processing elements, IP hardware cores, memory interfaces and interconnects together. Xilinx recently released PYNQ, an open-source framework to enable interactive testing, rapid design iteration, and fast prototyping on SoC FPGAs. In this paper we present Lynq, Lua for Lynq, a lightweight software layer for rapid SoC FPGA prototyping on Xilinx Zynq devices. We evaluate the performance and energy efficiency of the new software and assess hardware integration efficiency versus competing approaches. It is shown that we outperform Python implementations with PYNQ even when a JITed version of Python is available. Run-time speedups between 3.2× and 4.9× are shown with an energy improvement of 2.5× to 4.8× versus PYNQ. System bootup is achieved in less than 10 ms which fits time-critical application requirements.
Jonathan Déchelotte, Russell Tessier, Dominique Dallet, Jérémie Crenne
FPL2
2018 Hybrid Obfuscation to Protect Against Disclosure Attacks on Embedded Microprocessors
abstract
The risk of code reverse-engineering is particularly acute for embedded processors which often have limited available resources to protect program information. Previous efforts involving code obfuscation provide some additional security against reverse- engineering of programs, but the security benefits are typically limited and not quantifiable. Hence, new approaches to code protection and creation of associated metrics are highly desirable. This paper has two main contributions. We propose the first hybrid diversification approach for protecting embedded software and we provide statistical metrics to evaluate the protection. Diversification is achieved by combining hardware obfuscation at the microarchitecture level and the use of software-level obfuscation techniques tailored to embedded systems. Both measures are based on a compiler which generates obfuscated programs, and an embedded processor implemented in an FPGA with a randomized Instruction Set Architecture (ISA) encoding to execute the hybrid obfuscated program. We employ a fine-grained, hardware-enforced access control mechanism for information exchange with the processor and hardware-assisted booby traps to actively counteract manipulation attacks. It is shown that our approach is effective against a wide variety of possible information disclosure attacks in case of a physically present adversary. Moreover, we propose a novel statistical evaluation methodology that provides a security metric for hybrid-obfuscated programs.
Marc Fyrbiak, Simon Rokicki, Nicolai Bissantz, Russell Tessier, Christof Paar
IEEE Trans. Computers4
2017 Hardware support for embedded operating system security
abstract
Internet-connected embedded systems have limited capabilities to defend themselves against remote hacking attacks. The potential effects of such attacks, however, can have a significant impact in the context of the Internet of Things, industrial control systems, smart health systems, etc. Embedded systems cannot effectively utilize existing software-based protection mechanisms due to limited processing capabilities and energy resources. We propose a novel hardware-based monitoring technique that can detect if the embedded operating system or any running application deviates from the originally programmed behavior due to an attack. We present an FPGA-based prototype implementation that shows the effectiveness of such a security approach.
Arman Pouraghily, Tilman Wolf, Russell Tessier
ASAP3
2017 Energy Efficient Loop Unrolling for Low-Cost FPGAs
abstract
Many FPGA computations, including block ciphers, require repetitive loop operations that are difficult to parallelize. Sequential loop implementation leads to significant clock power while loop unrolling can lead to significant glitch power. In this paper, we provide a low overhead approach to unroll block ciphers and other loops in low-cost FPGAs to reduce energy consumption. A latch-based glitch filter is introduced for unrolled loops that reduces loop energy per operation by over an order of magnitude. Our filters and associated control for unrolled loops can be automatically instantiated as a macro for FPGA designs, allowing for easy designer use. We demonstrate our approach for SIMON-128 and AES-256 block ciphers implemented on a Xilinx Artix-7 FPGA.
Naveen Kumar Dumpala, Shivukumar B. Patil, Daniel E. Holcomb, Russell Tessier
FCCM4
2017 Scalable Network Function Virtualization for Heterogeneous Middleboxes
abstract
Over the past decade, a wide-ranging collection of network functions in middleboxes has been used to accommodate the needs of network users. Although the use of general-purpose processors has been shown to be feasible for this purpose, the serial nature of microprocessors limits network functional virtualization (NFV) performance. In this paper, we describe a new heterogeneous hardware-software approach to NFV construction that provides scalability and programmability, while supporting significant hardware-level parallelism and reconfiguration. Our computing platform uses both field-programmable gate arrays (FPGA) and microprocessors to implement numerous NFV operations that can be dynamically customized to specific network flow needs. As the number of required functions and their characteristics change, the hardware in the FPGA is automatically reconfigured to support the updated requirements. Traffic management and hardware reconfiguration functions are performed by a global coordinator which allows for the rapid sharing of middlebox state and continuous evaluation of network function needs. To evaluate our approach, a series of software tools and NFV modules have been implemented. Our system is shown to be scalable for collections of network functions exceeding one million shared states.
Xuzhi Zhang, Xiaozhe Shao, George Provelengios, Naveen Kumar Dumpala, Lixin Gao 0001, Russell Tessier
FCCM6
2017 A High-Speed Accelerator for Homomorphic Encryption using the Karatsuba Algorithm
abstract
Somewhat Homomorphic Encryption (SHE) schemes can be used to carry out operations on ciphered data. In a cloud computing scenario, personal information can be processed secretly, inferring a high level of confidentiality. The principle limitation of SHE is the size of ciphertext compared to the size of the message. This issue can be addressed by using a batching technique that “packs” several messages into one ciphertext. However, this method leads to important drawbacks in standard implementations. This paper presents a fast hardware/software co-design implementation of an encryption procedure using the Karatsuba algorithm. Our hardware accelerator is 1.5 times faster than the state of the art for 1 encryption and 4 times faster for 4 encryptions.
Vincent Migliore, Cédric Seguin, Maria Mendez Real, Vianney Lapotre, Arnaud Tisserand, Caroline Fontaine, Guy Gogniat, Russell Tessier
ACM Trans. Embed. Comput. Syst.8
2016 Effects of I/O routing through column interfaces in embedded FPGA fabrics
abstract
The emergence of 2.5D and 3D packaging technologies enables the integration of FPGA dice into more complex systems. Both heterogeneous manycore designs, which include an FPGA layer, and interposer-based multi-FPGA systems support the inclusion of reconfigurable hardware in 3D-stacked integrated circuits. In these architectures, the communication between FPGA dice or between FPGA and fixed-function layers often takes place through dedicated communication interfaces spread over the FPGA logic fabric, as opposed to an I/O ring around the fabric. In this paper, we investigate the effect of organizing FPGA fabric I/O into coarse-grained interface blocks distributed throughout the FPGA fabric. Specifically, we consider the quality of results for the placement and routing phases of the FPGA physical design flow. We evaluate the routing of I/O signals of large applications through dedicated interface blocks at various granularities in the logic fabric, and study its implications on the critical path delay of routed designs. We show that the impact of such I/O routing is limited and can improve chip routability and circuit delay in many cases.
Christophe Huriaux, Olivier Sentieys, Russell Tessier
FPL3
2016 Improving the efficiency of PUF-based key generation in FPGAs using variation-aware placement
abstract
Reconfigurable systems often require secret keys to encrypt and decrypt data. Applications requiring high security commonly generate keys based on physical unclonable functions (PUFs), circuits which use random manufacturing variations to produce secret keys that are unique to each device. The security of PUF-based keys comes at a high hardware cost. Due to the need for error correction to extract reliable keys from noisy PUFs, the total cost of an n-bit key far exceeds just the cost of producing n bits of PUF output. In this work, we propose variation-aware intra-FPGA PUF placement to reduce the area cost of PUF-based keys on FPGAs. We show that placing PUF instances according to the random variations of each chip instance reduces the bit error rate of the PUFs and consequently greatly reduces the overall cost of key generation. The proposed variation-aware placement approach is applicable to any PUF-based system implemented in reconfigurable logic. We demonstrate our approach on a Xilinx Zynq-7000 Programmable SoC using FPGA-specific PUFs with code-offset error correction based on BCH codes. We quantify the effectiveness of our approach by comparing the implementation costs of the same system when using the default approach of variation-agnostic placement and our proposed variation-aware placement. It is shown that our approach reduces the area required for PUF and error-correction circuitry by about 50% while achieving equivalent reliability.
Shrikant Vyas, Naveen Kumar Dumpala, Russell Tessier, Daniel E. Holcomb
FPL3
2016 Hybrid hard NoCs for efficient FPGA communication
abstract
Recent research has shown that the integration of a custom-silicon network-on-chip (NoC) into an FPGA fabric can significantly help on-chip communication bandwidth. Although not appropriate for all communication, FPGA hard NoCs provide a scalable infrastructure that will increase in importance as FPGA sizes increase. To date, most FPGA hard NoC implementation has focused on packet-switched routing which requires dynamic run-time decision making to transfer data from source to destination. In this paper we explore expanding hard NoC routers to support both packet-switched and prescheduled time-multiplexed communication. By limiting the use of energy-hungry routing buffers, time-multiplexed routing allows for throughput-predictable data transport with reduced energy consumption versus packet-switching for a variety of traffic patterns. The area overhead required to convert a packet-switched router to a hybrid packet-switched/time-multiplexed version is minimal (about 9%). In this research, we show significant NoC energy improvement (about 33%) while maintaining or improving packet latency values versus packet switching for select data traffic patterns.
Naveen Kumar Dumpala, Russell Tessier
FPT3
2016 Dynamic Hardware Monitors for Network Processor Protection
abstract
The importance of the Internet for society is increasing. To ensure a functional Internet, its routers need to operate correctly. However, the need for router flexibility has led to the use of software-programmable network processors in routers, which exposes these systems to data plane attacks. Recently, hardware monitors have been introduced into network processors to verify the expected behavior of processor cores at run time. If instruction-level execution deviates from the expected sequence, an attack is identified, triggering processor core recovery efforts. In this manuscript, we describe a scalable network processor monitoring system that supports the reallocation of hardware monitors to processor cores in response to workload changes. The scalability of our monitoring architecture is demonstrated using theoretical models, simulation, and router system-level experiments implemented on an FPGAbased hardware platform. For a system with four processor cores and six monitors, the monitors result in a 6 percent logic and 38 percent memory bit overhead versus the processor's core logic and instruction storage. No slowdown of system throughput due to monitoring is reported.
Kekai Hu, Harikrishnan Chandrikakutty, Zachary Goodman, Russell Tessier, Tilman Wolf
IEEE Trans. Computers4
2015 Multi-task support for security-enabled embedded processors
abstract
Embedded systems require low overhead security approaches to ensure that they are protected from attacks. In this paper, we propose a hardware-based approach to secure the operation of an embedded processor instruction-by-instruction, where deviations from expected program behavior are detected within the execution of an instruction. These security-enabled embedded processors provide effective defenses against common attacks, such as stack smashing. Previous work in this area has focused on monitoring a single task on a CPU while here we present a novel hardware monitoring system that can monitor multiple active tasks in an operating-system-based platform. The hardware monitor is able to track context switches that occur in the operating system and ensure that monitoring is performed continuously, thus ensuring system security. We present the design of our system and results obtained from a prototype implementation of the system on an Altera DE4 FPGA board. We demonstrate in hardware that applications can be monitored at the instruction level without execution slowdown and stack smashing attacks can be defeated using our system.
Tedy Thomas, Arman Pouraghily, Kekai Hu, Russell Tessier, Tilman Wolf
ASAP4
2015 Hardware-assisted code obfuscation for FPGA soft microprocessors
Meha Kainth, Lekshmi Krishnan, Chaitra Narayana, Sandesh Gubbi Virupaksha, Russell Tessier
DATE5
2015 Protecting against Cryptographic Trojans in FPGAs
abstract
In contrast to ASICs, hardware Trojans can potentially be injected into FPGA designs post-manufacturing by bit stream alteration. Hardware Trojans which target cryptographic primitives are particularly interesting for an adversary because a weakened primitive can lead to a complete loss of system security. One problem an attacker has to overcome is the identification of cryptographic primitives in a large bit stream with unknown semantics. As the first contribution, we demonstrate that AES can be algorithmically identified in a look-up table-level design for a variety of implementation styles. Our graph-based approach considers AES implementations which are created using several synthesis and technology mapping options. As the second contribution, we present and discuss the drawbacks of a dynamic obfuscation countermeasure which allows for the configuration of certain crucial parts of a cryptographic primitive after the algorithm has been loaded into the FPGA. As a result, reverse-engineering and modifying a primitive in the bit stream is more challenging.
Pawel Swierczynski, Marc Fyrbiak, Christof Paar, Christophe Huriaux, Russell Tessier
FCCM5
2015 Adaptive MRAM-based CGRAs
abstract
In this paper, we describe the use of magnetic RAM (MRAM) in coarse-grained reconfigurable arrays (CGRAs) as a configuration cache to allow for bulk low-energy storage and rapid device reconfiguration. If an energy-saving configuration update for an application is needed, a new configuration can be quickly swapped into compute blocks and interconnect switchboxes to minimize system down time. The determination of when to configure and an analysis of the CGRA architectural impact of MRAM is evaluated via system-level emulation. Our experiments show that the use of MRAM reduces overall application energy consumption by nearly 30% when dynamic reconfiguration is used.
Tedy Thomas, Alan Boguslawski, Russell Tessier
FPL4
2015 Reinforcement Learning for Thermal-aware Many-core Task Allocation
abstract
To maintain reliable operation, task allocation for many-core processors must consider the heat interaction of processor cores and network-on-chip routers in performing task assignment. Our approach employs reinforcement learning, machine learning algorithm that performs task allocation based on current core and router temperatures and a prediction of which assignment will minimize maximum temperature in the future. The algorithm updates prediction models after each allocation based on feedback regarding the accuracy of previous predictions. Our new algorithm is verified via detailed many-core simulation which includes on-chip routing. Our results show that the proposed technique is fast (scheduling performed in <1 ms) and can efficiently reduce peak temperature by up to 8°C in a 49-core processor (4.3°C on average) versus a competing task allocation approach for a series of SPLASH-2 benchmarks.
Shiting (Justin) Lu, Russell Tessier, Wayne P. Burleson
ACM Great Lakes Symposium on VLSI2
2015 Reconfigurable Computing Architectures
abstract
Reconfigurable architectures can bring unique capabilities to computational tasks. They offer the performance and energy efficiency of hardware with the flexibility of software. In some domains, they are the only way to achieve the required, real-time performance without fabricating custom integrated circuits. Their functionality can be upgraded and repaired during their operational lifecycle and specialized to the particular instance of a task. We survey the field of reconfigurable computing, providing a guide to the body-of-knowledge accumulated in architecture, compute models, tools, run-time reconfiguration, and applications.
Russell Tessier, Kenneth L. Pocek, André DeHon
Proc. IEEE1
2015 Securing Network Processors with High-Performance Hardware Monitors
abstract
As the Internet becomes integrated into nearly all aspects of everyday life, its reliability grows in importance. This vital communication resource, which has become an inviting target for attackers, must be protected with the same vigor as the end-systems it interconnects. Recent trends in network router architecture towards programmability and flexibility have increased the susceptibility of communication hardware to software attacks which modify intended data processing and forwarding functions. Contemporary routers typically feature network processors, whose protocol processing functions are determined via software. Prior work has shown that these general-purpose software-based processing systems can be attacked with data packets sent through the Internet. As a defense mechanism, the correct functionality of a network processor can be verified by a hardware monitor that observes processor operation and compares it to expected behavior. In the event of an attack, the monitor can interrupt the network processor, suppress malicious behavior, and reset the processor to a usable state for processing of subsequent traffic. In this work, we present several significant advances in hardware monitoring for network processors. A low-overhead monitor architecture that evaluates correct network processor operation in real-time on an instruction-by-instruction basis is described and tested. The monitor is shown to effectively prevent stack smashing attacks on processors that use a Harvard architecture, a widely used network processor configuration. Through experimentation, we show that our approach to hardware monitoring does not affect data plane packet throughput. In the event of an attack, malicious packets are dropped while packets of regular network traffic proceed through the network unaffected. A full evaluation of monitor architectural parameters is provided to create an optimized monitor design.
Tilman Wolf, Harikrishnan Chandrikakutty, Kekai Hu, Deepak Unnikrishnan, Russell Tessier
IEEE Trans. Dependable Secur. Comput.5
2014 System-Level Security for Network Processors with Hardware Monitors
abstract
New attacks are emerging that target the Internet infrastructure. Modern routers use programmable network processors that may be exploited by merely sending suitably crafted data packets into a network. Hardware monitors that are co-located with processor cores can detect attacks that change processor behavior with high probability. In this paper, we present a solution to the problem of secure, dynamic installation of hardware monitoring graphs on these devices. We also address the problem of how to overcome the homogeneity of a network with many identical devices, where a successful attack, albeit possible only with small probability, may have devastating effects.
Kekai Hu, Tilman Wolf, Thiago Teixeira, Russell Tessier
DAC4
2014 FPGA Architecture Enhancements to Support Heterogeneous Partially Reconfigurable Regions
abstract
In this work the author develop an FPGA architecture which allows for the placement of a partial FPGA design on the logic fabric even if the relative placement of heterogeneous blocks within the target region is not identical to the placement used to generate the bitstream for the partial design. This work has been conducted in the context of the European FP7 FlexTiles project in which a dynamically reconfigurable logic fabric is embedded in a 3-D stacked chip along with a manycore architecture. The reconfigurable logic fabric is used to load hardware-accelerated functions whose use is scheduled at run time. All communication between the fabric and manycore is made via dedicated I/O interface blocks in the fabric. This communication configuration increases the need for a flexible architecture which can handle the placement of a single application bitstream in multiple locations on the logic fabric.
Christophe Huriaux, Olivier Sentieys, Russell Tessier
FCCM3
2014 FPGA architecture support for heterogeneous, relocatable partial bitstreams
abstract
The use of partial dynamic reconfiguration in FPGA-based systems has grown in recent years as the spectrum of applications which use this feature has increased. For these systems, it is desirable to create a series of partial bitstreams which represent tasks which can be located in multiple regions in the FPGA fabric. While the transferal of homogeneous collections of lookup-table based logic blocks from region to region has been shown to be relatively straightforward, it is more difficult to transfer partial bitstreams which contain fixed-function resources, such as block RAMs and DSP blocks. In this paper we consider FPGA architecture enhancements which allow for the migration of partial bitstreams including fixed-function resources from region to region even if these resources are not located in the same position in each region. Our approach does not require significant, time-consuming place-and-route during the migration process. We quantify the cost of inserting additional routing resources into the FPGA architecture to allow for easy migration of heterogeneous, fixed-function resources. Our experiments show that this flexibility can be added for a relatively low overhead and performance penalty.
Christophe Huriaux, Olivier Sentieys, Russell Tessier
FPL3
2014 Dynamic On-Chip Thermal Sensor Calibration Using Performance Counters
abstract
Numerous sensors are currently deployed in modern processors to collect thermal information for fine-grained dynamic thermal management (DTM). Due to process variation and silicon aging, on-chip thermal sensors require periodic calibration before use in DTM. However, the calibration cost for thermal sensors can be prohibitively high as the number of on-chip sensors increases. In this paper, a model which is suitable for online calculation is employed to estimate the temperatures of multiple sensor locations on the silicon die. The estimated sensor and actual sensor thermal profile show a very high similarity with correlation coefficient${\sim}{\rm 0.9}$for most tested benchmarks. Our calibration approach combines potentially inaccurate temperature values obtained from two sources: temperature readings from thermal sensors and temperature estimations using system performance counters. A data fusion strategy based on Bayesian inference, which combines information from these two sources, is demonstrated along with a temperature estimation approach using performance counters. The average absolute error of the corrected sensor temperature readings is${<}{1.5}^{\circ}{\rm C}$and the standard deviation of error is less than${<}{\rm 0.5}^{\circ}{\rm C}$for tested benchmarks.
Shiting (Justin) Lu, Russell Tessier, Wayne P. Burleson
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 High-performance hardware monitors to protect network processors from data plane attacks
abstract
The Internet represents an essential communication infrastructure that needs to be protected from malicious attacks. Modern network routers are typically implemented using embedded multi-core network processors that are inherently vulnerable to attack. Hardware monitor subsystems, which can verify the behavior of a router's packet processing system at runtime, can be used to identify and respond to an ever-changing range of attacks. While hardware monitors have primarily been described in the context of general-purpose computing, our work focuses on two important aspects that are relevant to the embedded networking domain: We present the design and prototype implementation of a high-performance monitor that can track each processor instruction with low memory overhead. Additionally, our monitor is capable of defending against attacks on processors with a Harvard architecture, the dominant contemporary network processor organization. We demonstrate that our monitor architecture provides no network slowdown in the absence of an attack and provides the capability to drop attack packets without otherwise affecting regular network traffic when an attack occurs.
Harikrishnan Chandrikakutty, Deepak Unnikrishnan, Russell Tessier, Tilman Wolf
DAC3
2013 FPGA latency optimization using system-level transformations and DFG restructuring
abstract
This paper describes a system-level approach to improve the latency of FPGA designs by performing optimization of the design specification on a functional level prior to high-level synthesis. The approach uses Taylor Expansion Diagrams (TEDs), a functional graph-based design representation, as a vehicle to optimize the dataflow graph (DFG) used as input to the subsequent synthesis. The optimization focuses on critical path compaction in the functional representation before translating it into a structural DFG representation. Our approach engages several passes of a traditional high-level synthesis (HLS) process in a simulated annealing-based loop to make efficient cost tradeoffs. The algorithm is time efficient and can be used for fast design space exploration. The results indicate a latency performance improvement of 22% on average versus HLS with the initial DFG for a series of designs mapped to Altera Stratix II devices.
Daniel Gomez-Prado, Maciej J. Ciesielski, Russell Tessier
DATE3
2013 Run-time probabilistic detection of miscalibrated thermal sensors in many-core systems
abstract
Many-core architectures use large numbers of small temperature sensors to detect thermal gradients and guide thermal management schemes. In this paper a technique to identify thermal sensors which are operating outside a required accuracy is described. Unlike previous on-chip temperature estimation approaches, our algorithms are optimized to run on-line while thermal management decisions are being made. The accuracy of a sensor is determined by comparing its readings to expected values from a probability distribution function determined from surrounding sensors. Experiments show that a sensor operating outside a desired accuracy can be identified with a detection rate of over 90% and an average false alarm rate of < 6%, with a confidence level of 90%. The run time of our method is shown to be around 3× lower than a recently-published temperature estimation method, enhancing its suitability for run-time implementation.
Shiting (Justin) Lu, Wayne P. Burleson, Russell Tessier
DATE4
2013 FlexGrip: A soft GPGPU for FPGAs
abstract
Over the past decade, soft microprocessors and vector processors have been extensively used in FPGAs for a wide variety of applications. However, it is difficult to straightforwardly extend their functionality to support conditional and thread-based execution characteristic of general-purpose graphics processing units (GPGPUs) without recompiling FPGA hardware for each application. In this paper, we describe the implementation of FlexGrip, a soft GPGPU architecture which has been optimized for FPGA implementation. This architecture supports direct CUDA compilation to a binary which is executable on the FPGA-based GPGPU without hardware recompilation. Our architecture is customizable, thus providing the FPGA designer with a selection of GPGPU cores which display performance versus area tradeoffs. The benefits of our architecture are evaluated for a collection of five standard CUDA benchmarks which are compiled using standard GPGPU compilation tools. Speedups of up to 30× versus a MicroBlaze microprocessor are achieved for designs which take advantage of the conditional execution capabilities offered by FlexGrip.
Kevin Andryc, Murtaza Merchant, Russell Tessier
FPT3
2013 An open-source SATA core for Virtex-4 FPGAs
abstract
In this demonstration, we present an open-source Serial ATA core designed for Virtex-4 FPGAs. This core utilizes the RocketIO Multi-Gigabit Transceiver (MGT) of the Virtex-4 to interface with hard drives at SATA Generation 1 (SATA I, 1.5 Gb/s) and Generation 2 (SATA II, 3.0 Gb/s) speeds. A full design hierarchy from host software to the physical layer is provided with the distribution to facilitate design use. A simple, FIFO interface allows for easy integration with other FPGA modules. The demonstration illustrates the correct write and read behavior of the core using a Xilinx ML405 board and a solid state disk. The peak transfer rate of the core for SATA I (130 MB/s) is demonstrated. Our goal for the demonstration is to educate the reconfigurable computing community regarding the availability of the core and to illustrate its capabilities.
Cory Gorman, Paul Siqueira, Russell Tessier
FPT3
2013 Accelerating iterative algorithms with asynchronous accumulative updates on FPGAs
abstract
Iterative algorithms represent a pervasive class of data mining, web search and scientific computing applications. In iterative algorithms, a final result is derived by performing repetitive computations on an input data set. Existing techniques to parallelize such algorithms typically use software frameworks such as MapReduce and Hadoop to distribute data for an iteration across multiple CPU-based workstations in a cluster and collect per-iteration results. These platforms are marked by the need to synchronize data computations at iteration boundaries, impeding system performance. In this paper, we demonstrate that FPGAs in distributed computing systems can serve a vital role in breaking this synchronization barrier with the help of asynchronous accumulative updates. These updates allow for the accumulation of intermediate results for numerous data points without the need for iteration-based barriers allowing individual nodes in a cluster to independently make progress towards the final outcome. Computation is dynamically prioritized to accelerate algorithm convergence. A general-class of iterative algorithms have been implemented on a cluster of four FPGAs. A speedup of 7× is achieved over an implementation of asynchronous accumulative updates on a general-purpose CPU. The system offers up to 154× speedup versus a standard Hadoop-based CPU-workstation. Improved performance is achieved by clusters of FPGAs.
Deepak Unnikrishnan, Sandesh Gubbi Virupaksha, Lekshmi Krishnan, Lixin Gao 0001, Russell Tessier
FPT5
2013 Reconfigurable Data Planes for Scalable Network Virtualization
abstract
Network virtualization presents a powerful approach to share physical network infrastructure among multiple virtual networks. Recent advances in network virtualization advocate the use of field-programmable gate arrays (FPGAs) as flexible high performance alternatives to conventional host virtualization techniques. However, the limited on-chip logic and memory resources in FPGAs severely restrict the scalability of the virtualization platform and necessitate the implementation of efficient forwarding structures in hardware. The research described in this manuscript explores the implementation of a scalable heterogeneous network virtualization platform that integrates virtual data planes implemented in FPGAs with software data planes created using host virtualization techniques. The system exploits data plane heterogeneity to cater to the dynamic service requirements of virtual networks by migrating networks between software and hardware data planes. We demonstrate data plane migration as an effective technique to limit the impact of traffic on unmodified data planes during FPGA reconfiguration. Our system implements forwarding tables in a shared fashion using inexpensive off-chip memories and supports both Internet Protocol (IP) and non-IP-based data planes. Experimental results show that FPGA-based data planes can offer two orders of magnitude better throughput than their software counterparts, and FPGA reconfiguration can facilitate data plane customization within 12 seconds. An integrated system that supports up to 15 virtual networks has been validated on the NetFPGA platform.
Deepak Unnikrishnan, Ramakrishna Vadlamani, Jérémie Crenne, Lixin Gao 0001, Russell Tessier
IEEE Trans. Computers6
2013 Configurable memory security in embedded systems
abstract
System security is an increasingly important design criterion for many embedded systems. These systems are often portable and more easily attacked than traditional desktop and server computing systems. Key requirements for system security include defenses against physical attacks and lightweight support in terms of area and power consumption. Our new approach to embedded system security focuses on the protection of application loading and secure application execution. During secure application loading, an encrypted application is transferred from on-board flash memory to external double data rate synchronous dynamic random access memory (DDR-SDRAM) via a microprocessor. Following application loading, the core-based security technique provides both confidentiality and authentication for data stored in a microprocessor's system memory. The benefits of our low overhead memory protection approaches are demonstrated using four applications implemented in a field-programmable gate array (FPGA) in an embedded system prototyping platform. Each application requires a collection of tasks with varying memory security requirements. The configurable security core implemented on-chip inside the FPGA with the microprocessor allows for different memory security policies for different application tasks. An average memory saving of 63% is achieved for the four applications versus a uniform security approach. The lightweight circuitry included to support application loading from flash memory adds about 10% FPGA area overhead to the processor-based system and main memory security hardware.
Jérémie Crenne, Romain Vaslin, Guy Gogniat, Jean-Philippe Diguet, Russell Tessier, Deepak Unnikrishnan
ACM Trans. Embed. Comput. Syst.5
2012 Distributed sensor data processing for many-cores
abstract
Future many-core systems will rely heavily on a wide variety of sensors which provide run-time information about on-chip environment and workload. In this paper, a new dedicated infrastructure for distributed sensor processing for many-core systems is described. This infrastructure includes a sparse array of dedicated processors which evaluate on-chip sensor data and a two-level hierarchical network-on-chip (NoC) which allows for efficient sensor data collection. This design is evaluated using benchmark driven simulations for a three-dimensional (3D) stack, necessitating inter-layer sensor data communication. The experimental results for up to 1024 cores indicate that for typical sensor data collection rates, one sensor data processor (SDP) per 64 cores is optimal for sensor data latency. The use of a two-level NoC is shown to provide an average of 65% sensor data latency improvement versus a flat sensor data NoC structure for a 256-core system.
Russell Tessier, Wayne P. Burleson
ACM Great Lakes Symposium on VLSI2
2012 Saving energy and improving TCP throughput with rate adaptation in Ethernet
abstract
Reducing the power consumption of network interfaces contributes to lowering the overall power needs of the compute and communication infrastructure. Most modern Ethernet interfaces can operate at one of several data rates. In this paper, we present Queue Length Based Rate Adaptation (QLBRA), which can dynamically adapt the link rate for Ethernet interfaces at runtime using existing Ethernet standards. An implementation of the proposed rate adaptation functionality is demonstrated at runtime on a NetFPGA platform. Our results show that the rate adaptation approach can achieve significant energy savings and at the same time improve the throughput of TCP traffic due to the effect of packet pacing.
Y. Sinan Hanay, Russell Tessier, Tilman Wolf
ICC3
2012 Collaborative calibration of on-chip thermal sensors using performance counters
abstract
Thermal sensors are currently deployed in processors to collect thermal information for dynamic thermal management (DTM). The calibration cost for thermal sensors can be prohibitively high as the number of on-chip sensors increases. We propose an on-line multi-sensor calibration method which combines potentially inaccurate temperature values obtained from two sources: temperature readings from thermal sensors and temperature estimations using system performance counters. A data fusion strategy based on Bayesian inference, which combines information from these two sources, is demonstrated along with a temperature estimation approach using performance counters. The approaches are verified via simulation for an AMD Athlon 64 processor with 24 on-chip temperature sensors scaled to a 45nm technology node. Our results show that the standard deviation of temperature sensor measurement errors can be reduced from 3 ~ 4 °C to ≤ 1 °C using the proposed method. Additionally, our MATLAB implementation shows that the new approach runs at least 67x faster than competing approaches based on Kalman filtering making it highly appropriate for run-time use.
Shiting (Justin) Lu, Russell Tessier, Wayne P. Burleson
ICCAD2
2011 ReClick - A Modular Dataplane Design Framework for FPGA-Based Network Virtualization
abstract
Network virtualization has emerged as a powerful technique to deploy novel services and experimental protocols over shared network infrastructures. Although recent research has highlighted field programmable gate arrays (FPGAs) as attractive platforms for high performance network virtualization, these devices remain inaccessible to the larger networking research community due to the absence of user-friendly programming models. A programming model that can abstract the intricacies of the hardware platform while being aware of the underlying resource constraints is highly desirable. In this paper, we present ReClick, a framework to efficiently design and deploy reconfigurable data planes for FPGA-based network virtualization systems. A hardware-agnostic programming model is described that allows developers to focus on the virtual data plane semantics rather than the implementation details. The framework exposes interfaces similar to the popular software router development framework, Click, and promotes design reuse. Optimization strategies are included in ReClick which use similarities between virtual data plane configurations to implement multiple planes in an area-efficient manner. Data planes exhibiting up to 1 Gbps data rate have been automatically compiled and tested in hardware in a Net FPGA platform.
Deepak Unnikrishnan, Shiting (Justin) Lu, Lixin Gao 0001, Russell Tessier
ANCS4
2011 A Dynamically-Reconfigurable Phased Array Radar Processing System
abstract
Digital beam forming is an important radar processing technique used in many communication and radar sensing applications. This paper presents a low-cost digital beam forming system which takes advantage of four eight channel analog-to-digital (A-to-D) converter chips and dynamic FPGA reconfiguration. A full digital beam forming algorithm capable of forming up to 24 beams from 64 antenna input signals is described. FPGA reconfiguration is performed in 400 ms allowing for FPGA ret asking of the associated radar for weather and aircraft tracking. Beam forming performance of 64.2 GOPs per second for weather tracking and 72.2 GOPs per second for aircraft tracking is reported. The complete low cost digital beam forming board, including parts and assembly, costs less than $3,000.
Emmanuel Seguin, Russell Tessier, Eric J. Knapp, Robert W. Jackson
FPL2
2011 Efficient key-dependent message authentication in reconfigurable hardware
abstract
Cryptographic message authentication is a growing need for FPGA-based embedded systems. In this paper a customized FPGA implementation of a GHASH function that is used in AES-GCM, a widely-used message authentication protocol, is described. The implementation limits GHASH logic utilization by specializing the hardware implementation on a per-key basis. The implemented module can generate a 128bit message authentication code in both pipelined and unpipelined versions. The pipelined GHASH version achieves an authentication throughput of more than 14 Gbit/s on a Spartan-3 FPGA and 292 Gbit/s on a Virtex-6 device. To promote adoption in the field, the complete source code for this work has been made publically-available.
Jérémie Crenne, Pascal Cotret, Guy Gogniat, Russell Tessier, Jean-Philippe Diguet
FPT4
2011 Evaluation of the Universal Geocast Scheme for VANETs
abstract
Recently, a number of communications schemes have been proposed for Vehicular Ad hoc Networks (VANETs). A promising approach, the Universal Geocast Scheme (UGS), provides for a diverse variety of VANET-specific characteristics such as time-varying topology, protocol variation based on road congestion, and support for non line-of-sight communication. In this paper, the UGS protocol is extended to consider inter-vehicle multi-hop connections in intersections with surrounding obstructions. Since UGS is a probabilistic, repetition-based scheme, it supports the capacity-delay tradeoffs crucial for periodic safety message exchange. The approach is shown to support both vehicle-to-vehicle and vehicle-to-infrastructure communication. This research accurately evaluates this scheme using network (NS-2) and mobility (SUMO) simulators, verifying two crucial elements of successful VANETs, received packet ratio and message delay. A contemporary wireless radio propagation model is used to augment accuracy. Our results show a 6% improvement in received packet ratio combined with a decrease in average packet delay versus a previous, well-known inter vehicle communication protocol.
Ben Bovee, Mohammad Nekoui, Hossein Pishro-Nik, Russell Tessier
VTC Fall4
2011 Securing the data path of next-generation router systems
Tilman Wolf, Russell Tessier, Gayatri Prabhu
Comput. Commun.2
2011 A Dedicated Monitoring Infrastructure for Multicore Processors
abstract
On-chip monitoring of environmental information, such as temperature, voltage, and error data, is becoming increasingly important. To address this need, a low-overhead architectural approach to monitor data collection and use in multicore systems is described. A key aspect of our standalone monitoring subsystem is a low-complexity, on-chip network designed to transport monitor data with multiple priority levels. Collected monitor information is evaluated by a dedicated processor. Experimental results using architectural and interconnect simulators show that the new low-overhead subsystem facilitates employment of thermal and delay-aware dynamic voltage and frequency scaling. In contrast to using existing on-chip interconnect resources to communicate monitor data, the new subsystem provides necessary bandwidth for monitor data traffic without impacting application data traffic. Synthesis results show that our dedicated monitoring approach consumes about 0.2% of multicore area and power resources for an 8-core system based on AMD Athlon 64 processor cores.
Sailaja Madduri, Ramakrishna Vadlamani, Wayne P. Burleson, Russell Tessier
IEEE Trans. Very Large Scale Integr. Syst.5
2010 Multicore soft error rate stabilization using adaptive dual modular redundancy
abstract
The use of dynamic voltage and frequency scaling (DVFS) in contemporary multicores provides significant protection from unpredictable thermal events. A side effect of DVFS can be an increased processor exposure to soft errors. To address this issue, a flexible fault prevention mechanism has been developed to selectively enable a small amount of per-core dual modular redundancy (DMR) in response to increased vulnerability, as measured by the processor architectural vulnerability factor (AVF). Our new algorithm for DMR deployment aims to provide a stable effective soft error rate (SER) by using DMR in response to DVFS caused by thermal events. The algorithm is implemented in real-time on the multicore using a dedicated monitor network-on-chip and controller which evaluates thermal information and multicore performance statistics. Experiments with a multicore simulator using standard benchmarks show an average 6% improvement in overall power consumption and a stable SER by using selective DMR versus continuous DMR deployment.
Ramakrishna Vadlamani, Wayne P. Burleson, Russell Tessier
DATE4
2010 Scalable network virtualization using FPGAs
abstract
Recent virtual network implementations have shown the capability to implement multiple network data planes using a shared hardware substrate. In this project, a new scalable virtual networking data plane is demonstrated which combines the performance efficiency of FPGA hardware with the flexibility of software running on a commodity PC. Multiple virtual router data planes are implemented using a Virtex II-based NetFPGA card to accommodate virtual networks requiring superior packet forwarding performance. Numerous additional data planes for virtual networks which require less bandwidth and slower forwarding speeds are implemented on a commodity PC server via software routers. Through experimentation, we determine that a throughput improvement of up to two orders of magnitude can be achieved for FPGA-based virtual routers versus a software-based virtual router implementation. Dynamic FPGA reconfiguration is supported to adapt to changing networking needs. System scalability is demonstrated for up to 15 virtual routers.
Deepak Unnikrishnan, Ramakrishna Vadlamani, Abhishek Dwaraki, Jérémie Crenne, Lixin Gao 0001, Russell Tessier
FPGA7
2010 Thermal-aware voltage droop compensation for multi-core architectures
abstract
As the rated performance of microprocessors increases, voltage droop emergencies become a significant problem. In this paper, two new techniques to combat voltage droop emergencies are explored. First, a direct connection between temperature and processor clock frequency modulation during voltage droops is established. In general, a higher temperature leads to a lower voltage droop with the same processor activity. Thus, processor frequencies can be reduced less at high temperature in an effort to prevent voltage emergencies. Through experimentation, the benefits of temperature-flexible frequency scaling are explored. Second, processor signatures consisting of performance statistics are used to identify when voltage droop compensation is needed in a multicore environment. The use of an independent on-chip interconnect network allows for the sharing of signatures across cores at run time. Signature sharing in combination with frequency throttling is shown to provide an improvement in average run-time performance in a number of cases for an eight-core multiprocessor.
Basab Datta, Wayne P. Burleson, Russell Tessier
ACM Great Lakes Symposium on VLSI4
2009 A monitor interconnect and support subsystem for multicore processors
abstract
In many current SoCs, the architectural interface to on-chip monitors is ad hoc and inefficient. In this paper, a new architectural approach which advocates the use of a separate low-overhead subsystem for monitors is described. A key aspect of this approach is an on-chip interconnect specifically designed for monitor data with different priority levels. The efficiency of our monitor interconnect is assessed for a multicore system using both an interconnect and a system-level simulator. Collected monitor information is used by a dedicated processor to control the frequency and voltage of individual multicore processors. Experimental results show that the new low-overhead subsystem facilitates employment of thermal and delay-aware dynamic voltage and frequency scaling.
Sailaja Madduri, Ramakrishna Vadlamani, Wayne P. Burleson, Russell Tessier
DATE4
2009 Application Specific Customization and Scalability of Soft Multiprocessors
abstract
Although soft microprocessors are widely used in FPGAs, limited work has been performed regarding how to automatically and efficiently generate soft multiprocessors. In this paper, an automated parallel compilation environment for multiple soft processors which incorporates parallel compilation and inter-processor communication structures is described. A total of eight previously-developed parallel processing benchmarks have been automatically mapped to a varying number of synthesized soft microprocessors in commercial FPGAs. The new automated infrastructure allows for an evaluation of area, performance, and power tradeoffs for a range of architectural choices. Experiments show that our soft-multiprocessor systems consisting of up to 16 processors can offer up to 5× improvement in application performance against their uniprocessor counterparts.
Deepak Unnikrishnan, Russell Tessier
FCCM3
2009 CMOS vs Nano: comrades or rivals?
abstract
No abstract available.
Deming Chen, Russell Tessier, Kaustav Banerjee, Mojy C. Chian, André DeHon, Shinobu Fujita, James Hutchby, Steven Trimberger
FPGA2
2009 Design of a Secure Router System for Next-Generation Networks
abstract
Computer networks are vulnerable to attacks, where the network infrastructure itself is targeted. Emerging router designs, which use software-programmable embedded processors, increase the vulnerability to such attacks. We present the design of a secure packet processing platform (SPPP) that can protect these router systems. We use an instruction-level monitoring system to detect deviations in processing behavior. If such deviations are detected, a recovery system is invoked to restore the system into an operational state. Our preliminary results show that most attacks can be detected within a single instruction. The system overhead for secure monitoring is limited to a fraction of the overall space, memory, and power budget.
Tilman Wolf, Russell Tessier
NSS2
2009 Tetris-XL: A performance-driven spill reduction technique for embedded VLIW processors
abstract
As technology has advanced, the application space of Very Long Instruction Word (VLIW) processors has grown to include a variety of embedded platforms. Due to cost and power consumption constraints, many embedded VLIW processors contain limited resources, including registers. As a result, a VLIW compiler that maximizes instruction level parallelism (ILP) without considering register constraints may generate excessive register spills, leading to reduced overall system performance. To address this issue, this article presents a new spill reduction technique that improves VLIW runtime performance by reordering operations prior to register allocation and instruction scheduling. Unlike earlier algorithms, our approach explicitly considers both register reduction and data dependency in performing operation reordering. Data dependency control limits unexpected schedule length increases during subsequent instruction scheduling. Our technique has been evaluated using Trimaran, an academic VLIW compiler, and evaluated using a set of embedded systems benchmarks. Experimental results show that, on average, this technique improves VLIW performance by 10% for VLIW processors with 32 registers and 8 functional units compared with previous spill reduction techniques. Limited improvement is seen versus prior approaches for VLIW processors with 64 registers and 8 functional units.
Russell Tessier
ACM Trans. Archit. Code Optim.2
2008 Memory security management for reconfigurable embedded systems
abstract
The constrained operating environments of many FPGA-based embedded systems require flexible security that can be configured to minimize the impact on FPGA area and power consumption. In this paper, a security approach for external memory in FPGA-based embedded systems that exploits FPGA configurability is presented. Our FPGA-based security core provides both confidentiality and integrity for data stored externally to an FPGA which is accessed by a processor on the FPGA chip. The benefits of our security core are demonstrated using four embedded applications implemented on a Stratix II device. Each application requires a collection of tasks with varying memory security requirements. Our security core is used in conjunction with a NIOS II soft processor running the MicroC/OS II operating system. An average memory and energy savings of about 64%and 16%, respectively, is achieved for the four applications versus a non-configurable, uniform security approach.
Romain Vaslin, Guy Gogniat, Jean-Philippe Diguet, Russell Tessier, Deepak Unnikrishnan, Kris Gaj
FPT4
2007 Establishing Chain of Trust in Reconfigurable Hardware
abstract
Facing ubiquitous threats like computer viruses, trojans and theft of intellectual property, Trusted computing (TC) is an emerging technology towards building trustworthy computing platforms. A recent initiative by the trusted computing group (TCG) specifies the use of trusted platform modules (TPM), currently implemented as dedicated, cost-effective crypto-chips mounted on the main board of computer systems. In this paper we propose implementations for TC functionalities based on more flexible and versatile approaches for reconfigurable and embedded architectures. Our approach allows for (i) a scalable design and update of TPM functionalities in embedded systems, (ii) the integration of the TPM hardware in the chain of trust to bind applications to the underlying TPM and the reconfigurable hardware, and (iii) the design of vendor independent TPMs.
Thomas Eisenbarth 0001, Tim Güneysu, Christof Paar, Ahmad-Reza Sadeghi, Marko Wolf, Russell Tessier
FCCM6
2007 Power-aware FPGA logic synthesis using binary decision diagrams
abstract
Power consumption in field programmable gate arrays (FPGAs) has become an important issue as the FPGA market has grown to include mobile platforms. In this work we present a power-aware logic optimization tool that is specialized to facilitate subsequent power-aware technology mapping. Our synthesis framework uses binary decision diagram (BDD) based collapsing and decomposition techniques in conjunction with signal switching estimates to achieve power-efficient circuit networks. The results of synthesis and subsequent power-aware technology mapping are evaluated using two distinct physical design platforms: academic VPR and Altera Quartus II. Our approach achieves an average energy reduction of 13% for Altera Cyclone II devices versus synthesis with SIS-based algebraic optimization at the cost of 11% average circuit performance if performance-optimal technology mapping is performed after synthesis. If technology mapping is tuned to achieve the same average delay for both SIS and BDD-based flows, a 3% average energy reduction is achieved by our new synthesis approach.
Kevin Oo Tinmaung, David Howland, Russell Tessier
FPGA3
2007 Tetris: a new register pressure control technique for VLIW processors
abstract
113-122
Russell Tessier
LCTES2
2007 Power-Efficient RAM Mapping Algorithms for FPGA Embedded Memory Blocks
abstract
Contemporary field-programmable gate array (FPGA) design requires a spectrum of available physical resources. As FPGA logic capacity has grown, locally accessed FPGA embedded memory blocks have increased in importance. When targeting FPGAs, application designers often specify high-level memory functions, which exhibit a range of sizes and control structures. These logical memories must be mapped to FPGA embedded memory resources such that physical design objectives are met. In this paper, a set of power-efficient logical-to-physical RAM mapping algorithms is described, which converts user-defined memory specifications to on-chip FPGA memory block resources. These algorithms minimize RAM dynamic power by evaluating a range of possible embedded memory block mappings and selecting the most power-efficient choice. Our automated approach has been validated with both simulation of power dissipation and measurements of power dissipation on FPGA hardware. A comparison of measured power reductions to values determined via simulation confirms the accuracy of our simulation approach. Our power-aware RAM mapping algorithms have been integrated into a commercial FPGA compiler and tested with 34 large FPGA benchmarks. Through experimentation, we show that, on average, embedded memory dynamic power can be reduced by 26% and overall core dynamic power can be reduced by 6% with a minimal loss (1%) in design performance. In addition, it is shown that the availability of multiple embedded memory block sizes in an FPGA reduces embedded memory dynamic power by an additional 9.6% by giving more choices to the computer-aided design algorithms
Russell Tessier, Vaughn Betz, David Neto, Aaron Egier, Thiagaraja Gopalsamy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2006 An adaptive Reed-Solomon errors-and-erasures decoder
abstract
The development of Reed-Solomon (RS) codes has allowed for improved data transmission over a variety of communication media. Although Reed-Solomon decoding provides a powerful defense against burst data errors, the significant circuit area and power consumption of customized RS decoder hardware can be limiting for embedded computing environments. To support enhanced performance decoding with minimal power consumption, a dynamically-reconfigurable FPGA-based Reed-Solomon decoder has been developed. Our errors-and-erasures decoding system uses multiple erasure blocks to identify the location of likely corrupted data and multiple decoders to attempt error correction. The RS decoder design is implemented in reconfigurable hardware to leverage architectural parallelism and specialization. Run-time dynamic reconfiguration of the decoding system is used in response to variations in channel conditions to support the fastest possible data rate while, as a secondary metric, minimizing decoder power consumption. Algorithm parameters for the decoding system have been determined via simulation and the design has been implemented in Altera Stratix FPGAs. Through experimentation using an Altera 1S40 Stratix FPGA, we show that dynamic reconfiguration can result in an 14% performance improvement versus a non-reconfigurable decoder implementation. Comparisons with a Pentium IV microprocessor illustrate five orders of magnitude performance improvement.
Lilian Atieno, Jonathan Allen, Dennis Goeckel, Russell Tessier
FPGA4
2006 Power-aware RAM mapping for FPGA embedded memory blocks
abstract
Embedded memory blocks are important resources in contemporary FPGA devices. When targeting FPGAs, application designers often specify high-level memory functions which exhibit a range of sizes and control structures. These logical memories must be mapped to FPGA embedded memory resources such that physical design objectives are met. In this work a set of power-aware logical-to-physical RAM mapping algorithms are described which convert user-defined memory specifications to on-chip FPGA memory block resources. These algorithms minimize RAM dynamic power by evaluating a range of possible embedded memory block mappings and selecting the most power-efficient choice. Our automated approach has been integrated into a commercial FPGA compiler and tested with 40 large FPGA benchmarks. Through experimentation, we show that, on average, embedded memory dynamic power can be reduced by 21% and overall core dynamic power can be reduced by 7% with a minimal loss (1%) in design performance.
Russell Tessier, Vaughn Betz, David Neto, Thiagaraja Gopalsamy
FPGA1
2006 Design-specific path delay testing in lookup-table-based FPGAs
abstract
Due to the increased use of field-programmable gate arrays (FPGAs) in production circuits with high reliability requirements, the design-specific testing of FPGAs has become an important topic for research. Path delay testing of FPGAs is especially important since path delay faults can render an otherwise fault-free FPGA unusable for a given design layout. This paper presents a new approach for FPGA path delay testing, which partitions target paths into subsets that are tested in the same test configuration. Each path is tested for all combinations of signal inversions along the path length. Each configuration consists of a sequence generator, response analyzer, and circuitry for controlling inversions along tested paths, all of which are formed from FPGA resources not currently under test. Two algorithms are presented for target-path partitioning to determine the number of required test configurations. The test circuitry associated with these methods is also described. The results of applying the methods indicate that our path-delay-testing approach requires seconds per design to cover all paths with delay within 10% of the critical path delay. The approach has been validated using Xilinx Virtex devices.
Prem R. Menon, Russell Tessier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Salient features of radar nodes of the first generation NetRad System
abstract
The recently established National Science Foundation Engineering Research Center for Collaborative Adaptive Sensing of the Atmosphere (CASA) will be deploying the first generation of an automated network of four low-power, short-range, X-band, polarimetric, Doppler radars, known as NetRad, in central Oklahoma in late 2005. This network is developed with the goal of tracking tornadoes with high spatial and temporal resolution as well as mapping severe weather events in the lowest 2 km of the troposphere. Each radar node has been developed to accomplish this system goal through the coordinated interaction with other radars in the network via a real-time, closed-loop software control system. This paper will describe the characteristics of the individual radar nodes in the system, with emphasis on those aspects of the design that lend themselves toward operation as a coordinated network. Calibration results and performance characteristics of the single node radar of the first generation system will also be presented.
Francesc Junyent, V. Chandrasekar 0001, David McLaughlin, Stephen J. Frasier, Edin Insanic, Razi Ahmed, Nitin Bharadwaj, Eric J. Knapp, Luko Krnan, Russell Tessier
IGARSS10
2005 An energy-aware active smart card
abstract
Despite recent advances in smart card technology, most modern smart cards continue to rely on card readers for power and clocking, creating a potential security gap. In this paper, we present an energy-aware smart card architecture that operates using an embedded battery and crystal. This low-power VLSI system is continually active and provides enhanced security through periodic internal update when the card is detached from a reader. Our architecture achieves reduced power consumption by deactivating the majority of its circuitry, including an embedded microcontroller, for the vast majority of the card's lifetime. A proof-of-concept prototype implementation of the architecture has been developed including register-transfer-level and gate-level designs which have been synthesized to silicon. To permit extended operation for up to 18 months, critical design logic has been implemented using ultralow-power (adiabatic) circuit techniques.
Russell Tessier, David Jasinski, Atul Maheshwari, Aiyappan Natarajan, Wayne P. Burleson
IEEE Trans. Very Large Scale Integr. Syst.1
2005 A reconfigurable, power-efficient adaptive Viterbi decoder
abstract
Error-correcting convolutional codes provide a proven mechanism to limit the effects of noise in digital data transmission. Although hardware implementations of decoding algorithms, such as the Viterbi algorithm, have shown good noise tolerance for error-correcting codes, these implementations require an exponential increase in very large scale integration area and power consumption to achieve increased decoding accuracy. To achieve reduced decoder power consumption, we have examined and implemented decoders based on the reduced-complexity adaptive Viterbi algorithm (AVA). Run-time dynamic reconfiguration is performed in response to varying communication channel-noise conditions to match minimized power consumption to required error-correction capabilities. Experimental calculations indicate that the use of dynamic reconfiguration leads to a 69% reduction in decoder power consumption over a nonreconfigurable field-programmable gate array implementation with no loss of decode accuracy.
Russell Tessier, Sriram Swaminathan, Ramaswamy Ramaswamy, Dennis Goeckel, Wayne P. Burleson
IEEE Trans. Very Large Scale Integr. Syst.1
2004 A Dynamically-Reconfigurable, Power-Efficient Turbo Decoder
abstract
The development of turbo codes has allowed for near-Shannon limit information transfer in modern communication systems. Although turbo decoding is viewed as superior to alternate decoding techniques, the circuit complexity and power consumption of turbo decoder implementations can often be prohibitive for power-constrained systems. To address these issues, we have developed a reduced-complexity turbo decoder specifically optimized for contemporary FPGA devices. Our key power-saving technique is the use of decoder run-time dynamic reconfiguration in response to variations in the channel conditions. If less favorable channel conditions are detected, a more powerful, less power-efficient decoder is swapped into the FPGA hardware to maintain a fixed bit error rate. More favorable channel conditions result in the opposite effect. Through experimentation on a stratix-based NIOS development board, we show that dynamic reconfiguration can result in a 52% power reduction versus a static decoder implementation. Comparisons with contemporary microprocessors illustrate a 100/spl times/ performance improvement.
Russell Tessier, Dennis Goeckel
FCCM2
2004 A pod-based dual-beam SAR
abstract
A dual-beam along-track interferometric synthetic aperture radar that is entirely self-contained within an aircraft pod has been developed by the University of Massachusetts to study sea surface processes in coastal regions. The radar operates at 5.3 GHz with a bandwidth of up to 25 MHz. System hardware is described. Initial test flights aboard the National Oceanic and Atmospheric Administration's WP-3D research aircraft were performed to evaluate system performance over land and water surfaces. Imagery were collected for fore and aft squinted beams, though no interferometric data were collected. Notable look-angle dependences are observed in the sea surface normalized radar cross section under very low wind conditions.
Gordon Farquharson, William N. Junek, Arun Ramanathan, Stephen J. Frasier, Russell Tessier, David McLaughlin, Mark A. Sletten, Jakov V. Toporkov
IEEE Geosci. Remote. Sens. Lett.5
2004 An architecture and compiler for scalable on-chip communication
abstract
A dramatic increase in single chip capacity has led to a revolution in on-chip integration. Design reuse and ease of implementation have became important aspects of the design process. This paper describes a new scalable single-chip communication architecture for heterogeneous resources, adaptive system-on-a-chip (aSOC) and supporting software for application mapping. This architecture exhibits hardware simplicity and optimized support for compile-time scheduled communication. To illustrate the benefits of the architecture, four high-bandwidth signal processing applications including an MPEG-2 video encoder and a Doppler radar processor have been mapped to a prototype aSOC device using our design mapping technology. Through experimentation it is shown that aSOC communication outperforms a hierarchical bus-based system-on-chip (SoC) approach by up to a factor of five. A VLSI implementation of the communication architecture indicates clock rates of 400 MHz in 0.18-/spl mu/m technology for sustained on-chip communication. In comparison to previously-published results for an MPEG-2 decoder, our on-chip interconnect shows a runtime improvement of over a factor of four.
Andrew Laffely, Russell Tessier
IEEE Trans. Very Large Scale Integr. Syst.4
2004 Trading off transient fault tolerance and power consumption in deep submicron (DSM) VLSI circuits
abstract
High fault tolerance for transient faults and low-power consumption are key objectives in the design of critical embedded systems. Systems like smart cards, PDAs, wearable computers, pacemakers, defibrillators, and other electronic gadgets must not only be designed for fault tolerance but also for ultra-low-power consumption due to limited battery life. In this paper, a highly accurate method of estimating fault tolerance in terms of mean time to failure (MTTF) is presented. The estimation is based on circuit-level simulations (HSPICE) and uses a double exponential current-source fault model. Using counters, it is shown that the transient fault tolerance and power dissipation of low-power circuits are at odds and allow for a power fault-tolerance tradeoff. Architecture and circuit level fault tolerance and low-power techniques are used to demonstrate and quantify this tradeoff. Estimates show that incorporation of these techniques results either in a design with an MTTF of 36 years and power consumption of 102 /spl mu/W or a design with an MTTF of 12 years and power consumption of 20 /spl mu/W. Depending on the criticality of the system and the power budget, certain techniques might be preferred over others, resulting in either a more fault tolerant or a lower power design, at the sacrifice of the alternative objective.
Atul Maheshwari, Wayne P. Burleson, Russell Tessier
IEEE Trans. Very Large Scale Integr. Syst.3
2003 Floating Point Unit Generation and Evaluation for FPGAs
abstract
Most commercial and academic floating point libraries for FPGAs (field programmable gate arrays) provide only a small fraction of all possible floating point units. In contrast, the floating point unit generation approach outlined in this paper allows for the creation of a vast collection of floating point units with differing throughput, latency, and area characteristics. Given performance requirements, our generation tool automatically chooses the proper implementation algorithm and architecture to create a compliant floating point unit. Our approach is fully integrated into standard C++ using ASC, a stream compiler for FPGAs, and the PAM-Blox II module generation environment. The floating point units created by our approach exhibit a factor of two latency improvement versus commercial FPGA floating point units, while consuming only half of the FPGA logic area.
Russell Tessier, Oskar Mencer
FCCM2
2003 Adaptive Fault Recovery for Networked Reconfigurable Systems
abstract
The device-level size and complexity of reconfigurable architectures makes fault tolerance an important concern in system design. In this paper, we introduce a fully automated fault recovery system for networked systems, which contain FPGAs (field programmable gate arrays). If a fault is detected hat cannot be addressed locally, fault information is transferred to a reconfiguration server. Following design recompilation to avoid the fault, a new FPGA configuration is returned to the remote system and computation is reinitiated. To illustrate the benefit of this approach, we have implemented a complete fault recovery system, which requires no manual intervention. An important part of the system is a timing-driven incremental router for Xilinx Virtex devices. This router is directly interfaced to Xilinx JBits and uses no CAD tools from the standard Xilinx Alliance tool flow. Our completed system has been applied to three benchmark designs and exhibits complete fault recovery in up to 12x less time than the standard incremental Xilinx PAR flow.
Ramshankar Ramanarayanan, Russell Tessier
FCCM3
2003 A hybrid adiabatic content addressable memory for ultra low-power applications
abstract
This paper presents a hybrid adiabatic content addressable memory (CAM). The CAM uses an adiabatic switching technique to reduce the energy consumption in the match line while keeping the performance for the read/write operation. The adiabatic CAM is suitable for ultra low-power, low performance applications such as smart cards and portable devices. This CAM uses a clocked power supply for the match line while the rest of the circuit is the same as the basic CAM. A novel smart card application which uses the adiabatic CAM is illustrated. The circuit simulations for a 16x16 and 32x32 CAM were done in Hspice using 0.18 μm Berkeley models and the energy dissipation was compared with a basic CAM. The results show three orders of magnitude in energy savings for the 16x16 CAM and one order of magnitude savings for the 32x32 CAM when operated at 2Mhz. The maximum frequency of operation for which there was considerable energy savings was found to be 200 Mhz with a 20% and 45% energy savings for 16x16 and 32x32 CAM respectively.
Aiyappan Natarajan, David Jasinski, Wayne P. Burleson, Russell Tessier
ACM Great Lakes Symposium on VLSI4
2003 Adaptive system on a chip (ASOC): a backbone for power-aware signal processing cores
abstract
For motion estimation (ME) and discrete cosine transform (DCT) of MPEG video encoding, content variation and perceptual tolerance in video signals can be exploited to gracefully trade quality for low power. As a result, power-aware hardware cores have been proposed for these video encoding subsystems. Adaptive system-on-a-chip, aSoC, supports power-aware cores by providing an on-chip communications framework designed to promote scalability and flexibility in system-on-a-chip designs. This paper describes aSoC's ability to dynamically control voltage and frequency scaling through a simple voltage and frequency selection scheme. A small demonstration system is tested and shows up to 90% reduction in core power when the aSoC voltage scaling features are enabled.
Andrew Laffely, Russell Tessier, Wayne P. Burleson
ICIP (3)3
2003 First observations with the UMass dual-beam InSAR
abstract
A dual-beam along-track interferometric synthetic aperture radar which is self contained within an aircraft pod has been developed to study coastal regions. System hardware is described. Initial test flights aboard the NOAA WP-3D research aircraft were performed to evaluate system performance over land and water surfaces. Notable look-angle dependencies are observed in the sea surface NRCS under very low wind conditions.
William N. Junek, Arun Ramanathan, Gordon Farquharson, Stephen J. Frasier, Russell Tessier, David McLaughlin, Mark A. Sletten, Jakov V. Toporkov
IGARSS5
2003 Technology mapping algorithms for hybrid FPGAs containing lookup tables and PLAs
abstract
Programmable devices containing lookup tables (LUTs) and programmable logic arrays (PLAs) provide a heterogeneous target platform for user designs. Present commercial tools, which target these hybrid devices, require hand partitioning of user designs to isolate logic for each type of logic resource. In this paper, an automated technology mapping tool, hybridmap , is presented that identifies design logic partitions as suitable for either LUT or PLA implementation. A breadth-first search-based subgraph extraction and evaluation heuristic is integrated with product term (Pterm) count, area, and delay estimators to guide the technology mapping process. Hybridmap can be adapted to target a variety of PLA architectures and can accommodate user-provided timing constraints. It is shown that when timing constrained, hybridmap reduces LUT consumption for Apex20KE devices by 8% and when unconstrained by 14% by migrating logic from LUTs to Pterm structures. Hybridmap is shown to outperform previous mapping approaches for Apex20KE-type devices by up to 22%.
Srini Krishnamoorthy, Russell Tessier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 A dynamically reconfigurable adaptive viterbi decoder
abstract
The use of error-correcting codes has proven to be an effective way to overcome data corruption in digital communication channels. Although widely-used, the most popular communications decoding algorithm, the Viterbi algorithm, requires an exponential increase in hardware complexity to achieve greater decode accuracy. In this paper, we describe the analysis and implementation of a reduced-complexity decode approach, the adaptive Viterbi algorithm (AVA). Our AVA design is implemented in reconfigurable hardware to take full advantage of algorithm parallelism and specialization. Run-time dynamic reconfiguration is used in response to changing channel noise conditions to achieve improved decoder performance. Implementation parameters for the decoder have been determined through simulation and the decoder has been implemented on a Xilinx XC4036-based PCI board. An overall decode performance improvement of 7.5X for AVA has been achieved versus algorithm implementation on a Celeron-processor based system. The use of dynamic reconfiguration leads to a 20% performance improvement over a static implementation with no loss of decode accuracy.
Sriram Swaminathan, Russell Tessier, Dennis Goeckel, Wayne P. Burleson
FPGA2
2002 The Integration of SystemC and Hardware-Assisted Verification
Ramaswamy Ramaswamy, Russell Tessier
FPL2
2002 Testing and diagnosis of interconnect faults in cluster-based FPGA architectures
abstract
As IC densities are increasing, cluster-based field programmable gate arrays (FPGA) architectures are becoming the architecture of choice for major FPGA manufacturers. A cluster-base architecture is one in which several logic blocks are grouped together into a coarse-grained logic block. While the high-density local interconnect often found within clusters serves to improve FPGA utilization, it also greatly complicates the FPGA interconnect testing problem. To address this issue, we have developed a hierarchical approach to define a set of FPGA configurations which enable interconnect fault detection and diagnosis. This technique enables the detection of bridging faults involving intracluster interconnect and extracluster interconnect. The hierarchical structure of a cluster-based tile is exploited to define intracluster configurations separately from extracluster configurations, thereby improving the efficiency of the configuration definition process. The cornerstone of this work is the concise expression of the detectability conditions of each fault and the distinguishability conditions of each fault pair. By guaranteeing that both intracluster and extracluster configurations have several test transparency properties, hierarchical fault detectability is ensured.
Ian G. Harris, Russell Tessier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Static scheduling of multidomain circuits for fast functional verification
abstract
With the advent of system-on-a-chip design, many application specific integrated circuits (ASICs) now require multiple design clocks that operate asynchronously to each other. This design characteristic presents a significant challenge when these ASIC designs are mapped to parallel verification hardware such as parallel cycle-based simulators and logic emulators. In general, these systems require all computation and communication to be synchronized to a global system clock. As a result, the undefined relationship between design clocks can make it difficult to determine hold times for synchronous storage elements. and causality relationships along reconvergent communication paths. This paper presents new scheduling and synchronization techniques to support accurate mapping of designs with multiple asynchronous clocks to parallel verification hardware. Through analysis, it is shown that this approach is scalable to an unlimited number of domains and supports increasingly large design sizes. To prove the effectiveness of the authors' approach, developed algorithms have been integrated into the compilation system for a commercial multi-FPGA logic emulation system. For three designs mapped to a logic emulator using this software environment, modeling fidelity is maintained and performance is enhanced versus previous manual mapping approaches. A theoretical analysis based on Rent's rule validates the scalability of the approach as device sizes increase.
Murali Kudlugi, Russell Tessier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Fast placement approaches for FPGAs
abstract
Recent trends in FPGA development indicate a strong shift toward design reuse through the use of intellectual property (IP). This design shift has motivated the development of Frontier, a timing-driven FPGA placement system that uses design macroblocks in conjunction with a series of placement algorithms to achieve highly routable and high-performance layouts quickly. In the first stage of design placement, a macro-based floorplanner is used to quickly identify an initial layout based on intermacro connectivity. Next, FPGA routability and performance metrics are used to evaluate the quality of the initial placement. Finally, if the floorplan is determined to be insufficient from a routability or performance standpoint, a feedback-driven placement perturbation step is employed to achieve a lower cost placement. For a collection of large reconfigurable computing benchmark circuits our timing-driven placement system exhibits a 2.6× speedup in combined place and route time versus commercial FPGA CAD software with improved design performance for most designs. It is shown that floorplanning, placement evaluation, and backend optimization are all necessary to achieve high-performance placement solutions.
Russell Tessier
ACM Trans. Design Autom. Electr. Syst.1
2002 BDD-based logic synthesis for LUT-based FPGAs
abstract
Contemporary FPGA synthesis is a multiphase process that involves technology-independent logic optimization followed by FPGA-specific mapping to a target FPGA technology. Conventional technology-independent transformations target standard cells and are unable to optimize circuits with constraints and goals specific to FPGA architectures. This article describes an FPGA-specific logic synthesis approach, which unites multilevel logic transformation, decomposition, and optimization techniques into a single synthesis framework. This system performs network transformation, decomposition, and optimization at an early stage to generate a network that can be directly mapped onto FPGAs. Our techniques are built upon a BDD-based logic decomposition system. With this system, both AND-OR decompositions and AND-XOR decompositions can be identified, resulting in large area savings for synthesized XOR-intensive circuits. To induce good decompositions, a maximum fanout free cone (MFFC) -based partial clustering and collapsing technique is used. This step is followed by an area-minimizing variable partitioning heuristic that decomposes collapsed nodes into LUT-feasible subfunctions. As a postprocessing step, a performance-driven resynthesis phase is performed to alleviate increased delay caused by excessive logic sharing. We compare the quality of results obtained using our techniques with those of academic (BoolMap, SIS) and industry (Altera Quartus) FPGA synthesis tools. Experimental results indicate that the circuits generated by our techniques are not only smaller, but are also significantly faster than those synthesized by conventional FPGA synthesis tools. Furthermore, the computation times required by our techniques are significantly smaller than those of previous techniques.
Navin Vemuri, Priyank Kalla, Russell Tessier
ACM Trans. Design Autom. Electr. Syst.3
2002 Incremental compilation for parallel logic verification systems
abstract
Although simulation remains an important part of application-specific integrated circuit (ASIC) validation, hardware-assisted parallel verification is becoming a larger part of the overall ASIC verification flow. In this paper, we describe and analyze a set of incremental compilation steps that can be directly applied to a range of parallel logic verification hardware, including logic emulators. Important aspects of this work include the formulation and analysis of two incremental design mapping steps: the partitioning of newly added design logic onto multiple logic processors and the communication scheduling of newly added design signals between logic processors. To validate our incremental compilation techniques, the developed mapping heuristics have been integrated into the compilation flow for a field-programmable gate-array-based Ikos VirtuaLogic emulator . The modified compiler has been applied to five large benchmark circuits that have been synthesized from register-transfer level and mapped to the emulator. It is shown that our incremental approach reduces verification compile time for modified designs by up to a factor of five versus complete design recompilation for benchmarks of over 100 000 gates. In most cases, verification run-time following incremental compilation of a modified design matches the performance achieved with complete design recompilation.
Russell Tessier, Snigdha Jana
IEEE Trans. Very Large Scale Integr. Syst.1
2001 Static Scheduling of Multiple Asynchronous Domains For Functional Verification
abstract
While ASIC devices of a decade ago primarily contained synchro-nous circuitry triggered with a single clock, many contemporary architectures require multiple clocks that operate asynchronously to each other. This multi-clock domain behavior presents significant functional verification challenges for large parallel verification sys-tems such as distributed parallel simulators and logic emulators. In particular, multiple asynchronous design clocks make it difficult to verify that design hold times are met during logic evaluation and causality along reconvergent fanout paths is preserved during signal communication. In this paper, we describe scheduling and synchro-nization techniques to maintain modeling fidelity for designs with multiple asynchronous clock domains that are mapped to parallel verification systems. It is shown that when our approach is applied to an FPGA-based logic emulator, evaluation fidelity is maintained and increased design evaluation performance can be achieved for large benchmark designs with multiple asynchronous clock domains.
Murali Kudlugi, Charles Selvidge, Russell Tessier
DAC3
2001 Dynamically parameterized algorithms and architectures to exploit signal variations for improved performance and reduced power
abstract
Signal processing algorithms and architectures can use dynamic reconfiguration to exploit variations in signal statistics with the objectives of improved performance and reduced power consumption. Parameters provide a simple and formal way to characterize incremental changes to a computation and its computing mechanism. This paper examines five parameterized computations which are typically implemented in hardware for a wireless multimedia terminal: (1) motion estimation, (2) discrete cosine transform, (3) Lempel-Ziv lossless compression, (4) 3D graphics light rendering and (5) Viterbi decoding. Each computation is examined for the capability of dynamically adapting the algorithm and architecture parameters to variations in their respective input signals. Dynamically reconfigurable low-power implementations of each computation are currently underway.
Wayne P. Burleson, Russell Tessier, Dennis Goeckel, Sriram Swaminathan, Prashant Jain, Jeongseon Euh, Subramanian Venkatraman, Vidhya Thyagarajan
ICASSP2
2001 Static Scheduling of Multi-Domain Memories For Functional Verification
abstract
The presence of multiple clock domains presents significant challenges for large parallel verification systems such as parallel simulators and logic emulators that model both design logic and memory. Specifically, multiple asynchronous design clocks make it difficult to verify that design hold times are met during memory model execution and causality along memory data/control paths is preserved during signal communication. We describe new scheduling heuristics for memory-based designs with multiple asynchronous clock domains that are mapped to parallel verification systems. The scheduling approach scales to an unlimited number of clock domains and converges quickly to a feasible solution if one exists. It is shown that when the technique is applied to an FPGA-based emulator containing 48MB of SRAM, evaluation fidelity. is maintained and increased verification performance is achieved for large, memory-intensive circuits with multiple asynchronous clock domains.
Murali Kudlugi, Charles Selvidge, Russell Tessier
ICCAD3
2001 BIST-based delay path testing in FPGA architectures
abstract
The widespread use of field programmable gate arrays (FPGAs) as components in high-performance systems has increased the significance of path delay faults in FPGAs. We present a technique for FPGA path delay fault detection which integrates test insertion with the FPGA placement and routing stages to accomplish testing with low test application time. An accurate static timing analyzer is used to identify critical paths and built-in self-test (BIST) hardware is inserted using a placement and routing tool. Initial experimental results show that testing is accomplished with low test application time for several benchmark designs.
Ian G. Harris, Prem R. Menon, Russell Tessier
ITC3
2000 Interconnect testing in cluster-based FPGA architectures
abstract
As IC densities are increasing, cluster-based FPGA architectures are becoming the architecture of choice for major FPGA manufacturers. A cluster-based architecture is one in which several logic blocks are grouped together into a coarse-grained logic block. While the high density local interconnect often found within clusters serves to improve FPGA utilization, it also greatly complicates the FPGA interconnect testing problem. To address this issue, we have developed a hierarchical approach to define a set of FPGA configurations which enable interconnect faults to be detected. This technique enables the detection of bridging faults involving intra-cluster interconnect and extra-cluster interconnect. The hierarchical structure of a cluster-based tile is exploited to define intra-cluster configurations separately from extra-cluster configurations, thereby improving the efficiency of the configuration definition process. By guaranteeing that both intra-cluster and extra-cluster configurations have several test transparency properties, hierarchical fault detectability is ensured.
Ian G. Harris, Russell Tessier
DAC2
2000 Tolerating operational faults in cluster-based FPGAs
abstract
In recent years the application space of reconfigurable devices has grown to include many platforms with a strong need for fault tolerance. While these systems frequently contain hardware redundancy to allow for continued operation in the presence of operational faults, the need to recover faulty hardware and return it to full functionality quickly and efficiently is great. In addition to providing functional density, FPGAs provide a level of fault tolerance generally not found in mask-programmable devices by including the capability to reconfigure around operational faults in the field. In this paper, incremental CAD techniques are described that allow functional recovery of FPGA design configurations in the presence of single or multiple operational faults. Our preferred approach to fault recovery takes advantage of device routing hierarchy in architectural families such as Xilinx Virtex [2] and Altera Apex [3] to quickly swap unused logic and routing resources in place of faulty ones within logic clusters. These algorithms allow for straight-forward implementation within a local fault-tolerant system without the need to access a remote processing location. If initial recovery attempts through localized swapping fail, an incremental router based on the widely-used PathFinder maze routing algorithm [10] can be applied remotely in an attempt to form connections between newly-allocated logic and interconnect based on the history of the initial design route.
Vijay Lakamraju, Russell Tessier
FPGA2
2000 Diagnosis of Interconnect Faults in Cluster-Based FPGA Architectures
abstract
Fault diagnosis has particular importance in the context of field programmable gate arrays (FPGAs) because faults can be avoided by reconfiguration at almost no real cost. Cluster-based FPGA architectures, in which several logic blocks are grouped together into a coarse-grained logic block, are rapidly becoming the architecture of choice for major FPGA manufacturers. The high density interconnect found within clusters greatly complicates the problem of FPGA diagnosis. We propose a technique for the testing and diagnosis of cluster-based FPGA architectures. We present a hierarchical approach to define a set of FPGA configurations in which each fault is detectable, and each fault pair is differentiable. The cornerstone of this work is the concise expression of the distinguishing conditions of each fault pair. Experimental results demonstrate that nearly 100% fault coverage and diagnostic resolution are achieved with a low number of test configurations.
Ian G. Harris, Russell Tessier
ICCAD2
1997 Logic emulation with virtual wires
abstract
Logic emulation enables designers to functionally verify complex integrated circuits prior to chip fabrication. However, traditional FPGA-based logic emulators have poor inter-chip communication bandwidth, commonly limiting gate utilization to less than 20%. Global routing contention mandates the use of expensive crossbar and PC-board technology in a system of otherwise low-cost commodity parts. Even with crossbar technology, current emulators only use a fraction of potential communication bandwidth because they dedicate each FPGA pin (physical wire) to a single emulated signal (logical wire). Virtual wires overcome pin limitations by intelligently multiplexing each physical wire among multiple logical wires, and pipelining these connections at the maximum clocking frequency of the FPGA. The resulting increase in bandwidth allows effective use of low-dimension direct interconnect. The size of the FPGA array can be decreased as well, resulting in low-cost logic emulation. This paper covers major contributions of the MIT Virtual Wires project. In the context of a complete emulation system, we analyze phase-based static scheduling and routing algorithms, present virtual wires synthesis methodologies, and overview an operational prototype with 20 K-gate boards. Results, including in-circuit emulation of a SPARC microprocessor, indicate that virtual wires eliminate the need for expensive crossbar technology while increasing FPGA utilization beyond 45%. Theoretical analysis predicts that virtual wires emulation scales with FPGA size and average routing distance, while traditional emulation does not.
Jonathan Babb, Russell Tessier, Matthew Dahl, Silvina Hanono, David M. Hoki, Anant Agarwal
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1993 The NuMesh: A Modular, Scalable Communications Substrate
abstract
Many standardized hardware communication interfaces offer runtime flexibility and configurability at the cost of efficiency.An alternate approach is the use of a highly-efficien~minimal communication element with as much communication decision-making as possible done at compile time.NuMesh is a packaging and interconnect technology supporting high-bandwidth systolic communications on a 3D nearest-neighbor lattice; our goal is to combine Lego-like modularity with supercomputer performance.To date, the primary focus of the project has been the class of applications whose static communication patterns can be precompiled into independent and carefully choreographed finite state machines running on each node.Several extensions of the NuMesh to more general communication paradigms have been implemented, and the issues involved are under active exploration.This paper presents an overview of our approach, as well as an introduction to our current-generation prototype.We also discuss our software environment and simulation technology, and enumerate some of the applications and programming models we have developed to make full use of the capabdities of the NuMesh.
Steve Ward, Karim Abdalla, Rajeev Dujari, Michael Fetterman, Frank Honoré, Ricardo Jenez, Philippe Laffont, Kenneth Mackenzie, Chris Metcalf, Milan Minsky, John Nguyen, John Pezaris, Gill A. Pratt, Russell Tessier
International Conference on Supercomputing14