Roger F. Woods

dblp:w/RogerWoods · DBLP profile ↗
← Back
67ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-6201-4270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-authorComputer networks · 7 · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Predictive Quality of In-Fabrication Products in Smart Manufacturing Using Graph-Based Deep Learning
Peter Davison, Muhammad Fahim, Roger F. Woods, Scott Fischaber, Marcus Haron, Cormac McAteer
ICINCO (1)3
2025 Efficient Integer-Only-Inference of Gradient Boosting Decision Trees on Low-Power Devices
abstract
There is increasingly interest in developing embedded machine learning hardware as it can offer better performance in terms of privacy, bandwidth efficiency, and scalability. Gradient-boosted decision trees (GBDT) represent a strong candidate as they employ less complex logic, but their efficient implementation in field programmable gate array (FPGA) needs to be explored in detail. In this paper, we propose sophisticated quantisation approaches to balance the dual goals of efficiency and performance. In particular, we introduce quantisation-aware training of GBDT for integer-only and binary arithmetic. Results are presented for implementations on a Zynq UltraScale+ MPSoC FPGA with the best design using only 170 Look-up Tables and 233 flip-flops at a clock speed of 724 MHz. Implementations focused on network intrusion detection and jet substructure classification for large-scale physics experiments are explored. An order of magnitude less FPGA resources are used whilst offering extremely high throughput rate and maintaining accuracy. Code is available athttps://github.com/malsharari/QATGBDT.
Majed Alsharari, Son T. Mai, Roger F. Woods, Carlos Reaño
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Adaptive approximate computing in edge AI and IoT applications: A review
abstract
Recent advancements in hardware and software systems have been driven by the deployment of emerging smart health and mobility applications. These developments have modernized the traditional approaches by replacing conventional computing systems with cyber-physical and intelligent systems combining the Internet of Things (IoT) with Edge Artificial Intelligence. Despite the many advantages and opportunities of these systems within various application domains, the scarcity of energy, extensive computing needs, and limited communication must be considered when orchestrating their deployment. Inducing savings in these directions is central to the Approximate Computing (AxC) paradigm, in which the accuracy of some operations is traded off with energy, latency, and/or communication reductions. Unfortunately, the dynamics of the environments in which AxC-equipped IoT systems operate have been paid little attention. We bridge this gap by surveying adaptive AxC techniques applied to three emerging application domains, namely autonomous driving, smart sensing and wearables, and positioning, paying special attention to hardware acceleration. We discuss the challenges of such applications, how adaptive AxC can aid their deployment, and which savings it can bring based on traits of the data and devices involved. Insights arising thereof may serve as inspiration to researchers, engineers, and students active within the considered domains.
Hans Jakob Damsgaard, Antoine Grenier, Dewant Katare, Zain Taufique, Salar Shakibhamedan, Tiago Troccoli, Georgios Chatzitsompanis, Anil Kanduri, Aleksandr Ometov, Aaron Yi Ding, Nima Taherinejad, Georgios Karakonstantis, Roger F. Woods, Jari Nurmi
J. Syst. Archit.13
2024 FPAX: A Fast Prior Knowledge-Based Framework for DSE in Approximate Configurations
abstract
Current artificial intelligence and data science applications typically require complex computations and massive amounts of data handling, presenting unprecedented challenges for embedded platforms. Approximate computing has emerged as the most promising design technique to address this issue, by providing a potential performance increase, while sacrificing accuracy within an acceptable range. Approximate arithmetic units require the creation of design space exploration techniques that can swiftly and automatically form an approximate configuration in fault-tolerant systems. Existing methods, however, use iterative design space sampling, resulting in a large amount of redundant computation. In this work, we propose the efficient FPAX automatic search framework which can learn from prior knowledge regarding the exploration process of known applications and use it to guide design exploration. This avoids excessive redundant computation and quickly provides an impressive approximate configuration. Compared with the Jump Search algorithm known for its efficiency, FPAX can also achieve faster convergence speed and better exploration quality. Even compared to our previous ENAP framework, it exhibits an 18x faster performance while achieving almost identical exploration quality for several commonly used fault-tolerant applications.
Yuqin Dou, Chenghua Wang, Haroon Waris, Roger F. Woods, Weiqiang Liu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 CACTUS: A Comprehensive Abstraction and Classification Tool for Uncovering Structures
abstract
The availability of large datasets is providing the impetus for driving many current artificial intelligent developments. However, specific challenges arise in developing solutions that exploit small datasets, mainly due to practical and cost-effective deployment issues, as well as the opacity of deep learning models. To address this, the Comprehensive Abstraction and Classification Tool for Uncovering Structures (CACTUS) is presented as a means of improving secure analytics by effectively employing explainable artificial intelligence. CACTUS achieves this by providing additional support for categorical attributes, preserving their original meaning, optimising memory usage, and speeding up the computation through parallelisation. It exposes to the user the frequency of the attributes in each class and ranks them by their discriminative power. Performance is assessed by applying it to various domains, including Wisconsin Diagnostic Breast Cancer, Thyroid0387, Mushroom, Cleveland Heart Disease, and Adult Income datasets.
Luca Gherardini, Varun Ravi Varma, Karol Capala, Roger F. Woods, José L. R. Sousa
ACM Trans. Intell. Syst. Technol.4
2024 Towards Receiver-Agnostic and Collaborative Radio Frequency Fingerprint Identification
abstract
Radio frequency fingerprint identification (RFFI) is an emerging device authentication technique, which exploits the hardware characteristics of the RF front-end as device identifiers. The receiver hardware impairments interfere with the feature extraction of transmitter impairments, but their effect and mitigation have not been comprehensively studied. In this paper, we propose a receiver-agnostic RFFI system by employing adversarial training to learn the receiver-independent features. Moreover, when there are multiple receivers, collaborative inference are designed to enhance classification accuracy. Finally, we show how it is possible to leverage fine-tuning for further improvement with fewer collected signals. To validate the approach, we have conducted extensive experimental evaluation by applying the approach to a LoRaWAN case study involving ten LoRa devices and 20 software-defined radio (SDR) receivers. The results show that receiver-agnostic training enables the trained neural network to become robust to changes in receiver characteristics. The collaborative inference improves classification accuracy by up to 20% beyond a single-receiver RFFI system and fine-tuning can bring a 40% improvement for underperforming receivers. The system is further evaluated on a more practical testbed. By making additional use of online augmentation and multi-packet inference, the identification accuracy is improved from 50% to 90% at 10 dB.
Guanxiong Shen, Junqing Zhang, Alan Marshall 0001, Roger F. Woods, Joseph R. Cavallaro, Liquan Chen
IEEE Trans. Mob. Comput.4
2023 ENAP: An Efficient Number-Aware Pruning Framework for Design Space Exploration of Approximate Configurations
abstract
Approximate computing has emerged as a new computing architecture paradigm that trades off necessary numerical accuracy for performance. Various approximation operation units such as adders and multipliers have been created and provide the basis for improving system efficiency, but it is clear, that a design space exploration (DSE) is needed if improved performance is to be systematically achieved. The challenge is to determine a suitable configuration among approximation units with different error characteristics to ensure a minimization of resources while not exceeding user-defined error constraints. In this paper, we propose the efficient number-aware pruning (ENAP) technique that can compress the search space size. Using common fault-tolerant applications, we demonstrate a compression rate up to 0.0008%, meaning that 99.9992% of invalid designs can remain unsearched. An improved genetic algorithm (GA) is subsequently proposed to improve ENAP, allowing the creation of the optimal configuration in only 2 to 3 iterations, thereby greatly improving search efficiency compared to the initial 9 iterations. We integrate these two approaches into the proposed framework, demonstrating how we can achieve better exploration results compared to state-of-the-artwork.
Yuqin Dou, Chenghua Wang, Roger F. Woods, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Efficient, Dynamic Multi-Task Execution on FPGA-Based Computing Systems
abstract
With growing Field Programmable Gate Array (FPGA) device sizes and their integration in environments enabling sharing of computing resources such as cloud and edge computing, there is a requirement to share the FPGA area between multiple tasks. The resource sharing typically involves partitioning the FPGA space into fix-sized slots. This results in suboptimal resource utilisation and relatively poor performance, particularly as the number of tasks increase. Using OpenCL's exploration capabilities, we employ clever clustering and custom, task-specific partitioning and mapping to create a novel, area sharing methodology where task resource requirements are more effectively managed. Using models with varying resource/throughput profiles, we select the most appropriate distribution based on the runtime, workload needs to enhance temporal compute density. The approach is enabled in the system stack by a corresponding task-based virtualisation model. Using 11 high performance tasks from graph analysis, linear algebra and media streaming, we demonstrate an average 2.8× higher system throughput at 2.3× better energy efficiency over existing approaches.
Umar Ibrahim Minhas, Roger F. Woods, Dimitrios S. Nikolopoulos, Georgios Karakonstantis
IEEE Trans. Parallel Distributed Syst.2
2021 TOD: Transprecise Object Detection to Maximise Real-Time Accuracy on the Edge
abstract
Real-time video analytics on the edge is challenging as the computationally constrained resources typically cannot analyse video streams at full fidelity and frame rate, which results in loss of accuracy. This paper proposes a Transprecise Object Detector (TOD) which maximises the real-time object detection accuracy on an edge device by selecting an appropriate Deep Neural Network (DNN) on the fly with negligible computational overhead. TOD makes two key contributions over the state of the art: (1) TOD leverages characteristics of the video stream such as object size and speed of movement to identify networks with high prediction accuracy for the current frames; (2) it selects the best-performing network based on projected accuracy and computational demand using an effective and low-overhead decision mechanism. Experimental evaluation on a Jetson Nano demonstrates that TOD improves the average object detection precision by 34.7 % over the YOLOv4-tiny-288 model on average over the MOT17Det dataset. In the MOT17-05 test dataset, TOD utilises only 45.1 % of GPU resource and 62.7 % of the GPU board power without losing accuracy, compared to YOLOv4-416 model. We expect that TOD will maximise the application of edge devices to real-time object detection, since TOD maximises real-time object detection accuracy given edge devices according to dynamic input features without increasing inference latency in practice.
Blesson Varghese, Roger F. Woods, Hans Vandierendonck
ICFEC3
2021 Radio Frequency Fingerprint Identification for Narrowband Systems, Modelling and Classification
abstract
Device authentication is essential for securing Internet of things. Radio frequency fingerprint identification (RFFI) is an emerging technique that exploits intrinsic and unique hardware impairments as the device identifier. The existing RFFI literature focuses on experimental exploration but comprehensive modelling is missing. This paper systematically models impairments of transmitter and receiver in narrowband systems and carries out extensive experiments and simulations to evaluate their effects on RFFI. The modelled impairments include oscillator imperfections, imbalance of inphase (I) and quadrature (Q) branches of mixers and power amplifier (PA) nonlinearity. We then propose a convolutional neural network-based RFFI protocol. We carry out experimental measurements over three months and demonstrate that oscillator imperfections are not suitable for RFFI due to their unpredictable time variation caused by temperature change. Our simulation results show that our protocol can classify 50 and 200 devices with uniformly and randomly distributed IQ imbalances and PA nonlinearities with high accuracy, namely 99% and 89%, respectively. We also show that the RFFI has some tolerance on different receiver imbalances during training and classification. Specifically, the accuracy is shown to degrade less than 20% when the residual receiver's gain and phase imbalances are small. Based on the experimental and simulation results, we made recommendations for designing a robust RFFI protocol, namely compensate carrier frequency offset and calibrate IQ imbalances of receivers.
Junqing Zhang, Roger F. Woods, Magnus Sandell, Mikko Valkama, Alan Marshall 0001, Joseph R. Cavallaro
IEEE Trans. Inf. Forensics Secur.2
2019 An Investigation of Using Loop-Back Mechanism for Channel Reciprocity Enhancement in Secret Key Generation
abstract
Physical layer security key generation exploits unpredictable features from wireless channels to achieve high security, which requires high reciprocity in order to set up symmetric keys between two users. This paper investigates enhancing the channel reciprocity using a loop-back scheme with multiple frequency bands in time-division duplex (TDD) communication systems, in order to mitigate the effect of hardware fingerprint interference and synchronization offset. The scheme is evaluated to be robust to passive eavesdropping and active Man-in-the-Middle attack through both theoretical analyses and practical measurements. A secret key generation protocol is subsequently designed. The performance of the proposed secret key generation method is then evaluated through both numerical simulation and experiments. Results demonstrate that the proposed scheme can effectively mitigate non-reciprocity and outperforms the classical TDD scheme in both key disagreement rate and key generation rate.
Linning Peng, Guyue Li, Junqing Zhang, Roger F. Woods, Ming Liu 0010, Aiqun Hu
IEEE Trans. Mob. Comput.4
2018 Facilitating Easier Access to FPGAs in the Heterogeneous Cloud Ecosystems
abstract
With FPGAs being increasingly integrated into existing software-based heterogeneous cloud environments, novel evaluation mechanisms are required to reveal the energy-performance trade-offs of accelerators (FPGAs, GPUs, etc) in high-level heterogeneous programming environments. For FPGAs, this involves also a reconsideration of scheduling policies and reconfiguration methods with an aim of integrating software-based approaches as well as performance optimizations for wider workload sizes. The approaches are evaluated using various reconfiguration methodologies for a number of applications.
Umar Ibrahim Minhas, Roger F. Woods, Georgios Karakonstantis
FPL2
2018 Opportunistic Non-Orthogonal Multiple Access Scheme with Unreliable Wireless Backhauls
abstract
The demand for increased connectivity and reliability of devices in the fifth generation (5G) of wireless communications requires new technology for ensuring massive connectivity and high spectral efficiency. In addition, wireless backhauls with guaranteed reliability are being considered to improve the overall system performance. In this paper, we investigate an opportunistic non-orthogonal multiple access (NOMA) system with unreliable wireless backhauls. In particular, we develop two opportunistic selection rules which allow the selection of the best among either near or far-away group transmitters, considering both the unreliability of wireless backhauls and fading effects of fronthauls. In order to analyze the performance, new exact and approximated closed-form expressions for the outage probabilities of the grouped receivers are derived, thus providing an insight into the impact of unreliable random backhauls and opportunistic NOMA. We show that the proposed scheme gives an outage performance gain of more than 3dB gains to a dominant receiver in the selection rules and improvement in receiver fairness when compared to the orthogonal multiple access (OMA) with an unreliable wireless backhaul. In addition, our results clearly reveal that unreliability levels of wireless backhaul links are responsible for the outage floors.
Sunyoung Lee, Trung Quang Duong, Roger F. Woods
PIMRC3
2018 Security Optimization of Exposure Region-Based Beamforming With a Uniform Circular Array
abstract
This paper investigates the impact of a uniform circular array (UCA) in the context of wireless security via exposure region-based beamforming. An improvement is demonstrated for the security metric proposed in our previous paper, namely, the spatial secrecy outage probability (SSOP), by optimizing the configuration of the UCA. Our previous paper focused on formalizing the SSOP concept and exploring its applicability using a uniform linear array example. This paper proposes the UCA as a superior candidate because it is more robust against the effects of mutual coupling. The UCA's SSOP configuration is explored and a special expression is derived from the general expression for the first time, and a closed-form upper bound is then generated to facilitate analysis. By carefully designing the UCA structure particularly the radius, an SSOP optimization algorithm is derived and explored for mutual coupling. It is shown that the information leakage to eavesdroppers is reduced while the legitimate user's received signal quality is enhanced due to the use of beamforming.
Roger F. Woods, Youngwook Ko, Alan Marshall 0001, Junqing Zhang
IEEE Trans. Commun.2
2017 Defining Spatial Secrecy Outage Probability for Exposure Region-Based Beamforming
abstract
With the increasing number of antennae in base stations, there is considerable interest in using beamforming to improve physical layer security, by creating an “exposure region” that enhances the received signal quality for a legitimate user and reduces the possibility of leaking information to a randomly located passive eavesdropper. This paper formalizes this concept by proposing a novel definition for the security level of such a legitimate transmission, called the spatial secrecy outage probability (SSOP). By performing a theoretical and numerical analysis, it is shown how the antenna array parameters can affect the SSOP and its analytic upper bound. While this approach may be applied to any array type and any fading channel model, it is shown here how the security performance of a uniform linear array varies in a Rician fading channel by examining the analytic SSOP upper bound.
Youngwook Ko, Roger F. Woods, Alan Marshall 0001
IEEE Trans. Wirel. Commun.3
2016 Proposing the Deep Dynamic Bayesian Network as a Future Computer Based Medical System
abstract
The development of new learning models has been of great importance throughout recent years, with a focus on creating advances in the area of deep learning. Deep learning was first noted in 2006, and has since become a major area of research in a number of disciplines. This paper will delve into the area of deep learning to present its current limitations and provide a new idea for a fully integrated deep and dynamic probabilistic system. The new model will be applicable to a vast number of areas initially focusing on applications into medical image analysis with an overall goal of utilising this approach for prediction purposes in computer based medical systems.
Caoimhe M. Carbery, Adele H. Marshall, Roger F. Woods
CBMS3
2016 Efficient Key Generation by Exploiting Randomness From Channel Responses of Individual OFDM Subcarriers
abstract
Key generation from the randomness of wireless channels is a promising technique to establish a secret cryptographic key securely between legitimate users. This paper proposes a new approach to extract keys efficiently from the channel responses of individual orthogonal frequency-division multiplexing (OFDM) subcarriers. The efficiency is achieved by: 1) fully exploiting randomness from time and frequency domains and 2) improving the cross-correlation of the channel measurements. Through the theoretical modeling of the time and frequency autocorrelation relationship of the OFDM subcarrier's channel responses, we can obtain the optimal probing rate and use multiple uncorrelated subcarriers as random sources. We also study the effects of non-simultaneous measurements and noise on the cross-correlation of the channel measurements. We find that the cross-correlation is mainly impacted by noise effects in a slow fading channel and use a low-pass filter to reduce the key disagreement rate and extend the system's working signal-to-noise ratio range. The system is evaluated in terms of randomness, key generation rate, and key disagreement rate, verifying that it is feasible to extract randomness from both time and frequency domains of the OFDM subcarrier's channel responses.
Junqing Zhang, Alan Marshall 0001, Roger F. Woods, Trung Quang Duong
IEEE Trans. Commun.3
2015 An effective key generation system using improved channel reciprocity
abstract
In physical layer security systems there is a clear need to exploit the radio link characteristics to automatically generate an encryption key between two end points. The success of the key generation depends on the channel reciprocity, which is impacted by the non-simultaneous measurements and the white nature of the noise. In this paper, an OFDM subcarriers' channel responses based key generation system with enhanced channel reciprocity is proposed. By theoretically modelling the OFDM subcarriers' channel responses, the channel reciprocity is modelled and analyzed. A low pass filter is accordingly designed to improve the channel reciprocity by suppressing the noise. This feature is essential in low SNR environments in order to reduce the risk of the failure of the information reconciliation phase during key generation. The simulation results show that the low pass filter improves the channel reciprocity, decreases the key disagreement, and effectively increases the success of the key generation.
Junqing Zhang, Roger F. Woods, Alan Marshall 0001, Trung Quang Duong
ICASSP2
2014 Power modelling and capping for heterogeneous ARM/FPGA SoCs
abstract
Low-power processors and accelerators that were originally designed for the embedded systems market are emerging as building blocks for servers. Power capping has been actively explored as a technique to reduce the energy footprint of high-performance processors. The opportunities and limitations of power capping on the new low-power processor and accelerator ecosystem are less understood. This paper presents an efficient power capping and management infrastructure for heterogeneous SoCs based on hybrid ARM/FPGA designs. The infrastructure coordinates dynamic voltage and frequency scaling with task allocation on a customised Linux system for the Xilinx Zynq SoC. We present a compiler-assisted power model to guide voltage and frequency scaling, in conjunction with workload allocation between the ARM cores and the FPGA, under given power caps. The model achieves less than 5% estimation bias to mean power consumption. In an FFT case study, the proposed power capping schemes achieve on average 97.5% of the performance of the optimal execution and match the optimal execution in 87.5% of the cases, while always meeting power constraints.
Yun Wu 0003, José L. Núñez-Yáñez, Roger F. Woods, Dimitrios S. Nikolopoulos
FPT3
2014 Creating secure wireless regions using configurable beamforming
abstract
We present a novel approach to network security against passive eavesdroppers by employing a configurable beam-forming technique to create tightly defined regions of coverage for targeted users. In contrast to conventional encryption methods, our security scheme is developed at the physical layer by configuring antenna array beam patterns to transmit the data to specific regions. It is shown that this technique can effectively reduce vulnerability of the physical regions to eavesdropping by adapting the antenna configuration according to the intended user's channel state information. In this paper we present the application of our concept to 802.11n networks where an antenna array is employed at the access point, and consider the issue of minimizing the coverage area of the region surrounding the targeted user. A metric termed the exposure region is formally defined and used to evaluate the level of security offered by this technique. A range of antenna array configurations are examined through analysis and simulation, and these are subsequently used to obtain the optimum array configuration for a user traversing a coverage area.
Alan Marshall 0001, Roger F. Woods, Youngwook Ko
PIMRC3
2013 Optimization of Weighted Finite State Transducer for Speech Recognition
abstract
There is considerable interest in creating embedded, speech recognition hardware using the weighted finite state transducer (WFST) technique but there are performance and memory usage challenges. Two system optimization techniques are presented to address this; one approach improves token propagation by removing the WFST epsilon input arcs; another one-pass, adaptive pruning algorithm gives a dramatic reduction in active nodes to be computed. Results for memory and bandwidth are given for a 5,000 word vocabulary giving a better practical performance than conventional WFST; this is then exploited in an adaptive pruning algorithm that reduces the active nodes from 30,000 down to 4,000 with only a 2 percent sacrifice in speech recognition accuracy; these optimizations lead to a more simplified design with deterministic performance.
Louis-Marie Aubert, Roger F. Woods, Scott Fischaber, Richard Veitch
IEEE Trans. Computers2
2013 Power Efficient, FPGA Implementations of Transform Algorithms for Radar-Based Digital Receiver Applications
abstract
A key challenge in defense and security systems is to implement functionality within a power budget. We show how data bandwidth redundancy and the need to change performance is exploited to achieve power efficient, field programmable gate array realizations with improved sampling rates. A unified methodology is given for the implementation of a key function, the fast Fourier transform, for a Radar-based digital receiver. Locality of data, temporal and spatial resource usage are examined from first principles, leading to an algorithmic approach that demonstrates substantial industrial benefits in terms of power, performance and resource usage. A power saving of 18% is achieved over a Cooley Tukey design with a 100% speed improvement;the work is extended to other cyclical fast algorithms.
Stephen McKeown, Roger F. Woods
IEEE Trans. Ind. Informatics2
2012 Novel Application of Genetic Sequencing Algorithms to Optimization of Hardware Resource Sharing for DSP
abstract
Field programmable gate array (FPGA) technology is a powerful platform for implementing computationally complex, digital signal processing (DSP) systems. Applications that are multi-modal, however, are designed for worse case conditions. In this paper, genetic sequencing techniques are applied to give a more sophisticated decomposition of the algorithmic variations, thus allowing an unified hardware architecture which gives a 10-25% area saving and 15% power saving for a digital radar receiver.
Stephen McKeown, Roger F. Woods
ASAP2
2012 An On-Demand Queue Management Architecture for a Programmable Traffic Manager
abstract
A queue manager (QM) is a core traffic management (TM) function used to provide per-flow queuing in access and metro networks; however current designs have limited scalability. An on-demand QM (OD-QM) which is part of a new modular field-programmable gate-array (FPGA)-based TM is presented that dynamically maps active flows to the available physical resources; its scalability is derived from exploiting the observation that there are only a few hundred active flows in a high speed network. Simulations with real traffic show that it is a scalable, cost-effective approach that enhances per-flow queuing performance, thereby allowing per-flow QM without the need for extra external memory at speeds up to 10 Gbps. It utilizes 2.3%-16.3% of a Xilinx XC5VSX50t FPGA and works at 111 MHz.
Qi Zhang 0026, Roger F. Woods, Alan Marshall 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2011 A Scalable and Programmable Modular Traffic Manager Architecture
abstract
A key issue in the design of next-generation Internet routers and switches will be provision of Traffic Manager (TM) functionality in the datapaths of their high-speed switching fabrics. A new architecture that allows dynamic deployment of different TM functions is presented. By considering the processing requirements of operations such as policing and congestion, queuing, shaping, and scheduling, a solution has been derived that is scalable with a consistent programmable interface. Programmability is achieved using a function computation unit which determines the action (e.g., drop, queue, remark, forward) based on the packet attribute information and a memory storage part. Results of a Xilinx Virtex-5 FPGA reference design are presented.
Shane O'Neill, Roger F. Woods, Alan Marshall 0001, Qi Zhang 0026
ACM Trans. Reconfigurable Technol. Syst.2
2010 Adapting noisy speech models - Extended uncertainty decoding
abstract
Most conventional techniques for noise adaptation assume a clean initial speech model which is adapted to a specific noise condition using adaptation data accumulated from the condition. In this paper, a different problem is considered, i.e. adapting a noisy speech model to a specific noise condition. For example, the initial noisy model may be a multi-condition model which is used to provide more accurate transcripts for the adaptation data than could be provided by a clean model, thereby obtaining a more accurate adaptation. We develop the formulation for this new problem by combining and extending maximum likelihood linear regression (MLLR), constrained MLLR (CMLLR) and uncertainty decoding techniques. We also present an implementation which has been tested on the Aurora 4 database, assuming an initial multi-condition model trained using white noise corrupted data. Significant word error rate (WER) reductions are achieved in comparison with other approaches.
Jianhua Lu, Ji Ming, Roger F. Woods
ICASSP3
2010 Guest Editorial ARC 2009
abstract
No abstract available.
Roger F. Woods, Jürgen Becker 0001, Peter M. Athanas, Fearghal Morgan
ACM Trans. Reconfigurable Technol. Syst.1
2009 Replacing uncertainty decoding with subband re-estimation for large vocabulary speech recognition in noise
Jianhua Lu, Ji Ming, Roger F. Woods
INTERSPEECH3
2009 Introduction to the Special Issue ARC'08
abstract
No abstract available.
Katherine Compton, Roger F. Woods, Christos-Savvas Bouganis, Pedro C. Diniz
ACM Trans. Reconfigurable Technol. Syst.2
2008 Power efficient DSP datapath configuration methodology for FPGA
abstract
Exploiting the underutilisation of variable-length DSP algorithms during normal operation is vital, when seeking to maximise the achievable functionality of an application within peak power budget. A system level, low power design methodology for FPGA-based, variable length DSP IP cores is presented. Algorithmic commonality is identified and resources mapped with a configurable datapath, to increase achievable functionality. It is applied to a digital receiver application where a 100% increase in operational capacity is achieved in certain modes without significant power or area budget increases. Measured results show resulting architectures requires 19% less peak power, 33% fewer multipliers and 12% fewer slices than existing architectures.
Stephen McKeown, Roger F. Woods, John McAllister
FPL2
2008 Combining noise compensation and missing-feature decoding for large vocabulary speech recognition in noise
Jianhua Lu, Ji Ming, Roger F. Woods
INTERSPEECH3
2008 A real-time flow monitor architecture encompassing on-demand monitoring functions
abstract
This paper introduces a new flow monitoring architecture that is based on FPGA technology offering high speed and programmability for monitoring in the next generation Internet. Two monitoring functions are also presented that can be implemented on the proposed architecture. The functions are based on Internet traffic characteristics and can increase monitoring efficiency and compute additional network statistics. The two functions illustrate how the programmable nature of the architecture allows for flexible, intelligent and on-demand monitoring. The first function classifies flows based on size and speed and the second function, measures flow traffic bursts.
John McGlone, Alan Marshall 0001, Roger F. Woods
NOMS3
2007 Soft IP core implementation of recursive least squares filter using only multplicative and additive operators
abstract
Soft IP cores can be realized as parameterisable HDL descriptions of circuit architecture where the performance comes from efficiently mapping system functionality. However, special arithmetic operations e.g. division, reciprocal, can restrict this mapping. An approach is presented that maps the system onto foundation operations, multiplication and addition, thereby giving a freer mapping of the full system. The methodology and results are given for a QR-based recursive least squares filter design on a Xilinx Virtex 4 FPGA giving a 5 GFLOPS performance.
Gaye Lightbody, Roger F. Woods, Jonathan Francey
FPL2
2007 Rapid implementation and optimisation of DSP systems on FPGA-centric heterogeneous platforms
John McAllister, Roger F. Woods, Scott Fischaber, E. Malins
J. Syst. Archit.2
2006 From Bit Level Systolic Arrays to HDTV Processor Chips
abstract
The paper starts presents the work initially carried out by Queen's University and RSRE (now Qinetiq) in the development of advanced architectures and microchips based on systolic array architectures. The paper outlines how this has led to the development of highly complex designs for high definition TV and highlights work both on advanced signal processing architectures and tool flows for advanced systems.
John V. McCanny, Roger F. Woods, John G. McWhirter
ASAP2
2006 Providing Input-Output Throughput Guarantees in a Buffered Crossbar Switch
abstract
Space division switching fabrics, like the buffered crossbar, are required to meet the forwarding capacity requirements of Gigabit and Terabit packet switches. These fabrics use fast-operating arbitration algorithms to achieve high utilization under heavy traffic loads. However, Next Generation Network (NGN) switches also need to provide throughput guarantees. A new scheduling system is proposed in this paper that can achieve both of these goals. Primarily, backlogged traffic is serviced in a manner that meets throughput guarantees, after which, the excess capacity is serviced in a manner that achieves high switch utilization.
Shane O'Neill, Alan Marshall 0001, Roger F. Woods
ISCC3
2006 Hierarchical synthesis of complex DSP functions using IRIS
abstract
A "white box" design methodology, which deals with hierarchical synthesis issues by providing a bridge between a high-level algorithm representation and lower level design tools, is presented. It illustrates tradeoffs when dealing with designs in a hierarchical and flattened manner. An enhanced Minnesota architectural synthesis scheduling algorithm is given, which gives highly efficient field programmable gate array solutions by providing access to parameterized expressions for datapath latencies and is applied to normalized lattice and delayed least mean square filter examples.
Roger F. Woods
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2005 FPGA Core Network Implementation and Optimization: A Case Study
Scott Fischaber, R. Hasson, John McAllister, Roger F. Woods
FPT4
2005 FPGA-Based Hardware for Physical Modelling Sound Synthesis by Finite Difference Schemes
Erdem Motuk, Roger F. Woods, Stefan Bilbao
FPT2
2005 Implementation of finite difference schemes for the wave equation on FPGA
abstract
The computational requirements of finite difference schemes for the solution of the wave equation for physical modelling can be huge. Field programmable gate arrays (FPGAs) provide an ideal platform for performing highly parallel DSP computations, but the challenge is to be able to implement complex systems quickly and efficiently on FPGA platforms. The paper presents a system level design approach based on a dataflow model of computation using a particular finite difference scheme for the solution of a 2+1D wave equation. The results suggest that 84000 nodes could be accommodated on a single Virtex II FPGA.
Erdem Motuk, Roger F. Woods, Stefan Bilbao
ICASSP (3)2
2005 Rapid generation of hardware functionality in heterogeneous platforms [FPGA implementation applications]
abstract
One of the key problems in complex digital system design is the rapid generation of efficient hardware functionality. The paper introduces an architecture template for targeting FPGA implementations as part of a dataflow based design flow for heterogeneous platforms, thereby allowing a designer to perform system level optimizations for consistent FPGA performance. The architecture provides scalable capabilities in both communications and processing allowing the core to be scaled to the problem size. Matrix multiplication is used to demonstrate the capabilities of this methodology giving speeds ranging from 121.4 MHz to 188.3 MHz without optimization.
Darren Gerard Reilly, Roger F. Woods, John McAllister, Richard L. Walke
ICASSP (5)2
2005 A Novel Packet Marking Function for Real-Time Interactive MPEG-4 Video Applications in a Differentiated Services Network
Shane O'Neill, Alan Marshall 0001, Roger F. Woods
NETWORKING3
2005 Virtex FPGA implementation of a pipelined adaptive LMS predictor for electronic support measures receivers
abstract
High-speed field-programmable gate array (FPGA) implementations of an adaptive least mean square (LMS) filter with application in an electronic support measures (ESM) digital receiver, are presented. They employ "fine-grained" pipelining, i.e., pipelining within the processor and result in an increased output latency when used in the LMS recursive system. Therefore, the major challenge is to maintain a low latency output whilst increasing the pipeline stage in the filter for higher speeds. Using the delayed LMS (DLMS) algorithm, fine-grained pipelined FPGA implementations using both the direct form (DF) and the transposed form (TF) are considered and compared. It is shown that the direct form LMS filter utilizes the FPGA resources more efficiently thereby allowing a 120 MHz sampling rate.
Lok-Kee Ting, Roger F. Woods, Colin Cowan
IEEE Trans. Very Large Scale Integr. Syst.2
2004 Highly efficient, limited range multipliers for LUT-based FPGA architectures
abstract
A novel design technique for deriving highly efficient multipliers that operate on a limited range of multiplier values is presented. Using the technique, Xilinx Virtex field programmable gate array (FPGA) implementations for a discrete cosine transform and poly-phase filter were derived with area reductions of 31%-70% and speed increases of 5%-35% when compared to designs using general-purpose multipliers. The technique gives superior results over other fixed coefficient methods and is applicable to a range of FPGA technologies.
Richard H. Turner, Roger F. Woods
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Design Flow for Efficient FPGA Reconfiguration
Richard H. Turner, Roger F. Woods
FPL2
2003 Design of a parameterizable silicon intellectual property core for QR-based RLS filtering
abstract
The availability of an intellectual property core for recursive least squares (RLS) filtering could enable the RLS algorithm to replace the least mean squares algorithm in a wide range of applications. The goal of this study is to develop a parameterizable generic architecture for RLS filtering in the form of a hardware description language (HDL) description, which can be used to generate highly efficient silicon layout. The key issue is to develop a family of circuit architectures that are 100% efficient and locally connected. This paper presents a generic mapping for RLS filtering and circuit architectures that can be mapped to a range of application requirements. It outlines the transition from array to architecture covering detailed design issues such as timing and control generation. The result is a family of QR designs, which are parameterized in terms of architecture size, wordlength, performance, and arithmetic processor timing.
Gaye Lightbody, Roger F. Woods, Richard L. Walke
IEEE Trans. Very Large Scale Integr. Syst.2
2002 Mapping Multi-Mode Circuits to LUT-Based FPGA Using Embedded MUXes
abstract
For some systems, a general-purpose FPGA solution tends to be large and slow. A reconfigurable solution is smaller and faster but has a delay associated with the reconfiguration. In this paper, embedded MUXes are used to achieve the performance of reconfiguration without the time penalty. For a CRC circuit an area reduction of 93% compared to a general-purpose solution and a reduction of 17-34% compared to similar software compiled systems is achieved.
Tim Courtney, Richard H. Turner, Roger F. Woods
FCCM3
2002 Multiplier-less Realization of a Poly-phase Filter Using LUT-based FPGAs
Richard H. Turner, Roger F. Woods, Tim Courtney
FPL2
2002 FPGA-based system-level design framework based on the IRIS synthesis tool and System Generator
abstract
A system level design framework for FPGA-based DSP design is presented. The design flow utilizes System Generator, a system level tool developed by Xilinx, and links it to an "in-house" architectural synthesis tool, IRIS. Whilst System Generator allows FPGA-based Intellectual Property (IP) cores to be incorporated into the design flow, it does not address the timing and latency problems introduced by the cores which can be considerable, particularly when the cores are pipelined. These problems are addressed by the IRIS synthesis tool. The paper describes the tools, their interaction and illustrates the flow using an 8-tap Transpose-Form Retimed Delayed LMS (TF-RDLMS) adaptive filter.
Roger F. Woods
FPT2
2001 Virtex Implementation of Pipelined Adaptive LMS Predictor in Electronic Support Measures Receiver
Lok-Kee Ting, Roger F. Woods, Colin Cowan
FPL2
2001 Implementation of fixed DSP functions using the reduced coefficient multiplier
abstract
Distributed arithmetic (DA) has been successfully applied to the design of area efficient multipliers on FPGAs for DSP applications. Whilst DA is efficient in applications where the coefficients are fixed, there is little option for applications with a limited range of coefficient values. This paper describes a technique for developing area efficient multipliers for a range of DSP applications that fall into this category. This is accomplished by employing multiplexers at no extra cost to increase the functionality of existing fixed coefficient multipliers. The technique has been applied to a DCT FPGA implementation where an area decrease of up to 50% and a speed increase of 33% was achieved over the conventional route.
Richard H. Turner, Tim Courtney, Roger F. Woods
ICASSP3
2000 An Investigation of Reconfigurable Multipliers for Use in Adaptive Signal Processing
abstract
This paper looks at various XC6200 multiplier architectures for use within adaptive signal processing systems. It compares data throughput, block utilisation and reconfiguration times. A number of approaches are compared including fully programmable multipliers and three separate ways of implementing reconfigurable multipliers. The paper shows how fixed coefficient multipliers can be used to increase the number of multipliers from 4 to 10. Many of the ideas and rules can be extended to more recent fine grain architectures.
Tim Courtney, Richard H. Turner, Roger F. Woods
FCCM3
1999 Accelerating Run-Time Reconfiguration on FCCMs
abstract
The paper describes the implementation of the arithmetic operations of multiplication, division and square root on a Xilinx XC6200 FPGA. By using a design approach to enhance similarities across circuits, partial reconfiguration has been used to allow reductions in reconfiguration times of up to 75% on trials using the VCC HOTWorks board.
Jean-Paul Heron, Roger F. Woods
FCCM2
1999 A Virtual Hardware Handler for RTR Systems
abstract
The design of a Virtual Hardware Handler for run-time reconfiguration is presented. A windows-based system that works with the VCC Hotworks board has been implemented and results are presented.
Richard H. Turner, Roger F. Woods, Sakir Sezer, Jean-Paul Heron
FCCM2
1999 Novel mapping of a linear QR architecture
abstract
This paper presents a novel architecture mapping technique which was essential in the design of a QR array which forms the core processor of a single chip adaptive beamforming system. The mapping technique assigns a QR triangular array of 2m/sup 2/+3m+1 cells down onto a linear architecture of m+1 processors. The mapping results in a linear systolic architecture with one hundred percent hardware utilisation, local interconnects and individual processors for boundary and internal cell operations. In addition, this paper highlights the effect latency has on the validity of the linear architecture.
Gaye Lightbody, Richard L. Walke, Roger F. Woods, John V. McCanny
ICASSP3
1998 Fast Partial Reconfiguration for FCCMs
abstract
The emergence of new FPGA families such as the Xilinx 6200 FPGA family and the Atmel 40000 series has been an important development in the FPGAs for Custom Computing Machines (FCCMs). These devices have number of appealing features when compared to other technologies such as the Xilinx 4000 series SRAM technology. These can be characterised as follows: faster reconfiguration (typically m/spl mu/ s or /spl mu/s), support for partial reconfiguration, dedicated microprocessor interface. An approach for run-time reconfiguration can be achieved by considering a range of functions collectively and developing the specific circuit architectures for each so that a high degree of commonality exists between them in terms of their structure, wiring and cell function. This is done by representing the functions or algorithms using Signal Flow Graphs (SFGs) and manipulating them to produce similar graphs for different functions. This basic concept can only be exploited through the development of an efficient hardware system. This revolves around the concept of virtual hardware which is integrated within the operating system and is supported by programming languages such as C and C++. The reconfigurable designs which allow partial re-configuration, are stored within a configuration data graph. Whilst this allows the configuration data to be efficiently stored, reconfiguration state graphs are used for high speed reconfiguration. The entire software hardware system for fast partial reconfiguration is illustrated.
Sakir Sezer, Roger F. Woods, Jean-Paul Heron, Alan Marshall 0001
FCCM2
1998 The impact of data characteristics and hardware topology on hardware selection for low power DSP
abstract
Adders and multipliers are key operations in DSP systems. The power consumption of adders is well understood but there are few detailed results on the choice of multipliers available. This paper considers how the power consumption of a number of multiplier structures such as Carry-Save array and Wallace Tree multipliers varies with data wordlengths and different layout strategies. In all cases, results were obtained from EPIC PowerMill™ simulations of actual synthesised circuit layouts. Analysis of the results highlights the effects of routing and interconnect optimization for low power operation and gives clear indications on choice of multiplier structure and design flow for the rapid design of DSP systems.
Gareth Keane, Jonathan Robert Spanier, Roger F. Woods
ISLPED3
1997 FPGA synthesis on the XC6200 using IRIS and Trianus/Hades (or from heaven to hell and back again)
abstract
The implementation of a number of FIR filter structures in the Xilinx XC6200 technology is presented. The designs have been implemented using a combination of IRIS, an architectural synthesis tool and Trianus/Hades a set of integrated tools for implementing algorithms on Custom Computing Machines. The main attraction of this approach is that it allows algorithms to be compiled quickly allowing performance changes to be made at the architectural level in IRIS rather than at the FPGA layout level.
Roger F. Woods, Stefan H.-M. Ludwig, Jean-Paul Heron, David W. Trainor, Stephan W. Gehring
FCCM1
1996 A New FFT Architecture and Chip Design for Motion Compensation based on Phase Correlation
abstract
Details of a new low power FFT processor for use in digital television applications are presented. This has been fabricated using a 0.6 /spl mu/m CMOS technology and can perform a 64 point complex forward or inverse FFT on real-time video at up to 18 Megasamples per second. It comprises 0.5 million transistors in a die area of 7.8/spl times/8 mm/sup 2/ and dissipates 1 W. Its performance, in terms of computational rate per area per watt, is significantly higher than previously reported devices, leading to a cost-effective silicon solution for high quality video processing applications. This is the result of using a novel VLSI architecture which has been derived from a first principles factorisation of the DFT matrix and tailored to a direct silicon implementation.
Colin Chiu Wing Hui, Tiong Jiu Ding, John V. McCanny, Roger F. Woods
ASAP4
1996 VLSI architectures for field programmable gate arrays: a case study
abstract
The ability to achieve highly efficient hardware implementations of algorithms will form a key aspect in the success of custom computing, Developments in VLSI architectures where regularity, simple design and locality of connections appears to be an ideal approach in the development of efficient field programmable gate array (FPGA) designs. The authors present a case study namely the implementation of the majority of a two dimensional (2D) discrete cosine transform (DCT) in a XilinK XC6216 device. The design has a 70% hardware utilisation figure and operates at 25 Mega pixels per second or over 30 frames per second for standard NTSC. The paper clearly demonstrates the hardware efficiency of such an approach. The paper also describes work on architectural synthesis framework to automatically generate highly regular designs for a wider range of complex algorithms.
Roger F. Woods, A. Cassidy, J. Gray
FCCM1
1994 A high performance IIR filter chip and its evaluation system
abstract
A highly flexible programmable IIR filter chip has been designed and fabricated to commercial requirements within a collaborative project involving several industrial partners. The device uses 8 highly regular 16 bit array multiplier-accumulators which have been pipelined to achieve an overall computational rate of 30 MHz using a 1 micron gate array process. Most significant bit first arithmetic has been employed to achieve the target 15 MHz sample rate whilst implementing an 8th order filter. The paper reviews the principles behind the filter chip and its architecture, and describes a modular system which has been built to facilitate its demonstration and evaluation.>
Richard L. Walke, Roger Evans, Roger F. Woods, G. Floyd, K. W. Wood
ASAP3
1992 The systematic design of high performance digital filters
abstract
A modular design methodology for practically realizable fine-grain pipelined digital circuits is presented. This scheme differs from previously suggested methodologies in that the effects of practical design issues such as truncation and overflow can be readily dealt with. The design of a parameterized high performance IIR filter circuit is used to demonstrate the new scheme.>
B. P. McGovern, Roger F. Woods, John V. McCanny
ICASSP2
1991 A 40 megasample IIR filter chip
abstract
The design of a high performance bit parallel second order IIR filter chip is described. The chip in question is highly pipelined, uses most significant bit first arithmetic and consists mainly of arrays of simple carry save adders. It has been fabricated in 1.5 um double level metal CMOS technology, accepts 12 bit input data and coefficient values and can operate at up to 40 megasamples per second. All data inputs and outputs are in two's complement form, and the chip power consumption is 1 W. The highly regular nature of the architecture has been exploited for test pattern generation. It is shown how small, but important modifications to the basic architecture, can significantly improve testing. As a result, 100% fault coverage can be achieved using less than 1000 test vectors. 'The chip may be used in a cascade realisation to form a general n/sup th/ order filter.>
O. C. McNally, John V. McCanny, Roger F. Woods
ASAP3
1990 Pipelined two-port adaptor for wave digital filtering
abstract
The application of fine grain pipelining techniques in the design of high performance wave digital filters (WDFs) is described. The problems of latency in feedback loops can be significantly reduced if computations are organized most significant, as opposed to least significant, bit first and if the results are fed back as soon as they are formed. The result is that chips can be designed which offer significantly higher sampling rates than otherwise can be obtained using conventional methods. How these concepts can be extended to the more challenging problem of WDFs is discussed. It is shown that significant increases in the sampling rate of bit-parallel circuits can be achieved using most significant bit first arithmetic.>
Rajinder Jit Singh, John V. McCanny, Roger F. Woods
ICASSP3
1990 Optimized bit level architectures for IIR filtering
abstract
Optimized circuits for implementing high-performance bit-parallel IIR filters are presented. Circuits constructed mainly from simple carry save adders and based on most-significant-bit (MSB) first arithmetic are described. Two methods resulting in systems which are 100% efficient in that they are capable of sampling data every cycle are presented. In the first approach the basic circuit is modified so that the level of pipelining used is compatible with the small, but fixed, latency associated with the computation in question. This is achieved through insertion of pipeline delays (half latches) one every second row of cells. This produces an area-efficient solution in which the throughput rate is determined by a critical path of 76 gate delays. A second approach combines the MSB first arithmetic methods with the scattered look-ahead methods. Important design issues are addressed, including wordlength truncation, overflow detection, and saturation.>
O. C. McNally, John V. McCanny, Roger F. Woods
ICCD3
1989 A bit-level systolic architecture for very high performance IIR filters
abstract
A novel bit-level systolic array architecture for implementing bit-parallel IIR filter sections is presented. The authors have shown previously how the fundamental obstacle of pipeline latency in recursive structures can be overcome by the use of redundant arithmetic in combination with bit-level feedback. These ideas are extended by optimizing the degree of redundancy used in different parts of the circuit and combining redundant circuit techniques with those of conventional arithmetic. The resultant architecture offers significant improvements in hardware complexity and throughput rate.>
Simon C. Knowles, John G. McWhirter, Roger F. Woods, John V. McCanny
ICASSP3
1988 Systolic IIR filters with bit level pipelining
abstract
A novel bit-level systolic array architecture for implementing first-order IIR filter sections is presented. A latency of only two clock cycles is achieved by using a radix-4 redundant number representation, performing the recursive computation most-significant-digit first, and feeding back each digit of the result as soon as it is available.>
Roger F. Woods, Simon C. Knowles, John V. McCanny, John G. McWhirter
ICASSP1