John McAllister

dblp:58/4124 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-4017-115XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Co-Approximate Parallel Compressive Depth Reconstruction on FPGA
abstract
Efficient depth image reconstruction from sparse samples is crucial for machine perception applications, such as robotics, vehicle assistance and autonomy. It demands fast processing speed with low power consumption for sensing quality and safety, as well as cost reduction for FPGA and solid state implementations, within constrained resource budgets on edge devices. A new co-approximate framework of parallel approximate compressive depth reconstruction engine on FPGA is proposed using ℓ1solvers, proximal gradient decent (PGD), with instrumented frequency and voltage scaling during the iterative optimization process. By evaluating various number of parallel approximate processing units for the depth image reconstruction engine, up to 51% further power saving is achieved, and 421× speed up of parallel processing compared to the baseline, henceforth the efficiency is elevated over 43×.
Yun Wu 0003, John McAllister
ICASSP2
2024 Quantum Circuit Cutting Minimising Loss of Qubit Entanglement
abstract
Quantum circuit cutting allows an arbitrary circuit to be executed on a quantum computer with fewer qubits by fragmenting it into subcircuits, each with fewer qubits. If such a method is to be employed, it is critical that the quantum properties of the original circuit, such as entanglement, are not violated. This paper demonstrates, using circuit cutting tools, how to select circuit bipartitions while prioritising the preservation of underlying entanglement between the qubits. We present two entanglement-based criteria. The former minimises losses of reduced-state pairwise qubit entanglement, and has polynomial complexity in the number of qubits. The latter minimises loss of multipartite entanglement across qubit bipartitions, with exponential complexity. Violations of entanglement with frequencies up to 72% are observed in current circuit cutting software. This paper describes how to prevent such violations, or to minimise qubit entanglement losses, fragmenting circuits with pairwise qubit negativity losses of only 5%, and for certain circuits 0% von Neumann entropy losses across qubit bipartitions.
Michael Hart, John McAllister
CF2
2024 Reconstructing Cut Quantum Circuits Maximising Fidelity between Quantum States
abstract
Quantum circuit cutting is an emerging field of quantum computing allowing quantum circuits requiring a relatively large number of qubits to be executed on a quantum computer which has a smaller number of qubits. Current works focus mainly on choosing the optimal position to cut circuits, in order to minimise classical reconstruction costs. They do not consider the detrimental effects of cutting circuits on the entanglement shared between the circuit qubits. This paper presents, so far as the authors are aware, the first works to globally reconstruct circuit cutting fragments by maximising the underlying fidelity between quantum states used within circuit computation. Consequently, the original entanglement present that was broken, is reformed. Gradient-based techniques are used to design an optimisation protocol for how to maximise the fidelity between highly multipartite-entangled states. Using adaptive gradient descent this work presents how unit fidelities can be achieved with a relatively small number of optimisation iterations.
Michael Hart, John McAllister
CF2
2021 An Emulation of Quantum Error-Correction on an FPGA device
abstract
Practical fault-tolerant quantum computation faces enormous engineering challenges and emulation is a key support for quantum algorithm development in the meantime. Field-Programmable Gate Array are promising hosts for these emulators but have seen almost no application to this problem so far. This paper presents emulation of an error-correct qubit via a 9-qubit Shor code on Xilinx FPGA. It shows how by exploiting sparsity linked to the errors, operational complexity can be reduced by more than 99% and high-fidelity operation enabled by 11-bit floating point computation.
Michael Hart, John McAllister, Leo Rogers, Charles Gillan
FPL2
2021 Configurable Quasi-Optimal Sphere Decoding for Scalable MIMO Communications
abstract
Sphere Decoding (SD) enables real-time quasi-optimal symbol detection for Multiple-Input Multiple-Output (MIMO) communication systems via custom circuit accelerators. Configurable SDs allow accelerator cost to be balanced with detection accuracy for the most constrained MIMO environments, such as power-constrained Internet-of-Things (IoT) scenarios. However this high detection accuracy comes at high accelerator cost. This paper proposes a novel configurable SD which addresses this issue. A Robust Bounded Spanning with Fast Enumeration (R-BSFE) approach employs novel strategies for channel matrix pre-processing and symbol enumeration to maintain quasi-ML accuracy whilst reducing complexity by up to 74%. This enables accelerators for 802.11n on Xilinx FPGA with significantly lower cost and higher throughput. To the best of the authors' knowledge, the accelerators produced are the highest performance, lowest cost quasi-ML SD accelerators on record.
Yun Wu 0003, John McAllister
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 Programmable Dataflow Accelerators: A 5G OFDM Modulation/Demodulation Case Study
abstract
Via OFDM technology, FFT and Inverse FFT (IFFT) operators enable the latest 5G radio standards. In these latests standards, the behaviour of FFT and IFFT needs to be flexible, supporting sub-carrier spacings from 15kHz to 480kHz and point sizes of up to 4096 point. An FFT or IFFT accelerator for 5G can take any configuration inside this spectrum both at design time, or potentially at run-time, under the control of a system control plane. This necessitates accelerators which combine high levels of flexibility and performance. This paper describes an FFT accelerator for such a context. Specifically, a novel data-driven programmable softcore processor is presented which enables run-time variable workloads and respond to data as provided by control processors in an MPSoC operating architecture. It is the first such accelerator to enable real-time IFFT/FFT for 5G, providing up to 3.89 times greater data rate than comparable accelerators.
Yun Wu 0003, Peng Wang 0091, John McAllister
ICASSP3
2019 EMG Wrist-hand Motion Recognition System for Real-time Embedded Platform
abstract
Electromyography (EMG) signal analysis is a popular method for controlling prosthetic and gesture control equipment. For portable systems, such as prosthetic limbs, real-time low-power operation on embedded processors is critical, but to date there has been no record of how existing EMG analysis approaches support such deployments. This paper presents a novel approach to time-domain classification of multichannel EMG signals harnessed from randomly-placed sensors according to the wrist-hand movements which caused their occurrence. It shows how, by employing a very small set of time-domain features, Kernel Fisher discriminant feature projection and Radial Bias Function neural network classifiers, nine wrist-hand movements can be detected with accuracy exceeding 99% - surpassing the state-of-the-art on record. It also shows how, when deployed on ARM Cortex-A53, the processing time is not only sufficient to enable real-time processing but is also a factor 50 shorter than the leading time-frequency techniques on record.
Sumit A. Raurale, John McAllister, Jesús Martínez del Rincón
ICASSP2
2019 On Modified Squared Givens Rotations for Sphere Decoder Preprocessing
abstract
Sphere Decoding for Multiple-Input Multiple-Output (MIMO) wireless systems is a complex operation, usually demanding custom accelerators in order to support real-time performance. The cost of these accelerators is disproportionately influenced by channel matrix preprocessing, which represents a relatively small fraction of the overall computational cost of detecting an OFDM MIMO frame in standards such as 802.11n, but consumes a very large amount of hardware resource. Modified Squared Givens' Rotations has been proposed to resolve this issue and shown to dramatically reduce accelerator cost. However, there is no analysis on the record of the complexity of this algorithm, nor its detection performance. This paper shows that, despite offering modest reductions in operational complexity, MFSD-SQRD enables dramatic cost reductions by explicitly addressing the overhead of matrix permutation steps. Further, it shows that for most SNR values of practical interest, the performance of MFSD-SQRD is not appreciably diminished relative to the standard SQRD approach to preprocessing. To the best of the authors' knowledge, the proposed modified SQRD preprocessing approach is the highest performance sub-optimal preprocessing approach on record.
Yun Wu 0003, John McAllister
ICASSP2
2019 Window Size Estimation for Nearest Neighbour Compliant Quantum Circuit Mapping
abstract
In general, quantum circuits permit any two logical qubits to be combined via a quantum gate. When these are to be deployed on locally-connected grid-shaped quantum processors, swap operations are required to make adjacent interacting qubits. Swap insertion algorithms are complex and time-consuming. Windowing limits this complexity by considering only a subset of the circuit's gates when inserting swaps, but has not been applied to the latest generation of swap insertion algorithms, nor has any systematic method been proposed for determining the appropriate window length. This paper shows how off-line analysis of swap density across the circuit identifies thresholds for window length which limit increases in the number of swap gates. When adopted, speed-ups in the swap insertion process asymptotically approach 100% with the length of the circuit, whilst maintaining swap costs at comparable levels to non-windowed algorithms.
Leo Rogers, John McAllister
ISCAS2
2018 Emg Acquisition and Hand Pose Classification for Bionic Hands from Randomly-Placed Sensors
abstract
This paper presents a unique real-time motion recognition system for Electromyographic (EMG) signal acquisition and classification. It is the first approach which can classify hand poses from multi-channel EMG signals gathered from randomly placed arm sensors as accurately as current placed-sensor EMG acquisition approaches. It combines time-domain feature extraction, Linear Discriminant Analysis (LDA) feature projection and Multilayer Perceptron (MLP) classification to allow nine distinct poses to be correctly identified more than 95% of the time. This is comparable to state-of-the-art placed-sensor EMG acquisition systems. Processing times of 11.70 ms also make this a viable candidate approach for real-time EMG acquisition and processing in practical prosthesis applications.
Sumit A. Raurale, John McAllister, Jesús Martínez del Rincón
ICASSP2
2018 Architectural Synthesis of Multi-SIMD Dataflow Accelerators for FPGA
abstract
Field Programmable Gate Array (FPGA) boast abundant resources with which to realise high-performance accelerators for computationally demanding operations. Highly efficient accelerators may be automatically derived from Signal Flow Graph (SFG) models by using architectural synthesis techniques, but in practical design scenarios, these currently operate under two important limitations - they cannot efficiently harness the programmable datapath components which make up an increasing proportion of the computational capacity of modern FPGA and they are unable to automatically derive accelerators to meet a prescribed throughput or latency requirement. This paper addresses these limitations. SFG synthesis is enabled which derives software-programmable multicore single-instruction, multiple-data (SIMD) accelerators which, via combined offline characterisation of multicore performance and compile-time program analysis, meet prescribed throughput requirements. The effectiveness of these techniques is demonstrated on tree-search and linear algebraic accelerators for 802.11n WiFi transceivers, an application for which satisfying real-time performance requirements has, to this point, proven challenging for even manually-derived architectures.
Yun Wu 0003, John McAllister
IEEE Trans. Parallel Distributed Syst.2
2017 Multicore distributed dictionary learning: A microarray gene expression biclustering case study
abstract
The increasing pervasion and scale of machine learning technologies is posing fundamental challenges for their realisation. In the main, current algorithms are centralised, with a large number of processing agents, distributed across parallel processing resources, accessing a single, very large data object. This creates bottlenecks as a result of limited memory access rates. Distributed learning has the potential to resolve this problem by employing networks of co-operating agents each operating on subsets of the data, but as yet their suitability for realisation on parallel architectures such as multicore are unknown. This paper presents the results of a case study deploying distributed dictionary learning for microarray gene expression bi-clustering on a 16-core Epiphany multicore. It shows that distributed learning approaches can enable near-linear speed-up with the number of processing resources and, via the use of DMA-based communication, a 50% increase in throughput can be enabled.
Stephen Laide, John McAllister
ICASSP2
2016 Streaming Elements for FPGA Signal and Image Processing Accelerators
abstract
Field-programmable gate array (FPGA) devices boast abundant resources with which custom accelerator components for signal, image, and data processing may be realized; however, realizing high-performance, low-cost accelerators currently demands manual register transfer level design. Software-programmable soft processors have been proposed as a way to reduce this design burden, but they are unable to support performance and cost comparable to custom circuits. This paper proposes a new soft processing approach for FPGA that promises to overcome this barrier. A high-performance, fine-grained streaming processor, known as a streaming accelerator element, is proposed, which realizes accelerators as large-scale custom multicore networks. By adopting a streaming execution approach with advanced program control and memory addressing capabilities, typical program inefficiencies can be almost completely eliminated to enable performance and cost, which are unprecedented among software-programmable solutions. When used to realize accelerators for fast Fourier transform, motion estimation, matrix multiplication, and sobel edge detection, it is shown how the proposed architecture enables real-time performance and with performance and cost comparable with hand-crafted custom circuit accelerators and up to two orders of magnitude beyond existing soft processors.
Peng Wang 0091, John McAllister
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Soft-core stream processing on FPGA: An FFT case study
abstract
The increasing design complexity associated with modern Field Programmable Gate Array (FPGA) has prompted the emergence of 'soft'-programmable processors which attempt to replace at least part of the custom circuit design problem with a problem of programming parallel processors. Despite substantial advances in this technology, its performance and resource efficiency for computationally complex operations remains in doubt. In this paper we present the first recorded implementation of a softcore Fast-Fourier Transform (FFT) on Xilinx Virtex FPGA technology. By employing a streaming processing architecture, we show how it is possible to achieve architectures which offer 1.1 GSamples/s throughput and up to 19 times speed-up against the Xilinx Radix-2 FFT dedicated circuit with comparable cost.
Peng Wang 0091, John McAllister, Yun Wu 0003
ICASSP2
2012 Valved dataflow for FPGA memory hierarchy synthesis
abstract
For modern FPGA, implementation of memory intensive processing applications such as high end image and video processing systems necessitates manual design of complex multilevel memory hierarchies incorporating off-chip DDR and on-chip BRAM and LUT RAM. In fact, automated synthesis of multi-level memory hierarchies is an open problem facing high level synthesis technologies for FPGA devices. In this paper we describe the first automated solution to this problem. By exploiting a novel dataflow application modelling dialect, known as Valved Dataflow, we show for the first time how, not only can such architectures be automatically derived, but also that the resulting implementations support real-time processing for current image processing application standards such as H.264. We demonstrate the viability of this approach by reporting the performance and cost of hierarchies automatically generated for Motion Estimation, Matrix Multiplication and Sobel Edge Detection applications on Virtex-5 FPGA.
Matthew Milford, John McAllister
ICASSP2
2010 FPGA based soft-core SIMD processing: A MIMO-OFDM Fixed-Complexity Sphere Decoder case study
abstract
To enable reliable data transfer in next generation Multiple-Input Multiple-Output (MIMO) communication systems, terminals must be able to react to fluctuating channel conditions by having flexible modulation schemes and antenna configurations. This creates a challenging real-time implementation problem: to provide the high performance required of cutting edge MIMO standards, such as 802.11n, with the flexibility for this behavioural variability. FPGA softcore processors offer a solution to this problem, and in this paper we show how heterogeneous SISD/SIMD/MIMD architectures can enable programmable multicore architectures on FPGA with similar performance and cost as traditional dedicated circuit-based architectures. When applied to a 4×4 16-QAM Fixed-Complexity Sphere Decoder (FSD) detector we present the first soft-processor based solution for real-time 802.11n MIMO.
Xuezheng Chu, John McAllister
FPT2
2008 Power efficient DSP datapath configuration methodology for FPGA
abstract
Exploiting the underutilisation of variable-length DSP algorithms during normal operation is vital, when seeking to maximise the achievable functionality of an application within peak power budget. A system level, low power design methodology for FPGA-based, variable length DSP IP cores is presented. Algorithmic commonality is identified and resources mapped with a configurable datapath, to increase achievable functionality. It is applied to a digital receiver application where a 100% increase in operational capacity is achieved in certain modes without significant power or area budget increases. Measured results show resulting architectures requires 19% less peak power, 33% fewer multipliers and 12% fewer slices than existing architectures.
Stephen McKeown, Roger F. Woods, John McAllister
FPL3
2008 Modified givens rotations and their application to matrix inversion
abstract
Complex wireless communication systems such as MIMO require high-performance real-time implementations of operations such as matrix inversion. This paper presents two novel algorithms for this application. A novel modified conventional Givens rotations (MCGR) method has been derived which offers high-performance implementation since it avoids high-latency angle-based architectures, such as CORDIC. Furthermore, a novel modified squared Givens rotations (MSGR) method has been proposed which extends the original SGR method for complex valued data, and also corrects erroneous results in the original SGR method when zeros occur on the diagonal of the matrix either initially or during processing. In addition, both of the proposed methods avoid complex dividers in the matrix inversion, thus minimising the complexity of potential real-time implementations.
Lei Ma 0012, Kevin Dickson, John McAllister, John V. McCanny
ICASSP3
2007 Rapid implementation and optimisation of DSP systems on FPGA-centric heterogeneous platforms
John McAllister, Roger F. Woods, Scott Fischaber, E. Malins
J. Syst. Archit.1
2005 Core-Based Methodology: An Automated Approach for Implementing a Complete System from Algorithms to a Heterogeneous Network including FPGAs
abstract
In this paper, we present a methodology for implementing a complete digital signal processing (DSP) system onto a heterogeneous network including field programmable gate arrays (FPGAs) automatically. The methodology aims to allow design refinement and real time verification at the system level. The DSP application is constructed in the form of a data flow graph (DFG) which provides an entry point to the methodology. The netlist for parts that are mapped onto the FPGA(s) together with the corresponding software and hardware application protocol interface (API) are also generated. Using a set of case studies, we demonstrate that the design and development time can be significantly reduced using the methodology developed.
Jasmine Lam, John McAllister, Jennifer Dudley
FCCM2
2005 FPGA Core Network Implementation and Optimization: A Case Study
Scott Fischaber, R. Hasson, John McAllister, Roger F. Woods
FPT3
2005 Rapid generation of hardware functionality in heterogeneous platforms [FPGA implementation applications]
abstract
One of the key problems in complex digital system design is the rapid generation of efficient hardware functionality. The paper introduces an architecture template for targeting FPGA implementations as part of a dataflow based design flow for heterogeneous platforms, thereby allowing a designer to perform system level optimizations for consistent FPGA performance. The architecture provides scalable capabilities in both communications and processing allowing the core to be scaled to the problem size. Matrix multiplication is used to demonstrate the capabilities of this methodology giving speeds ranging from 121.4 MHz to 188.3 MHz without optimization.
Darren Gerard Reilly, Roger F. Woods, John McAllister, Richard L. Walke
ICASSP (5)3