Klaus D. McDonald-Maier

dblp:78/4433 · also Klaus D. Maier 0001, Klaus Dieter McDonald-Maier · DBLP profile ↗
← Back
62ranked-venue papers
1as first author
32since 2021 · last 2026
0000-0002-6412-8519ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 21 since 2021Software engineering, systems software and programming languages · 13 · 3 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 Mitigating Scalability Challenges in LUT-Based Neural Networks via Pruning Optimisations
abstract
Modern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computational cost. To address this, Look-Up Table (LUT)-based matrix multiplications have emerged as a promising alternative for reducing the computational cost and time of the multiply-accumulate operations in a neural network. However, the LUT-based neural network still faces the scalability challenge due to the inherent limitations of LUT-based matrix multiplication. To mitigate these scalability limitations, this paper proposes a scalable and energy-efficient LUT-based approximate matrix multiplication unit (LUT-MU) constituting the basic component of the neural networks by integrating a pruning strategy on the MADDNESS algorithm, a LUT-based matrix multiplication methodology. With increasing problem size and precision demands in matrix multiplication, our proposed LUT-MU architecture effectively constrains resource expansion. The case study shows that deploying our LUT-MU in neural network architectures, including fully connected layers (MNIST) and ResNets (CIFAR-10, ImageNet)—on XCZU7EV and XCZU19EG FPGAs, produces up to 1.6× throughput improvement and 4.2× energy efficiency gains over mainstream CUDA-based network implementations, and 1.8× energy efficiency compared to leading quantised neural network implementations, with moderate impact on accuracy. Compared to original MADDNESS-based neural networks, our LUT-MU shows 1.3 to 2.6× resource savings based on various resolution configuration settings of MADDNESS.
Xuqi Zhu, Huaizhi Zhang, Chandrajit Pal, Sangeet Saha, Klaus D. McDonald-Maier, Xiaojun Zhai
IEEE Trans. Computers7
2025 Late Breaking Results: Approximated LUT-Based Neural Networks for FPGA Accelerated Inference
abstract
This work presents LUT-MU, an approximated LUT-based Matrix Multiplication (MM) architecture designed for FPGA-based Neural Network (NN) inference across. The proposed architecture maximises the utilisation of on-chip memory bandwidth through dedicated memory distribution and pipeline design, addressing performance limitations inherent to LUT-based MM. Experimental evaluation demonstrates that LUT-MU achieves a four-fold improvement in NN inference throughput whilst reducing hardware resource consumption by 80% with only a 5% decline in accuracy. These results validate that our optimisation approach successfully resolves the performance constraints caused by the limited arithmetic intensity and memory bandwidth, enabling the LUT-MU to serve as a foundation for efficient NN acceleration systems.
Xuqi Zhu, Huaizhi Zhang, Tamim M. Al-Hasan, Klaus D. McDonald-Maier, Xiaojun Zhai
DATE5
2025 Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic Systems
abstract
Integrating Large Language Models (LLMs) into modern robotic systems presents significant computational and energy constraint challenges, particularly for human-centered robotic applications. This paper presents a novel hardware optimization technique for deploying LLMs on resource-constrained embedded devices, achieving an up to 77% reduction in computational latency through an FPGA implementation in comparison to other popular embedded computing devices (e.g., CPU and GPUs). Additionally, we demonstrate our methodology by deploying a LLaMA 2-7B model on a Unitree Go2 robotic dog integrated with the proposed FPGA platform. The proposed optimization framework preserves real-time interaction capabilities while significantly reducing computational and energy overhead, facilitating efficient natural language processing for human-robot interaction in safety-critical and dynamic environments. Experimental results demonstrate that the FPGA-based LLaMA 2-7B implementation achieves up to 6.06-fold and 1.95-fold higher throughput compared to baseline CPU and GPU implementations while maintaining comparable inference accuracy. Furthermore, the proposed FPGA design surpasses existing state-of-the-art FPGA implementations, delivering a 30% improvement in computational efficiency.
Huaizhi Zhang, Tamim M. Al-Hasan, Xuqi Zhu, Weiyong Si, Klaus D. McDonald-Maier, Xiaojun Zhai
IROS6
2025 A Privacy-aware Quantilisation Approach for Efficient Edge Deep Learning Accelerator
abstract
Data privacy is one of the key concerns in machine learning model applications at the edge, especially in sensitive domains such as the healthcare sector. Here adversaries may exploit Membership Inference Attacks (MIAs) to determine if particular data points were used as part of datasets from the model’s training set, potentially leading to further data leakage issues. Although privacy preservation techniques like differential privacy (DP) can mitigate such risks during the training phase, this often results in degradation of model accuracy, making them less suitable for cloud training and edge deployment paradigms. For AI edge applications, existing research works for designing neural network accelerators primarily prioritize computational performance and power efficiency as their design target. In this paper, we introduce a novel privacy-aware quantilisation approach for deep learning accelerators and analyse the tradeoff between computational efficiency and privacy protection. The proposed system allows to adjust privacy constraints through tunable parameters, enabling flexible deployment on edge devices while meeting privacy and performance constraints. We have evaluated the proposed design on an AMD VCK190 board using a range of hypothetical MIA benchmarks. The results demonstrate that the proposed approach can effectively reduce the success rate of MIA attacks across multiple performance metrics.
Huaizhi Zhang, Xuqi Zhu, Klaus D. McDonald-Maier, Xiaojun Zhai
ISCAS4
2025 RENOWNED: A Real-Time Anomaly Detection and Mitigation Framework in Edge-Enabled IoV
abstract
The rapid adoption of smart vehicles and their interconnection through the Internet of Vehicles (IoV) has increased the use of electronic control units (ECUs) in cars. These ECUs, while enabling advanced features, also present a larger target for cyberattacks, which can disrupt critical functions and jeopardize safety. The time-sensitive nature of automotive systems necessitates swift responses, making the protection of ECUs crucial. The imprecise computation (IC) task model can mitigate the risk of task completion failures by generating acceptable approximation results within deadlines when achieving absolute accuracy becomes difficult within fixed deadlines and energy budgets. This article introduces RENOWNED, a solution that ensures the normal functioning of these controller area networks (CAN) controlled ECUs even in the face of anomalies. It combines anomaly detection and mitigation through the HEALING module to maintain the desired performance. The anomaly detection module uses graph attention networks (GAT) to identify unusual processor behavior. If an anomaly is detected, the HEALING module takes over, reallocating tasks based on the available resources to guarantee that deadlines are met and energy constraints are not exceeded. Experiments have shown that RENOWNED delivers a Quality of Service (QoS) of 25% to 64% when system utilisation is varied in the range from 40% to 90%. It exhibits an excelling performance in detecting anomalies, achieving a 97.6% accuracy even when the magnitude mixed anomaly signals are very minute. Thus, our proposed RENOWNED offers a robust way to enhance the reliability and energy efficiency of safety-critical automotive applications prevalent in IoV.
Chandrajit Pal, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier
IEEE Internet Things J.4
2025 PRECIOUS: Approximate Real-Time Computing in MLC-MRAM Based Heterogeneous CMPs
abstract
Enhancing quality of service (QoS) in approximate-computing (AC) based real-time systems, without violating power limits is becoming increasingly challenging due to contradictory constraints, i.e., power consumption and time criticality, as multicore computing platforms are becoming heterogeneous. To fulfill these constraints and optimise system QoS, AC tasks should be judiciously mapped on such platforms. However, prior approaches rarely considered the problem of AC task deployment on heterogeneous platforms. Moreover, the majority of prior approaches typically neglect the runtime architectural phenomena, which can be accounted for along with the approximation tolerance of the applications to enhance the QoS. We presentPRECIOUS, a novel hybrid offline-online approach that firstschedules AC real-timetasks on aheterogeneous multicorewith an objective to maximise QoS and determines the appropriate cluster for each task constrained by a system-wide power limit, deadline, and task-dependency. At runtime,PRECIOUSintroduces novel architectural techniques for the AC tasks, where tasks are executed on a heterogeneous platform equipped withmultilevel-cell (MLC)-MRAMbased last-level cache to improve energy efficiency and performance by prudentially leveraging storage density of MLC-MRAM while ameliorating associated high write latency and write energy. Our novel block management for the MLC-MRAM cache further improves performance of the system, which we exploit opportunistically to enhance system QoS, and turn off processor cores during the dynamically generated slacks.PRECIOUS-Offlineachieves up to 76% QoS for a specific task-set, surpassing prior art, whereasPRECIOUS-Onlineenhances QoS by 9.0% by reducing cache miss-rate by 19% on a 64-core heterogeneous system without incurring any energy overhead over a conventional MRAM based cache design.
Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier
IEEE Trans. Computers5
2025 MESSI: Task Mapping and Scheduling Strategy for FPGA-based Heterogeneous Real-Time Systems
abstract
Continuous demands for improved performance within constrained resource budgets are driving a move from homogeneous to heterogeneous processing platforms for the implementation of today’s Real-Time (RT) embedded systems. The applications executing on such systems are typically represented as a Precedence Task Graph (PTG), where a node represents a task or algorithm for one functionality and edges represent the complex interactions between multiple functionalities. Due to RT constraints, the task graph needs to be executed within a specified deadline. Although some existing studies have looked into solving this challenge, comprehensive studies that combine the theoretical features of RT task-graph mapping and scheduling with practical runtime architectural characteristics have mostly been ignored to date. Hence, in this article, we consider the challenge of scheduling an RT application modeled as a single PTG, with the objective of minimizing the overall execution time under Hardware (HW) resource and deadline constraints for heterogeneous Central Processing Unit (CPU) + Field Programmable Gate Array (FPGA) architectures. First, we introduce an optimal solution using Integer Linear Programming (ILP). However, this ILP-based optimal solution suffers from computational complexity and does not scale well even for moderately large problem sizes. Hence, we additionally propose heuristic algorithms for task mapping and scheduling. The efficiency of the proposed scheme, named MESSI, has been evaluated through experiments using PTG on a practical CPU+FPGA system regarding current technology restrictions. Our experiments demonstrate that performance gains of 55.6% and area usage reductions of 46.3% are possible compared to full Software (SW) and HW execution, respectively.
Sallar Ahmadi-Pour, Sangeet Saha, Klaus D. McDonald-Maier, Rolf Drechsler
ACM Trans. Design Autom. Electr. Syst.3
2025 APPARENT: AI-Powered Platform Anomaly Detection in Edge Computing
abstract
Embedded systems serving as IoT nodes are often vulnerable to malicious and unknown runtime software that could compromise the system, steal sensitive data, and cause undesirable system behaviour. Commercially available embedded systems used in automation, medical equipment, and automotive industries, are especially exposed to this vulnerability since they lack the resources to incorporate conventional safety features and are challenging to mitigate through conventional approaches. We propose a novel system design coined as APPARENT which can identify program characteristics by monitoring and counting the maximum possible low-level hardware events from Hardware Performance Counters (HPCs) that occur during the program's execution and analyse the correlation among the counts of various monitored events. To further utilise these captured events as features we propose a self-supervised machine learning algorithm that combines a Graph Attention Network GAT and a Generative Topographic Mapping GTM to detect unusual program behaviour as anomalies to enhance the system security. Our proposed methodology takes advantage of attributes like program counter, cycles per instruction, and physical and virtual timers at various exception levels of the embedded processor to identify abnormal activity. APPARENT identifies unknown program behaviours not present in the training phase with an accuracy of over 98.46% on Autobench EEMBC benchmarks.
Chandrajit Pal, Sangeet Saha, Xiaojun Zhai, Gareth Howells 0001, Klaus D. McDonald-Maier
IEEE Trans. Sustain. Comput.5
2024 Evaluating Lightweight GAN- and Adapted CTGAN-Based Data Synthesis for Predictive Maintenance in High-Radiation Environments
abstract
This paper presents a comparative analysis of two developed Generative Adversarial Network (GAN) architectures for synthesizing sensor data in predictive maintenance (PdM) applications within high-radiation environments. The study ad-dresses the challenge of data scarcity in such settings, where experimental runs are constrained by the risk of device failure and economic considerations. The two GAN models: GAN-1 uses the Conditional Tabular GAN (CTGAN) architecture, and GAN-2 employs a custom network. These models generated synthetic datasets that were used to train and evaluate three machine learning algorithms: Random Forest, k-Nearest Neighbours, and eXtreme Gradient Boosting. The performance of these PdM models trained on synthetic data was compared against models trained on the original limited dataset. Results demonstrate that GAN-1 produced synthetic data closely mirroring the characteristics of the original dataset, enabling PdM models to achieve comparable performance levels. This study highlights the potential of GAN-based data synthesis in enhancing PdM model development for high-radiation environments, offering a viable solution to the challenges of limited data availability in such harsh settings. The findings have significant implications for improving operational reliability and safety in nuclear and other extreme environments where electronic systems are deployed.
Tamim M. Al-Hasan, Xiaojun Zhai, Klaus D. McDonald-Maier, Faycal Bensaali, Alice Cryer
BDCAT3
2024 MAFin: Maximizing Accuracy in FinFET based Approximated Real-Time Computing
abstract
We propose MAFin that exploits the unique temperature effect inversion (TEI) property of a FinFET based multicore platform, where processing speed increases with temperature, in the context of approximate real-time computing. In approximate real-time computing platforms, the execution of each task can be divided into two parts: (i) the mandatory part, execution of which provides a result of acceptable quality, followed by (ii) the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy (QoS) without violating deadlines. With an objective to maximize the QoS for a FinFET based multicore system, MAFin, our proposed real-time scheduler first derives a task-to-core allocation, while respecting system-wide constraints and prepares a schedule. During execution, MAFin further increases the achieved QoS, while balancing the performance and temperature on-the-fly by incorporating a prudential temperature cognizant frequency management mechanism and guarantees imposed constraints. Specifically, MAFin exploits the TEI property of FinFET based processors, where processor-speed is enhanced at the increased temperature, to reduce the execution time of the individual tasks. This reduced execution-time is then traded off either to enhance QoS by executing more from the tasks' optional parts or to improve energy efficiency by turning off the core. While surpassing prior art, MAFin achieves 70% QoS, which is further enhanced by 8.3% in online, with a maximum EDP gain of up to 12%, based on benchmark based evaluation on a 4-core based system.
Shounak Chakraborty 0001, Sangeet Saha, Magnus Själander, Klaus D. McDonald-Maier
DAC4
2024 ARCTIC: Approximate Real-Time Computing in a Cache-Conscious Multicore Environment
abstract
Improving result-accuracy in approximate computing (AC) based time-critical systems, without violating power constraints of the underlying circuitry, is gradually becoming challenging with the rapid progress in technology scaling. The execution span of each AC real-time tasks can be split into a couple of parts: (i) the mandatory part, execution of which offers a result of acceptable quality, followed by (ii) the optional part, which can be executed partially or completely to refine the initially obtained result in order to increase the result-accuracy, while respecting the time-constraint. In this article, we introduce a novel hybrid offline-online scheduling strategy, for AC real-time tasks. The goal of real-time scheduler of is to maximise the results-accuracy (QoS) of the task-set with opportunistic shedding of the optional part, while respecting system-wide constraints. During execution, retains exclusive copy of the private cache blocks only in the local caches in a multi-core system and no copies of these blocks are maintained at the other caches, and improves performance (i.e., reduces execution-time) by accumulating more live blocks on-chip. Combining offline scheduling with the online cache optimization improves both QoS and energy efficiency. While surpassing prior arts, our proposed strategy reduces the task-rejection-rate by up to 25%, whereas enhances QoS by 10%, with an average energy-delay-product gain of up to 9.1%, on an 8-core system.
Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Bayesian Optimization for Efficient Heterogeneous MPSoC Based DNN Accelerator Runtime Tuning
abstract
With the explosive growth of Internet of Things (IoT) devices and applications, deploying Deep Neural Networks (DNNs) on resource-constrained embedded edge devices has become a popular research trend. Because such systems have limited resources, they need to rely on optimising resource utilisation to meet performance requirements. However, for scenarios where the DNN application and workloads are dynamically changing, the offline system optimisation technique cannot achieve optimal runtime performance in practical environments. Hence, in this PhD project, we propose a Bayesian Optimisation (BO)-based runtime tuning scheme for improving energy efficiency of heterogeneous MPSoC-based DNN accelerator in the context of DNN applications. By seeking suitable hardware configurations of the accelerator for dynamic DNN inference workloads ranging from 200 M to 600 M FLOPs (floating-point operations) at runtime, the recommended configuration can averagely save up to 15.33% energy consumption from a random configuration setting.
Xuqi Zhu, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier
FPL5
2023 A Complementarity-Based Switch-Fuse System for Improved Visual Place Recognition
abstract
Recently several fusion and switching based approaches have been presented to solve the problem of Visual Place Recognition. In spite of these systems demonstrating significant boost in VPR performance they each have their own set of limitations. The multi-process fusion systems usually involve employing brute force and running all available VPR techniques simultaneously while the switching method attempts to negate this practise by only selecting the best suited VPR technique for given query image. But switching does fail at times when no available suitable technique can be identified. An innovative solution would be an amalgamation of the two otherwise discrete approaches to combine their competitive advantages while negating their shortcomings. The proposed, Switch-Fuse system, is an interesting way to combine both the robustness of switching VPR techniques based on complementarity and the force of fusing the carefully selected techniques to significantly improve performance. Our system holds a structure superior to the basic fusion methods as instead of simply fusing all or any random techniques, it is structured to first select the best possible VPR techniques for fusion, according to the query image. The system combines two significant processes, switching and fusing VPR techniques, which together as a hybrid model substantially improve performance on all major VPR data sets illustrated using PR curves.
Maria Waheed, Sania Waheed, Michael Milford, Klaus D. McDonald-Maier, Shoaib Ehsan
IROS4
2023 DELICIOUS: Deadline-Aware Approximate Computing in Cache-Conscious Multicore
abstract
Enhancing result-accuracy in approximate computing (AC) based real-time systems, without violating power constraints of the underlying hardware, is a challenging problem. Execution of such AC real-time applications can be split into two parts: (i)the mandatory part, execution of which provides a result of acceptable quality, followed by (ii)the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy, without violating the time-constraint. This article introducesDELICIOUS, a novel hybrid offline-onlinescheduling strategyfor AC real-time dependent tasks. By employing an efficientheuristic algorithm,DELICIOUSfirst generates a schedule for a task-set with an objective to maximize the results-accuracy, while respecting system-wide constraints. During execution,DELICIOUSthen introduces aprudential cache resizingthat reduces temperature of the adjacent cores, by generating thermal buffers at the turned off cache ways.DELICIOUSfurther trades off this thermal benefits by enhancing the processing speed of the cores for a stipulated duration, calledV/F Spiking, without violating the power budget of the core, to shorten the execution length of the tasks. This reduced runtime is exploited either to enhance result-accuracy by dynamically adjusting the optional part, or to reduce temperature by enabling sleep mode at the cores. While surpassing the prior art,DELICIOUSoffers 80% result-accuracy with its scheduling strategy, which is further enhanced by 8.3% in online, while reducing runtime peak temperature by 5.8°C on average, as shown by benchmark based evaluation on a 4-core based multicore.
Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Rahul Gangopadhyay, Magnus Själander, Klaus D. McDonald-Maier
IEEE Trans. Parallel Distributed Syst.6
2022 Task Mapping and Scheduling in FPGA-based Heterogeneous Real-time Systems: A RISC-V Case-Study
abstract
Heterogeneous platforms, that integrate CPU and FPGA-based processing units, are emerging as a promising solution for accelerating various applications in the embedded system domain. However, in this context, so far, comprehensive studies that combine theoretical features of real-time task scheduling with practical runtime architectural characteristics have mostly been ignored. To fill this gap, in this paper we propose a real-time scheduling algorithm with the objective of minimizing the overall execution time under hardware resource constraints for heterogeneous CPU+FPGA architectures. In particular, we propose an Integer Linear Programming (ILP) based technique for task allocation and scheduling. We then show how to implement a given scheduling on a practical CPU+FPGA system regarding current technology restrictions and validate our methodology using a practical RISC-V case-study. Our experiments demonstrate that performance gains of 40 % and area usage reductions of 67 % are possible compared to a full software and hardware execution, respectively.
Sallar Ahmadi-Pour, Sangeet Saha, Vladimir Herdt, Rolf Drechsler, Klaus D. McDonald-Maier
DSD5
2022 OpenSceneVLAD: Appearance Invariant, Open Set Scene Classification
abstract
Scene classification is a well-established area of computer vision research that aims to classify a scene image into pre-defined categories such as playground, beach and airport. Recent work has focused on increasing the variety of pre-defined categories for classification, but so far failed to consider two major challenges: changes in scene appearance due to lighting and open set classification (the ability to classify unknown scene data as not belonging to the trained classes). Our first contribution, SceneVLAD, fuses scene classification and visual place recognition CNNs for appearance invariant scene classification that outperforms state-of-the-art scene classification by a mean F1 score of up to 0.1. Our second contribution, OpenSceneVLAD, extends the first to an open set classification scenario using intra-class splitting to achieve a mean increase in F1 scores of up to 0.06 compared to using state-of-the-art openmax layer. We achieve these results on three scene class datasets extracted from large scale outdoor visual localisation datasets, one of which we collected ourselves.
William H. B. Smith, Michael Milford, Klaus D. McDonald-Maier, Shoaib Ehsan, Robert B. Fisher
ICRA3
2022 SENAS: Security driven ENergy Aware Scheduler for Real Time Approximate Computing Tasks on Multi-Processor Systems
abstract
Present day real time approximate computing applications like image and video processing involves execution of a set of tasks before a certain amount of time or deadline. In addition to this, present day systems are associated with strict energy budget that cannot be changed post deployment. The tasks comprises of a mandatory and optional part. Completion of all mandatory portions of all tasks before deadline is much more important than result accuracy in such real time approximate computing applications. Based on the energy budget, the optional portions can be executed that determines the quality of service (QoS) of the system. In ideal scenario, sufficient energy budget is present that ensures completion of both mandatory and optional portions in a system with a pre-determined number of processors. However, if fault or malware attack occurs on one or more processors, then the system will cease to work and results may be fatal. In this work, we consider such a scenario where the processors may be faulty and stop functioning in post deployment phases or some malware may cause unexpected delays in processing or may cause unexpected power draining at runtime that will prevent the system from meeting its deadline. We propose a Security driven ENergy Aware Scheduler (SENAS) that works as a self aware agent. Initially, based on the available energy budget, SENAS determines which task is to be executed in which processor of a system. At runtime, SENAS constantly monitors the working of the processors and on detecting any anomaly in any of the processors, it reschedules its tasks at runtime by reducing execution of the optional portions of the tasks and ensuring completion before deadline with high QoS.
Krishnendu Guha, Sangeet Saha, Klaus D. McDonald-Maier
IOLTS3
2022 Highly-Efficient Binary Neural Networks for Visual Place Recognition
abstract
VPR is a fundamental task for autonomous navigation as it enables a robot to localize itself in the workspace when a known location is detected. Although accuracy is an essential requirement for a VPR technique, computational and energy efficiency are not less important for real-world applications. CNN-based techniques archive state-of-the-art VPR performance but are computationally intensive and energy demanding. Binary neural networks (BNN) have been recently proposed to address VPR efficiently. Although a typical BNN is an order of magnitude more efficient than a CNN, its processing time and energy usage can be further improved. In a typical BNN, the first convolution is not completely binarized for the sake of accuracy. Consequently, the first layer is the slowest network stage, requiring a large share of the entire computational effort. This paper presents a class of BNNs for VPR that combines depthwise separable factorization and binarization to replace the first convolutional layer to improve computational and energy efficiency. Our best model achieves higher VPR performance while spending considerably less time and energy to process an image than a BNN using a non-binary convolution as a first stage.
Bruno Ferrarini, Michael Milford, Klaus D. McDonald-Maier, Shoaib Ehsan
IROS3
2022 SwitchHit: A Probabilistic, Complementarity-Based Switching System for Improved Visual Place Recognition in Changing Environments
abstract
Visual place recognition (VPR) - a fundamental task in computer vision and robotics - is the problem of identifying a place mainly based on visual information. View-point and appearance changes, such as due to weather and seasonal variations, make this task challenging. Currently, there is no universal VPR technique that can work in all types of environments, on a variety of robotic platforms, and under a wide range of viewpoint and appearance changes. Recent work has shown the potential of combining different VPR methods intelligently by evaluating complementarity for some specific VPR datasets to achieve better performance. This, however, requires ground truth information (correct matches) which is not available when a robot is deployed in a real-world scenario. Moreover, running multiple VPR techniques in parallel may be prohibitive for resource-constrained embedded platforms. To overcome these limitations, this paper presents a probabilistic complementarity-based switching VPR system, SwitchHit. Our proposed system consists of multiple VPR techniques, however, it does not simply run all techniques at once, rather predicts the probability of correct match for an incoming query image and dynamically switches to another complementary technique if the probability of correctly matching the query is below a certain threshold. This innovative use of multiple VPR techniques allow our system to be more efficient and robust than other combined VPR approaches employing brute force and running multiple VPR techniques at once. Thus making it more suitable for resource constrained embedded systems and achieving an overall superior performance from what any individual VPR method in the system could have by achieved running independently.
Maria Waheed, Michael Milford, Klaus D. McDonald-Maier, Shoaib Ehsan
IROS3
2022 Benchmark Tool for Detecting Anomalous Program Behaviour on Embedded Devices
abstract
This paper presents an open-source benchmark tool for anomaly detection in program behaviour, using program counter (PC) and instruction type information. It is introducing anomalies in artificial way, allowing for fine-grained evaluation with adjustable sliding window sizes and preprocessing configuration. The usage of the benchmark, including demonstrated data collection, does not require any additional hardware other than a standard computer. The benchmark uses the output of llvm-objdump program to focus on non-library code which allows for rapid evaluation of various detection methods with different configurations. The proposed tool extracts features derived from processor’s PC and instruction type information and then utilizes the features to identify abnormal behavior using 4 different anomaly detection algorithms. New detection methods can be easily incorporated into the benchmark, which provides a solid foundation for evaluating novel, previously unseen methods against methods we selected for our experiment.
Michal Borowski, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier
TrustCom4
2022 InterpolatedXY: a two-step strategy to normalize DNA methylation microarray data avoiding sex bias
abstract
MOTIVATION: Data normalization is an essential step to reduce technical variation within and between arrays. Due to the different karyotypes and the effects of X chromosome inactivation, females and males exhibit distinct methylation patterns on sex chromosomes; thus, it poses a significant challenge to normalize sex chromosome data without introducing bias. Currently, existing methods do not provide unbiased solutions to normalize sex chromosome data, usually, they just process autosomal and sex chromosomes indiscriminately. RESULTS: Here, we demonstrate that ignoring this sex difference will lead to introducing artificial sex bias, especially for thousands of autosomal CpGs. We present a novel two-step strategy (interpolatedXY) to address this issue, which is applicable to all quantile-based normalization methods. By this new strategy, the autosomal CpGs are first normalized independently by conventional methods, such as funnorm or dasen; then the corrected methylation values of sex chromosome-linked CpGs are estimated as the weighted average of their nearest neighbors on autosomes. The proposed two-step strategy can also be applied to other non-quantile-based normalization methods, as well as other array-based data types. Moreover, we propose a useful concept: the sex explained fraction of variance, to quantitatively measure the normalization effect. AVAILABILITY AND IMPLEMENTATION: The proposed methods are available by calling the function 'adjustedDasen' or 'adjustedFunnorm' in the latest wateRmelon package (https://github.com/schalkwyk/wateRmelon), with methods compatible with all the major workflows, including minfi. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yucheng Wang 0008, Tyler J. Gorrie-Stone, Olivia A. Grant, Alexandria D. Andrayas, Xiaojun Zhai, Klaus D. McDonald-Maier, Leonard C. Schalkwyk
Bioinform.6
2022 ACCURATE: Accuracy Maximization for Real-Time Multicore Systems With Energy-Efficient Way-Sharing Caches
abstract
Improving result accuracy in approximate computing (AC)-based real-time applications without violating deadlines has recently become an active research domain. Execution time of AC real-time tasks can individually be separated into: execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve the result accuracy of the initial result within a given deadline. However, obtaining higher result accuracy at the cost of enhanced execution time may lead to deadline violation, along with higher energy usage. We present ACCURATE, a novel hybrid offline–online approximate real-time scheduling approach that first schedules AC-based tasks on multicore with an objective to maximize result accuracy and determines operational processing speeds for each task constrained by system-wide power limit, deadline, and task dependency. At runtime, by employing a way-sharing technique (WH_LLC) at the last level cache (LLC), ACCURATE improves performance, which is further leveraged, to enhance result accuracy by executing more from the optional part and to improve the energy efficiency of the cache by turning off a controlled number of cache ways. ACCURATE also exploits the slacks either to improve the result accuracy of the tasks or to enhance the energy efficiency of the underlying system, or both. ACCURATE achieves 85% QoS with 36% average reduction in cache leakage consumption with a 24% average gain in energy-delay product (EDP) for a 4-core-based chip multiprocessor (CMP) with 6.4% average improvement in performance.
Sangeet Saha, Shounak Chakraborty 0001, Xiaojun Zhai, Shoaib Ehsan, Klaus D. McDonald-Maier
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Effects of Non-Driving Related Tasks During Self-Driving Mode
abstract
Perception reaction time and mental workload have proven to be crucial in manual driving. Moreover, in highly automated cars, where most of the research is focusing on Level 4 Autonomous driving, take-over performance is also a key factor when taking road safety into account. This study aims to investigate how the immersion in non-driving related tasks affects the take-over performance of drivers in given scenarios. The paper also highlights the use of virtual simulators to gather efficient data that can be crucial in easing the transition between manual and autonomous driving scenarios. The use of Computer Aided Simulations is of absolute importance in this day and age since the automotive industry is rapidly moving towards Autonomous technology. An experiment comprising of 40 subjects was performed to examine the reaction times of driver and the influence of other variables in the success of take-over performance in highly automated driving under different circumstances within a highway virtual environment. The results reflect the relationship between reaction times under different scenarios that the drivers might face under the circumstances stated above as well as the importance of variables such as velocity in the success on regaining car control after automated driving. The implications of the results acquired are important for understanding the criteria needed for designing Human Machine Interfaces specifically aimed towards automated driving conditions. Understanding the need to keep drivers in the loop during automation, whilst allowing drivers to safely engage in other non-driving related tasks is an important research area which can be aided by the proposed study.
Saad Minhas, Aura Hernández-Sabaté, Shoaib Ehsan, Klaus D. McDonald-Maier
IEEE Trans. Intell. Transp. Syst.4
2022 Binary Neural Networks for Memory-Efficient and Effective Visual Place Recognition in Changing Environments
abstract
Visual place recognition (VPR) is a robot’s ability to determine whether a place was visited before using visual data. While conventional handcrafted methods for VPR fail under extreme environmental appearance changes, those based on convolutional neural networks (CNNs) achieve state-of-the-art performance but result in heavy runtime processes and model sizes that demand a large amount of memory. Hence, CNN-based approaches are unsuitable for resource-constrained platforms, such as small robots and drones. In this article, we take a multistep approach of decreasing the precision of model parameters, combining it with network depth reduction and fewer neurons in the classifier stage to propose a new class of highly compact models that drastically reduces the memory requirements and computational effort while maintaining state-of-the-art VPR performance. To the best of our knowledge, this is the first attempt to propose binary neural networks for solving the VPR problem effectively under changing conditions and with significantly reduced resource requirements. Our best-performing binary neural network, dubbed FloppyNet, achieves comparable VPR performance when considered against its full-precision and deeper counterparts while consuming 99% less memory and increasing the inference speed by seven times.
Bruno Ferrarini, Michael Milford, Klaus D. McDonald-Maier, Shoaib Ehsan
IEEE Trans. Robotics3
2022 RASA: Reliability-Aware Scheduling Approach for FPGA-Based Resilient Embedded Systems in Extreme Environments
abstract
Field-programmable gate arrays (FPGAs) offer the flexibility of general-purpose processors along with the performance efficiency of dedicated hardware that essentially renders it as a platform of choice for modern-day robotic systems for achieving real-time performance. Such robotic systems when deployed in harsh environments often get plagued by faults due to extreme conditions. Consequently, the real-time applications running on FPGA become susceptible to errors which call for a reliability-aware task scheduling approach, the focus of this article. We attempt to address this challenge using a hybrid offline-online approach. Given a set of periodic real-time tasks that require to be executed, the offline component generates a feasible preemptive schedule with specific preemption points. At runtime, these preemption events are utilized for fault detection. Upon detecting any faulty execution at such distinct points, the reliability-aware scheduling approach, RASA, orchestrates the recovery mechanism to remediate the scenario without jeopardizing the predefined schedule. Effectiveness of the proposed strategy has been verified through simulation-based experiments and we observed that the RASA is able to achieve 72% of task acceptance rate even under 70% of system workloads with high fault occurrence rates.
Sangeet Saha, Xiaojun Zhai, Shoaib Ehsan, Shakaiba Majeed, Klaus D. McDonald-Maier
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Energy-Aware Real-Time Tasks Processing for FPGA-Based Heterogeneous Cloud
abstract
Cloud computing is becoming a popular model of computing. Due to the increasing complexity of the cloud service request, it often exploits heterogeneous architecture. Moreover, some service requests (SRs)/tasks exhibit real-time features, which are required to be handled within a specified duration. Along with the stipulated temporal management, the strategy should also be energy efficient, as energy consumption in cloud computing is challenging. In this paper, we have proposed a strategy, called “Efficient Resource Allocation of Service Request” (ERASER) for energy efficient allocation and scheduling of periodic real-time SRs on cloud platform. The cloud platform is consists of Field Programmable Gate Arrays (FPGAs) as Processing Elements (PEs) along with the General Purpose Processors (GPP). We have further proposed, an SR migration technique to reduce the tasks rejection by serving maximum SRs. Simulation based experimental results demonstrate that the proposed methodology is capable to achieve upto 90 percent resource utilization with only 26 percent SR rejection rate over different experimental scenarios. Comparison results with other state-of-the-art techniques reveal that the proposed strategy outperforms the existing technique with 17 percent reduction in SR rejection rate and 21 percent reduction in energy consumption. Further, the simulation outcomes have been validated on real FPGA test-bed based on Xilinx Zynq SoC with standard benchmark tasks.
Atanu Majumder, Sangeet Saha, Amlan Chakrabarti, Klaus D. McDonald-Maier
IEEE Trans. Sustain. Comput.4
2021 EnSuRe: Energy & Accuracy Aware Fault-tolerant Scheduling on Real-time Heterogeneous Systems
abstract
This paper proposes an energy efficient real-time scheduling strategy called EnSuRe, which (i) executes real-time tasks on low power consuming primary processors to enhance the system accuracy by maintaining the deadline and (ii) provides reliability against a fixed number of transient faults by selectively executing backup tasks on high power consuming backup processor. Simulation results reveal that EnSuRe consumes nearly 25% less energy, compared to existing techniques, while satisfying the fault tolerance requirements. EnSuRe is also able to achieve 75% system accuracy with 50% system utilisation. Further, the obtained simulation outcomes are validated on benchmark tasks via a fault injection framework on Xilinx ZYNQ APSoC heterogeneous dual core platform.
Sangeet Saha, Adewale Adetomi, Xiaojun Zhai, Server Kasap, Shoaib Ehsan, Tughrul Arslan, Klaus D. McDonald-Maier
IOLTS7
2021 Design and Implementation of a RISC V Processor on FPGA
abstract
The RISC-V ISA is becoming one of the leading instruction sets for the Internet-of-Things and System-on-Chip applications. Due to its strong security features and open-source nature, it is becoming a competitor to the popular ARM architecture. This paper describes the design of a light weight, open-source implementation of a RISCV processor using modern hardware design teclmiques, the implementation of the design onto a Field Programmable Gate Array (FPGA), and its testing. We wanted to create a RISC-V processor that is easy for beginners to learn from and lightweight enough to be implemented on even small FPGAs. While there are existing opensource implementations of RISC-V processors, none are intuitive enough for a beginner to follow. For this reason, in this paper we have minimised the use of conventions and components in modern processors that are not strictly necessary for a barebones implementation. For example, the processor does not include pipelining and uses a simple Harvard architecture. The barebones nature of the design allows for a lot of potential for upgradability. The implementation of each component, and the corresponding test benches, are written in concise and conventional System Verilog. The project produced a RISC-V processor with files for targeting Basys 3 Artix-7 FPGA. Performance was tested using the Dhyrstone benchmark and achieved a strong 2276 DMIPs/MHz, even outperforming the ARM Cortex-A9, while maintaining very low resource utilization on the FPGA.
Ludovico Poli, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier
MSN4
2021 VPR-Bench: An Open-Source Visual Place Recognition Evaluation Framework with Quantifiable Viewpoint and Appearance Change
abstract
Abstract Visual place recognition (VPR) is the process of recognising a previously visited place using visual information, often under varying appearance conditions and viewpoint changes and with computational constraints. VPR is related to the concepts of localisation, loop closure, image retrieval and is a critical component of many autonomous navigation systems ranging from autonomous vehicles to drones and computer vision systems. While the concept of place recognition has been around for many years, VPR research has grown rapidly as a field over the past decade due to improving camera hardware and its potential for deep learning-based techniques, and has become a widely studied topic in both the computer vision and robotics communities. This growth however has led to fragmentation and a lack of standardisation in the field, especially concerning performance evaluation. Moreover, the notion of viewpoint and illumination invariance of VPR techniques has largely been assessed qualitatively and hence ambiguously in the past. In this paper, we address these gaps through a new comprehensive open-source framework for assessing the performance of VPR techniques, dubbed “VPR-Bench”. VPR-Bench (Open-sourced at: https://github.com/MubarizZaffar/VPR-Bench ) introduces two much-needed capabilities for VPR researchers: firstly, it contains a benchmark of 12 fully-integrated datasets and 10 VPR techniques, and secondly, it integrates a comprehensive variation-quantified dataset for quantifying viewpoint and illumination invariance. We apply and analyse popular evaluation metrics for VPR from both the computer vision and robotics communities, and discuss how these different metrics complement and/or replace each other, depending upon the underlying applications and system requirements. Our analysis reveals that no universal SOTA VPR technique exists, since: (a) state-of-the-art (SOTA) performance is achieved by 8 out of the 10 techniques on at least one dataset, (b) SOTA technique in one community does not necessarily yield SOTA performance in the other given the differences in datasets and metrics. Furthermore, we identify key open challenges since: (c) all 10 techniques suffer greatly in perceptually-aliased and less-structured environments, (d) all techniques suffer from viewpoint variance where lateral change has less effect than 3D change, and (e) directional illumination change has more adverse effects on matching confidence than uniform illumination change. We also present detailed meta-analyses regarding the roles of varying ground-truths, platforms, application requirements and technique parameters. Finally, VPR-Bench provides a unified implementation to deploy these VPR techniques, metrics and datasets, and is extensible through templates.
Mubariz Zaffar, Sourav Garg, Michael Milford, Julian F. P. Kooij, David Flynn, Klaus D. McDonald-Maier, Shoaib Ehsan
Int. J. Comput. Vis.6
2021 A novel ICMetric public key framework for secure communication
Ruhma Tahir, Shahzaib Tahir, Hasan Tahir, Klaus D. McDonald-Maier, Gareth Howells 0001, Ali Sajjad
J. Netw. Comput. Appl.4
2021 Prepare: Power-Aware Approximate Real-time Task Scheduling for Energy-Adaptive QoS Maximization
abstract
Achieving high result-accuracy in approximate computing (AC) based real-time applications without violating power constraints of the underlying hardware is a challenging problem. Execution of such AC real-time tasks can be divided into the execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve accuracy of the initially obtained result within the given time-limit. However, enhancing result-accuracy at the cost of increased execution length might lead to deadline violations with higher energy usage. We propose Prepare , a novel hybrid offline-online approximate real-time task-scheduling approach, that first schedules AC-based tasks and determines operational processing speeds for each individual task constrained by system-wide power limit, deadline, and task-dependency. At runtime, by employing fine-grained DVFS, the energy-adaptive processing speed governing mechanism of Prepare reduces processing speed during each last level cache miss induced stall and scales up the processing speed once the stall finishes to a higher value than the predetermined one. To ensure on-chip thermal safety, this higher processing speed is maintained only for a short time-span after each stall, however, this reduces execution times of the individual task and generates slacks. Prepare exploits the slacks either to enhance result-accuracy of the tasks, or to improve thermal and energy efficiency of the underlying hardware, or both. With a 70 - 80% workload, Prepare offers 75% result-accuracy with its constrained scheduling, which is enhanced by 5.3% for our benchmark based evaluation of the online energy-adaptive mechanism on a 4-core based homogeneous chip multi-processor, while meeting the deadline constraint. Overall, while maintaining runtime thermal safety, Prepare reduces peak temperature by up to 8.6 °C for our baseline system. Our empirical evaluation shows that constrained scheduling of Prepare outperforms a state-of-the-art scheduling policy, whereas our runtime energy-adaptive mechanism surpasses two current DVFS based thermal management techniques.
Shounak Chakraborty 0001, Sangeet Saha, Magnus Själander, Klaus D. McDonald-Maier
ACM Trans. Embed. Comput. Syst.4
2021 Memorable Maps: A Framework for Re-Defining Places in Visual Place Recognition
abstract
This paper presents a cognition-inspired agnostic framework for building a map for Visual Place Recognition. This framework draws inspiration from human-memorability, utilizes the traditional image entropy concept and computes the static content in an image; thereby presenting a tri-folded criteria to assess the ‘memorability’ of an image for visual place recognition. A dataset namely ‘ESSEX3IN1’ is created, composed of highly confusing images from indoor, outdoor and natural scenes for analysis. When used in conjunction with state-of-the-art visual place recognition methods, the proposed framework provides significant performance boost to these techniques, as evidenced by results on ESSEX3IN1 and other public datasets.
Mubariz Zaffar, Shoaib Ehsan, Michael Milford, Klaus D. McDonald-Maier
IEEE Trans. Intell. Transp. Syst.4
2020 Temporal Motionless Analysis of Video using CNN in MPSoC
abstract
This paper proposes a novel human-inspired methodology called IRON-MAN (Integrated RatiONal prediction and Motionless ANalysis of videos) on mobile multi-processor systems-on-chips (MPSoCs). The methodology integrates analysis of the previous image frames of the video to represent the analysis of the current frame in order to perform Temporal Motionless Analysis of the Video (TMAV). This is the first work on TMAV using Convolutional Neural Network (CNN) for scene prediction in MPSoCs. Experimental results show that our methodology outperforms state-of-the-art. We also introduce a metric named, Energy Consumption per Training Image (ECTI) to assess the suitability of using a CNN model in mobile MPSoCs with a focus on energy consumption of the device.
Somdip Dey, Amit Kumar Singh 0002, Dilip K. Prasad, Klaus D. McDonald-Maier
ASAP4
2020 User Interaction Aware Reinforcement Learning for Power and Thermal Efficiency of CPU-GPU Mobile MPSoCs
abstract
Mobile user’s usage behaviour changes throughout the day and the desirable Quality of Service (QoS) could thus change for each session. In this paper, we propose a QoS aware agent to monitor mobile user’s usage behaviour to find the target frame rate, which satisfies the desired user’s QoS, and applies reinforcement learning based DVFS on a CPU-GPU MPSoC to satisfy the frame rate requirement. Experimental study on a real Exynos hardware platform shows that our proposed agent is able to achieve a maximum of 50% power saving and 29% reduction in peak temperature compared to stock Android’s power saving scheme. It also outperforms the existing state-of-the-art power and thermal management scheme by 41% and 19%, respectively.
Somdip Dey, Amit Kumar Singh 0002, Xiaohang Wang 0001, Klaus D. McDonald-Maier
DATE4
2020 A self-scrubbing scheme for embedded systems in radiation environments
abstract
As one of the most important components in the embedded systems, the SRAM are sensitive to radiation effects. When the embedded systems working in the extreme radiation environments, the bit flips could occur frequently and decrease the reliability of the systems significantly. In this paper, the self-scrubbing RAM scheme is proposed for light wight embedded systems in the extreme radiation environments. In the scheme, both scrubbing and ECC are used to mitigate the large number of the errors in the RAMs. The separately scrubber is designed to scrub the RAM separately. Therefore it is can be able to operating the scrubbing, when the CPUs are busy. In addition, the scrubber is a portable modules and the hardware costs do not grow with the size of the available RAM. The results of the real world radiation experiments show that it can correct most errors in the RAM under neutron radiation where the errors rates in unhardened RAMs is approximately 1.2bit/(KB·h). The results of the 6 hours radiation experiments show that the error rates of in the conventional ECC RAM is approximately 4.3×10-4bit/(KB·h), while the self-scrubbing RAMs is less than 8.7×10-5bit/(KB·h).
Yufan Lu, Xiaojun Zhai, Sangeet Saha, Shoaib Ehsan, Klaus D. McDonald-Maier
IOLTS5
2020 A Framework and Protocol for Dynamic Management of Fault Tolerant Systems in Harsh Environments
abstract
Robots can be used to deal with hazardous materials like nuclear waste. Unfortunately, electronic components are also susceptible to radiation effects. Current proposals to tackle this issue solve only parts of the problems for the specific scenarios and the specific types of radiation. At the same time, current computational devices should provide run-time capabilities to monitor and adapt to different situations. In this paper, we target a possible solution presenting a framework which provides the flexibility to employ fault-tolerant techniques on distributed systems. As proof of concept, we target a fault-tolerant technique to extend the operating time of systems in harsh environments. Results show a very low overhead, of few microseconds, to execute a majority voter with replicated tasks.
Eduardo Wächter, Server Kasap, Xiaojun Zhai, Shoaib Ehsan, Klaus D. McDonald-Maier
IOLTS5
2020 Proxy Circuits for Fault-Tolerant Primitive Interfacing in Reconfigurable Devices Targeting Extreme Environments
abstract
Continuous interface access to device-level primitives in reconfigurable devices in extreme environments is key to reliable operation. However, it is possible for a primitive's interface controller, which is static to be rendered non-operational by a permanent damage in the controller's circuitry. In order to mitigate this, this paper proposes the use of relocatable proxy circuits to provide remote interfacing capability to primitives from anywhere on a reconfigurable device. A demonstration with device register read controller shows that an improvement in fault-tolerance can be achieved.
Adewale Adetomi, Sangeet Saha, Klaus D. McDonald-Maier, Tughrul Arslan
ISCAS3
2020 A Holistic Visual Place Recognition Approach Using Lightweight CNNs for Significant ViewPoint and Appearance Changes
abstract
This article presents a lightweight visual place recognition approach, capable of achieving high performance with low computational cost, and feasible for mobile robotics under significant viewpoint and appearance changes. Results on several benchmark datasets confirm an average boost of 13% in accuracy, and 12x average speedup relative to state-of-the-art methods.
Ahmad Khaliq, Shoaib Ehsan, Zetao Chen, Michael Milford, Klaus D. McDonald-Maier
IEEE Trans. Robotics5
2019 TEEM: Online Thermal- and Energy-Efficiency Management on CPU-GPU MPSoCs
abstract
Heterogeneous Multiprocessor System-on-Chip (MPSoC) are progressively becoming predominant in most modern mobile devices. These devices are required to perform processing of applications within thermal, energy and performance constraints. However, most stock power and thermal management mechanisms either neglect some of these constraints or rely on frequency scaling to achieve energy-efficiency and temperature reduction on the device. Although this inefficient technique can reduce temporal thermal gradient, but at the same time hurts the performance of the executing task. In this paper, we propose a thermal and energy management mechanism which achieves reduction in thermal gradient as well as energy-efficiency through resource mapping and thread-partitioning of applications with online optimization in heterogeneous MPSoCs. The efficacy of the proposed approach is experimentally appraised using different applications from Polybench benchmark suite on Odroid-XU4 developmental platform. Results show 28% performance improvement, 28.32% energy saving and reduced thermal variance of over 76% when compared to the existing approaches. Additionally, the method is able to free more than 90% in memory storage on the MPSoC, which would have been previously utilized to store several task-to-thread mapping configurations.
Samuel Isuwa, Somdip Dey, Amit Kumar Singh 0002, Klaus D. McDonald-Maier
DATE4
2019 Security and PrIvacy foR the Internet of Things: an overview of the project
abstract
As the adoption of digital technologies expands, it becomes vital to build trust and confidence in the integrity of such technology. The SPIRIT project investigates the proof of concept of employing novel secure and privacy-ensuring techniques in services set-up in the Internet of Things (IoT) environment, aiming to increase the trust of users in IoTbased systems. The proposed system integrates three highly novel technology concepts developed by the consortium partners. Specifically, a technology, ermed ICMetrics, for deriving encryption keys directly from the operating characteristics of digital devices; secondly, a technology based on a contentbased signature of user data in order to ensure the integrity of sentdata upon arrival; a third technology, termed semantic firewall, which is able to allow or deny the transmission of data derived from an IoT device according to the information contained within the data and the information gathered about the requester.
Sabrine Aroua, Julian Murphy, Mourad Rabah, Kais Rouis, Nicolas Sidere, Nouredine Tamani, Ronan Champagnat, Mickaël Coustaty, Gilles Falquet, Sami Ghadfi, Yacine Ghamri-Doudane, Petra Gomez-Krämer, Gareth Howells 0001, Klaus D. McDonald-Maier
SMC14
2019 Profi-Load: An FPGA-Based Solution for Generating Network Load in Profinet Communication
abstract
Industrial automation has received a considerable attention in the last few years with the rise of Internet of Things (IoT). Specifically, industrial communication network technology such as Profinet has proved to be a major game changer for such automation. However, industrial automation devices often have to exhibit robustness to dynamically changing network conditions and thus, demand a rigorous testing environment to avoid any safety-critical failures. Hence, in this paper, we have proposed an FPGA-based novel framework called “Profi-Load” to generate Profinet traffic with specific intensities for a specified duration of time. The proposed Profi-Load intends to facilitate the performance testing of the industrial automated devices under various network conditions. By using the advantage of inherent hardware parallelism and re-configurable features of FPGA, Profi-Load is able to generate Profinet traffic efficiently. Moreover, it can be reconfigured on the fly as per the specific requirements. We have developed our proposed Profi-Load framework by employing the Xilinx-based “NetJury” device which belongs to Zynq-7000 FPGA family. A series of experiments have been conducted to evaluate the effectiveness of Profi-Load and it has been observed that Profi-Load is able to generate precise load at a constant rate for stringent timing requirements. Furthermore, a suitable Human Machine Interface (HMI) has also been developed for quick access to our framework. The HMI at the client side can directly communicate with the NetJury device and parameters such as, required load amount, number of packet(s) to be sent or desired time duration can be selected using the HMI.
Ahmad Khaliq, Sangeet Saha, Bina Bhatt, Dongbing Gu, Klaus D. McDonald-Maier
SMC5
2017 Classification of Graphomotor Impressions Using Convolutional Neural Networks: An Application to Automated Neuro-Psychological Screening Tests
abstract
Graphomotor impressions are a product of complex cognitive, perceptual and motor skills and are widely used as psychometric tools for the diagnosis of a variety of neuro-psychological disorders. Apparent deformations in these responses are quantified as errors and are used are indicators of various conditions. Contrary to conventional assessment methods where manual analysis of impressions is carried out by trained clinicians, an automated scoring system is marked by several challenges. Prior to analysis, such computerized systems need to extract and recognize individual shapes drawn by subjects on a sheet of paper as an important pre-processing step. The aim of this study is to apply deep learning methods to recognize visual structures of interest produced by subjects. Experiments on figures of Bender Gestalt Test (BGT), a screening test for visuo-spatial and visuo-constructive disorders, produced by 120 subjects, demonstrate that deep feature representation brings significant improvements over classical approaches. The study is intended to be extended to discriminate coherent visual structures between produced figures and expected prototypes.
Haris Bin Nazar, Momina Moetesum, Shoaib Ehsan, Imran Siddiqi, Khurram Khurshid, Nicole Vincent, Klaus D. McDonald-Maier
ICDAR7
2015 An intrusion detection system against malicious attacks on the communication network of driverless cars
abstract
Vehicular ad hoc networking (VANET) have become a significant technology in the current years because of the emerging generation of self-driving cars such as Google driverless cars. VANET have more vulnerabilities compared to other networks such as wired networks, because these networks are an autonomous collection of mobile vehicles and there is no fixed security infrastructure, no high dynamic topology and the open wireless medium makes them more vulnerable to attacks. It is important to design new approaches and mechanisms to rise the security these networks and protect them from attacks. In this paper, we design an intrusion detection mechanism for the VANETs using Artificial Neural Networks (ANNs) to detect Denial of Service (DoS) attacks. The main role of IDS is to detect the attack using a data generated from the network behavior such as a trace file. The IDSs use the features extracted from the trace file as auditable data. In this paper, we propose anomaly and misuse detection to detect the malicious attack.
Khattab M. Ali Alheeti, Anna Gruebler, Klaus D. McDonald-Maier
CCNC3
2015 Exploring ICMetrics to detect abnormal program behaviour on embedded devices
Xiaojun Zhai, Kofi Appiah, Shoaib Ehsan, Gareth Howells 0001, Huosheng Hu, Dongbing Gu, Klaus D. McDonald-Maier
J. Syst. Archit.7
2015 A Method for Detecting Abnormal Program Behavior on Embedded Devices
abstract
A potential threat to embedded systems is the execution of unknown or malicious software capable of triggering harmful system behavior, aimed at theft of sensitive data or causing damage to the system. Commercial off-the-shelf embedded devices, such as embedded medical equipment, are more vulnerable as these type of products cannot be amended conventionally or have limited resources to implement protection mechanisms. In this paper, we present a self-organizing map (SOM)-based approach to enhance embedded system security by detecting abnormal program behavior. The proposed method extracts features derived from processor's program counter and cycles per instruction, and then utilises the features to identify abnormal behavior using the SOM. Results achieved in our experiment show that the proposed method can identify unknown program behaviors not included in the training set with over 98.4% accuracy.
Xiaojun Zhai, Kofi Appiah, Shoaib Ehsan, Gareth Howells 0001, Huosheng Hu, Dongbing Gu, Klaus D. McDonald-Maier
IEEE Trans. Inf. Forensics Secur.7
2014 Image fusion using multivariate and multidimensional EMD
abstract
We present a novel methodology for the fusion of multiple (two or more) images using the multivariate extension of empirical mode decomposition (MEMD). Empirical mode decomposition (EMD) is a data-driven method which decomposes input data into its intrinsic oscillatory modes, known as intrinsic mode functions (IMFs), without making a priori assumptions regarding the data. We show that the multivariate and multidimensional extensions of EMD are suitable for image fusion purposes. We further demonstrate that while multidimensional extensions, by design, may seem more appropriate for tasks related to image processing, the proposed multivariate extension outperforms these in image fusion applications owing to its mode-alignment property for IMFs. Case studies involving multi-focus image fusion and pan-sharpening of multi-spectral images are presented to demonstrate the effectiveness of the proposed method.
Naveed ur Rehman, Muhammad Murtaza Khan, M. I. Sohaib, M. Jehanzaib, Shoaib Ehsan, Klaus D. McDonald-Maier
ICIP6
2012 An Algorithm for the Contextual Adaption of SURF Octave Selection With Good Matching Performance: Best Octaves
abstract
Speeded-Up Robust Features is a feature extraction algorithm designed for real-time execution, although this is rarely achievable on low-power hardware such as that in mobile robots. One way to reduce the computation is to discard some of the scale-space octaves, and previous research has simply discarded the higher octaves. This paper shows that this approach is not always the most sensible and presents an algorithm for choosing which octaves to discard based on the properties of the imagery. Results obtained with this best octaves algorithm show that it is able to achieve a significant reduction in computation without compromising matching performance.
Shoaib Ehsan, Nadia Kanwal, Adrian F. Clark, Klaus D. McDonald-Maier
IEEE Trans. Image Process.4
2010 A Fuzzy Logic Reconfiguration Engine for Symmetric Chip Multiprocessors
abstract
Recent developments in reconfigurable multiprocessor system on chip (MPSoC) have offered system designers a great amount of flexibility to exploit task concurrency with higher throughput and less energy consumption. This paper presents a novel fuzzy logic reconfiguration engine (FLRE) for coarse grain MPSoC reconfiguration that facilitates to identify an optimum balance between power and performance of the system. The FLRE is composed on two levels of abstraction layers. The system selects an optimal configuration of Level 1 / Level 2 cache size and Associativity, processor operating frequency and voltage, the number of cores based on miss rate, and energy and throughput information of the system both at core and SoC level. An 8-core symmetric chip multiprocessor has been used to evaluate the proposed scheme. The results show an overall decrease of energy consumption with not more than 30% decrease in the throughput.
Muhammad Yasir Qadri, Klaus D. McDonald-Maier
CISIS2
2010 Model Learning from Weights by Adaptive Enhanced Probabilistic Convergent Network
Pierre Lorrentz, Gareth Howells 0001, Klaus D. McDonald-Maier
ESANN3
2010 A Novel Weightless Artificial Neural Based Multi-Classifier for Complex Classifications
Pierre Lorrentz, Gareth Howells 0001, Klaus D. McDonald-Maier
Neural Process. Lett.3
2009 FPGA-based enhanced probabilistic convergent weightless Network for human Iris recognition
Pierre Lorrentz, Gareth Howells 0001, Klaus D. McDonald-Maier
ESANN3
2008 Dynamic Scheduling of Test Routines for Efficient Online Self-Testing of Embedded Microprocessors
abstract
This paper presents a self-testing framework targeting the LEON3 embedded microprocessor with built-in test-scheduling features. The proposed design exploits existing post production test sets, designed for software-based testing of embedded microprocessors. The framework also includes a constraint-based approach of test-routine scheduling. The initial results show that the test execution time could be dynamically scaled by the test selection algorithm.
Nikolaos G. Bartzoudis, Vasileios Tantsios, Klaus D. McDonald-Maier
IOLTS3
2008 A Model-Driven Development Approach to Mapping UML State Diagrams to Synthesizable VHDL
abstract
With the continuing rise in the complexity of embedded systems, there is an emerging need for a higher level modelling environment that facilitates efficient handling of this complexity. The aim here is to produce such a high level environment using Model Driven Development (MDD) techniques that maps a high level abstract description of an electronic embedded system into its low level implementation details. The Unified Modelling Language (UML) is a high level graphical based language that is broad enough in scope to model embedded systems hardware circuits. The authors have developed a framework for deriving Very High Speed Integrated Circuits Hardware Description Language (VHDL) code from UML state diagrams and defined a set of rules that enable automated generation of synthesisable VHDL code from UML specifications using MDD techniques. By adopting the techniques and tools described in this paper the design and implementation of complex state-based systems is greatly simplified.
Stephen Wood, David H. Akehurst, O. Uzenkov, Gareth Howells 0001, Klaus D. McDonald-Maier
IEEE Trans. Computers5
2007 Integrating Multi-Modal Circuit Features within an Efficient Encryption System
abstract
The problem of the incorporation of pattern features with unusual distributions is well known within pattern recognition systems even if not easily addressed. The problem is more acute when features are derived from characteristics of given integrated electronic circuits. The current paper introduces novel efficient techniques for normalising sets of features which are highly multi-modal in nature, so as to allow them to be incorporated within a single encryption key generation system based primarily on measured hardware characteristics. The utility of the proposed system lies in the observation that the need for data sent to and from remote network nodes to be secure and verified is substantial. Security can be improved by using encryption techniques based on keys, which are based on unique properties of the individual nodes within the network. This will serve both to minimize the need for key storage and sharing as well as to validate the initiator node of a message.
Evangelos Papoutsis, Gareth Howells 0001, Andrew B. T. Hopkins, Klaus D. McDonald-Maier
IAS4
2007 Compiling UML State Diagrams into VHDL: An Experiment in Using Model Driven Development
David H. Akehurst, Gareth Howells 0001, Klaus D. McDonald-Maier, Behzad Bordbar
FDL3
2007 Online monitoring of FPGA-based co-processing engines embedded in dependable workstations
abstract
An assertion-based monitoring system was implemented to enforce the operation of a FPGA coprocessing engine, which is part of a dependable workstation. The monitor was built in VHDL using simple state machines. Concurrent error detection is an important aspect in dependable workstations helping to prevent the propagation of errors within the system. The functionality of the implemented monitoring component does not interfere with the performance of the workstation and ensures high system availability. The monitor checks for PCI protocol and application errors based on the forensic analysis of a dependable workstation with FPGA- based co-processing support.
Nikolaos G. Bartzoudis, Klaus D. McDonald-Maier
IOLTS2
2007 Implementing associations: UML 2.0 to Java 5
David H. Akehurst, Gareth Howells 0001, Klaus D. McDonald-Maier
Softw. Syst. Model.3
2006 Debug support for embedded processor reuse
abstract
This research presents a test-bench implementation of a novel debug support system that targets the needs of hard real-time embedded systems. The solution provides over 70 percent combined program and data trace compression using a low complexity messaging framework and subtraction based differential compression. The test-bench is based on an open source multi-processor system-on-chip design, where novel debug support has been attached through defined interfaces and close integration of the debug adapter with the processor cores. Synthesis to an FPGA platform shows that the approach of using core adapters to implement debug support with a defined core generic interface is practical, and that the resulting circuitry is compact
Andrew B. T. Hopkins, Klaus D. McDonald-Maier
ISCAS2
2006 SiTra: Simple Transformations in Java
David H. Akehurst, Behzad Bordbar, M. J. Evans, Gareth Howells 0001, Klaus D. McDonald-Maier
MoDELS5
2006 Debug Support Strategy for Systems-on-Chips with Multiple Processor Cores
abstract
On-chip program and data tracing is now an essential part of any system level development platform for system-on-chip (SoC). Current debug support solutions are platform specific and incompatible with processors and active peripherals from other sources, restricting effective design reuse. In order to overcome this reuse challenge, this paper defines interfaces to decouple the debug support from processor cores and other active data accessing units. The on-chip debug support infrastructure is also decoupled from each core's debug support and from the trace port or trace memory, using an additional interface. As a result, this decoupling of the debug support infrastructure provides freedom from a specific SoC platform. These interfaces are applied through a reference design modeled using VHDL that is based on a novel low overhead trace message framework. Compared with a leading implementation of a relevant standard, the reference design is 50 percent more compact while providing improvements in trace compression of 8.4 percent for program trace messages and almost 24 percent for data trace messages. This reference design is a multiple core solution that is compatible with most SoC architectures, including those based on emerging network-on-chip architectures.
Andrew B. T. Hopkins, Klaus D. McDonald-Maier
IEEE Trans. Computers2
2005 Debug Support, Calibration and Emulation for Multiple Processor and Powertrain Control SoCs
abstract
The introduction of complex SoCs with multiple processor cores presents new development challenges, such that development support is now a decisive factor when choosing a system-on-chip (SoC). The presented development support strategy addresses the challenges using both architecture and technology approaches. The multi-core debug support (MCDS) architecture provides flexible triggering using cross triggers and a multiple core break and suspend switch. Temporal trace ordering is guaranteed down to cycle level by on-chip time stamping. The package sized-ICE (PSI) approach is a novel method of including trace buffers, overlay memories, processing resources and communication interfaces without changing device behavior. PSI requires no external emulation box, as the debug host interfaces directly with the SoC using a standard interface.
Albrecht Mayer, Harry Siebert, Klaus D. McDonald-Maier
DATE3
2000 Controlling fast spring-legged locomotion with artificial neural networks
Klaus D. McDonald-Maier, Volkmar Glauche, Clemens Beckstein, Reinhard Blickhan
Soft Comput.1