EDBT 2026 Demo / reviewers in the wild / expert
Geoff V. Merrett
dblp:79/2874
· DBLP profile ↗
74ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-4980-3894ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 2 first-author · 14 since 2021Software engineering, systems software and programming languages · 20 · 1 first-author · 5 since 2021Computer networks · 15 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: Adaptive Ensembles of Dynamic DNNs for Collaborative Edge InferenceabstractEdge computing enables low-latency and privacy-preserving DNN inference, yet heterogeneous and dynamically changing device resources make it difficult to satisfy real-time constraints. In this paper, we present AdaEnsemble, an adaptive and collaborative ensemble inference framework that integrates Dynamic DNNs with deadline-aware scheduling. The system profiles accuracy and latency offline and selects both model widths and participating devices at runtime to maximize accuracy under a given deadline. Experiments on heterogeneous edge devices show that AdaEnsemble adapts effectively to different latency requirements and consistently outperforms the state-of-the-art. Mingyu Hu, Amit Kumar Singh 0002, Jonathon S. Hare, Geoff V. Merrett |
DATE | 4 |
| 2025 | Power- and Deadline-Aware Dynamic Inference on Intermittent Computing SystemsabstractIn energy-harvesting intermittent computing systems, balancing power constraints with the need for timely and accurate inference remains a critical challenge. Existing methods often sacrifice significant accuracy or fail to adapt effectively to fluctuating power conditions. This paper presents DualAdaptNet, a power- and deadline-aware neural network architecture that dynamically adapts both its width and depth to ensure reliable inference under variable power conditions. Additionally, a runtime scheduling method is introduced to select an appropriate sub-network configuration based on real-time energy-harvesting conditions and system deadlines. Experimental results on the MNIST dataset demonstrate that our approach completes up to 7.0% more inference tasks within a specified deadline, while also improving average accuracy by 15.4% compared to the state-of-the-art. Hengrui Zhao, Lei Xun, Jagmohan Chauhan, Geoff V. Merrett |
DATE | 4 |
| 2025 | FedTMOS: Efficient One-Shot Federated Learning with Tsetlin MachineabstractOne-Shot Federated Learning (OFL) is a promising approach that reduce communication to a single round, minimizing latency and resource consumption. However, existing OFL methods often rely on Knowledge Distillation, which introduce server-side training, increasing latency. While neuron matching and model fusion techniques bypass server-side training, they struggle with alignment when heterogeneous data is present. To address these challenges, we proposed One-Shot Federated Learning with Tsetlin Machine (FedTMOS), a novel data-free OFL framework built upon the low-complexity and class-adaptive properties of the Tsetlin Machine. FedTMOS first clusters then reassigns class-specific weights to form models using an inter-class maximization approach, efficiently generating balanced server models without requiring additional training. Our extensive experiments demonstrate that FedTMOS significantly outperforms its ensemble counterpart by an average of $6.16$%, and the leading state-of-the-art OFL baselines by $7.22$% across various OFL settings. Moreover, FedTMOS achieves at least a $2.3\times$ reduction in upload communication costs and a $75\times$ reduction in server latency compared to methods requiring server-side training. These results establish FedTMOS as a highly efficient and practical solution for OFL scenarios. Shannon How Shi Qi, Jagmohan Chauhan, Geoff V. Merrett, Jonathon S. Hare |
ICLR | 3 |
| 2025 | Realization of Early-Exit Dynamic Neural Networks on Reconfigurable HardwareabstractEarly-exiting is a strategy that is becoming popular in deep neural networks (DNNs), as it can lead to faster execution and a reduction in the computational intensity of inference. To achieve this, intermediate classifiers abstract information from the input samples to strategically stop forward propagation and generate an output at an earlier stage. Confidence criteria are used to identify easier-to-recognize samples over the ones that need further filtering. However, such dynamic DNNs have only been realized in conventional computing systems (CPU+GPU) using libraries designed for static networks. In this article, we first explore the feasibility and benefits of realizing early-exit dynamic DNNs on field-programmable gate arrays (FPGAs), a platform already proven to be highly effective for neural network applications. We consider two approaches for implementing and executing the intermediate classifiers: 1) pipeline, which uses existing hardware and 2) parallel, which uses additional dedicated modules. We model their energy needs and execution time and explore their performance using the BranchyNet early-exit approach on LeNet-5, AlexNet, VGG19, and ResNet32, and a Xilinx ZCU106 Evaluation Board. We found that the dynamic approaches are at least 24% faster than a static network executed on an FPGA, consuming a minimum of$1.32\times $lower energy. We further observe that FPGAs can enhance the performance of early-exit dynamic DNNs by minimizing the complexities introduced by the decision intermediate classifiers through parallel execution. Finally, we compare the two approaches and identify which is best for different network types and confidence levels. Anastasios Dimitriou, Lei Xun, Jonathon S. Hare, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Enhancing State Retention With Energy-Efficient Memory Tracing in Intermittent SystemsabstractIntermittent systems powered by harvested energy frequently encounter power outages, requiring efficient mechanisms to save and restore their internal state, to ensure computational progress. In these systems, minimising the overhead of state retention, comprising CPU core registers and main volatile memory contents (a snapshot), is essential to maximise computational progress within the constraints of limited energy availability. This paper introduces a hardware module,MeTra(MEmory TRAcing), designed to enhance energy-efficient state retention in intermittent systems. This is achieved by tracing and selectively saving the modified parts of the main volatile memory (RAM) to non-volatile memory (NVM), i.e. FRAM. Additionally,MeTradynamically adjusts the voltage threshold that initiates state saving, optimizing energy usage for each snapshot and enabling the system to dedicate more harvested energy to useful computations.MeTrawas integrated with an Arm Cortex-M1 processor on an FPGA and evaluated using benchmarks including matrix multiplication, array sorting, and advanced encryption standard (AES), demonstrating its effectiveness in reducing state-saving time and improving energy efficiency by selectively saving modified RAM regions instead of the entire memory. Experimental results demonstrate that snapshot time can be reduced by 48.34% to 57.56% in FRAM-based systems, leading to an improvement in energy efficiency of 65.24% to 77.76%. These gains are achieved by selectively saving 5.66% to 19.82% of RAM, usingMeTra, compared to saving entire RAM. Osama Bin Tariq, Theodoros D. Verykios, Geoff V. Merrett, Domenico Balsamo |
IEEE Trans. Sustain. Comput. | 3 |
| 2024 | Fluid Dynamic DNNs for Reliable and Adaptive Distributed Inference on Edge DevicesabstractDistributed inference is a popular approach for efficient DNN inference at the edge. However, traditional Static and Dynamic DNNs are not distribution-friendly, causing system reliability and adaptability issues. In this paper, we introduce Fluid Dynamic DNNs (Fluid DyDNNs), tailored for distributed inference. Distinct from Static and Dynamic DNNs, Fluid DyDNNs utilize a novel nested incremental training algorithm to enable independent and combined operation of its sub-networks, enhancing system reliability and adaptability. Evaluation on embedded Arm CPUs with a DNN model and the MNIST dataset, shows that in scenarios of single device failure, Fluid Dy DNNs ensure continued inference, whereas Static and Dynamic DNNs fail. When devices are fully operational, Fluid DyDNNs can operate in either a High-Accuracy mode and achieve comparable accuracy with Static DNNs, or in a High-Throughput mode and achieve 2.5x and 2x throughput compared with Static and Dynamic DNNs, respectively. Lei Xun, Mingyu Hu, Hengrui Zhao, Amit Kumar Singh 0002, Jonathon S. Hare, Geoff V. Merrett |
DATE | 6 |
| 2024 | Enabling ImageNet-Scale Deep Learning on MCUs for Accurate and Efficient InferenceabstractConventional approaches to Tiny Machine Learning (TinyML) achieve high accuracy by deploying the largest deep learning model with the highest input resolutions that fit within the size constraints imposed by the microcontroller’s (MCUs) fast internal storage and memory. In this article, we perform an in-depth analysis of prior works to show that models derived within these constraints suffer from low accuracy and, surprisingly, high latency. We propose an alternative approach that enables the deployment of efficient models with low inference latency, but free from the constraints of internal memory. We take a holistic view of typical MCU architectures and utilize plentiful but slower external memories to relax internal storage and memory constraints. To avoid the lower speed of external memory impacting inference latency, we build on the TinyOps inference framework, which performs operation partitioning and uses overlays via DMA, to accelerate the latency. Using insights from our study, we deploy efficient models from the TinyOps design space onto a range of embedded MCUs achieving record performance on TinyML ImageNet classification with up to 6.7% higher accuracy and$1.4\times $faster latency compared to state-of-the-art internal memory approaches. Sulaiman Sadiq, Jonathon S. Hare, Simon Craske, Partha Maji, Geoff V. Merrett |
IEEE Internet Things J. | 5 |
| 2023 | Exploration of Decision Sub-Network Architectures for FPGA-based Dynamic DNNsabstractDynamic Deep Neural Networks (DNNs) can achieve faster execution and less computationally intensive inference by spending fewer resources on easy to recognise or less informative parts of an input. They make data-dependent decisions, which strategically deactivate a model's components, e.g. layers, channels or sub-networks. However, dynamic DNNs have only been explored and applied on conventional computing systems ($\text{CPU} +\text{GPU}$)) and programmed with libraries designed for static networks, limiting their effects. In this paper, we propose and explore two approaches for efficiently realising the sub-networks that make these decisions on FPGAs. A pipeline approach targets the use of the existing hardware to execute the sub-network, while a parallel approach uses dedicated circuitry for it. We explore the performance of each using the BranchyNet early exit approach on LeNet-5, and evaluate on a Xilinx ZCU106. The pipeline approach is 36% faster than a desktop CPU. It consumes 0.51 mJ per inference, 16x lower than a non-dynamic network on the same platform and 8x lower than an Nvidia Jetson Xavier NX. The parallel approach executes 17% faster than the pipeline approach when on dynamic inference no early exits are taken, but incurs an increase in energy consumption of 28%. Anastasios Dimitriou, Mingyu Hu, Jonathon S. Hare, Geoff V. Merrett |
DATE | 4 |
| 2023 | Content- and Lighting-Aware Adaptive Brightness Scaling for Improved Mobile User ExperienceabstractFor an improved user experience, the display sub-system is expected to provide superior resolution and optimal brightness despite its impact on battery life. Existing brightness scaling approaches set the display brightness statically or adaptively in response to predefined events such as low-battery or ambient light of the environment, which are independent of the displayed content. Approaches that consider the displayed content are either limited to video content or do not account for the user's expected battery life, thereby failing to maximise the user experience. This paper proposes Content- and ambient Lighting-aware Adaptive Brightness Scaling in mobile devices that maximises user experience while meeting battery life expectations. The approach employs a content- and ambient lighting-aware profiler that learns and classifies each sample into predefined clusters at runtime by leveraging insights on user perceptions of content and ambient luminance variations. We maximise user experience through adaptive scaling of the display's brightness using an energy prediction model that determines appropriate brightness levels while meeting expected battery life. The evaluation of the proposed approach on a commercial smartphone improves Quality of Experience (QoE) by up to 24.5 % compared to state-of-art. Samuel Isuwa, David Amos, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 5 |
| 2023 | Maximising mobile user experience through self-adaptive content- and ambient-aware display brightness scalingabstractDisplay subsystems have become the predominant user interface on mobile devices, serving as both input and output interfaces. For a better quality of user experience (QoE), the display subsystem is expected to provide appropriate resolution and brightness despite its impact on battery life. Existing display brightness approaches either consider content- and ambient-light in isolation or do not account for the user’s expected battery life, thereby failing to maximise the QoE. This paper proposes aCADS, a self-Adaptive Content- and Ambient-aware Display brightness Scaling in mobile devices that maximises QoE while meeting battery life expectations. The approach employs a content- and ambient lighting-aware profiler that learns and classifies each sample into predefined clusters at runtime by leveraging insights on user perceptions of content and ambient luminances variations. We maximise QoE through adaptive scaling of the display’s brightness using an energy model that determines appropriate brightness levels while meeting expected battery life. The evaluation on a commercial smartphone shows that aCADS improves QoE by up to 32.5 % compared to state-of-the-art. Samuel Isuwa, David Amos, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
J. Syst. Archit. | 5 |
| 2023 | Pragmatic Memory-System Support for Intermittent Computing Using Emerging Nonvolatile MemoryabstractIntermittent computing (IC) is a key enabler for the vision of a trillion Internet of Things devices. By harvesting energy from the environment and leveraging nonvolatile memory (NVM) to retain computational progress across power cycles, IC enables untethered and battery-free devices to perform computation whenever ambient energy is available. The backbone of state retention is NVM, and recent advances in energy-efficient NVM have the potential to expand the application domain of IC significantly. Utilizing emerging NVM at the level of bit cells, researchers have proposed nonvolatile processors. However, these do not leverage hardware–software co-design, which can be used to overcome hardware limitations and to provide support for application-level constraints such as atomicity. In this article, we propose MEMIC, a memory architecture tailored for IC devices with byte-addressable NVM. A core focus of MEMIC is to combine volatile and NVM in such a way that the operations of IC are as efficient as possible, while also maximizing computational performance per joule. MEMIC uses volatile memory for energy efficiency and NVM for data retention. To avoid double-buffered checkpoints and costly roll backs when code needs to be reexecuted, MEMIC is designed to track and minimize writes to NVM during failure-atomic sections. Our evaluation shows that MEMIC’s instruction cache reduces workload completion time under intermittent operation by 41%–70% and its data cache provides a further reduction of 13%–39%. Sivert T. Sliper, Nikos Nikoleris, Alex S. Weddell, Anand Savanth, Pranay Prabhat, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | Multihop Networking for Intermittent DevicesabstractEnergy harvesting (EH) devices without batteries can enable the Internet of Things (IoT) to reach new and challenging scenarios. Multihop routing is needed to extend the range but, when low EH causes intermittency, it has been overlooked and is not possible with existing protocols. Also, whilst wake-up receivers (WuRxs) have been used to enable star networks, the cost of another EH node sending wake-ups, required for multihop communication, has not been considered. This paper adapts the opportunistic RPL (ORPL) protocol to make possible multihop routing between intermittently-powered devices. Furthermore, the benefit of using WuRx to enable networks is measured, considering different sensitivity devices and associated range. Comparing ORPL to RPL, we show that opportunistic routing enables multihop communication where RPL cannot. If WuRx are used for routing towards a central hub, the more sensitive WuRx perform better, but routing cross-network benefits from lower sensitivity, lower power WuRx. Edward Longman, Mohammed El-Hajjar, Geoff V. Merrett |
SenSys | 3 |
| 2022 | Similarity-Aware CNN for Efficient Video Recognition at the EdgeabstractConvolutional neural networks (CNNs) often extract similar features from successive video frames due to having identical appearances. In contrast, conventional CNNs for video recognition process individual frames with a fixed computational effort. Each video frame is independently processed, resulting in numerous redundant computations and an inefficient use of limited energy resources, particularly for edge computing applications. To alleviate the high energy requirements associated with video frame processing, this article presented similarity-aware CNNs that recognize similar feature pixels across frames and avoid computations on them. First, with a loss of less than 1% in recognition accuracy, a proposed similarity-aware quantization technique increases the average number of unchanged feature pixels across frame pairs by up to 85%. Then, a proposed similarity-aware dataflow improves energy consumption by minimizing redundant computations and memory accesses across frame pairs. According to simulation experiments, the proposed dataflow decreases the energy consumed by video frame processing by up to 30%. Amin Sabet, Jonathon S. Hare, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | A High-Level Approach for Energy Efficiency Improvement of FPGAs by Voltage TrimmingabstractChip manufacturers define voltage margins on top of the “best-case” operational voltage of their chips to ensure reliable functioning in the worst-case settings. The margins guarantee correctness of operation, but at the cost of performance and power efficiency. Violating the margins is tempting to save energy, but might lead to timing errors. This article proposes an algorithmic solution that enables reliable removal of the margins by detecting errors on the fly. In contrast to previous approaches that require special hardware to detect timing errors, the proposed method is fully implementable using high-level synthesis tools without reliance on additional hardware. The approach is demonstrated using a$32 \times 32$matrix-matrix multiplication and a simple multilayer neural network implemented on two Xilinx ZC702 field-programmable gate array (FPGA) System-on-Chip (SoC) platforms, showcasing its utility in detecting errors that may originate from different sources of logic circuits, clock tree, or memory. Results show that the energy dissipation is halved, while the implementation is clocked at 2.5x faster than specified by the design tool of the vendor. Mehdi Safarpour, Lei Xun, Geoff V. Merrett, Olli Silvén |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Exploring the Effect of Energy Storage Sizing on Intermittent Computing System PerformanceabstractBatteryless energy-harvesting devices promise to deliver a sustainable Internet of Things. Intermittent computing is an emerging area, where the application forward progress, i.e., computation beneficial to the progress of the active application, is maintained by saving the volatile computing state into nonvolatile memory before power interruptions, and restored afterward. Conventional intermittent computing approaches typically minimize energy storage to reduce device dimensions and interruption periods, but this can result in high state-saving and -restoring overheads and impede forward progress. In this article, we argue that adding a small amount of energy storage can significantly improve the forward progress. We develop an intermittent computing model that accurately estimates the forward progress, with an experimentally validated mean error of 0.5%. Using this model, we show that sizing energy storage can improve the forward progress by up to 65% with a constant current supply, and 43% with real-world photovoltaic sources. An extension to this approach, which uses a cost function to trade off the energy storage size against forward progress, can save 83% of capacitor volume and 91% of interruption periods while maintaining 93% of the maximum forward progress. Jie Zhan, Geoff V. Merrett, Alex S. Weddell |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | QUAREM: Maximising QoE Through Adaptive Resource Management in Mobile MPSoC PlatformsabstractHeterogeneous multi-processor system-on-chip (MPSoC) smartphones are required to offer increasing performance and user quality-of-experience (QoE) , despite comparatively slow advances in battery technology. Approaches to balance instantaneous power consumption, performance and QoE have been reported, but little research has considered how to perform longer-term budgeting of resources across a complete battery discharge cycle. Approaches that have considered this are oblivious to the daily variability in the user’s desired charging time-of-day (plug-in time), resulting in a failure to meet the user’s battery life expectations, or else an unnecessarily over-constrained QoE. This paper proposes QUAREM, an adaptive resource management approach in mobile MPSoC platforms that maximises QoE while meeting battery life expectations. The proposed approach utilises a model that learns and then predicts the dynamics of the energy usage pattern and plug-in times. Unlike state-of-the-art approaches, we maximise the QoE through the adaptive balancing of the battery life and the quality of service (QoS) for the duration of the battery discharge. Our model achieves a good degree of accuracy with a mean absolute percentage error of 3.47% and 2.48% for the energy demand and plug-in times, respectively. Experimental evaluation on an off-the-shelf commercial smartphone shows that QUAREM achieves the expected battery life of the user within 20–25% energy demand variation with little or no QoE degradation. Samuel Isuwa, Somdip Dey, Andre P. Ortega, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2021 | GhostShiftAddNet: More Features from Energy-Efficient Operations
Jia Bi, Jonathon S. Hare, Geoff V. Merrett |
BMVC | 3 |
| 2021 | Wake-up Radio-enabled Intermittently-powered Devices for Mesh Networking: A Power AnalysisabstractThis paper analyzes the successful communication probability between two intermittently-powered nodes in a homogeneous energy harvesting (EH) mesh network. Powering devices using EH can enable networks to operate indefinitely; however, with limited energy storage and in scarce EH conditions, nodes may only be intermittently-powered. This reduces the effectiveness of conventional networking techniques, where listening modes of radios deplete the storage too quickly, rendering nodes useless. This is particularly problematic in the deployment of mesh networks, where there is no provision of a high-power coordinator. To counter this, wake-up receivers (WuRxs) provide extended listening time without the need for a high power conventional radio, but with a cost to sensitivity. Therefore, listening time must be balanced with transmission and wake-to-receive cost, where if all the harvested energy is spent listening none remains to transmit, and vice versa. From stochastic analysis and simulation of the energy usage in mesh nodes, we obtain the optimum transmission load to maximize goodput, which is the rate of successful communications. We include the cost of each wake-up based on the network size in our goodput analysis. Simulations for a fixed number of homogeneous nodes verify this. Furthermore, we model and evaluate the energy consumption trade-off between transmit power and WuRx sensitivity to enable the maximum goodput. Edward Longman, Oktay Cetinkaya, Mohammed El-Hajjar, Geoff V. Merrett |
CCNC | 4 |
| 2021 | Partner selection in self-organised wireless sensor networks for opportunistic energy negotiation: A multi-armed bandit based approach
Andre P. Ortega, Sarvapali D. Ramchurn, Long Tran-Thanh, Geoff V. Merrett |
Ad Hoc Networks | 4 |
| 2021 | Improving the Forward Progress of Transient SystemsabstractEmerging applications for Internet of Things (IoT) devices demand smaller mass, size, and cost whilst increasing capability and reliability. Energy harvesting can provide power to these ultra-constrained devices, but introduces unreliability, unpredictability, and intermittency. Schemes for wireless sensors without batteries or supercapacitors overcome intermittency through saving system state into nonvolatile memory before the supply drops below the minimum operating voltage, termed transient, or intermittent computing. However, this introduces significant time and energy overheads. This article presents two schemes that significantly reduce these overheads: entering a sleep mode to avoid saving state and utilizing direct memory access (DMA) when state saves are required. Time and energy previously wasted on state saves can instead be used to perform useful computation, termed “forward progress.” We practically validate the proposed approaches across a range of energy sources and IoT benchmarks and demonstrate up to 46.8% and 40.3% increase in forward progress and up to 91.1% and 85.6% reduction in overheads for each scheme, respectively. Tim Daulby, Anand Savanth, Geoff V. Merrett, Alex S. Weddell |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Optimising Resource Management for Embedded Machine LearningabstractMachine learning inference is increasingly being executed locally on mobile and embedded platforms, due to the clear advantages in latency, privacy and connectivity. In this paper, we present approaches for online resource management in heterogeneous multi-core systems and show how they can be applied to optimise the performance of machine learning work-loads. Performance can be defined using platform-dependent (e.g. speed, energy) and platform-independent (accuracy, confidence) metrics. In particular, we show how a Deep Neural Network (DNN) can be dynamically scalable to trade-off these various performance metrics. Achieving consistent performance when executing on different platforms is necessary yet challenging, due to the different resources provided and their capability, and their time-varying availability when executing alongside other workloads. Managing the interface between available hardware resources (often numerous and heterogeneous in nature), software requirements, and user experience is increasingly complex. Lei Xun, Long Tran-Thanh, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 4 |
| 2020 | Fused: Closed-Loop Performance and Energy Simulation of Embedded SystemsabstractEnergy-driven computing is an emerging paradigm that aims to fuel the proliferation of tiny and low-cost IoT sensing and monitoring devices. Energy-driven computers are generally powered by energy harvesting sources, and adapt their operation at runtime according to energy availability; thus, they must be designed and tested according to the expected dynamics of their power source. However, today's processor simulators and debuggers typically assume that power is always available, so they are unable to correctly model the interactions between power supply, power consumption and energy-driven execution. To address this shortcoming, we propose Fused, an open source full-system simulator for energy-driven computers. Fused models execution, power consumption, and power supply in a closed loop, thus correctly models the interaction between them. It targets energy-driven embedded systems, and employs SystemC for digital and mixed-signal simulation to model a microcontroller and mixed-signal circuitry, enabling hardware-software codesign and design space exploration. Fused includes a high-level power modelling methodology, whereby events recorded during simulation are correlated to power measurements of real hardware to extract features for power modelling. Results show that Fused can model the execution time and power consumption of a commercially available microcontroller with a geometric mean error of 0.2% and 3.4% respectively, across a wide range of workloads. Through a case-study, we demonstrate that Fused can accurately model a state-of-the art intermittent computing system, where execution is heavily dependent on energy availability: although up to 70 power cycles were needed to complete the tested workload on the constrained energy supply, Fused modelled the completion time with less than 7% error. Sivert T. Sliper, Nikos Nikoleris, Alex S. Weddell, Geoff V. Merrett |
ISPASS | 5 |
| 2020 | Efficient Deployment of UAV-powered Sensors for Optimal Coverage and ConnectivityabstractThe Internet of Things (IoT) digitizes the physical world with wireless devices sensing their surroundings and delivering periodic notifications of parameters they are monitoring. However, this operation is bound by finite-capacity batteries, in which replenishment is practically infeasible due to the envisioned size of the IoT networks. By also considering the autonomous and self-sufficient service vision of the IoT paradigm, the need for novel approaches overcoming the energy constraints is evident. Here, unmanned aerial vehicles (UAVs) come into prominence. The UAVs can remotely energize wireless devices, via wireless power transfer (WPT), and thus guarantee reliable sensing coverage as well as longevity in the IoT domain. However, this can be only achieved by the precise alignment of both UAVs and wireless devices. Thus, this paper presents an efficient deployment strategy based on the circle packing problem, in which a lower bound for the required number of wireless devices achieving optimal coverage is derived. The analysis, based on empirical measurements, reveals the design considerations for an energy harvesting (EH)-aided UAV scenario with regard to Federal Communications Commission (FCC) regulations, power consumption of wireless devices, and reporting frequency requirements of the IoT applications. Our results elaborate on a number of trade-offs, based on UAV, device, and medium characteristics, and provide realistic guidelines, achieving optimal coverage while meeting application requirements. Oktay Cetinkaya, Geoff V. Merrett |
WCNC | 2 |
| 2020 | Collaborative Adaptation for Energy-Efficient Heterogeneous Mobile SoCsabstractHeterogeneous Mobile System-on-Chips (SoCs) containing CPU and GPU cores are becoming prevalent in embedded computing, and they need to execute applications concurrently. However, existing run-time management approaches do not perform adaptive mapping and thread-partitioning of applications while exploiting both CPU and GPU cores at the same time. In this paper, we propose an adaptive mapping and thread-partitioning approach for energy-efficient execution of concurrent OpenCL applications on both CPU and GPU cores while satisfying performance requirements. To start execution of concurrent applications, the approach makes mapping (number of cores and operating frequencies) and partitioning (distribution of threads between CPU and GPU) decisions to satisfy performance requirements for each application. The mapping and partitioning decisions are made by having a collaboration between the CPU and GPU cores' processing capabilities such that balanced execution can be performed. During execution, adaptation is triggered when new application(s) arrive, or an executing one finishes, that frees cores. The adaptation process identifies a new mapping and thread-partitioning in a similar collaborative manner for remaining applications provided it leads to an improvement in energy efficiency. The proposed approach is experimentally validated on the Odroid-XU3 hardware platform with varying set of applications. Results show an average energy saving of 37%, compared to existing approaches while satisfying the performance requirements. Amit Kumar Singh 0002, Basireddy Karunakar Reddy, Alok Prakash, Geoff V. Merrett, Bashir M. Al-Hashimi |
IEEE Trans. Computers | 4 |
| 2020 | AdaMD: Adaptive Mapping and DVFS for Energy-Efficient Heterogeneous MulticoresabstractModern heterogeneous multicore systems, containing various types of cores, are increasingly dealing with concurrent execution of dynamic application workloads. Moreover, the performance constraints of each application vary, and applications enter/exit the system at any time. Existing approaches are not efficient in such dynamic scenarios, especially if applications are unknown, as they require extensive offline application analysis and do not consider the runtime execution scenarios (application arrival/completion, and workload and performance variations) for runtime management. To address this, we present AdaMD, an adaptive mapping and dynamic voltage and frequency scaling (DVFS) approach for improving energy consumption and performance. The key feature of the proposed approach is the elimination of dependency on offline profiled results while making runtime decisions. This is achieved through a performance prediction model having a maximum error of 7.9% lower than the previously reported model and a mapping approach that allocates processing cores to applications while respecting performance constraints. Furthermore, AdaMD adapts to runtime execution scenarios efficiently by monitoring the application status, and performance/workload variations to adjust the previous DVFS settings and thread-to-core mappings. The proposed approach is experimentally validated on the Odroid-XU3, with various combinations of diverse multithreaded applications from PARSEC and SPLASH benchmarks. Results show energy savings of up to 28% compared to the recently proposed approach while meeting performance constraints. Basireddy Karunakar Reddy, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Efficient State Retention through Paged Memory Management for Reactive Transient ComputingabstractReactive transient computing systems preserve computational progress despite frequent power failures by suspending (saving state to nonvolatile memory) when detecting a power failure, and restoring once power returns. Existing methods inefficiently save and restore all allocated memory. We propose lightweight memory management that applies the concept of paging to load pages only when needed, and save only modified pages. We then develop a model that maximises available execution time by dynamically adjusting the suspend and restore voltage thresholds. Experiments on an MSP430FR5994 microcontroller show that our method reduces state retention overheads by up to 86.9% and executes algorithms up to 5.3× faster than the state-of-the-art. Sivert T. Sliper, Domenico Balsamo, Nikos Nikoleris, Alex S. Weddell, Geoff V. Merrett |
DAC | 6 |
| 2019 | BRB: Mitigating Branch Predictor Side-ChannelsabstractModern processors use branch prediction as an optimization to improve processor performance. Predictors have become larger and increasingly more sophisticated in order to achieve higher accuracies which are needed in high performance cores. However, branch prediction can also be a source of side channel exploits, as one context can deliberately change the branch predictor state and alter the instruction flow of another context. Current mitigation techniques either sacrifice performance for security, or fail to guarantee isolation when retaining the accuracy. Achieving both has proven to be challenging. In this work we address this by, (1) introducing the notions of steady-state and transient branch predictor accuracy, and (2) showing that current predictors increase their misprediction rate by as much as 90% on average when forced to flush branch prediction state to remain secure. To solve this, (3) we introduce the branch retention buffer, a novel mechanism that partitions only the most useful branch predictor components to isolate separate contexts. Our mechanism makes thread isolation practical, as it stops the predictor from executing cold with little if any added area and no warm-up overheads. At the same time our results show that, compared to the state-of-the-art, average misprediction rates are reduced by 15-20% without increasing area, leading to a 2% performance increase. Ilias Vougioukas, Nikos Nikoleris, Andreas Sandberg, Stephan Diestelhorst, Bashir M. Al-Hashimi, Geoff V. Merrett |
HPCA | 6 |
| 2019 | The Circuit Breaker Pattern Targeted to Future IoT Applications
Gibeon Aquino, Rafael Fernandes de Queiroz, Geoff V. Merrett, Bashir M. Al-Hashimi |
ICSOC | 3 |
| 2019 | A Digital In-Analogue Out Logic Gate Based on Metal-Oxide Memristor DevicesabstractAn important cornerstone of data processing is the ability to efficiently capture structure in data and perform data classification. More recently, memristive technologies enabled the incorporation of continuous tuneable resistive elements directly in hardware, thus increasing the efficiency of reconfigurable systems power and area-wise. Memristors are a promising candidate for reconfigurable circuits capable of carrying out classification with physical computing, such as dot-product vector multiplication and accumulation technique. In this work, we demonstrate a novel proof-of-concept memristor-based Digital-In-Analogue-Out logic circuit and present preliminary results highlighting the effect of non-uniform non-linear memristor IV characteristics that result in device-to-device behavioural variation. Georgios Papandroulidakis, Loukas Michalas, Alexander Serb, Ali Khiat, Geoff V. Merrett, Themistoklis Prodromakis |
ISCAS | 5 |
| 2019 | Momentum: Power-neutral Performance Scaling with Intrinsic MPPT for Energy Harvesting Computing SystemsabstractRecent research has looked to supplement or even replace the batteries in embedded computing systems with energy harvesting, where energy is derived from the device’s environment. However, such supplies are generally unpredictable and highly variable, and hence systems typically incorporate large external energy buffers (e.g., supercapacitors) to sustain computation; however, these pose environmental issues and increase system size and cost. This article proposes Momentum , a general power-neutral methodology, with intrinsic system-wide maximum power point tracking, that can be applied to a wide range of different computing systems, where the system dynamically scales its performance (and hence power consumption) to optimize computational progress depending on the power availability. Momentum enables the system to operate around an efficient operating voltage, maximizing forward application execution, without adding any external tracking or control units. This methodology combines at runtime (1) a hierarchical control strategy that utilizes available power management controls (such as dynamic voltage and frequency scaling, and core hot-plugging) to achieve efficient power-neutral operation; (2) a software-based maximum power point tracking scheme (unlike existing approaches, this does not require any additional hardware), which adapts the system power consumption so that it can work at the optimal operating voltage, considering the efficiency of the entire system rather than just the energy harvester; and (3) experimental validation on two different scales of computing system: a low power microcontroller (operating from the already-present 4.7μF decoupling capacitance) and a multi-processor system-on-chip (operating from 15.4mF added capacitance). Experimental results from both a controlled supply and energy harvesting source show that Momentum operates correctly on both platforms and exhibits improvements in forward application execution of up to 11% when compared to existing power-neutral approaches and 46% compared to existing static approaches. Domenico Balsamo, Benjamin J. Fletcher, Alex S. Weddell, Giorgos Karatziolas, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2019 | Predictive Thermal Management for Energy-Efficient Execution of Concurrent Applications on Heterogeneous MulticoresabstractCurrent multicore platforms contain different types of cores, organized in clusters (e.g., ARM's big.LITTLE). These platforms deal with concurrently executing applications, having varying workload profiles and performance requirements. Runtime management is imperative for adapting to such performance requirements and workload variabilities and to increase energy and temperature efficiency. Temperature has also become a critical parameter since it affects reliability, power consumption, and performance and, hence, must be managed. This paper proposes an accurate temperature prediction scheme coupled with a runtime energy management approach to proactively avoid exceeding temperature thresholds while maintaining performance targets. Experiments show up to 20% energy savings while maintaining high-temperature averages and peaks below the threshold. Compared with state-of-the-art temperature predictors, this paper predicts 35% faster and reduces the mean absolute error from 3.25 to 1.15 °C for the evaluated applications' scenarios. Eduardo Wächter, Cedric de Bellefroid, Basireddy Karunakar Reddy, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | Online concurrent workload classification for multi-core energy managementabstractModern embedded multi-core processors are organized as clusters of cores, where all cores in each cluster operate at a common Voltage-frequency (V-f). Such processors often need to execute applications concurrently, exhibiting varying and mixed workloads (e.g. compute- and memory-intensive) depending on the instruction mix and resource sharing. Runtime adaptation is key to achieving energy savings without trading-off application performance with such workload variabilities. In this paper, we propose an online energy management technique that performs concurrent workload classification using the metric Memory Reads Per Instruction (MRPI) and pro-actively selects an appropriate V-fsetting through workload prediction. Subsequently, it monitors the workload prediction error and performance loss, quantified by Instructions Per Second (IPS) at runtime and adjusts the chosen V-fto compensate. We validate the proposed technique on an Odroid-XU3 with various combinations of benchmark applications. Results show an improvement in energy efficiency of up to 69% compared to existing approaches. Basireddy Karunakar Reddy, Geoff V. Merrett, Bashir M. Al-Hashimi, Amit Kumar Singh 0002 |
DATE | 2 |
| 2018 | Hardware-Validated CPU Performance and Energy ModellingabstractFull-system simulation frameworks such as gem5 are used extensively to evaluate research ideas and for design-space exploration. Moreover, energy-efficiency has become the key design constraint in recent years and many works use a separate power modelling framework to evaluate energy consumption. While such tools are convenient and flexible, they are known to contain sources of error which are often not fully understood and potentially impact the conclusions drawn from investigations. This work enables accurate, hardware-validated performance, power, and energy modelling of CPUs by first presenting a methodology to evaluate and identify sources of error in CPU performance models, and secondly developing empirical power models optimised for use with such performance models. Hierarchical clustering, correlation analysis, and regression techniques are used to identify sources of error without requiring detailed CPU specifications and enable existing models to be improved, new models to be developed, validation of simulator changes, and testing of model suitability for specific use-cases. Furthermore, the GemStone open-source software tool is presented, which automates the process of characterising hardware platforms, identifying sources of error in gem5 models, applying power analysis, and quantifying the effect of errors on the performance, power, and energy estimations. In addition, the mean percentage error in execution time was found to swing from -51% to +10% between two versions of the same gem5 model, underlining the need for an automated tool to validate models against reference hardware, ensuring accuracy and consistency. Matthew J. Walker, Sascha Bischoff, Stephan Diestelhorst, Geoff V. Merrett, Bashir M. Al-Hashimi |
ISPASS | 4 |
| 2018 | An Energy-driven Wireless Bicycle Trip Counter with Zero Energy StorageabstractThis paper presents the implementation of a bicycle trip counter, which measures cycling speed, traveled distance, and cycling time, that is directly powered from tiny periodic pulses of energy with only the intrinsically present decoupling capacitance as an energy buffer. To cope with the highly variable amount of energy generated during each pulse, an energy-driven approach is used. The core principles in this approach are to dynamically adjust operational mode according to energy availability, to scale performance, for example sensing accuracy, proportional to energy harvested, and to perform intermittent or transient computing to enable computation across multiple power cycles. The device presented is able to start operation from energy supply pulses as low as 4 uJ, where a rough estimate of the sensing parameters is done, and perform increasingly complex and time-consuming tasks such as additional more accurate measurements, sensor fusion, and filtering computations as more energy becomes available. Samuel Chang Bing Wong, Domenico Balsamo, Geoff V. Merrett |
SenSys | 3 |
| 2018 | Internet-of-Things and big data for smarter healthcare: From device to architecture, applications and analytics
Farshad Firouzi, Amir-Mohammad Rahmani, Kunal Mankodiya, Mustafa Badaroglu, Geoff V. Merrett, Bahareh J. Farahani |
Future Gener. Comput. Syst. | 5 |
| 2018 | A model-based framework for software portability and verification in embedded power management systems
Asieh Salehi Fathabadi, Michael J. Butler, Sheng Yang 0003, Luis Alfonso Maeda-Nunez, James R. B. Bantock, Bashir M. Al-Hashimi, Geoff V. Merrett |
J. Syst. Archit. | 7 |
| 2018 | Runtime Performance and Power Optimization of Parallel Disparity Estimation on Many-Core PlatformsabstractThis article investigates the use of many-core systems to execute the disparity estimation algorithm, used in stereo vision applications, as these systems can provide flexibility between performance scaling and power consumption. We present a learning-based runtime management approach that achieves a required performance threshold while minimizing power consumption through dynamic control of frequency and core allocation. Experimental results are obtained from a 61-core Intel Xeon Phi platform for the aforementioned investigation. The same performance can be achieved with an average reduction in power consumption of 27.8% and increased energy efficiency by 30.04% when compared to Dynamic Voltage and Frequency Scaling control alone without runtime management. Charles Leech, Charan Kumar Vala, Amit Acharyya, Sheng Yang 0003, Geoff V. Merrett, Bashir M. Al-Hashimi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | Machine learning for run-time energy optimisation in many-core systemsabstractIn recent years, the focus of computing has moved away from performance-centric serial computation to energy-efficient parallel computation. This necessitates run-time optimisation techniques to address the dynamic resource requirements of different applications on many-core architectures. In this paper, we report on intelligent run-time algorithms which have been experimentally validated for managing energy and application performance in many-core embedded system. The algorithms are underpinned by a cross-layer system approach where the hardware, system software and application layers work together to optimise the energy-performance trade-off. Algorithm development is motivated by the biological process of how a human brain (acting as an agent) interacts with the external environment (system) changing their respective states over time. This leads to a pay-off for the action taken, and the agent eventually learns to take the optimal/best decisions in future. In particular, our online approach uses a model-free reinforcement learning algorithm that suitably selects the appropriate voltage-frequency scaling based on workload prediction to meet the applications' performance requirements and achieve energy savings of up to 16% in comparison to state-of-the-art-techniques, when tested on four ARM A15 cores of an ODROID-XU3 platform. Dwaipayan Biswas, Vibishna Balagopal, Rishad A. Shafik, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 5 |
| 2017 | Power neutral performance scaling for energy harvesting MP-SoCsabstractUsing energy `harvested' from the environment to power autonomous embedded systems is an attractive ideal, alleviating the burden of periodic battery replacement. However, such energy sources are typically low-current and transient, with high temporal and spatial variability. To overcome this, large energy buffers such as supercapacitors or batteries are typically incorporated to achieve energy neutral operation, where the energy consumed over a certain period of time is equal to the energy harvested. Large energy buffers, however, pose environmental issues in addition to increasing the size and cost of systems. In this paper we propose a novel power neutral performance scaling approach for multiprocessor system-on-chips (MP-SoCs) powered by energy harvesting. Under power neutral operation, the system's performance is dynamically scaled through DVFS and DPM such that the instantaneous power consumption is approximately equal to the instantaneous harvested power. Power neutrality means that large energy buffers are no longer required, while performance scaling ensures that available power is effectively utilised. The approach is experimentally validated using the Samsung Exynos5422 big.LITTLE SoC directly coupled to a monocrystalline photovoltaic array, with only 47mF of intermediate energy storage. Results show that the proposed approach is successful in tracking harvested power, stabilising the supply voltage to within 5% of the target value for over 93% of the test duration, resulting in the execution of 69% more instructions compared to existing static approaches. Benjamin J. Fletcher, Domenico Balsamo, Geoff V. Merrett |
DATE | 3 |
| 2017 | Energy-driven computing: Rethinking the design of energy harvesting systemsabstractEnergy harvesting computing has been gaining increasing traction over the past decade, fueled by technological developments and rising demand for autonomous and battery-free systems. Energy harvesting introduces numerous challenges to embedded systems but, arguably the greatest, is the required transition from an energy source that typically provides virtually unlimited power for a reasonable period of time until it becomes exhausted, to a power source that is highly unpredictable and dynamic (both spatially and temporally, and with a range spanning many orders of magnitude). The typical approach to overcome this is the addition of intermediate energy storage/buffering to smooth out the temporal dynamics of both power supply and consumption. This has the advantage that, if correctly sized, the system `looks like' a battery-powered system; however, it also adds volume, mass, cost and complexity and, if not sized correctly, unreliability. In this paper, we consider energy-driven computing, where systems are designed from the outset to operate from an energy harvesting source. Such systems typically contain little or no additional energy storage (instead relying on tiny parasitic and decoupling capacitance), alleviating the aforementioned issues. Examples of energy-driven computing include transient systems (which power down when the supply disappears and efficiently continue execution when it returns) and power-neutral systems (which operate directly from the instantaneous power harvested, gracefully modulating their consumption and performance to match the supply). In this paper, we introduce a taxonomy of energy-driven computing, articulating how power-neutral, transient, and energy-driven systems present a different class of computing to conventional approaches. Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 1 |
| 2017 | Online tuning of Dynamic Power Management for efficient execution of interactive workloadsabstractModern mobile devices contain powerful Multi-Processor System-on-Chips (MPSoCs) that are performance throttled by Dynamic Power Management (DPM) runtime systems to extend battery lifetime. Applications on mobile devices commonly generate highly interactive workloads, dependent on interaction between the processor cores, peripherals, external resources and the user, such as touch input during web-browsing. Inevitably, a subset of interactive workloads are affected by delays caused by data unavailability, e.g. loss or delay of data packets during voice-over-IP. At the same time, the system is required to respond quickly upon data retrieval to ensure that the user Quality of Experience (QoE) metrics (frame-rate, latency, etc.) are not degraded. Traditionally, operating systems have mitigated this problem with periodic sampling or event-driven approaches. Through experimentation using a mobile MPSoC platform, however, we demonstrate that improving the tuning of DPM parameters for certain interactive user inputs can provide energy savings of up to 21% or QoE improvements of up to 36%, when compared with the traditional approach. To capture these improvements, we propose a dynamic modeling of user input and data resource access times (e.g. mobile network bandwidth and latency) for interactive workloads, which is based on workload profiling and which we refer to herein as inelasticity analysis. The proposed approach is implemented through online tuning of a DPM runtime in the Android operating system and is validated through a Monte Carlo simulation of interactive workloads. In comparison to the default DPM tuning, the proposed approach achieves energy savings of 13% or QoE improvement of 27% or a selectable trade-off, e.g. 9% energy savings and 15% QoE improvement. James R. B. Bantock, Vasileios Tenentes, Bashir M. Al-Hashimi, Geoff V. Merrett |
ISLPED | 4 |
| 2017 | Applications of Energy-Driven and Transient Computing: A Wireless Bicycle Trip CounterabstractEnergy harvesting is an efficient solution to power embedded systems instead of using batteries. However, it has been traditionally coupled with large energy buffers to tackle the temporal variation of the source. These buffers require time to charge and introduce a cost, size and weight overhead. Energy-driven and transiently-powered systems can operate from an energy harvesting source, while containing little or no additional energy storage. However, few real-life applications have been considered for such systems to demonstrate that they can actually be realised. This poster presents a transiently-powered wireless bicycle trip counter which measures distance, speed and active cycling time, and transmits data wirelessly. The system sustains operation by harvesting energy from the rotation of the wheel, operating from 6kph. Uvis Senkans, Domenico Balsamo, Theodoros D. Verykios, Geoff V. Merrett |
SenSys | 4 |
| 2017 | Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUsabstractModern mobile and embedded devices are required to be increasingly energy-efficient while running more sophisticated tasks, causing the CPU design to become more complex and employ more energy-saving techniques. This has created a greater need for fast and accurate power estimation frameworks for both run-time CPU energy management and design-space exploration. We present a statistically rigorous and novel methodology for building accurate run-time power models using performance monitoring counters (PMCs) for mobile and embedded devices, and demonstrate how our models make more efficient use of limited training data and better adapt to unseen scenarios by uniquely considering stability. Our robust model formulation reduces multicollinearity, allows separation of static and dynamic power, and allows a 100× reduction in experiment time while sacrificing only 0.6% accuracy. We present a statistically detailed evaluation of our model, highlighting and addressing the problem of heteroscedasticity in power modeling. We present software implementing our methodology and build power models for ARM Cortex-A7 and Cortex-A15 CPUs, with 3.8% and 2.8% average error, respectively. We model the behavior of the nonideal CPU voltage regulator under dynamic CPU activity to improve modeling accuracy by up to 5.5% in situations where the voltage cannot be measured. To address the lack of research utilizing PMC data from real mobile devices, we also present our data acquisition method and experimental platform software. We support this paper with online resources including software tools, documentation, raw data and further results. Matthew J. Walker, Stephan Diestelhorst, Andreas Hansson 0001, Anup Das 0001, Sheng Yang 0003, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2017 | Energy-Efficient Run-Time Mapping and Thread Partitioning of Concurrent OpenCL Applications on CPU-GPU MPSoCsabstractHeterogeneous Multi-Processor Systems-on-Chips (MPSoCs) containing CPU and GPU cores are typically required to execute applications concurrently. However, as will be shown in this paper, existing approaches are not well suited for concurrent applications as they are developed either by considering only a single application or they do not exploit both CPU and GPU cores at the same time. In this paper, we propose an energy-efficient run-time mapping and thread partitioning approach for executing concurrent OpenCL applications on both GPU and GPU cores while satisfying performance requirements. Depending upon the performance requirements, for each concurrently executing application, the mapping process finds the appropriate number of CPU cores and operating frequencies of CPU and GPU cores, and the partitioning process identifies an efficient partitioning of the applications’ threads between CPU and GPU cores. We validate the proposed approach experimentally on the Odroid-XU3 hardware platform with various mixes of applications from the Polybench benchmark suite. Additionally, a case-study is performed with a real-world application SLAMBench. Results show an average energy saving of 32% compared to existing approaches while still satisfying the performance requirements. Amit Kumar Singh 0002, Alok Prakash, Basireddy Karunakar Reddy, Geoff V. Merrett, Bashir M. Al-Hashimi |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | Nucleus: Finding the Sharing Limit of Heterogeneous CoresabstractHeterogeneous multi-processors are designed to bridge the gap between performance and energy efficiency in modern embedded systems. This is achieved by pairing Out-of-Order (OoO) cores, yielding performance through aggressive speculation and latency masking, with In-Order (InO) cores, that preserve energy through simpler design. By leveraging migrations between them, workloads can therefore select the best setting for any given energy/delay envelope. However, migrations introduce execution overheads that can hurt performance if they happen too frequently. Finding the optimal migration frequency is critical to maximize energy savings while maintaining acceptable performance. We develop a simulation methodology that can 1) isolate the hardware effects of migrations from the software, 2) directly compare the performance of different core types, 3) quantify the performance degradation and 4) calculate the cost of migrations for each case. To showcase our methodology we run mibench, a microbenchmark suite, and show that migrations can happen as fast as every 100k instructions with little performance loss. We also show that, contrary to numerous recent studies, hypothetical designs do not need to share all of their internal components to be able to migrate at that frequency. Instead, we propose a feasible system that shares level 2 caches and a translation lookaside buffer that matches performance and efficiency. Our results show that there are phases comprising up to 10% that a migration to the OoO core leads to performance benefits without any additional energy cost when running on the InO core, and up to 6% of phases where a migration to the InO core can save energy without affecting performance. When considering a policy that focuses on improving the energy-delay product, results show that on average 66% of the phases can be migrated to deliver equal or better system operation without having to aggressively share the entire memory system or to revert to migration periods finer than 100k instructions. Ilias Vougioukas, Andreas Sandberg, Stephan Diestelhorst, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | Invited - Energy harvesting and transient computing: a paradigm shift for embedded systems?abstractEmbedded systems powered from time-varying energy harvesting sources traditionally operate using the principles of energy-neutral computing: over a certain period of time, the energy that they consume equals the energy that they harvest. This has the significant advantage of making the system 'look like' a battery-powered system, yet typically results in large, complex and expensive power conversion circuitry and introduces numerous challenges including fast and reliable cold-start. In recent years, the concept of transient computing has emerged to challenge this traditional approach, whereby low-power embedded systems are enabled to operate as usual while energy is available but, after loss of supply, can quickly regain state and continue where they left off. This paper provides a summary of these different approaches. Geoff V. Merrett |
DAC | 1 |
| 2016 | The slowdown or race-to-idle question: Workload-aware energy optimization of SMT multicore platforms under process variation
Anup Das 0001, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 2 |
| 2016 | Workload Change Point Detection for Runtime Thermal Management of Embedded SystemsabstractApplications executed on multicore embedded systems interact with system software [such as the operating system (OS)] and hardware, leading to widely varying thermal profiles which accelerate some aging mechanisms, reducing the lifetime reliability. Effectively managing the temperature therefore requires: 1) autonomous detection of changes in application workload and 2) appropriate selection of control levers to manage thermal profiles of these workloads. In this paper, we propose a technique for workload change detection using density ratio-based statistical divergence between overlapping sliding windows of CPU performance statistics. This is integrated in a runtime approach for thermal management, which uses reinforcement learning to select workload-specific thermal control levers by sampling on-board thermal sensors. Identified control levers override the OSs native thread allocation decision and scale hardware voltage-frequency to improve average temperature, peak temperature, and thermal cycling. The proposed approach is validated through its implementation as a hierarchical runtime manager for Linux, with heuristic-based thread affinity selected from the upper hierarchy to reduce thermal cycling and learningbased voltage-frequency selected from the lower hierarchy to reduce average and peak temperatures. Experiments conducted with mobile, embedded, and high performance applications on ARM-based embedded systems demonstrate that the proposed approach increases workload change detection accuracy by an average 3.4×, reducing the average temperature by 4 °C-25 °C, peak temperature by 6 °C-24 °C, and thermal cycling by 7%-35% over state-of-the-art approaches. Anup Das 0001, Geoff V. Merrett, Mirco Tribastone, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Graceful Performance Modulation for Power-Neutral Transient Computing SystemsabstractTransient computing systems do not have energy storage, and operate directly from energy harvesting. These systems are often faced with the inherent challenge of low-current or transient power supply. In this paper, we propose “power-neutral” operation, a new paradigm for such systems, whereby the instantaneous power consumption of the system must match the instantaneous harvested power. Power neutrality is achieved using a control algorithm for dynamic frequency scaling, modulating system performance gracefully in response to the incoming power. Detailed system model is used to determine design parameters for selecting the system voltage thresholds where the operating frequency will be raised or lowered, or the system will be hibernated. The proposed control algorithm for power-neutral operation is experimentally validated using a microcontroller incorporating voltage threshold-based interrupts for frequency scaling. The microcontroller is powered directly from real energy harvesters; results demonstrate that a power-neutral system sustains operation for 4%-88% longer with up to 21% speedup in application execution. Domenico Balsamo, Anup Das 0001, Alex S. Weddell, Davide Brunelli, Bashir M. Al-Hashimi, Geoff V. Merrett, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | Hibernus++: A Self-Calibrating and Adaptive System for Transiently-Powered Embedded DevicesabstractEnergy harvesters are being used to power autonomous systems, but their output power is variable and intermittent. To sustain computation, these systems integrate batteries or supercapacitors to smooth out rapid changes in harvester output. Energy storage devices require time for charging and increase the size, mass, and cost of systems. The field of transient computing moves away from this approach, by powering the system directly from the harvester output. To prevent an application from having to restart computation after a power outage, approaches such as Hibernus allow these systems to hibernate when supply failure is imminent. When the supply reaches the operating threshold, the last saved state is restored and the operation is continued from the point it was interrupted. This paper proposes Hibernus++ to intelligently adapt the hibernate and restore thresholds in response to source dynamics and system load properties. Specifically, capabilities are built into the system to autonomously characterize the hardware platform and its performance during hibernation in order to set the hibernation threshold at a point which minimizes wasted energy and maximizes computation time. Similarly, the system auto-calibrates the restore threshold depending on the balance of energy supply and consumption in order to maximize computation time. Hibernus++ is validated both theoretically and experimentally on microcontroller hardware using both synthesized and real energy harvesters. Results show that Hibernus++ provides an average 16% reduction in energy consumption and an improvement of 17% in application execution time over state-of-the-art approaches. Domenico Balsamo, Alex S. Weddell, Anup Das 0001, Alberto Rodriguez Arreola, Davide Brunelli, Bashir M. Al-Hashimi, Geoff V. Merrett, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2016 | Learning Transfer-Based Adaptive Energy Minimization in Embedded SystemsabstractEmbedded systems execute applications with varying performance requirements. These applications exercise the hardware differently depending on the computation task, generating varying workloads with time. Energy minimization with such workload and performance variations within (intra) and across (inter) applications is particularly challenging. To address this challenge, we propose an online approach, capable of minimizing energy through adaptation to these variations. At the core of this approach is a reinforcement learning algorithm that suitably selects the appropriate voltage/frequency scaling (VFS) based on workload predictions to meet the applications' performance requirements. The adaptation is then facilitated and expedited through learning transfer, which uses the interaction between the application, runtime, and hardware layers to adjust the VFS. The proposed approach is implemented as a power governor in Linux and extensively validated on an ARM Cortex-A8 running different benchmark applications. We show that with intra- and inter-application variations, our proposed approach can effectively minimize energy consumption by up to 33% compared to the existing approaches. Scaling the approach to multicore systems, we also demonstrate that it can minimize energy by up to 18% with 2× reduction in the learning time when compared with an existing approach. Rishad A. Shafik, Sheng Yang 0003, Anup Das 0001, Luis Alfonso Maeda-Nunez, Geoff V. Merrett, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Adaptive and Hierarchical Runtime Manager for Energy-Aware Thermal Management of Embedded SystemsabstractModern embedded systems execute applications, which interact with the operating system and hardware differently depending on the type of workload. These cross-layer interactions result in wide variations of the chip-wide thermal profile. In this article, a reinforcement learning-based runtime manager is proposed that guarantees application-specific performance requirements and controls the POSIX thread allocation and voltage/frequency scaling for energy-efficient thermal management. This controls three thermal aspects: peak temperature, average temperature, and thermal cycling. Contrary to existing learning-based runtime approaches that optimize energy and temperature individually, the proposed runtime manager is the first approach to combine the two objectives, simultaneously addressing all three thermal aspects. However, determining thread allocation and core frequencies to optimize energy and temperature is an NP-hard problem. This leads to exponential growth in the learning table (significant memory overhead) and a corresponding increase in the exploration time to learn the most appropriate thread allocation and core frequency for a particular application workload. To confine the learning space and to minimize the learning cost, the proposed runtime manager is implemented in a two-stage hierarchy: a heuristic-based thread allocation at a longer time interval to improve thermal cycling, followed by a learning-based hardware frequency selection at a much finer interval to improve average temperature, peak temperature, and energy consumption. This enables finer control on temperature in an energy-efficient manner while simultaneously addressing scalability, which is a crucial aspect for multi-/many-core embedded systems. The proposed hierarchical runtime manager is implemented for Linux running on nVidia’s Tegra SoC, featuring four ARM Cortex-A15 cores. Experiments conducted with a range of embedded and cpu-intensive applications demonstrate that the proposed runtime manager not only reduces energy consumption by an average 15% with respect to Linux but also improves all the thermal aspects—average temperature by 14°C, peak temperature by 16°C, and thermal cycling by 54%. Anup Das 0001, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Workload uncertainty characterization and adaptive frequency scaling for energy minimization of embedded systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 5 |
| 2015 | Hardware-software interaction for run-time power optimization: A case study of embedded Linux on multicore smartphonesabstractApplications running on smartphones interact with the hardware and the system software differently, resulting in widely varying power consumption and hence thermal profiles. Typically, these smartphone platforms expose some hardware power control features to users, controlled through software governors such as cpufreq for dynamic voltage-frequency scaling (DVFS) and cpuquiet for dynamic core selection (DCS). Operating systems on these platforms manage these governors conservatively, independent of application's performance requirement. To address this, we propose an alternative approach, which uses reinforcement learning to explore the trade-off between power saving opportunities using DVFS and DCS and application's performance at run-time. The objective is to reduce power consumption, taking into consideration dynamic power, leakage power, and the inter-dependency between temperature and power. The reinforcement learning-based control is validated as a case-study on ARM A15-based nvidia's tegra smartphone through its implementation as a run-time manager (RTM). This RTM interfaces with different hardware performance counters and the embedded Linux Operating System through (1) the cpuquiet API to select cores at run-time; and (2) the cpufreq API to scale the frequency of active cores. Experiments with mobile and high performance applications demonstrate that the proposed approach achieves an average 22% (7-40%) power reduction compared to existing techniques. Anup Das 0001, Matthew J. Walker, Andreas Hansson 0001, Bashir M. Al-Hashimi, Geoff V. Merrett |
ISLPED | 5 |
| 2015 | Application-specific memory protection policies for energy-efficient reliable designabstractIn this paper, we show that the vulnerability of memory components due to data retention in the presence of soft errors exhibit orders of magnitude variations with applications through extensive analysis of MiBench benchmarks. Underpinning such analysis, we propose a novel application-specific design flow for joint energy efficiency and reliability optimization. The energy efficiency is achieved through voltage/frequency scaling (VFS), while reliability is achieved through suitably choosing the appropriate protection policies (L1-Cache resizing and selective ECC) for hierarchical memory components. Fundamental to such joint optimization is a design analysis framework, which can analyze trade-off between memory protection policies considering the impact of VFS, and apply design optimization algorithm to provide with an energy-efficient design, while meeting a given reliability target. Using this framework the proposed design flow is validated through extensive number of application case studies based on ARMv7 processors modeled in GEM5. We show that the joint consideration of cache resizing and VFS can improve the L1-Cache reliability by up to 5x compared to VFS alone, while incurring <10% energy overhead. Additionally, using selective ECC for L2-Cache and DRAM, we show that energy consumption can be reduced by up to 40%. Sheng Yang 0003, Rishad A. Shafik, S. Saqib Khursheed, David Flynn, Geoff V. Merrett, Bashir M. Al-Hashimi |
RSP | 5 |
| 2015 | Poster: Solar-Powered Adaptive Street Lighting Evaluated with Real Traffic and Sunlight DataabstractStreet lighting is an important resource; it has been shown to reduce crime, improve road safety, and increase economic activity. These benefits, however, come with a cost: an annual emission of 64 million tonnes of CO2. Solar-powered street lighting is attractive for its use of renewable energy and its ease of installation (particularly in off-grid applications), but sizing and control is a non-trivial task. This paper describes TALiSMaN-Green, a traffic-aware street lighting scheme which takes account of road users as well as the available energy to dynamically adjust lighting levels. Simulations using real traffic and sunlight data illustrate that solar-powered streetlights can be managed to deliver consistent usefulness throughout the night. Sei Ping Lau, Alex S. Weddell, Neil M. White, Geoff V. Merrett |
SenSys | 4 |
| 2015 | ENSsys 2015: 3rd International Workshop on Energy Harvesting and Energy Neutral Sensing SystemsabstractComplementing the topics of ACM SenSys 2015, the 3rd International Workshop on Energy Harvesting and Energy Neutral Sensing Systems (ENSsys) 2015 brings together researchers to explore the challenges, issues and opportunities in the research, design, and engineering of energy-harvesting and energy-neutral sensing systems. These are an enabling technology for future applications in smart energy, transportation, environmental monitoring and smart cities. Innovative solutions in hardware for energy scavenging, adaptive algorithms, and power management policies are needed to enable uninterrupted operation. This one-day workshop features invited and peer reviewed talks on these areas, and provides a forum for feedback, discussion and networking. Geoff V. Merrett, Christian Renner, Davide Brunelli |
SenSys | 1 |
| 2015 | Poster: Enspect: Simplifying the Design of Energy Harvesting SystemsabstractThe design of sensing systems powered from energy harvesting can be complex. Design decisions are required concerning the properties and parameters of energy harvesting, conversion, and storage devices. The quantity and properties of environmental energy are typically both temporally and spatially variant, while the current consumption of the load electronics also changes dynamically. In this paper we describe Enspect, an open-source hardware/software tool which simplifies the design of energy harvesting sensing systems by assisting in the specification of harvesting and storage devices. It does this by enabling the long-term collection of data on energy availability, and modeling and simulating the performance of a complete system. Nick F. Tinsley, Stuart T. Witts, Jacob M. R. Ansell, Emily Barnes, Simeon M. Jenkins, Dhanushan Raveendran, Geoff V. Merrett, Alex S. Weddell |
SenSys | 7 |
| 2014 | Listening to the forest and its curators: lessons learnt from a bioacoustic smartphone application deploymentabstractOur natural environment is complex and sensitive, and is home to a number of species on the verge of extinction. Surveying is one approach to their preservation, and can be supported by technology. This paper presents the deployment of a smartphone-based citizen science biodiversity application. Our findings from interviews with members of the biodiversity community revealed a tension between the technology and their established working practices. From our experience, we present a series of general guidelines for those designing citizen science apps. Stuart Moran, Nadia Pantidi, Tom Rodden, Alan Chamberlain, Chloe Griffiths, Davide Zilli, Geoff V. Merrett, Alex Rogers |
CHI | 7 |
| 2014 | Reinforcement Learning-Based Inter- and Intra-Application Thermal Optimization for Lifetime Improvement of Multicore SystemsabstractThe thermal profile of multicore systems vary both within an application's execution (intra) and also when the system switches from one application to another (inter). In this paper, we propose an adaptive thermal management approach to improve the lifetime reliability of multicore systems by considering both inter- and intra-application thermal variations. Fundamental to this approach is a reinforcement learning algorithm, which learns the relationship between the mapping of threads to cores, the frequency of a core and its temperature (sampled from on-board thermal sensors). Action is provided by overriding the operating system's mapping decisions using affinity masks and dynamically changing CPU frequency using in-kernel governors. Lifetime improvement is achieved by controlling not only the peak and average temperatures but also thermal cycling, which is an emerging wear-out concern in modern systems. The proposed approach is validated experimentally using an Intel quad-core platform executing a diverse set of multimedia benchmarks. Results demonstrate that the proposed approach minimizes average temperature, peak temperature and thermal cycling, improving the mean-time-to-failure (MTTF) by an average of 2x for intra-application and 3x for inter-application scenarios when compared to existing thermal management techniques. Furthermore, the dynamic and static energy consumption are also reduced by an average 10% and 11% respectively. Anup Das 0001, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi, Akash Kumar 0001, Bharadwaj Veeravalli |
DAC | 3 |
| 2014 | A Hidden Markov Model-Based Acoustic Cicada Detector for Crowdsourced Smartphone Biodiversity MonitoringabstractIn recent years, the field of computational sustainability has striven to apply artificial intelligence techniques to solve ecological and environmental problems. In ecology, a key issue for the safeguarding of our planet is the monitoring of biodiversity. Automated acoustic recognition of species aims to provide a cost-effective method for biodiversity monitoring. This is particularly appealing for detecting endangered animals with a distinctive call, such as the New Forest cicada. To this end, we pursue a crowdsourcing approach, whereby the millions of visitors to the New Forest, where this insect was historically found, will help to monitor its presence by means of a smartphone app that can detect its mating call. Existing research in the field of acoustic insect detection has typically focused upon the classification of recordings collected from fixed field microphones. Such approaches segment a lengthy audio recording into individual segments of insect activity, which are independently classified using cepstral coefficients extracted from the recording as features. This paper reports on a contrasting approach, whereby we use crowdsourcing to collect recordings via a smartphone app, and present an immediate feedback to the users as to whether an insect has been found. Our classification approach does not remove silent parts of the recording via segmentation, but instead uses the temporal patterns throughout each recording to classify the insects present. We show that our approach can successfully discriminate between the call of the New Forest cicada and similar insects found in the New Forest, and is robust to common types of environment noise. A large scale trial deployment of our smartphone app collected over 6000 reports of insect activity from over 1000 users. Despite the cicada not having been rediscovered in the New Forest, the effectiveness of this approach was confirmed for both the detection algorithm, which successfully identified the same cicada through the app in countries where the same species is still present, and of the crowdsourcing methodology, which collected a vast number of recordings and involved thousands of contributors. Davide Zilli, Oliver Parson, Geoff V. Merrett, Alex Rogers |
J. Artif. Intell. Res. | 3 |
| 2014 | Active Mode Subclock Power GatingabstractThis paper presents a technique, called subclock power gating, for reducing leakage power during the active mode in low performance, energy-constrained applications. The proposed technique achieves power reduction through two mechanisms: 1) power gating the combinational logic within the clock period (subclock) and 2) reducing the virtual supply to less than Vth rather than shutting down completely as is the case in conventional power gating. To achieve this reduced voltage, a pair of nMOS and pMOS transistors are used at the head and foot of the power gated logic for symmetric virtual rail clamping of the power and ground supplies. The subclock power gating technique has been validated by incorporating it with an ARM Cortex-M0 microprocessor, which was fabricated in a 65-nm process. Two sets of experiments are done: the first experimentally validates the functionality of the proposed technique in the fabricated test chip and the second investigates the utility of the proposed technique in example applications. Measured results from the fabricated chip show 27% power saving during the active mode for an example wireless sensor node application when compared with the same microprocessor without subclock power gating. Jatin N. Mistry, James Myers, Bashir M. Al-Hashimi, David Flynn, John Biggs, Geoff V. Merrett |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2013 | DoE-based performance optimization of energy management in sensor nodes powered by tunable energy-harvesters
Tom J. Kazmierski, Leran Wang, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 4 |
| 2013 | A survey of multi-source energy harvesting systemsabstractEnergy harvesting allows low-power embedded devices to be powered from naturally-ocurring or unwanted environmental energy (e.g. light, vibration, or temperature difference). While a number of systems incorporating energy harvesters are now available commercially, they are specific to certain types of energy source. Energy availability can be a temporal as well as spatial effect. To address this issue, ‘hybrid’ energy harvesting systems combine multiple harvesters on the same platform, but the design of these systems is not straight-forward. This paper surveys their design, including trade-offs affecting their efficiency, applicability, and ease of deployment. This survey, and the taxonomy of multi-source energy harvesting systems that it presents, will be of benefit to designers of future systems. Furthermore, we identify and comment upon the current and future research directions in this field. Alex S. Weddell, Michele Magno, Geoff V. Merrett, Davide Brunelli, Bashir M. Al-Hashimi, Luca Benini |
DATE | 3 |
| 2013 | Opportunistic Direct Interconnection between Co-Located Wireless Sensor NetworksabstractWireless sensor networks are usually designed to avoid interaction with other networks. To share information, they are usually connected via a backbone network (e.g. the Internet) using gateways. The realization of visions for pervasive computing depends upon effective interconnection between individual networks. As the number of deployed sensor networks increases, the chance of any network having multiple neighbors also increases. In this paper, we argue that a paradigm shift towards 'opportunistic direct interconnection' is required. This enables one network to share information or resources with neighboring networks that it was unaware of at design-time. We present OI-MAC, which supports automatic neighbor discovery and cross- boundary data exchange without sacrificing the independence of each network. The effects of discovery and cross-boundary data injection are evaluated using both analytical models and network simulation. Initial results indicate that neighbor discovery has little effect on latency, while energy consumption increases insignificantly compared to ordinary operations of each node. If network traffic is doubled by packets 'injected' from a neighboring network, latency increases by around 7% while average power consumption increases by 20%. Teng Jiang, Geoff V. Merrett, Nick R. Harris |
ICCCN | 2 |
| 2013 | Energy and Accuracy Trade-Offs in Accelerometry-Based Activity RecognitionabstractDriven by real-world applications such as fitness, wellbeing and healthcare, accelerometry-based activity recognition has been widely studied to provide context-awareness to future pervasive technologies. Accurate recognition and energy efficiency are key issues in enabling long-term and unobtrusive monitoring. While the majority of accelerometry-based activity recognition systems stream data to a central point for processing, some solutions process data locally on the sensor node to save energy. In this paper, we investigate the trade-offs between classification accuracy and energy efficiency by comparing on- and off-node schemes. An empirical energy model is presented and used to evaluate the energy efficiency of both systems, and a practical case study (monitoring the physical activities of office workers) is developed to evaluate the effect on classification accuracy. The results show a 40% energy saving can be obtained with a 13% reduction in classification accuracy, but this performance depends heavily on the wearer's activity. Geoff V. Merrett, Robert G. Maunder, Alex Rogers |
ICCCN | 2 |
| 2013 | A Hidden Markov Model-Based Acoustic Cicada Detector for Crowdsourced Smartphone Biodiversity Monitoring
Davide Zilli, Oliver Parson, Geoff V. Merrett, Alex Rogers |
IJCAI | 3 |
| 2012 | An Explicit Linearized State-Space Technique for Accelerated Simulation of Electromagnetic Vibration Energy HarvestersabstractVibration energy harvesting systems pose significant modeling and design challenges due to their mixed-technology nature, extremely low levels of available energy and disparate time scales between different parts of a complete harvester. An energy harvester is a complex system of tightly coupled components modeled in the mechanical, magnetic, as well as electrical analog and digital domains. Currently available design tools are inadequate for simulating such systems due to prohibitive CPU times. This paper proposes a new technique to accelerate simulations of complete vibration energy harvesters by approximately two orders of magnitude. The proposed technique is to linearize the state equations of the system's analog components to obtain a fast estimate of the maximum step-size to guarantee the numerical stability of explicit integration based on the Adams-Bashforth formula. We show that the energy harvester's analog electronics can be efficiently and reliably simulated in this way with CPU times two orders of magnitude lower than those obtained from two state-of-the-art tools, VHDL-AMS and SystemC-A. As a case study, a practical, complex microgenerator with magnetic tuning and two types of power-processing circuits have been simulated using the proposed technique and verified experimentally. Tom J. Kazmierski, Leran Wang, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | Accelerated simulation of tunable vibration energy harvesting systems using a linearised state-space techniqueabstractThis paper proposes a linearised state-space technique to accelerate the simulation of tunable vibration energy harvesting systems by at least two orders of magnitude. The paper provides evidence that currently available simulation tools are inadequate for simulating complete energy harvesting systems where prohibitive CPU times are encountered due to disparate time scales. In the proposed technique, the model of a complete mixed-technology energy harvesting system is divided into component blocks whose mechanical and analogue electrical parts are modelled by local state equations and terminal variables while the digital electrical part is modelled as a digital process. Unlike existing simulation tools that use Newton-Raphson method, the proposed technique uses explicit integration such as Adams-Bashforth method to solve the state equations of the complete energy harvester model in short simulation time. Experimental measurements of a practical tunable energy harvester have been carried out to validate the proposed technique. Leran Wang, Tom J. Kazmierski, Bashir M. Al-Hashimi, Alex S. Weddell, Geoff V. Merrett, Ivo Netali Ayala-Garcia |
DATE | 5 |
| 2011 | Ultra low-power photovoltaic MPPT technique for indoor and outdoor wireless sensor nodesabstractPhotovoltaic (PV) energy harvesting is commonly used to power wireless sensor nodes. To optimise harvesting efficiency, maximum power point tracking (MPPT) techniques are often used. Recently-reported techniques focus solely on outdoor applications, being too power-hungry for use under indoor lighting. Additionally, some techniques have required light sensors (or pilot cells) to control their operating point. This paper describes an ultra low-power MPPT technique which is based on a novel system design and sample-and-hold arrangement, which enables MPPT across the range of light intensities found indoors and outdoors and is capable of cold-starting. The proposed sample-and-hold based technique has been validated through a prototype system. Its performance compares favourably against state-of-the-art systems, and does not require an additional pilot cell or photodiode. This represents an important contribution, in particular for sensors which may be exposed to different types of lighting (such as body-worn or mobile sensors). Alex S. Weddell, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 2 |
| 2009 | Energy-Aware Simulation for Wireless Sensor NetworksabstractEnergy-aware sensor nodes are usually tightly energy-constrained, execute energy-efficient algorithms, have the ability to interrogate and control the devices used for storing and consuming energy, and often feature one or more sources of energy harvesting. Due to the cost, time and expertise required to deploy a wireless sensor network (WSN), simulation is currently the most widely adopted evaluation method. Network simulation is well established for mobile ad hoc networks, using simulators such as the popular ns2. However, the differing characteristics and performance criteria of WSNs introduce additional simulation requirements, and this has resulted in a number of simulators and simulator extensions developed specifically for this purpose. This paper investigates the suitability of a number of state-of-the-art simulators for evaluating energy-aware WSNs, and subsequently proposes a novel structure for simulating energy-aware WSNs. The proposed structure provides diverse, flexible and extensible hardware and environment models, and integrates a structured architecture for embedded software to enhance the design of energy-aware sensor nodes. To illustrate an implementation of the structure, details of - and observations obtained using - an in-house simulator (WSNsim) are presented. Geoff V. Merrett, Neil M. White, Nick R. Harris, Bashir M. Al-Hashimi |
SECON | 1 |
| 2008 | Iterative Decoding for Redistributing Energy Consumption in Wireless Sensor NetworksabstractIn this paper, we propose a method for desirably redistributing a wireless sensor network's energy consumption from its sensor nodes (which may have scarce energy resources obtained through energy harvesting, for example) to its central node (which often has an abundant energy resource, such as the mains). At the cost of increasing the central node's decoding complexity, our method facilitates (1) a significant reduction in the number of times the sensor nodes are required to retransmit data owing to transmission errors and/or (2) a reduction of up to 3.99 dB in the sensor node's total transmit energy consumption. We show that our approach can reduce the overall energy consumption of transmitting sensor nodes by more than 20% in practice. Robert G. Maunder, Alex S. Weddell, Geoff V. Merrett, Bashir M. Al-Hashimi, Lajos Hanzo |
ICCCN | 3 |
| 2008 | A Structured Hardware/Software Architecture for Embedded Sensor NodesabstractOwing to the limited requirement for sensor processing in early networked sensor nodes, embedded software was generally built around the communication stack. Modern sensor nodes have evolved to contain significant on-board functionality in addition to communications, including sensor processing, energy management, actuation and locationing. The embedded software for this functionality, however, is often implemented in the application layer of the communications stack, resulting in an unstructured, top-heavy and complex stack. In this paper, we propose an embedded system architecture to formally specify multiple interfaces on a sensor node. This architecture differs from existing solutions by providing a sensor node with multiple stacks (each stack implements a separate node function), all linked by a shared application layer. This establishes a structured platform for the formal design, specification and implementation of modern sensor and wireless sensor nodes. We describe a practical prototype of an intelligent sensing, energy-aware, sensor node that has been developed using this architecture, implementing stacks for communications, sensing and energy management. The structure and operation of the intelligent sensing and energy management stacks are described in detail. The proposed architecture promotes structured and modular design, allowing for efficient code reuse and being suitable for future generations of sensor nodes featuring interchangeable components. Geoff V. Merrett, Alex S. Weddell, Nick R. Harris, Bashir M. Al-Hashimi, Neil M. White |
ICCCN | 1 |
| 2008 | An Empirical Energy Model for Supercapacitor Powered Wireless Sensor NodesabstractThe modeling of energy components in wireless sensor network (WSN) simulation is important for obtaining realistic lifetime predictions and ensuring the faithful operation of energy-aware algorithms. The use of supercapacitors as energy stores on WSN nodes is increasing, but their behavior differs from that of batteries. This paper proposes a model for a supercapacitor energy store based upon experimental results, and compares obtained simulation results to those using an 'ideal' energy store model. The proposed model also considers the variety and behavior of energy consumers, and finds that contrary to many existing models, the energy consumed depends on the store voltage (which varies considerably during supercapacitor discharge). Furthermore, energy models in a node's embedded firmware are shown to be paramount for providing energy-aware operation. Geoff V. Merrett, Alex S. Weddell, Adam P. Lewis, Nick R. Harris, Bashir M. Al-Hashimi, Neil M. White |
ICCCN | 1 |