EDBT 2026 Demo / reviewers in the wild / expert
Bashir M. Al-Hashimi
dblp:48/5139
· DBLP profile ↗
169ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-3591-1328ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 149 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 53 · 1 first-author · 1 since 2021Computer networks · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Backpropagation-Free On-Chip Training: Circuit-Algorithm Co-Design of a Goodness-Based Learning Tile
Qingchun Gong, Robert Bogdan Staszewski, Bashir M. Al-Hashimi, Kai Xu 0018 |
ISCAS | 3 |
| 2025 | Sparsity-Aware Optimization of In-Memory Bayesian Binary Neural Network AcceleratorsabstractBayesian Neural Networks (BNNs) provide principled estimates of model and data uncertainty by encoding parameters as distributions. This makes them key enablers for reliable AI that can be deployed on safety critical edge systems. These systems can be made resource efficient by restricting synapses to two synaptic states {−1, +1}, and using a memristive in-memory computing (IMC) paradigm. However, BNNs pose an additional challenge – they require multiple instantiations for ensembling, consuming extra resources in terms of energy and area. In this work, we propose a novel sparsity-aware optimization for Bayesian Binary Neural Network (BBNN) accelerators that exploits the inherent BBNN sampling sparsity – most of the network is made up of synapses that have a high probability of being fixed at ±1 and require no sampling. The optimization scheme proposed here exploits the sampling sparsity that exists both among layers, i.e., only a few layers of the network contain a majority of the probabilistic synapses, as well as the parameters i.e., a tiny fraction of parameters in these layers require sampling, reducing total sampled parameter count further by up to 86%. We demonstrate no loss in accuracy or uncertainty quantification performance for a VGGBinaryConnect network on CIFAR-100 dataset mapped on a custom sparsity-aware phase change memory (PCM) based IMC simulator. We also develop a simple drift compensation technique to demonstrate robustness to drift-induced degradation. Finally, we project latency, energy, and area for sparsity-aware BNN implementation in both pipelined and non-pipelined modes. With sparsity-aware implementation, we estimate upto 5.3× reduction in area and 8.8× reduction in energy compared to a non-sparsity-aware implementation. Our approach also results in 2.9× more power efficiency compared to the state-of-the-art BNN accelerator. Prabodh Katti, Bashir M. Al-Hashimi, Bipin Rajendran |
ISCAS | 2 |
| 2025 | Context-Aware Doubly-Robust Semi-Supervised LearningabstractThe widespread adoption of artificial intelligence (AI) in next-generation communication systems is challenged by the heterogeneity of traffic and network conditions, which call for the use of highly contextual, site-specific, data. A promising solution is to rely not only on real-world data, but also on synthetic pseudo-data generated by a network digital twin (NDT). However, the effectiveness of this approach hinges on the accuracy of the NDT, which can vary widely across different contexts. To address this problem, this paper introduces contextaware doubly-robust (CDR) learning, a novel semi-supervised scheme that adapts its reliance on the pseudo-data to the different levels of fidelity of the NDT across contexts. CDR is evaluated on the task of downlink beamforming where it outperforms previous state-of-the-art approaches, providing a 24% loss decrease when compared to doubly-robust (DR) semi-supervised learning in regimes with low labeled data availability. Clement Ruah, Houssem Sifaou, Osvaldo Simeone, Bashir M. Al-Hashimi |
IEEE Signal Process. Lett. | 4 |
| 2025 | Bayes2IMC: In-Memory Computing for Bayesian Binary Neural NetworksabstractBayesian Neural Networks (BNNs) generate an ensemble of possible models by treating model weights as random variables. This enables them to provide superior estimates of decision uncertainty. However, implementing Bayesian inference in hardware is resource-intensive, as it requires noise sources to generate the desired model weights. In this work, we introduce Bayes2IMC, an in-memory computing (IMC) architecture designed for binary BNNs that leverages the stochasticity inherent to nanoscale devices. Our novel design, based on Phase-Change Memory (PCM) crossbar arrays eliminates the necessity for Analog-to-Digital Converter (ADC) within the array, significantly improving power and area efficiency. Hardware-software co-optimized corrections are introduced to reduce device-induced accuracy variations across deployments on hardware, as well as to mitigate the effect of conductance drift of PCM devices. We validate the effectiveness of our approach on the CIFAR-10 dataset with a VGGBinaryConnect model containing 14 million parameters, achieving accuracy metrics comparable to ideal software implementations. We also present a complete core architecture, and compare its projected power, performance, and area efficiency against an equivalent SRAM baseline, showing a 3.8 to$9.6 \times $improvement in total efficiency (in GOPS/W/mm2) and a 2.2 to$5.6 \times $improvement in power efficiency (in GOPS/W). In addition, the projected hardware performance of Bayes2IMC surpasses most memristive BNN architectures reported in the literature, achieving up to 20% higher power efficiency compared to the state-of-the-art. Prabodh Katti, Clement Ruah, Osvaldo Simeone, Bashir M. Al-Hashimi, Bipin Rajendran |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Bayesian Inference Accelerator for Spiking Neural NetworksabstractBayesian neural networks offer better estimates of model uncertainty compared to frequentist networks. However, inference involving Bayesian models requires multiple instantiations or sampling of the network parameters, requiring significant computational resources. Compared to traditional deep learning networks, spiking neural networks (SNNs) have the potential to reduce computational area and power, thanks to their event-driven and spike-based computational framework. Most works in literature either address frequentist SNN models or non-spiking Bayesian neural networks. In this work, we demonstrate an optimization framework for developing and implementing efficient Bayesian SNNs in hardware by additionally restricting network weights to be binary-valued to further decrease power and area consumption. We demonstrate accuracies comparable to Bayesian binary networks with full-precision Bernoulli parameters, while requiring up to 25× less spikes than equivalent binary SNN implementations. We show the feasibility of the design by mapping it onto Zynq-7000, a lightweight SoC, and achieve a 6.5× improvement in GOPS/DSP while utilizing up to 30 times less power compared to the state-of-the-art. Prabodh Katti, Anagha Nimbekar, Amit Acharyya, Bashir M. Al-Hashimi, Bipin Rajendran |
ISCAS | 5 |
| 2023 | Content- and Lighting-Aware Adaptive Brightness Scaling for Improved Mobile User ExperienceabstractFor an improved user experience, the display sub-system is expected to provide superior resolution and optimal brightness despite its impact on battery life. Existing brightness scaling approaches set the display brightness statically or adaptively in response to predefined events such as low-battery or ambient light of the environment, which are independent of the displayed content. Approaches that consider the displayed content are either limited to video content or do not account for the user's expected battery life, thereby failing to maximise the user experience. This paper proposes Content- and ambient Lighting-aware Adaptive Brightness Scaling in mobile devices that maximises user experience while meeting battery life expectations. The approach employs a content- and ambient lighting-aware profiler that learns and classifies each sample into predefined clusters at runtime by leveraging insights on user perceptions of content and ambient luminance variations. We maximise user experience through adaptive scaling of the display's brightness using an energy prediction model that determines appropriate brightness levels while meeting expected battery life. The evaluation of the proposed approach on a commercial smartphone improves Quality of Experience (QoE) by up to 24.5 % compared to state-of-art. Samuel Isuwa, David Amos, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 4 |
| 2023 | Digital Twin-Based Multiple Access Optimization and Monitoring via Model-Driven Bayesian LearningabstractCommonly adopted in the manufacturing and aerospace sectors, digital twin (DT) platforms are increasingly seen as a promising paradigm to control and monitor software-based, “open”, communication systems, which play the role of the physical twin (PT). In the general framework presented in this work, the DT builds a Bayesian model of the communication system, which is leveraged to enable core DT functionalities such as control via multi-agent reinforcement learning (MARL) and monitoring of the PT for anomaly detection. We specifically investigate the application of the proposed framework to a simple case-study system encompassing multiple sensing devices that report to a common receiver. The Bayesian model trained at the DT has the key advantage of capturing epistemic uncertainty regarding the communication system, e.g., regarding current traffic conditions, which arise from limited PT-to-DT data transfer. Experimental results validate the effectiveness of the proposed Bayesian framework as compared to standard frequentist model-based solutions. Clement Ruah, Osvaldo Simeone, Bashir M. Al-Hashimi |
ICC | 3 |
| 2023 | Bayesian Inference on Binary Spiking Networks Leveraging Nanoscale Device StochasticityabstractBayesian Neural Networks (BNNs) can overcome the problem of overconfidence that plagues traditional frequentist deep neural networks, and are hence considered to be a key enabler for reliable AI systems. However, conventional hardware realizations of BNNs are resource intensive, requiring the imple-mentation of random number generators for synaptic sampling. Owing to their inherent stochasticity during programming and read operations, nanoscale memristive devices can be directly leveraged for sampling, without the need for additional hardware resources. In this paper, we introduce a novel Phase Change Memory (PCM)-based hardware implementation for BNNs with binary synapses. The proposed architecture consists of separate weight and noise planes, in which PCM cells are configured and operated to represent the nominal values of weights and to generate the required noise for sampling, respectively. Using experimentally observed PCM noise characteristics, for the ex-emplary Breast Cancer Dataset classification problem, we obtain hardware accuracy and expected calibration error matching that of an 8-bit fixed-point (FxP8) implementation, with projected savings of over$9\times$in terms of core area transistor count. Prabodh Katti, Nicolas Skatchkovsky, Osvaldo Simeone, Bipin Rajendran, Bashir M. Al-Hashimi |
ISCAS | 5 |
| 2023 | Maximising mobile user experience through self-adaptive content- and ambient-aware display brightness scalingabstractDisplay subsystems have become the predominant user interface on mobile devices, serving as both input and output interfaces. For a better quality of user experience (QoE), the display subsystem is expected to provide appropriate resolution and brightness despite its impact on battery life. Existing display brightness approaches either consider content- and ambient-light in isolation or do not account for the user’s expected battery life, thereby failing to maximise the QoE. This paper proposes aCADS, a self-Adaptive Content- and Ambient-aware Display brightness Scaling in mobile devices that maximises QoE while meeting battery life expectations. The approach employs a content- and ambient lighting-aware profiler that learns and classifies each sample into predefined clusters at runtime by leveraging insights on user perceptions of content and ambient luminances variations. We maximise QoE through adaptive scaling of the display’s brightness using an energy model that determines appropriate brightness levels while meeting expected battery life. The evaluation on a commercial smartphone shows that aCADS improves QoE by up to 32.5 % compared to state-of-the-art. Samuel Isuwa, David Amos, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
J. Syst. Archit. | 4 |
| 2023 | A Bayesian Framework for Digital Twin-Based Control, Monitoring, and Data Collection in Wireless SystemsabstractCommonly adopted in the manufacturing and aerospace sectors, digital twin (DT) platforms are increasingly seen as a promising paradigm to control, monitor, and analyze software-based, “open”, communication systems that are expected to dominate 6G deployments. Notably, DT platforms provide a sandbox in which to test artificial intelligence (AI) solutions for communication systems, potentially reducing the need to collect data and test algorithms in the field, i.e., on the physical twin (PT). A key challenge in the deployment of DT systems is to ensure that virtual control optimization, monitoring, and analysis at the DT are safe and reliable, avoiding incorrect decisions caused by “model exploitation”. To address this challenge, this paper presents a general Bayesian framework with the aim of quantifying and accounting for model uncertainty at the DT that is caused by limitations in the amount and quality of data available at the DT from the PT. In the proposed framework, the DT builds a Bayesian model of the communication system, which is leveraged to enable core DT functionalities such as control via multi-agent reinforcement learning (MARL), monitoring of the PT for anomaly detection, prediction, data-collection optimization, and counterfactual analysis. To exemplify the application of the proposed framework, we specifically investigate a case-study system encompassing multiple sensing devices that report to a common receiver. Experimental results validate the effectiveness of the proposed Bayesian framework as compared to standard frequentist model-based solutions. Clement Ruah, Osvaldo Simeone, Bashir M. Al-Hashimi |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Power-Efficient and Aging-Aware Primary/Backup Technique for Heterogeneous Embedded SystemsabstractOne of the essential requirements of embedded systems is a guaranteed level of reliability. In this regard, fault-tolerance techniques are broadly applied to these systems to enhance reliability. However, fault-tolerance techniques may increase power consumption due to their inherent redundancy. For this purpose, power management techniques are applied, along with fault-tolerance techniques, which generally prolong the system lifespan by decreasing the temperature and leading to an aging rate reduction. Yet, some power management techniques, such as Dynamic voltage and frequency scaling (DVFS), increase the transient fault rate and timing error. For this reason, heterogeneous multicore platforms have received much attention due to their ability to make a trade-off between power consumption and performance. Still, it is more complicated to map and schedule tasks in a heterogeneous multicore system. In this paper, for the first time, we propose a power management method for a heterogeneous multicore system that reduces power consumption and tolerates both transient and permanent faults through primary/backup technique while considering core-level power constraint, real-time constraint, and aging effect. Experimental evaluations demonstrate the efficiency of our proposed method in terms of reducing power consumption compared to the state-of-the-art schemes, together with guaranteeing reliability and considering the aging effect. Mohsen Ansari, Sepideh Safari, Nezam Rohbani, Alireza Ejlali, Bashir M. Al-Hashimi |
IEEE Trans. Sustain. Comput. | 5 |
| 2022 | Similarity-Aware CNN for Efficient Video Recognition at the EdgeabstractConvolutional neural networks (CNNs) often extract similar features from successive video frames due to having identical appearances. In contrast, conventional CNNs for video recognition process individual frames with a fixed computational effort. Each video frame is independently processed, resulting in numerous redundant computations and an inefficient use of limited energy resources, particularly for edge computing applications. To alleviate the high energy requirements associated with video frame processing, this article presented similarity-aware CNNs that recognize similar feature pixels across frames and avoid computations on them. First, with a loss of less than 1% in recognition accuracy, a proposed similarity-aware quantization technique increases the average number of unchanged feature pixels across frame pairs by up to 85%. Then, a proposed similarity-aware dataflow improves energy consumption by minimizing redundant computations and memory accesses across frame pairs. According to simulation experiments, the proposed dataflow decreases the energy consumed by video frame processing by up to 30%. Amin Sabet, Jonathon S. Hare, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | QUAREM: Maximising QoE Through Adaptive Resource Management in Mobile MPSoC PlatformsabstractHeterogeneous multi-processor system-on-chip (MPSoC) smartphones are required to offer increasing performance and user quality-of-experience (QoE) , despite comparatively slow advances in battery technology. Approaches to balance instantaneous power consumption, performance and QoE have been reported, but little research has considered how to perform longer-term budgeting of resources across a complete battery discharge cycle. Approaches that have considered this are oblivious to the daily variability in the user’s desired charging time-of-day (plug-in time), resulting in a failure to meet the user’s battery life expectations, or else an unnecessarily over-constrained QoE. This paper proposes QUAREM, an adaptive resource management approach in mobile MPSoC platforms that maximises QoE while meeting battery life expectations. The proposed approach utilises a model that learns and then predicts the dynamics of the energy usage pattern and plug-in times. Unlike state-of-the-art approaches, we maximise the QoE through the adaptive balancing of the battery life and the quality of service (QoS) for the duration of the battery discharge. Our model achieves a good degree of accuracy with a mean absolute percentage error of 3.47% and 2.48% for the energy demand and plug-in times, respectively. Experimental evaluation on an off-the-shelf commercial smartphone shows that QUAREM achieves the expected battery life of the user within 20–25% energy demand variation with little or no QoE degradation. Samuel Isuwa, Somdip Dey, Andre P. Ortega, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2020 | Optimising Resource Management for Embedded Machine LearningabstractMachine learning inference is increasingly being executed locally on mobile and embedded platforms, due to the clear advantages in latency, privacy and connectivity. In this paper, we present approaches for online resource management in heterogeneous multi-core systems and show how they can be applied to optimise the performance of machine learning work-loads. Performance can be defined using platform-dependent (e.g. speed, energy) and platform-independent (accuracy, confidence) metrics. In particular, we show how a Deep Neural Network (DNN) can be dynamically scalable to trade-off these various performance metrics. Achieving consistent performance when executing on different platforms is necessary yet challenging, due to the different resources provided and their capability, and their time-varying availability when executing alongside other workloads. Managing the interface between available hardware resources (often numerous and heterogeneous in nature), software requirements, and user experience is increasingly complex. Lei Xun, Long Tran-Thanh, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 3 |
| 2020 | Collaborative Adaptation for Energy-Efficient Heterogeneous Mobile SoCsabstractHeterogeneous Mobile System-on-Chips (SoCs) containing CPU and GPU cores are becoming prevalent in embedded computing, and they need to execute applications concurrently. However, existing run-time management approaches do not perform adaptive mapping and thread-partitioning of applications while exploiting both CPU and GPU cores at the same time. In this paper, we propose an adaptive mapping and thread-partitioning approach for energy-efficient execution of concurrent OpenCL applications on both CPU and GPU cores while satisfying performance requirements. To start execution of concurrent applications, the approach makes mapping (number of cores and operating frequencies) and partitioning (distribution of threads between CPU and GPU) decisions to satisfy performance requirements for each application. The mapping and partitioning decisions are made by having a collaboration between the CPU and GPU cores' processing capabilities such that balanced execution can be performed. During execution, adaptation is triggered when new application(s) arrive, or an executing one finishes, that frees cores. The adaptation process identifies a new mapping and thread-partitioning in a similar collaborative manner for remaining applications provided it leads to an improvement in energy efficiency. The proposed approach is experimentally validated on the Odroid-XU3 hardware platform with varying set of applications. Results show an average energy saving of 37%, compared to existing approaches while satisfying the performance requirements. Amit Kumar Singh 0002, Basireddy Karunakar Reddy, Alok Prakash, Geoff V. Merrett, Bashir M. Al-Hashimi |
IEEE Trans. Computers | 5 |
| 2020 | AdaMD: Adaptive Mapping and DVFS for Energy-Efficient Heterogeneous MulticoresabstractModern heterogeneous multicore systems, containing various types of cores, are increasingly dealing with concurrent execution of dynamic application workloads. Moreover, the performance constraints of each application vary, and applications enter/exit the system at any time. Existing approaches are not efficient in such dynamic scenarios, especially if applications are unknown, as they require extensive offline application analysis and do not consider the runtime execution scenarios (application arrival/completion, and workload and performance variations) for runtime management. To address this, we present AdaMD, an adaptive mapping and dynamic voltage and frequency scaling (DVFS) approach for improving energy consumption and performance. The key feature of the proposed approach is the elimination of dependency on offline profiled results while making runtime decisions. This is achieved through a performance prediction model having a maximum error of 7.9% lower than the previously reported model and a mapping approach that allocates processing cores to applications while respecting performance constraints. Furthermore, AdaMD adapts to runtime execution scenarios efficiently by monitoring the application status, and performance/workload variations to adjust the previous DVFS settings and thread-to-core mappings. The proposed approach is experimentally validated on the Odroid-XU3, with various combinations of diverse multithreaded applications from PARSEC and SPLASH benchmarks. Results show energy savings of up to 28% compared to the recently proposed approach while meeting performance constraints. Basireddy Karunakar Reddy, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | BRB: Mitigating Branch Predictor Side-ChannelsabstractModern processors use branch prediction as an optimization to improve processor performance. Predictors have become larger and increasingly more sophisticated in order to achieve higher accuracies which are needed in high performance cores. However, branch prediction can also be a source of side channel exploits, as one context can deliberately change the branch predictor state and alter the instruction flow of another context. Current mitigation techniques either sacrifice performance for security, or fail to guarantee isolation when retaining the accuracy. Achieving both has proven to be challenging. In this work we address this by, (1) introducing the notions of steady-state and transient branch predictor accuracy, and (2) showing that current predictors increase their misprediction rate by as much as 90% on average when forced to flush branch prediction state to remain secure. To solve this, (3) we introduce the branch retention buffer, a novel mechanism that partitions only the most useful branch predictor components to isolate separate contexts. Our mechanism makes thread isolation practical, as it stops the predictor from executing cold with little if any added area and no warm-up overheads. At the same time our results show that, compared to the state-of-the-art, average misprediction rates are reduced by 15-20% without increasing area, leading to a 2% performance increase. Ilias Vougioukas, Nikos Nikoleris, Andreas Sandberg, Stephan Diestelhorst, Bashir M. Al-Hashimi, Geoff V. Merrett |
HPCA | 5 |
| 2019 | The Circuit Breaker Pattern Targeted to Future IoT Applications
Gibeon Aquino, Rafael Fernandes de Queiroz, Geoff V. Merrett, Bashir M. Al-Hashimi |
ICSOC | 4 |
| 2019 | Run-time Detection and Mitigation of Power-Noise VirusesabstractPower-noise viruses can be used as denial-of-service attacks by causing voltage emergencies in multi-core microprocessors that may lead to data corruptions and system crashes. In this paper, we present a run-time system for detecting and mitigating power-noise viruses. We present voltage noise data from a power-noise virus and benchmarks collected from an Arm multi-core processor, and we observe that the frequency of voltage emergencies is dramatically increasing during the execution of power-noise attacks. Based on this observation, we propose a regression model that allows for a run-time estimation of the severity of voltage emergencies by monitoring the frequency of voltage emergencies and the operating frequency of the microprocessor. For mitigating the problem, during the execution of critical tasks that require protection, we propose a system which periodically evaluates the severity of voltage emergencies and adapts its operating frequency in order to honour a predefined severity constraint. We demonstrate the efficacy of the proposed run-time system. Vasileios Tenentes, Shidhartha Das, Daniele Rossi 0001, Bashir M. Al-Hashimi |
IOLTS | 4 |
| 2019 | Arbitrarily Parallel Turbo Decoding for Ultra-Reliable Low Latency Communication in 3GPP LTEabstractIn order to meet the latency requirements of the ultra-reliable low latency communication (URLLC) mode of the third-generation partnership project's long term evolution (LTE) mobile communication standard, this paper proposes a novel turbo decoding algorithm that supports an arbitrarily high degree of parallel processing, facilitating significantly higher processing throughputs and substantially lower processing latencies than the state-of-the-art (SOTA) LTE turbo decoder. As in conventional turbo decoding algorithms, the proposed Arbitrarily Parallel Turbo Decoder (APTD) decomposes each frame of information bits into a sequence of windows, where the bits within different windows are processed simultaneously using forward and backward recursions in a serial manner. However, in contrast to conventional turbo decoding algorithms, the APTD does not require different windows to be composed of an identical number of bits, which allows the use of an arbitrary number of windows and hence an arbitrary degree of parallelism, when decoding information bits of an arbitrary frame length. Furthermore, conventional turbo decoding algorithms alternate between simultaneously processing the windows in the upper decoder and those in the lower decoder. By contrast, the APTD processes the odd-indexed windows in the upper decoder at the same time as the even-indexed windows in the lower decoder and alternates between this and the reversed arrangement, hence further improving the decoding throughput and latency. Furthermore, the APTD achieves a reduced hardware resource requirement by calculating the extrinsic information based only on the outputs of the forward recursions, rather than based on both the forward and backward recursions of conventional turbo decoding algorithms. We demonstrate that the proposed APTD achieves superior latency, throughput, and computational efficiency than the SOTA LTE turbo decoder at all frame lengths, but particularly at the short frame lengths that are typically used in URLLC approaches. For example, at a frame length of N = 504 bits, the proposed APTD achieves an FER of 10-5at the same Eb/N0as I = 8 iterations of a conventional turbo decoder but with a computational efficiency that is 6 times higher than that of the SOTA turbo decoder, while achieving a latency and throughput that are 0.7 and 1.4 times those of the SOTA decoder, respectively. Luping Xiang, Matthew F. Brejza, Robert G. Maunder, Bashir M. Al-Hashimi, Lajos Hanzo |
IEEE J. Sel. Areas Commun. | 4 |
| 2019 | Momentum: Power-neutral Performance Scaling with Intrinsic MPPT for Energy Harvesting Computing SystemsabstractRecent research has looked to supplement or even replace the batteries in embedded computing systems with energy harvesting, where energy is derived from the device’s environment. However, such supplies are generally unpredictable and highly variable, and hence systems typically incorporate large external energy buffers (e.g., supercapacitors) to sustain computation; however, these pose environmental issues and increase system size and cost. This article proposes Momentum , a general power-neutral methodology, with intrinsic system-wide maximum power point tracking, that can be applied to a wide range of different computing systems, where the system dynamically scales its performance (and hence power consumption) to optimize computational progress depending on the power availability. Momentum enables the system to operate around an efficient operating voltage, maximizing forward application execution, without adding any external tracking or control units. This methodology combines at runtime (1) a hierarchical control strategy that utilizes available power management controls (such as dynamic voltage and frequency scaling, and core hot-plugging) to achieve efficient power-neutral operation; (2) a software-based maximum power point tracking scheme (unlike existing approaches, this does not require any additional hardware), which adapts the system power consumption so that it can work at the optimal operating voltage, considering the efficiency of the entire system rather than just the energy harvester; and (3) experimental validation on two different scales of computing system: a low power microcontroller (operating from the already-present 4.7μF decoupling capacitance) and a multi-processor system-on-chip (operating from 15.4mF added capacitance). Experimental results from both a controlled supply and energy harvesting source show that Momentum operates correctly on both platforms and exhibits improvements in forward application execution of up to 11% when compared to existing power-neutral approaches and 46% compared to existing static approaches. Domenico Balsamo, Benjamin J. Fletcher, Alex S. Weddell, Giorgos Karatziolas, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2019 | Predictive Thermal Management for Energy-Efficient Execution of Concurrent Applications on Heterogeneous MulticoresabstractCurrent multicore platforms contain different types of cores, organized in clusters (e.g., ARM's big.LITTLE). These platforms deal with concurrently executing applications, having varying workload profiles and performance requirements. Runtime management is imperative for adapting to such performance requirements and workload variabilities and to increase energy and temperature efficiency. Temperature has also become a critical parameter since it affects reliability, power consumption, and performance and, hence, must be managed. This paper proposes an accurate temperature prediction scheme coupled with a runtime energy management approach to proactively avoid exceeding temperature thresholds while maintaining performance targets. Experiments show up to 20% energy savings while maintaining high-temperature averages and peaks below the threshold. Compared with state-of-the-art temperature predictors, this paper predicts 35% faster and reduces the mean absolute error from 3.25 to 1.15 °C for the evaluated applications' scenarios. Eduardo Wächter, Cedric de Bellefroid, Basireddy Karunakar Reddy, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | Online concurrent workload classification for multi-core energy managementabstractModern embedded multi-core processors are organized as clusters of cores, where all cores in each cluster operate at a common Voltage-frequency (V-f). Such processors often need to execute applications concurrently, exhibiting varying and mixed workloads (e.g. compute- and memory-intensive) depending on the instruction mix and resource sharing. Runtime adaptation is key to achieving energy savings without trading-off application performance with such workload variabilities. In this paper, we propose an online energy management technique that performs concurrent workload classification using the metric Memory Reads Per Instruction (MRPI) and pro-actively selects an appropriate V-fsetting through workload prediction. Subsequently, it monitors the workload prediction error and performance loss, quantified by Instructions Per Second (IPS) at runtime and adjusts the chosen V-fto compensate. We validate the proposed technique on an Odroid-XU3 with various combinations of benchmark applications. Results show an improvement in energy efficiency of up to 69% compared to existing approaches. Basireddy Karunakar Reddy, Geoff V. Merrett, Bashir M. Al-Hashimi, Amit Kumar Singh 0002 |
DATE | 3 |
| 2018 | Collective-Aware System-on-Chips for Dependable IoT ApplicationsabstractIoT applications with low-budget connected nodes are emerging for a variety of domains, such as smart cities, geomonitoring, parking sensors, surveillance etc. These low-cost nodes contain System-on-Chips (SoCs) with networking capabil- ities. In this paper, we propose to exploit this feature for their dependability management. In particular, we propose collective- awareness, which is a run-time system that emerges when cloud resources are provided to the SoCs for IoT applications for storing information related to their in-the-field status, such as preferable operating modes and performance degradation. Periodically, a dynamic dependability model is constructed by the collected data and SoCs software is updated to meet user-defined lifetime, reliability and performance requirements. To evaluate the operations of the proposed system, we emulate the in-the- field performance degradation of a fleet of a 10K IoT nodes using Monte Carlo on temperature and workload conditions using the largest IWLS’05 benchmarks. During the first two years of system operation, the dynamically constructed model performs lifetime estimation with up to 57% higher accuracy, compared to a static model that considers data only from the design phase of the circuits, while after three years the dynamic model is always accurate for all the devices. Vasileios Tenentes, Daniele Rossi 0001, Bashir M. Al-Hashimi |
IOLTS | 3 |
| 2018 | Hardware-Validated CPU Performance and Energy ModellingabstractFull-system simulation frameworks such as gem5 are used extensively to evaluate research ideas and for design-space exploration. Moreover, energy-efficiency has become the key design constraint in recent years and many works use a separate power modelling framework to evaluate energy consumption. While such tools are convenient and flexible, they are known to contain sources of error which are often not fully understood and potentially impact the conclusions drawn from investigations. This work enables accurate, hardware-validated performance, power, and energy modelling of CPUs by first presenting a methodology to evaluate and identify sources of error in CPU performance models, and secondly developing empirical power models optimised for use with such performance models. Hierarchical clustering, correlation analysis, and regression techniques are used to identify sources of error without requiring detailed CPU specifications and enable existing models to be improved, new models to be developed, validation of simulator changes, and testing of model suitability for specific use-cases. Furthermore, the GemStone open-source software tool is presented, which automates the process of characterising hardware platforms, identifying sources of error in gem5 models, applying power analysis, and quantifying the effect of errors on the performance, power, and energy estimations. In addition, the mean percentage error in execution time was found to swing from -51% to +10% between two versions of the same gem5 model, underlining the need for an automated tool to validate models against reference hardware, ensuring accuracy and consistency. Matthew J. Walker, Sascha Bischoff, Stephan Diestelhorst, Geoff V. Merrett, Bashir M. Al-Hashimi |
ISPASS | 5 |
| 2018 | A model-based framework for software portability and verification in embedded power management systems
Asieh Salehi Fathabadi, Michael J. Butler, Sheng Yang 0003, Luis Alfonso Maeda-Nunez, James R. B. Bantock, Bashir M. Al-Hashimi, Geoff V. Merrett |
J. Syst. Archit. | 6 |
| 2018 | Exploiting Aging Benefits for the Design of Reliable Drowsy Cache MemoriesabstractIn this paper, we show how beneficial effects of aging on static power consumption can be exploited to design reliable drowsy cache memories adopting dynamic voltage scaling (DVS) to reduce static power. First, we develop an analytical model allowing designers to evaluate the long-term threshold voltage degradation induced by bias temperature instability (BTI) in a drowsy cache memory. Through HSPICE simulations, we demonstrate that, as drowsy memories age, static power reduction techniques based on DVS become more effective because of reduction in subthreshold current due to BTI aging. We develop a simulation framework to evaluate tradeoffs between static power and reliability, and a methodology to properly select the “drowsy” data retention voltage. We then propose different architectures of a drowsy cache memory allowing designers to meet different power and reliability constraints. The performed HSPICE simulations show a soft error rate and static noise margin improvement up to 20.8% and 22.7%, respectively, compared to standard aging unaware drowsy technique. This is achieved with a limited static power increase during the very early lifetime, and with static energy saving of up to 37% in 10 years of operation, at no or very limited hardware overhead. Daniele Rossi 0001, Vasileios Tenentes, Sudhakar M. Reddy, Bashir M. Al-Hashimi, Andrew D. Brown |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Leakage Current Analysis for Diagnosis of Bridge Defects in Power-Gating DesignsabstractManufacturing defects that do not affect the functional operation of low power integrated circuits (ICs) can nevertheless impact their power saving capability. We show that stuck-ON faults on the power switches and resistive bridges between the power networks can impair the power saving capability of power-gating designs. For quantifying the impact of such faults on the power savings of power-gating designs, we propose a diagnosis technique that targets bridges between the power networks. The proposed technique is based on the static power analysis of a power-gating design in stand-by mode and it utilizes a novel on-chip signature generation unit, which is sensitive to the voltage level between power rails, the measurements of which are processed off-line for the diagnosis of bridges that can adversely affect power savings. We explore, through SPICE simulation of the largest IWLS’05 benchmarks synthesized using a 32 nm CMOS technology, the tradeoffs achieved by the proposed technique between diagnosis accuracy and area cost and we evaluate its robustness against process variation. The proposed technique achieves a diagnosis resolution that is higher than 98.6% and 97.9% for bridges of${R}~{\gtrsim }~{10~{ M}\Omega }$(weak bridges) and bridges of${R~\lesssim~10~{ M}\Omega }$(strong bridges), respectively, and a diagnosis accuracy higher than 94.5% for all the examined defects. The area overhead is small and scalable: it is found to be 1.8% and 0.3% for designs with 27 K and 157 K gate equivalents, respectively. Vasileios Tenentes, Daniele Rossi 0001, S. Saqib Khursheed, Bashir M. Al-Hashimi, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Runtime Performance and Power Optimization of Parallel Disparity Estimation on Many-Core PlatformsabstractThis article investigates the use of many-core systems to execute the disparity estimation algorithm, used in stereo vision applications, as these systems can provide flexibility between performance scaling and power consumption. We present a learning-based runtime management approach that achieves a required performance threshold while minimizing power consumption through dynamic control of frequency and core allocation. Experimental results are obtained from a 61-core Intel Xeon Phi platform for the aforementioned investigation. The same performance can be achieved with an average reduction in power consumption of 27.8% and increased energy efficiency by 30.04% when compared to Dynamic Voltage and Frequency Scaling control alone without runtime management. Charles Leech, Charan Kumar Vala, Amit Acharyya, Sheng Yang 0003, Geoff V. Merrett, Bashir M. Al-Hashimi |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2017 | Machine learning for run-time energy optimisation in many-core systemsabstractIn recent years, the focus of computing has moved away from performance-centric serial computation to energy-efficient parallel computation. This necessitates run-time optimisation techniques to address the dynamic resource requirements of different applications on many-core architectures. In this paper, we report on intelligent run-time algorithms which have been experimentally validated for managing energy and application performance in many-core embedded system. The algorithms are underpinned by a cross-layer system approach where the hardware, system software and application layers work together to optimise the energy-performance trade-off. Algorithm development is motivated by the biological process of how a human brain (acting as an agent) interacts with the external environment (system) changing their respective states over time. This leads to a pay-off for the action taken, and the agent eventually learns to take the optimal/best decisions in future. In particular, our online approach uses a model-free reinforcement learning algorithm that suitably selects the appropriate voltage-frequency scaling based on workload prediction to meet the applications' performance requirements and achieve energy savings of up to 16% in comparison to state-of-the-art-techniques, when tested on four ARM A15 cores of an ODROID-XU3 platform. Dwaipayan Biswas, Vibishna Balagopal, Rishad A. Shafik, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 4 |
| 2017 | Energy-driven computing: Rethinking the design of energy harvesting systemsabstractEnergy harvesting computing has been gaining increasing traction over the past decade, fueled by technological developments and rising demand for autonomous and battery-free systems. Energy harvesting introduces numerous challenges to embedded systems but, arguably the greatest, is the required transition from an energy source that typically provides virtually unlimited power for a reasonable period of time until it becomes exhausted, to a power source that is highly unpredictable and dynamic (both spatially and temporally, and with a range spanning many orders of magnitude). The typical approach to overcome this is the addition of intermediate energy storage/buffering to smooth out the temporal dynamics of both power supply and consumption. This has the advantage that, if correctly sized, the system `looks like' a battery-powered system; however, it also adds volume, mass, cost and complexity and, if not sized correctly, unreliability. In this paper, we consider energy-driven computing, where systems are designed from the outset to operate from an energy harvesting source. Such systems typically contain little or no additional energy storage (instead relying on tiny parasitic and decoupling capacitance), alleviating the aforementioned issues. Examples of energy-driven computing include transient systems (which power down when the supply disappears and efficiently continue execution when it returns) and power-neutral systems (which operate directly from the instantaneous power harvested, gracefully modulating their consumption and performance to match the supply). In this paper, we introduce a taxonomy of energy-driven computing, articulating how power-neutral, transient, and energy-driven systems present a different class of computing to conventional approaches. Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 2 |
| 2017 | Online tuning of Dynamic Power Management for efficient execution of interactive workloadsabstractModern mobile devices contain powerful Multi-Processor System-on-Chips (MPSoCs) that are performance throttled by Dynamic Power Management (DPM) runtime systems to extend battery lifetime. Applications on mobile devices commonly generate highly interactive workloads, dependent on interaction between the processor cores, peripherals, external resources and the user, such as touch input during web-browsing. Inevitably, a subset of interactive workloads are affected by delays caused by data unavailability, e.g. loss or delay of data packets during voice-over-IP. At the same time, the system is required to respond quickly upon data retrieval to ensure that the user Quality of Experience (QoE) metrics (frame-rate, latency, etc.) are not degraded. Traditionally, operating systems have mitigated this problem with periodic sampling or event-driven approaches. Through experimentation using a mobile MPSoC platform, however, we demonstrate that improving the tuning of DPM parameters for certain interactive user inputs can provide energy savings of up to 21% or QoE improvements of up to 36%, when compared with the traditional approach. To capture these improvements, we propose a dynamic modeling of user input and data resource access times (e.g. mobile network bandwidth and latency) for interactive workloads, which is based on workload profiling and which we refer to herein as inelasticity analysis. The proposed approach is implemented through online tuning of a DPM runtime in the Android operating system and is validated through a Monte Carlo simulation of interactive workloads. In comparison to the default DPM tuning, the proposed approach achieves energy savings of 13% or QoE improvement of 27% or a selectable trade-off, e.g. 9% energy savings and 15% QoE improvement. James R. B. Bantock, Vasileios Tenentes, Bashir M. Al-Hashimi, Geoff V. Merrett |
ISLPED | 3 |
| 2017 | Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUsabstractModern mobile and embedded devices are required to be increasingly energy-efficient while running more sophisticated tasks, causing the CPU design to become more complex and employ more energy-saving techniques. This has created a greater need for fast and accurate power estimation frameworks for both run-time CPU energy management and design-space exploration. We present a statistically rigorous and novel methodology for building accurate run-time power models using performance monitoring counters (PMCs) for mobile and embedded devices, and demonstrate how our models make more efficient use of limited training data and better adapt to unseen scenarios by uniquely considering stability. Our robust model formulation reduces multicollinearity, allows separation of static and dynamic power, and allows a 100× reduction in experiment time while sacrificing only 0.6% accuracy. We present a statistically detailed evaluation of our model, highlighting and addressing the problem of heteroscedasticity in power modeling. We present software implementing our methodology and build power models for ARM Cortex-A7 and Cortex-A15 CPUs, with 3.8% and 2.8% average error, respectively. We model the behavior of the nonideal CPU voltage regulator under dynamic CPU activity to improve modeling accuracy by up to 5.5% in situations where the voltage cannot be measured. To address the lack of research utilizing PMC data from real mobile devices, we also present our data acquisition method and experimental platform software. We support this paper with online resources including software tools, documentation, raw data and further results. Matthew J. Walker, Stephan Diestelhorst, Andreas Hansson 0001, Anup Das 0001, Sheng Yang 0003, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | Energy-Efficient Run-Time Mapping and Thread Partitioning of Concurrent OpenCL Applications on CPU-GPU MPSoCsabstractHeterogeneous Multi-Processor Systems-on-Chips (MPSoCs) containing CPU and GPU cores are typically required to execute applications concurrently. However, as will be shown in this paper, existing approaches are not well suited for concurrent applications as they are developed either by considering only a single application or they do not exploit both CPU and GPU cores at the same time. In this paper, we propose an energy-efficient run-time mapping and thread partitioning approach for executing concurrent OpenCL applications on both GPU and GPU cores while satisfying performance requirements. Depending upon the performance requirements, for each concurrently executing application, the mapping process finds the appropriate number of CPU cores and operating frequencies of CPU and GPU cores, and the partitioning process identifies an efficient partitioning of the applications’ threads between CPU and GPU cores. We validate the proposed approach experimentally on the Odroid-XU3 hardware platform with various mixes of applications from the Polybench benchmark suite. Additionally, a case-study is performed with a real-world application SLAMBench. Results show an average energy saving of 32% compared to existing approaches while still satisfying the performance requirements. Amit Kumar Singh 0002, Alok Prakash, Basireddy Karunakar Reddy, Geoff V. Merrett, Bashir M. Al-Hashimi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | Nucleus: Finding the Sharing Limit of Heterogeneous CoresabstractHeterogeneous multi-processors are designed to bridge the gap between performance and energy efficiency in modern embedded systems. This is achieved by pairing Out-of-Order (OoO) cores, yielding performance through aggressive speculation and latency masking, with In-Order (InO) cores, that preserve energy through simpler design. By leveraging migrations between them, workloads can therefore select the best setting for any given energy/delay envelope. However, migrations introduce execution overheads that can hurt performance if they happen too frequently. Finding the optimal migration frequency is critical to maximize energy savings while maintaining acceptable performance. We develop a simulation methodology that can 1) isolate the hardware effects of migrations from the software, 2) directly compare the performance of different core types, 3) quantify the performance degradation and 4) calculate the cost of migrations for each case. To showcase our methodology we run mibench, a microbenchmark suite, and show that migrations can happen as fast as every 100k instructions with little performance loss. We also show that, contrary to numerous recent studies, hypothetical designs do not need to share all of their internal components to be able to migrate at that frequency. Instead, we propose a feasible system that shares level 2 caches and a translation lookaside buffer that matches performance and efficiency. Our results show that there are phases comprising up to 10% that a migration to the OoO core leads to performance benefits without any additional energy cost when running on the InO core, and up to 6% of phases where a migration to the InO core can save energy without affecting performance. When considering a policy that focuses on improving the energy-delay product, results show that on average 66% of the phases can be migrated to deliver equal or better system operation without having to aggressively share the entire memory system or to revert to migration periods finer than 100k instructions. Ilias Vougioukas, Andreas Sandberg, Stephan Diestelhorst, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2017 | Coarse-Grained Online Monitoring of BTI Aging by Reusing Power-Gating InfrastructureabstractIn this paper, we present a novel coarse-grained technique for monitoring online the bias temperature instability (BTI) aging of circuits by exploiting their power gating infrastructure. The proposed technique relies on monitoring the discharge time of the virtual-power-network during standby operations, the value of which depends on the threshold voltage of the CMOS devices in a power-gated design (PGD). It does not require any distributed sensors, because the virtual-power-network is already distributed in a PGD. It consists of a hardware block for measuring the discharge time concurrently with normal standby operations and a processing block for estimating the BTI aging status of the PGD according to collected measurements. Through SPICE simulation, we demonstrate that the BTI aging estimation error of the proposed technique is less than 1% and 6.2% for PGDs with static operating frequency and dynamic voltage and frequency scaling, respectively. Its area cost is also found negligible. The power gating minimum idle time (MIT) cost induced by the energy consumed for monitoring the discharge time is evaluated on two scalar machine models using either x86 or ARM instruction sets. It is found less than 1.3× and 1.45× the original power gating MIT, respectively. We validate the proposed technique through accelerated aging experiments conducted with five actual chips that contain an ARM cortex M0 processor, manufactured with a 65 nm CMOS technology. Vasileios Tenentes, Daniele Rossi 0001, Sheng Yang 0003, S. Saqib Khursheed, Bashir M. Al-Hashimi, Steve R. Gunn |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | The slowdown or race-to-idle question: Workload-aware energy optimization of SMT multicore platforms under process variation
Anup Das 0001, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 3 |
| 2016 | Co-optimization of fault tolerance, wirelength and temperature mitigation in TSV-based 3D ICsabstractTSV failures due to manufacturing defects and thermal-induced latent defects result in yield and reliability issues in 3D-ICs. Recent work has shown different temperature mitigation techniques and fault tolerant architectures for 3D-ICs. It is known that TSVs are effective in reducing temperature by providing thermal conductivity. This is the first work that jointly considers temperature mitigation and fault tolerance for TSV based 3D ICs without introducing additional redundant TSVs (also called dummy TSVs). By reusing and carefully placing spare TSVs that are frequently deployed for improving yield and reliability in 3D ICs, temperature is reduced without affecting fault tolerance capability. The proposed technique consists of two steps: first is TSV determination step, which provides optimised allocation of regular and spare TSVs in groups to achieve expected repair capability. The second step is TSV placement, where, for the first time, temperature mitigation is addressed when considering TSVs impact on both vertical and horizontal heat flow. Meanwhile routing difference and total wirelength can be co-optimized. Simulation results show that using the proposed technique, 100% repair capability is achieved across all (five) benchmarks with an average temperature reduction of 33% (best case is 58.2%), while the wirelength may slightly increase depending on assumed TSV fault rate. Yi Zhao 0001, S. Saqib Khursheed, Bashir M. Al-Hashimi, Zhiwen Zhao |
VLSI-SoC | 3 |
| 2016 | Workload Change Point Detection for Runtime Thermal Management of Embedded SystemsabstractApplications executed on multicore embedded systems interact with system software [such as the operating system (OS)] and hardware, leading to widely varying thermal profiles which accelerate some aging mechanisms, reducing the lifetime reliability. Effectively managing the temperature therefore requires: 1) autonomous detection of changes in application workload and 2) appropriate selection of control levers to manage thermal profiles of these workloads. In this paper, we propose a technique for workload change detection using density ratio-based statistical divergence between overlapping sliding windows of CPU performance statistics. This is integrated in a runtime approach for thermal management, which uses reinforcement learning to select workload-specific thermal control levers by sampling on-board thermal sensors. Identified control levers override the OSs native thread allocation decision and scale hardware voltage-frequency to improve average temperature, peak temperature, and thermal cycling. The proposed approach is validated through its implementation as a hierarchical runtime manager for Linux, with heuristic-based thread affinity selected from the upper hierarchy to reduce thermal cycling and learningbased voltage-frequency selected from the lower hierarchy to reduce average and peak temperatures. Experiments conducted with mobile, embedded, and high performance applications on ARM-based embedded systems demonstrate that the proposed approach increases workload change detection accuracy by an average 3.4×, reducing the average temperature by 4 °C-25 °C, peak temperature by 6 °C-24 °C, and thermal cycling by 7%-35% over state-of-the-art approaches. Anup Das 0001, Geoff V. Merrett, Mirco Tribastone, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Graceful Performance Modulation for Power-Neutral Transient Computing SystemsabstractTransient computing systems do not have energy storage, and operate directly from energy harvesting. These systems are often faced with the inherent challenge of low-current or transient power supply. In this paper, we propose “power-neutral” operation, a new paradigm for such systems, whereby the instantaneous power consumption of the system must match the instantaneous harvested power. Power neutrality is achieved using a control algorithm for dynamic frequency scaling, modulating system performance gracefully in response to the incoming power. Detailed system model is used to determine design parameters for selecting the system voltage thresholds where the operating frequency will be raised or lowered, or the system will be hibernated. The proposed control algorithm for power-neutral operation is experimentally validated using a microcontroller incorporating voltage threshold-based interrupts for frequency scaling. The microcontroller is powered directly from real energy harvesters; results demonstrate that a power-neutral system sustains operation for 4%-88% longer with up to 21% speedup in application execution. Domenico Balsamo, Anup Das 0001, Alex S. Weddell, Davide Brunelli, Bashir M. Al-Hashimi, Geoff V. Merrett, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Hibernus++: A Self-Calibrating and Adaptive System for Transiently-Powered Embedded DevicesabstractEnergy harvesters are being used to power autonomous systems, but their output power is variable and intermittent. To sustain computation, these systems integrate batteries or supercapacitors to smooth out rapid changes in harvester output. Energy storage devices require time for charging and increase the size, mass, and cost of systems. The field of transient computing moves away from this approach, by powering the system directly from the harvester output. To prevent an application from having to restart computation after a power outage, approaches such as Hibernus allow these systems to hibernate when supply failure is imminent. When the supply reaches the operating threshold, the last saved state is restored and the operation is continued from the point it was interrupted. This paper proposes Hibernus++ to intelligently adapt the hibernate and restore thresholds in response to source dynamics and system load properties. Specifically, capabilities are built into the system to autonomously characterize the hardware platform and its performance during hibernation in order to set the hibernation threshold at a point which minimizes wasted energy and maximizes computation time. Similarly, the system auto-calibrates the restore threshold depending on the balance of energy supply and consumption in order to maximize computation time. Hibernus++ is validated both theoretically and experimentally on microcontroller hardware using both synthesized and real energy harvesters. Results show that Hibernus++ provides an average 16% reduction in energy consumption and an improvement of 17% in application execution time over state-of-the-art approaches. Domenico Balsamo, Alex S. Weddell, Anup Das 0001, Alberto Rodriguez Arreola, Davide Brunelli, Bashir M. Al-Hashimi, Geoff V. Merrett, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | Learning Transfer-Based Adaptive Energy Minimization in Embedded SystemsabstractEmbedded systems execute applications with varying performance requirements. These applications exercise the hardware differently depending on the computation task, generating varying workloads with time. Energy minimization with such workload and performance variations within (intra) and across (inter) applications is particularly challenging. To address this challenge, we propose an online approach, capable of minimizing energy through adaptation to these variations. At the core of this approach is a reinforcement learning algorithm that suitably selects the appropriate voltage/frequency scaling (VFS) based on workload predictions to meet the applications' performance requirements. The adaptation is then facilitated and expedited through learning transfer, which uses the interaction between the application, runtime, and hardware layers to adjust the VFS. The proposed approach is implemented as a power governor in Linux and extensively validated on an ARM Cortex-A8 running different benchmark applications. We show that with intra- and inter-application variations, our proposed approach can effectively minimize energy consumption by up to 33% compared to the existing approaches. Scaling the approach to multicore systems, we also demonstrate that it can minimize energy by up to 18% with 2× reduction in the learning time when compared with an existing approach. Rishad A. Shafik, Sheng Yang 0003, Anup Das 0001, Luis Alfonso Maeda-Nunez, Geoff V. Merrett, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | Adaptive and Hierarchical Runtime Manager for Energy-Aware Thermal Management of Embedded SystemsabstractModern embedded systems execute applications, which interact with the operating system and hardware differently depending on the type of workload. These cross-layer interactions result in wide variations of the chip-wide thermal profile. In this article, a reinforcement learning-based runtime manager is proposed that guarantees application-specific performance requirements and controls the POSIX thread allocation and voltage/frequency scaling for energy-efficient thermal management. This controls three thermal aspects: peak temperature, average temperature, and thermal cycling. Contrary to existing learning-based runtime approaches that optimize energy and temperature individually, the proposed runtime manager is the first approach to combine the two objectives, simultaneously addressing all three thermal aspects. However, determining thread allocation and core frequencies to optimize energy and temperature is an NP-hard problem. This leads to exponential growth in the learning table (significant memory overhead) and a corresponding increase in the exploration time to learn the most appropriate thread allocation and core frequency for a particular application workload. To confine the learning space and to minimize the learning cost, the proposed runtime manager is implemented in a two-stage hierarchy: a heuristic-based thread allocation at a longer time interval to improve thermal cycling, followed by a learning-based hardware frequency selection at a much finer interval to improve average temperature, peak temperature, and energy consumption. This enables finer control on temperature in an energy-efficient manner while simultaneously addressing scalability, which is a crucial aspect for multi-/many-core embedded systems. The proposed hierarchical runtime manager is implemented for Linux running on nVidia’s Tegra SoC, featuring four ARM Cortex-A15 cores. Experiments conducted with a range of embedded and cpu-intensive applications demonstrate that the proposed runtime manager not only reduces energy consumption by an average 15% with respect to Linux but also improves all the thermal aspects—average temperature by 14°C, peak temperature by 16°C, and thermal cycling by 54%. Anup Das 0001, Bashir M. Al-Hashimi, Geoff V. Merrett |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Two-Phase Low-Energy N-Modular Redundancy for Hard Real-Time Multi-Core SystemsabstractThis paper proposes an N-modular redundancy (NMR) technique with low energy-overhead for hard real-time multi-core systems. NMR is well-suited for multi-core platforms as they provide multiple processing units and low-overhead communication for voting. However, it can impose considerable energy overhead and hence its energy overhead must be controlled, which is the primary consideration of this paper. For this purpose the system operation can be divided into two phases: indispensable phase and on-demand phase. In the indispensable phase only half-plus-one copies for each task are executed. When no fault occurs during this phase, the results must be identical and hence the remaining copies are not required. Otherwise, the remaining copies must be executed in the on-demand phase to perform a complete majority voting. In this paper, for such a two-phase NMR, an energy-management technique is developed where two new concepts have been considered:i) Block-partitioned scheduling that enables parallel task execution during on-demand phase, thereby leaving more slack for energy saving,ii) Pseudo-dynamic slack, that results when a task has no faulty execution during the indispensable phase and hence the time which is reserved for its copies in the on-demand phase is reclaimed for energy saving. The energy-management technique has an off-line part that manages static and pseudo-dynamic slacks at design time and an online part that mainly manages dynamic slacks at run-time. Experimental results show that the proposed NMR technique provides up to 29 percent energy saving and is 6 orders of magnitude higher reliable as compared to a recent previous work. Alireza Ejlali, Bashir M. Al-Hashimi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Sequence-Aware Watermark Design for Soft IP Embedded ProcessorsabstractThis paper describes a design approach for incorporating sequence-aware watermarks in soft intellectual property (IP) embedded processors. The influence of watermark sequence parameters on detection, area, and power overheads is examined, and consequently a method for incorporating sequence-aware watermarks in soft IP embedded processors is proposed. The intrinsic parameters of sequences, such as the activity factor and the overlapping factor, are introduced, and their impact on correlation results is demonstrated. Measurement and application-specified integrated circuits validate the design approach and demonstrate the resulting IP protection and subsequent costs for constrained embedded processors. Results presented in this paper show that the tradeoff occurs between the watermark robustness against third-party IP attacks and hardware implementation costs. The analysis of this tradeoff is provided, and an application specific watermark implementation is proposed. Jedrzej Kufel, Peter R. Wilson, Stephen Hill, Bashir M. Al-Hashimi, Paul N. Whatmough |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Reliable Power Gating With NBTI Aging BenefitsabstractIn this paper, we show that negative bias temperature instability (NBTI) aging of sleep transistors (STs), together with its detrimental effect for circuit performance and lifetime (LT), presents considerable benefits for power-gated circuits. Indeed, it reduces static power due to leakage current, and increases ST switch efficiency, making power gating more efficient and effective over time. The magnitude of these aging benefits depends on operating and environmental conditions. By means of HSPICE simulations, considering a 32-nm CMOS technology, we demonstrate that static power may reduce by more than 80% in 10 years of operation. Static power decrease over time due to NBTI aging is also proven experimentally, using a test chip manufactured with a 65-nm technology. We propose an ST design strategy for reliable power gating, in order to harvest the benefits offered by NBTI aging. It relies on the design of STs with a proper lower$V_{\textrm {th}}$compared with the standard STs. This can be achieved by either redesigning the STs with the identified$V_{\textrm {th}}$value or applying a proper forward body bias to the available power switching fabrics. Through the HSPICE simulations, we show LT extension up to$21.4\times $and average static power reduction up to 16.3% compared with the standard ST design approach, without additional area overhead. Finally, we show LT extension and several performance-cost tradeoffs when a target maximum LT is considered. Daniele Rossi 0001, Vasileios Tenentes, Sheng Yang 0003, S. Saqib Khursheed, Bashir M. Al-Hashimi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Workload uncertainty characterization and adaptive frequency scaling for energy minimization of embedded systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 6 |
| 2015 | NBTI and leakage aware sleep transistor design for reliable and energy efficient power gatingabstractIn this paper we show that power gating techniques become more effective during their lifetime, since the aging of sleep transistors (STs) due to negative bias temperature instability (NBTI) drastically reduces leakage power. Based on this property, we propose an NBTI and leakage aware ST design method for reliable and energy efficient power gating. Through SPICE simulations, we show lifetime extension up to 19.9x and average leakage power reduction up to 14.4% compared to standard STs design approach without additional area overhead. Finally, when a maximum 10-year lifetime target is considered, we show that the proposed method allows multiple beneficial options compared to a standard STs design method: either to improve circuit operating frequency up to 9.53% or to reduce ST area overhead up to 18.4%. Daniele Rossi 0001, Vasileios Tenentes, S. Saqib Khursheed, Bashir M. Al-Hashimi |
ETS | 4 |
| 2015 | Diagnosis of power switches with power-distribution-network considerationabstractThis paper examines diagnosis of power switches when the power-distribution-network (PDN) is considered as a high resolution distributed electrical model. The analysis shows that for a diagnosis method to perform high diagnosis accuracy and resolution, the distributed nature of PDN should not be simplified by a lumped model. For this reason, a PDN-aware diagnosis method for power switches fault grading is proposed. The proposed method utilizes a novel signature generation design-for-testability (DFT) unit, the signatures of which are processed by a novel diagnosis algorithm that grades the magnitude of faults. Through simulations of physical layout SPICE models, we explore the trade-offs of the proposed method between diagnosis accuracy and diagnosis resolution against area overhead and we show that 100% diagnosis accuracy and up to 98% diagnosis resolution can be achieved with negligible cost. Vasileios Tenentes, Daniele Rossi 0001, S. Saqib Khursheed, Bashir M. Al-Hashimi |
ETS | 4 |
| 2015 | Adaptive iterative detection for expediting the convergence of a serially concatenated Unary Error Correction decoder, turbo decoder and an iterative demodulatorabstractUnary Error Correction (UEC) codes constitute a recently proposed Joint Source and Channel Code (JSCC) family, conceived for alphabets having an infinite cardinality, whilst out-performing previously used Separate Source and Channel Codes (SSCCs). UEC based schemes rely on an iterative decoding process, which involves three decoding blocks when concatenated with a turbo code. Owing to this, following the activation of one of the three blocks, the next block to be activated must be chosen from the other two decoding block options. Furthermore, the UEC decoder offers a number of decoding options, allowing its complexity and error correction capability to be dynamically adjusted. It has been shown that iterative decoding convergence can be expedited by activating the specific decoding option that offers the highest Mutual Information (MI) improvement to computational complexity ratio. This paper introduces an iterative demodulator, which is shown to improve the associated error correction performance, while reducing the overall iterative decoding complexity. The challenge is that the iterative demodulator has to forward its soft-information to the other two iterative decoding blocks, and hence the corresponding MI improvements cannot be compared on a like-for-like basis. Additionally, we also propose a method of eliminating the logarithmic calculations from the adaptive iterative decoding algorithm, hence further reducing its implementational complexity without impacting its error correcting performance. Matthew F. Brejza, Wenbo Zhang 0011, Robert G. Maunder, Bashir M. Al-Hashimi, Lajos Hanzo |
ICC | 4 |
| 2015 | BTI and leakage aware dynamic voltage scaling for reliable low power cache memoriesabstractWe propose a novel dynamic voltage scaling (DVS) approach for reliable and energy efficient cache memories. First, we demonstrate that, as memories age, leakage power reduction techniques become more effective due to sub-threshold current reduction with aging. Then, we provide an analytical model and a design exploration framework to evaluate trade-offs between leakage power and reliability, and propose a BTI and leakage aware selection of the “drowsy” state retention voltage for DVS of cache memories. We propose three DVS policies, allowing us to achieve different power/reliability trade-offs. Through SPICE simulations, we show that a critical charge and a static noise margin increase up to 150% and 34.7%, respectively, is achieved compared to standard aging unaware drowsy technique, with a limited leakage power increase during the very early lifetime, and with leakage energy saving up to 37% in 10 years of operation. These improvements are attained at zero or negligible area cost. Daniele Rossi 0001, Vasileios Tenentes, S. Saqib Khursheed, Bashir M. Al-Hashimi |
IOLTS | 4 |
| 2015 | Hardware-software interaction for run-time power optimization: A case study of embedded Linux on multicore smartphonesabstractApplications running on smartphones interact with the hardware and the system software differently, resulting in widely varying power consumption and hence thermal profiles. Typically, these smartphone platforms expose some hardware power control features to users, controlled through software governors such as cpufreq for dynamic voltage-frequency scaling (DVFS) and cpuquiet for dynamic core selection (DCS). Operating systems on these platforms manage these governors conservatively, independent of application's performance requirement. To address this, we propose an alternative approach, which uses reinforcement learning to explore the trade-off between power saving opportunities using DVFS and DCS and application's performance at run-time. The objective is to reduce power consumption, taking into consideration dynamic power, leakage power, and the inter-dependency between temperature and power. The reinforcement learning-based control is validated as a case-study on ARM A15-based nvidia's tegra smartphone through its implementation as a run-time manager (RTM). This RTM interfaces with different hardware performance counters and the embedded Linux Operating System through (1) the cpuquiet API to select cores at run-time; and (2) the cpufreq API to scale the frequency of active cores. Experiments with mobile and high performance applications demonstrate that the proposed approach achieves an average 22% (7-40%) power reduction compared to existing techniques. Anup Das 0001, Matthew J. Walker, Andreas Hansson 0001, Bashir M. Al-Hashimi, Geoff V. Merrett |
ISLPED | 4 |
| 2015 | Application-specific memory protection policies for energy-efficient reliable designabstractIn this paper, we show that the vulnerability of memory components due to data retention in the presence of soft errors exhibit orders of magnitude variations with applications through extensive analysis of MiBench benchmarks. Underpinning such analysis, we propose a novel application-specific design flow for joint energy efficiency and reliability optimization. The energy efficiency is achieved through voltage/frequency scaling (VFS), while reliability is achieved through suitably choosing the appropriate protection policies (L1-Cache resizing and selective ECC) for hierarchical memory components. Fundamental to such joint optimization is a design analysis framework, which can analyze trade-off between memory protection policies considering the impact of VFS, and apply design optimization algorithm to provide with an energy-efficient design, while meeting a given reliability target. Using this framework the proposed design flow is validated through extensive number of application case studies based on ARMv7 processors modeled in GEM5. We show that the joint consideration of cache resizing and VFS can improve the L1-Cache reliability by up to 5x compared to VFS alone, while incurring <10% energy overhead. Additionally, using selective ECC for L2-Cache and DRAM, we show that energy consumption can be reduced by up to 40%. Sheng Yang 0003, Rishad A. Shafik, S. Saqib Khursheed, David Flynn, Geoff V. Merrett, Bashir M. Al-Hashimi |
RSP | 6 |
| 2015 | DFT Architecture With Power-Distribution-Network Consideration for Delay-Based Power Gating TestabstractThis paper shows that existing delay-based testing techniques for power gating exhibit both fault coverage and yield loss due to deviations at the charging delay introduced by the distributed nature of the power-distribution-networks (PDNs). To restore this test quality (TQ) loss, which could reach up to 67.7% of false passes and 25% of false fails due to stuck-open faults, we propose a design-for-testability logic that accounts for a distributed PDN. The proposed logic is optimized by an algorithm that also handles uncertainty due to process variations and offers tradeoff flexibility between test application time and area cost. A calibration process is proposed to bridge model-to-hardware discrepancies and increase TQ when considering systematic variations. Through SPICE simulations, we show complete recovery of the TQ lost due to PDNs. The proposed method is robust, sustaining 80.3%–98.6% of the achieved TQ under high random and systematic process variations. To the best of our knowledge, this paper presents the first analysis of the PDN impact on TQ and offers a unified test solution for both ring and grid power gating styles. Vasileios Tenentes, S. Saqib Khursheed, Daniele Rossi 0001, Sheng Yang 0003, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2015 | Online Fault Tolerance Technique for TSV-Based 3-D-ICabstractThis brief presents the design, validation, and evaluation of an efficient online fault tolerance technique for fault detection and recovery in presence of three through-silicon-vias (TSV) defects: 1) voids; 2) delamination between TSV and landing pad; and 3) TSV short-to-substrate. The technique employs transition delay test for TSV fault detection. Fault recovery is achieved by employing redundant TSVs and rerouting signals to fault-free TSVs. This technique is efficient because it requires a small (2× number of TSVs per group) number of clock cycles for fault detection and recovery. Synthesis results using 130-nm design library show that 100% repair capability can be achieved with low area overhead (4% for the best case). Yi Zhao 0001, S. Saqib Khursheed, Bashir M. Al-Hashimi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | High Quality Testing of Grid Style Power GatingabstractThis paper shows that existing delay-based testing techniques for power gating exhibit fault coverage loss due to unconsidered delays introduced by the structure of the virtual voltage power-distribution-network (VPDN). To restore this loss, which could reach up to 70.3% on stuck-open faults, we propose a design-for-testability (DFT) logic that considers the impact of VPDN on fault coverage in order to constitute the proper interface between the VPDN and the DFT. The proposed logic can be easily implemented on-top of existing DFT solutions and its overhead is optimized by an algorithm that offers trade-off flexibility between test-application-time and hardware overhead. Through physical layout SPICE simulations, we show complete fault coverage recovery on stuck-open faults and 43.2% test-application-time improvement compared to a previously proposed DFT technique. To the best of our knowledge, this paper presents the first analysis of the VPDN impact on test quality. Vasileios Tenentes, S. Saqib Khursheed, Bashir M. Al-Hashimi, Shida Zhong, Sheng Yang 0003 |
ATS | 3 |
| 2014 | Reinforcement Learning-Based Inter- and Intra-Application Thermal Optimization for Lifetime Improvement of Multicore SystemsabstractThe thermal profile of multicore systems vary both within an application's execution (intra) and also when the system switches from one application to another (inter). In this paper, we propose an adaptive thermal management approach to improve the lifetime reliability of multicore systems by considering both inter- and intra-application thermal variations. Fundamental to this approach is a reinforcement learning algorithm, which learns the relationship between the mapping of threads to cores, the frequency of a core and its temperature (sampled from on-board thermal sensors). Action is provided by overriding the operating system's mapping decisions using affinity masks and dynamically changing CPU frequency using in-kernel governors. Lifetime improvement is achieved by controlling not only the peak and average temperatures but also thermal cycling, which is an emerging wear-out concern in modern systems. The proposed approach is validated experimentally using an Intel quad-core platform executing a diverse set of multimedia benchmarks. Results demonstrate that the proposed approach minimizes average temperature, peak temperature and thermal cycling, improving the mean-time-to-failure (MTTF) by an average of 2x for intra-application and 3x for inter-application scenarios when compared to existing thermal management techniques. Furthermore, the dynamic and static energy consumption are also reduced by an average 10% and 11% respectively. Anup Das 0001, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi, Akash Kumar 0001, Bharadwaj Veeravalli |
DAC | 4 |
| 2014 | Advanced SIMD: Extending the reach of contemporary SIMD architecturesabstractSIMD extensions have gained widespread acceptance in modern microprocessors as a way to exploit data-level parallelism in general-purpose cores. Popular SIMD architectures (e.g. Intel SSE/AVX) have evolved by adding support for wider registers and datapaths, and advanced features like indexed memory accesses, per-lane predication and inter-lane instructions, at the cost of additional silicon area and design complexity. This paper evaluates the performance impact of such advanced features on a set of workloads considered hard to vectorize for traditional SIMD architectures. Their sensitivity to the most relevant design parameters (e.g. register/datapath width and L1 data cache configuration) is quantified and discussed. We developed an ARMv7 NEON based ISA extension (ARGON), augmented a cycle accurate simulation framework for it, and derived a set of benchmarks from the Berkeley dwarfs. Our analyses demonstrate how ARGON can, depending on the structure of an algorithm, achieve speedups of 1.5x to 16x. Matthias Boettcher, Bashir M. Al-Hashimi, Mbou Eyole, Giacomo Gabrielli, Alastair Reid 0001 |
DATE | 2 |
| 2014 | Clock-modulation based watermark for protection of embedded processorsabstractThis paper presents a novel watermark generation technique for the protection of embedded processors. In previous work, a load circuit is used to generate detectable watermark patterns in the ASIC power supply. This approach leads to hardware area overheads. We propose removing the dedicated load circuit entirely, instead to compensate the reduced power consumption the watermark power pattern is emulated by reusing existing clock gated sequential logic as a zero-overhead load circuit and modulating the clock-gating enable signal with the watermark sequence. The proposed technique has been validated through experiments using two ASICs in 65nm CMOS, one with an ARM Cortex-M0 microcontroller and one with a Cortex-A5 microprocessor. Silicon measurement results verify the viability of the technique for embedded processors. Furthermore, the proposed clock modulation technique demonstrates a significant area reduction, without compromising the detection performance. In our experiments an area overhead reduction of 98% was achieved. Through reuse of existing logic and reduction of watermark hardware implementation costs, the proposed clock modulation technique offers an improved robustness against removal attacks. Jedrzej Kufel, Peter R. Wilson, Stephen Hill, Bashir M. Al-Hashimi, Paul N. Whatmough, James Myers |
DATE | 4 |
| 2014 | Energy-Aware Streaming Multimedia Adaptation: An Educational PerspectiveabstractAs mobile devices are getting more powerful and more affordable the use of online educational multimedia is also getting very prevalent. Limited battery power is nevertheless a major restricting factor as streaming multimedia drains battery power quickly. Many battery efficient multimedia adaptation techniques have been proposed that achieve battery efficiency by lowering presentation quality of entire multimedia. Adaptation is usually done without considering any impact on the information contents of multimedia. In this paper, based on the results of an experimental study, we argue that without considering any negative impact on information contents of multimedia the adaptation may negatively impact the learning process. Some portions of the multimedia that require a higher visual quality for conveying learning information may lose their learning effectiveness in the adapted lowered quality. We report results of our experimental study that indicate that different parts of the same learning multimedia do not have same minimum acceptable quality. This strengthens the position that power-saving adaptation techniques for educational multimedia must be developed that lower the quality of multimedia based on the needs of its individual fragments for successfully conveying learning information. Asim Jalal, Nicholas Gibbins, David E. Millard, Bashir M. Al-Hashimi, Naif R. Aljohani |
MoMM | 4 |
| 2014 | Learner-battery interaction in energy-aware learning multimedia systemsabstractUsing online multimedia content on mobile devices is a power hungry activity and drains battery power very quickly. This poses a big challenge in using mobile devices with limited battery power for learning purposes using online educational multimedia. Multimedia adaptation techniques have been developed that preserve battery power by lowering multimedia quality. These adaptation techniques do not provide users with any power-saving options and the adaptation is done automatically without involvement of users. In this paper, we propose a Learner-Battery Interaction model that suggests involving learners in the adaptation process. The idea is to provide learners with power-saving options and relevant feedback about the form of adapted multimedia in advance. This will help leaners in making informed power-saving decisions for adaptation. We implemented the model in a prototype system and conducted an evaluation in the form of a user study. Asim Jalal, Nicholas Gibbins, David E. Millard, Bashir M. Al-Hashimi, Naif R. Aljohani |
MUM | 4 |
| 2014 | Efficient Variation-Aware Delay Fault Simulation Methodology for Resistive Open and Bridge DefectsabstractSPICE offers an accurate method of simulating defect behavior. However, as demonstrated by recent research, it requires long computation time to simulate defect behavior, when considering process variation. To the best of our knowledge, there is no efficient variation-aware delay fault simulation methodology for resistive opens and resistive bridges. This paper presents a fast and accurate delay fault simulation methodology for these two defects. It is fast because it speeds up delay fault computation time by employing two efficient algorithms. The first algorithm is used to calculate transient gate output voltage, which is a key variable needed to compute delay faults. It employs a three step strategy to accelerate the computation of transient gate output voltage without compromising accuracy. The second algorithm uses bisection method to efficiently compute delay fault behavior of a fault-site. The proposed methodology (PM) has been incorporated in an open-source SPICE (NGSPICE) with BSIM4.7 transistor model. The methodology has been validated by comparing results with HSPICE using industrial designs from IWLS 2005 benchmarks and realistic fault-sites have been extracted from synthesized designs. Simulations are carried out using a 65-nm gate library (for illustration). When compared with HSPICE, results show that the PM is on average up to 52-times faster with ≤ 4.2% error in accuracy for resistive open and 39-times faster with ≤ 5.2% error in accuracy for resistive bridge defects. Shida Zhong, S. Saqib Khursheed, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Delay Test for Diagnosis of Power SwitchesabstractPower switches are used as a part of the power-gating technique to reduce the leakage power of a design. To the best of our knowledge, this is the first report in open literature to show a systematic diagnosis method for accurately diagnosing power switches. The proposed diagnosis method utilizes the recently proposed design-for-test solution for efficient testing of power switches in the presence of process, voltage, and temperature variation. It divides power switches into segments such that any faulty power switch is detectable, thereby achieving high diagnosis accuracy. The proposed diagnosis method is validated through SPICE simulation using a number of ISCAS benchmarks synthesized with a 90-nm gate library. Simulation results show that, when considering the influence of process variation, the worst case loss of accuracy is less than 4.5%; it is less than 12% when considering VT variations. S. Saqib Khursheed, Kan Shi, Bashir M. Al-Hashimi, Peter R. Wilson, Krishnendu Chakrabarty |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Active Mode Subclock Power GatingabstractThis paper presents a technique, called subclock power gating, for reducing leakage power during the active mode in low performance, energy-constrained applications. The proposed technique achieves power reduction through two mechanisms: 1) power gating the combinational logic within the clock period (subclock) and 2) reducing the virtual supply to less than Vth rather than shutting down completely as is the case in conventional power gating. To achieve this reduced voltage, a pair of nMOS and pMOS transistors are used at the head and foot of the power gated logic for symmetric virtual rail clamping of the power and ground supplies. The subclock power gating technique has been validated by incorporating it with an ARM Cortex-M0 microprocessor, which was fabricated in a 65-nm process. Two sets of experiments are done: the first experimentally validates the functionality of the proposed technique in the fabricated test chip and the second investigates the utility of the proposed technique in example applications. Measured results from the fabricated chip show 27% power saving during the active mode for an example wireless sensor node application when compared with the same microprocessor without subclock power gating. Jatin N. Mistry, James Myers, Bashir M. Al-Hashimi, David Flynn, John Biggs, Geoff V. Merrett |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | MALEC: a multiple access low energy cacheabstractThis paper addresses the dynamic energy consumption in L1 data cache interfaces of out-of-order superscalar processors. The proposed Multiple Access Low Energy Cache (MALEC) is based on the observation that consecutive memory references tend to access the same page. It exhibits a performance level similar to state of the art caches, but consumes approximately 48% less energy. This is achieved by deliberately restricting accesses to only 1 page per cycle, allowing the utilization of single-ported TLBs and cache banks, and simplified lookup structures of Store and Merge Buffers. To mitigate performance penalties it shares memory address translation results between multiple memory references, and shares data among loads to the same cache line. In addition, it uses a Page-Based Way Determination scheme that holds way information of recently accessed cache lines in small storage structures called way tables that are closely coupled to TLB lookups and are able to simultaneously service all accesses to a particular page. Moreover, it removes the need for redundant tag-array accesses, usually required to confirm way predictions. For the analyzed workloads, MALEC achieves average energy savings of 48% in the L1 data memory subsystem over a high performance cache interface that supports up to 2 loads and 1 store in parallel. Comparing MALEC and the high performance interface against a low power configuration limited to only 1 load or 1 store per cycle reveals 14% and 15% performance gain requiring 22% less and 48% more energy, respectively. Furthermore, Page-Based Way Determination exhibits coverage of 94%, which is a 16% improvement over the originally proposed line-based way determination. Matthias Boettcher, Giacomo Gabrielli, Bashir M. Al-Hashimi, Danny Kershaw |
DATE | 3 |
| 2013 | DoE-based performance optimization of energy management in sensor nodes powered by tunable energy-harvesters
Tom J. Kazmierski, Leran Wang, Bashir M. Al-Hashimi, Geoff V. Merrett |
DATE | 3 |
| 2013 | A survey of multi-source energy harvesting systemsabstractEnergy harvesting allows low-power embedded devices to be powered from naturally-ocurring or unwanted environmental energy (e.g. light, vibration, or temperature difference). While a number of systems incorporating energy harvesters are now available commercially, they are specific to certain types of energy source. Energy availability can be a temporal as well as spatial effect. To address this issue, ‘hybrid’ energy harvesting systems combine multiple harvesters on the same platform, but the design of these systems is not straight-forward. This paper surveys their design, including trade-offs affecting their efficiency, applicability, and ease of deployment. This survey, and the taxonomy of multi-source energy harvesting systems that it presents, will be of benefit to designers of future systems. Furthermore, we identify and comment upon the current and future research directions in this field. Alex S. Weddell, Michele Magno, Geoff V. Merrett, Davide Brunelli, Bashir M. Al-Hashimi, Luca Benini |
DATE | 5 |
| 2013 | Energy-Aware Adaptation of Educational Multimedia in Mobile LearningabstractAs a result of tremendous enhancements in the capabilities of mobile devices and availability of higher data rate mobile internet, the use of online multimedia learning resources on mobile devices is increasingly becoming popular. Limited Battery Power of mobile devices, however, is still one big challenge in Mobile Learning. High Quality multimedia learning resources are power hungry and if used on mobile devices drain battery power rapidly limiting learning opportunities on the move. Lack of significant improvements in battery capacities has resulted in significant interest in battery power saving techniques. Existing power-saving streaming multimedia adaptation techniques tend to extend battery life by reducing quality of multimedia making them susceptible to information loss. This loss may affect the learning content efficacy and jeopardizes the learning process. To the best of our knowledge, no previous work has considered the learning content efficacy in multimedia streaming adaptation mechanism. In this paper, we present MoBELearn system, which is a prototype implementation of our proposed Content Aware Power Saving Educational Multimedia Adaptation (CAPS-EMA) approach. We demonstrate battery efficiency in educational multimedia streaming while keeping the adapted resource suitable for learning. We also describe our semantic metamodel for educational multimedia resource that support our energy efficient adaptation technique. Asim Jalal, Nicholas Gibbins, David E. Millard, Bashir M. Al-Hashimi, Naif R. Aljohani |
MoMM | 4 |
| 2013 | Analysis of voltage- and clock-scaling-induced timing errors in stochastic LDPC decodersabstractLow Density Parity Check (LDPC) decoders have an inherent capability of correcting the transmission errors that occur, when communicating over a hostile wireless channel. This capability allows LDPC-coded schemes to employ lower transmission energies than uncoded schemes, at the cost of introducing a significant processing energy consumption during LDPC decoding. Traditional energy-reduction techniques, such as voltage and clock scaling can be employed for reducing the LDPC decoder's energy consumption. However, these techniques may induce timing errors, which can degrade the LDPC decoder's error correction capability. Our previous work has demonstrated that in contrast to other types of LDPC decoders, stochastic decoders have an inherent tolerance to timing errors, allowing them to maintain a high error correction capability in clockscaling scenarios. In this paper, we investigate this timing error tolerance in voltage-scaling scenarios, by extending our previous model of timing errors using extensive SPICE simulations. Furthermore, we use these SPICE simulations to characterize the processing energy consumption of stochastic LDPC decoders for the first time. We demonstrate that a modified stochastic LDPC decoder can operate at 0.8 V and a clock period of 915.11 ps, while maintaining the error correction capability of a conventional stochastic decoder operating at 1 V and a clock period of 1019.2 ps, offering a 36.7% reduction in processing energy consumption. Isaac Perez-Andrade, Robert G. Maunder, Bashir M. Al-Hashimi, Lajos Hanzo |
WCNC | 4 |
| 2013 | A Low-Complexity Turbo Decoder Architecture for Energy-Efficient Wireless Sensor NetworksabstractTurbo codes have recently been considered for energy-constrained wireless communication applications, since they facilitate a low transmission energy consumption. However, in order to reduce the overall energy consumption, lookup table-log-BCJR (LUT-Log-BCJR) architectures having a low processing energy consumption are required. In this paper, we decompose the LUT-Log-BCJR architecture into its most fundamental add compare select (ACS) operations and perform them using a novel low-complexity ACS unit. We demonstrate that our architecture employs an order of magnitude fewer gates than the most recent LUT-Log-BCJR architectures, facilitating a 71% energy consumption reduction. Compared to state-of-the-art maximum logarithmic Bahl-Cocke-Jelinek-Raviv implementations, our approach facilitates a 10% reduction in the overall energy consumption at ranges above 58 m. Liang Li 0013, Robert G. Maunder, Bashir M. Al-Hashimi, Lajos Hanzo |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Accelerators and emulators: Can they become the platform of choice for hardware verification?
Bashir M. Al-Hashimi, Ronny Morad |
DATE | 1 |
| 2012 | Response-surface-based design space exploration and optimisation of wireless sensor nodes with tunable energy harvestersabstractIn an energy harvester powered wireless sensor node, the energy harvester is often the only energy source, therefore it is crucial to configure the microcontroller and the sensor node so that the harvested energy is used efficiently. This paper presents a response surface model (RSM) based design space exploration and optimisation of a complete wireless sensor node system. In our work the power consumption models of the microcontroller and the sensor node are defined based on their digital operations so that the parameters of the digital algorithms can be optimised to achieve the best energy efficiency. In the proposed technique, SystemC-A is used to model the system's analogue components as well as the digital control algorithms implemented in the microcontroller and the sensor node. A series of simulations are carried out and a response surface model is constructed from the simulation results. The RSM is then optimised using MATLAB's optimisation toolbox and the results show that the optimised system configuration can double the total number of wireless transmissions with fixed amount of harvested energy. The great improvement in the system performance validates the efficiency of our technique. Leran Wang, Tom J. Kazmierski, Bashir M. Al-Hashimi, Mansour Aloufi, Joseph Wenninger |
DATE | 3 |
| 2012 | Low-Energy Standby-Sparing for Hard Real-Time SystemsabstractTime-redundancy techniques are commonly used in real-time systems to achieve fault tolerance without incurring high energy overhead. However, reliability requirements of hard real-time systems that are used in safety-critical applications are so stringent that time-redundancy techniques are sometimes unable to achieve them. Standby sparing as a hardware-redundancy technique can be used to meet high reliability requirements of safety-critical applications. However, conventional standby-sparing techniques are not suitable for low-energy hard real-time systems as they either impose considerable energy overheads or are not proper for hard timing constraints. In this paper we provide a technique to use standby sparing for hard real-time systems with limited energy budgets. The principal contribution of this paper is an online energy-management technique which is specifically developed for standby-sparing systems that are used in hard real-time applications. This technique operates at runtime and exploits dynamic slacks to reduce the energy consumption while guaranteeing hard deadlines. We compared the low-energy standby-sparing (LESS) system with a low-energy time-redundancy system (from a previous work). The results show that for relaxed time constraints, the LESS system is more reliable and provides about 26% energy saving as compared to the time-redundancy system. For tight deadlines when the time-redundancy system is not sufficiently reliable (for safety-critical application), the LESS system preserves its reliability but with about 49% more energy consumption. Alireza Ejlali, Bashir M. Al-Hashimi, Petru Eles |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | An Explicit Linearized State-Space Technique for Accelerated Simulation of Electromagnetic Vibration Energy HarvestersabstractVibration energy harvesting systems pose significant modeling and design challenges due to their mixed-technology nature, extremely low levels of available energy and disparate time scales between different parts of a complete harvester. An energy harvester is a complex system of tightly coupled components modeled in the mechanical, magnetic, as well as electrical analog and digital domains. Currently available design tools are inadequate for simulating such systems due to prohibitive CPU times. This paper proposes a new technique to accelerate simulations of complete vibration energy harvesters by approximately two orders of magnitude. The proposed technique is to linearize the state equations of the system's analog components to obtain a fast estimate of the maximum step-size to guarantee the numerical stability of explicit integration based on the Adams-Bashforth formula. We show that the energy harvester's analog electronics can be efficiently and reliably simulated in this way with CPU times two orders of magnitude lower than those obtained from two state-of-the-art tools, VHDL-AMS and SystemC-A. As a case study, a practical, complex microgenerator with magnetic tuning and two types of power-processing circuits have been simulated using the proposed technique and verified experimentally. Tom J. Kazmierski, Leran Wang, Bashir M. Al-Hashimi, Geoff V. Merrett |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Cost-Effective TSV Grouping for Yield Improvement of 3D-ICsabstractThree-dimensional Integrated Circuits (3D-ICs) vertically stack multiple silicon dies to reduce overall wire length, power consumption, and allow integration of heterogeneous technologies. Through-silicon-vias (TSVs) which act as vertical links between layers pose challenges for 3D integration design. TSV defects can happen in fabrication process and bonding stage, which can reduce the yield and increase the cost. Recent work proposed the employment of redundant TSVs to improve the yield of 3D-ICs. This paper presents a redundant TSVs grouping technique, which partitions regular and redundant TSVs into groups. For each group, a set of multiplexers are used to select good signal paths away from defective TSVs. We investigate the impact of grouping ratio (regular-to-redundant TSVs in one group) on trade-off between yield and hardware overhead. We also show probabilistic models for yield analysis under the influence of independent and clustering defect distributions. Simulation results show that for a given number of TSVs and TSV failure rate, careful selection of grouping ratios lead to achieving 100% yield at minimal hardware cost (number of multiplexers and redundant TSVs) in comparison to a design that does not exploit TSV grouping ratios. Yi Zhao 0001, S. Saqib Khursheed, Bashir M. Al-Hashimi |
Asian Test Symposium | 3 |
| 2011 | Analysis of Resistive Bridge Defect Delay Behavior in the Presence of Process VariationabstractRecent research has shown that tests generated without taking process variation into account may lead to loss of test quality. Using transition delay test, this paper analyzes the behavior of resistive bridge defect under the influence of process variation. The effect of process variation is incorporated by using three transistor parameters: gate length (L), threshold voltage (Vth) and effective mobility (μeff), where each follows Gaussian distribution. Through HSPICE simulations using a 65-nm gate library, this paper brings the following two contributions: firstly, it analyzes the delay behavior of bridge defect using all three transition delay classes to determine the most effective class of transition test that achieves maximum coverage in the presence of process variation. Secondly, recent research has shown that low voltage testing improves detectability of bridge fault, this work compares bridge resistance coverage using logic test and delay test at multiple voltage settings to identify the best voltage setting and test type for detecting resistive bridge defects. Shida Zhong, S. Saqib Khursheed, Bashir M. Al-Hashimi, Sudhakar M. Reddy, Krishnendu Chakrabarty |
Asian Test Symposium | 3 |
| 2011 | Sub-clock power-gating technique for minimising leakage power during active modeabstractThis paper presents a new technique, called sub-clock power gating, for reducing leakage power in digital circuits. The proposed technique works concurrently with voltage and frequency scaling and power reduction is achieved by power gating within the clock cycle during active mode unlike traditional power gating which is applied during idle mode. The proposed technique can be implemented using standard EDA tools with simple modifications to the standard power gating design flow. Using a 90nm technology library, the technique is validated using two case studies: 16-bit parallel multiplier and ARM Cortex-M0™ microprocessor, provided by our industrial project partner. Compared to designs without sub-clock power gating, in a given power budget, we show that leakage power saved allows 45× and 2.5× improvements in energy efficiency in the case of multiplier and microprocessor, respectively. Jatin N. Mistry, Bashir M. Al-Hashimi, David Flynn, Stephen Hill |
DATE | 2 |
| 2011 | Accelerated simulation of tunable vibration energy harvesting systems using a linearised state-space techniqueabstractThis paper proposes a linearised state-space technique to accelerate the simulation of tunable vibration energy harvesting systems by at least two orders of magnitude. The paper provides evidence that currently available simulation tools are inadequate for simulating complete energy harvesting systems where prohibitive CPU times are encountered due to disparate time scales. In the proposed technique, the model of a complete mixed-technology energy harvesting system is divided into component blocks whose mechanical and analogue electrical parts are modelled by local state equations and terminal variables while the digital electrical part is modelled as a digital process. Unlike existing simulation tools that use Newton-Raphson method, the proposed technique uses explicit integration such as Adams-Bashforth method to solve the state equations of the complete energy harvester model in short simulation time. Experimental measurements of a practical tunable energy harvester have been carried out to validate the proposed technique. Leran Wang, Tom J. Kazmierski, Bashir M. Al-Hashimi, Alex S. Weddell, Geoff V. Merrett, Ivo Netali Ayala-Garcia |
DATE | 3 |
| 2011 | Ultra low-power photovoltaic MPPT technique for indoor and outdoor wireless sensor nodesabstractPhotovoltaic (PV) energy harvesting is commonly used to power wireless sensor nodes. To optimise harvesting efficiency, maximum power point tracking (MPPT) techniques are often used. Recently-reported techniques focus solely on outdoor applications, being too power-hungry for use under indoor lighting. Additionally, some techniques have required light sensors (or pilot cells) to control their operating point. This paper describes an ultra low-power MPPT technique which is based on a novel system design and sample-and-hold arrangement, which enables MPPT across the range of light intensities found indoors and outdoors and is capable of cold-starting. The proposed sample-and-hold based technique has been validated through a prototype system. Its performance compares favourably against state-of-the-art systems, and does not require an additional pilot cell or photodiode. This represents an important contribution, in particular for sensors which may be exposed to different types of lighting (such as body-worn or mobile sensors). Alex S. Weddell, Geoff V. Merrett, Bashir M. Al-Hashimi |
DATE | 3 |
| 2011 | Improved DFT for Testing Power SwitchesabstractPower switches are used as part of power-gating technique to reduce leakage power of a design. To the best of our knowledge this is the first study that analyzes recently proposed DFT solutions for testing power switches through SPICE simulations on a number of ISCAS benchmarks and presents the following contributions. It provides evidence of long discharge time when power switches are turned-off, when testing power switches using available DFT solutions. This may either lead to false test (false-fail or false-pass) or long test time. This problem is addressed through a simple and effective DFT solution to reduce the discharge time. The proposed DFT solution has been validated through SPICE simulation and shows an improvement in discharge time of at least 28-times, based on a number of ISCAS benchmarks synthesized with a 90-nm gate library. S. Saqib Khursheed, Sheng Yang 0003, Bashir M. Al-Hashimi, David Flynn |
ETS | 3 |
| 2011 | Simplified logic design methodology for fuzzy membership function based robust detection of maternal modulus maxima location: A low complexity Fetal ECG extraction architecture for mobile health monitoring systemsabstractThis paper proposes a simplified logic design methodology for the fuzzy membership function used for robust and reliable detection of modulus-maxima locations in wavelet domain for fetal ECG extraction from the abdominal composite ECG signal. This simplification is achieved by exploiting the inherent time-position information of the wavelet coefficients decomposed at different resolution levels. Subsequently, a low complexity VLSI architecture for Fetal ECG extraction is presented which is designed using the recently proposed memory- efficient, multiplierless Discrete Wavelet Transform method. The generic memory model within this architecture will provide the flexibility to configure the on-chip memory with any type of orthonormal wavelets suitable for different applications. Total synthesized cell area of the proposed architecture is 14.2 mm2and power consumption is 101.5 μW at 1.2 V @ 1 MHz frequency using 0.13 μm standard cell technology. The proposed architecture is targeted for the personalized health monitoring applications within a mobile home-care medical device in the resource constrained environment. Amit Acharyya, Koushik Maharatna, Bashir M. Al-Hashimi, Hasitha Tudugalle |
ISCAS | 3 |
| 2011 | Investigation into voltage and process variation-aware manufacturing testabstractTraditional test methods that use abstract fault models potentially results in low defect coverage and test escapes for ICs with multiple supply voltage (Vdd) settings for adaptive power management, and in the presence of process variation. In this paper, we address two important defect types, resistive bridge defects and full open defects, and present foundational work on variation-aware test methods. To test ICs with multiple Vdds, Multi-Vdd Test Generation (MVTG) produces Vdd-specific test sets, such that tests are applied using the most effective Vdd. For Process Variation-aware Test Generation (PVTG), we target the most significant test escapes, guided by a novel process variation-aware metric for test quality, called test robustness. We implemented our test methods, and integrated them into a flow of commercial EDA tools. Experimental results on benchmark designs and realistic defects, extracted from layout, show that our test methods achieve high defect coverage while keeping the test sets size low. This serves as proof-of-concept for variation-aware test. Urban Ingelsson, Bashir M. Al-Hashimi |
ITC | 2 |
| 2011 | Reliable State Retention-Based Embedded Processors Through Monitoring and RecoveryabstractState retention power gating and voltage-scaled state retention are two effective design techniques, commonly employed in embedded processors, for reducing idle circuit leakage power. This paper presents a methodology for improving the reliability of embedded processors in the presence of power supply noise and soft errors. A key feature of the method is low cost, which is achieved through reuse of the scan chain for state monitoring, and it is effective because it can correct single and multiple bit errors through hardware and software, respectively. To validate the methodology, ARM® Cortex™-M0 embedded microprocessor (provided by our industrial project partner) is implemented in field-programmable gate array and further synthesized using 65-nm technology to quantify the cost in terms of area, latency, and energy. It is shown that the proposed methodology has a small area overhead (8.6%) with less than 4% worst-case increase in critical path and is capable of detecting and correcting both single bit and multibit errors for a wide range of fault rates. Sheng Yang 0003, S. Saqib Khursheed, Bashir M. Al-Hashimi, David Flynn, Sachin Idgunji |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | A Fast and Accurate Process Variation-Aware Modeling Technique for Resistive Bridge DefectsabstractRecent research has shown that tests generated without taking process variation into account may lead to loss of test quality. At present, there is no efficient device-level modeling technique that models the effect of process variation on resistive bridge defects. This paper presents a fast and accurate technique to achieve this, including modeling the effect of voltage and temperature variation using the BSIM4 transistor model. To speed up the computation time and without compromising simulation accuracy (achieved through BSIM4), two efficient voltage approximation algorithms are proposed for calculating logic threshold of driven gates and voltages on bridged lines of a fault-site to calculate bridge critical resistance. Experiments are conducted on a 65 nm gate library (for illustration purposes), and results show that on average the proposed modeling technique is more than 53 times faster and in the worst case, error in bridge critical resistance is 2.64% when compared with HSPICE. Shida Zhong, S. Saqib Khursheed, Bashir M. Al-Hashimi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2010 | Evaluation and design exploration of solar harvested-energy prediction algorithmabstractTo respond to variations in solar energy, harvested-energy prediction is essential to harvested-energy management approaches. The effectiveness of such approaches is dependent on both the achievable accuracy and computation overhead of prediction algorithm implementation. This paper presents detailed evaluation of a recently reported solar energy prediction algorithm to determine empirical bounds on achievable accuracy and implementation overhead using an effective error evaluation technique. We evaluate the algorithm performance over varying prediction horizons and propose guidelines for algorithm parameter selection across different real solar energy profiles to simplify implementation. The prediction algorithm computation overhead is measured on actual hardware to demonstrate prediction accuracy-cost trade-off. Finally, we motivate the basis for dynamic prediction algorithm and show that more than 10% increase in prediction accuracy can be achieved compared to static algorithm. Mustafa Imran Ali, Bashir M. Al-Hashimi, Joaquín Recas, David Atienza 0001 |
DATE | 2 |
| 2010 | Soft error-aware design optimization of low power and time-constrained embedded systemsabstractIn this paper, we examine the impact of application task mapping on the reliability of MPSoC in the presence of single-event upsets (SEUs). We propose a novel soft error-aware design optimization using joint power minimization with voltage scaling and reliability improvement through application task mapping. The aim is to minimize the number of SEUs experienced by the MPSoC for a suitably identified voltage scaling of the system processing cores such that the power is reduced and the specified real-time constraint is met.We evaluate the effectiveness of the proposed optimization technique using an MPEG-2 decoder and random task graphs. We show that for an MPEG-2 decoder with four processing cores, our optimization technique produces a design that experiences 38% less SEUs than soft error-unaware design optimization for a soft error rate of 10-9, while consuMPEG-2 decoderming 9% less power and meeting a given real-time constraint. Furthermore, we investigate the impact of architecture allocation (varying the number of MPSoC cores) on the power consumption and SEUs experienced. We show that for an MPSoC with six processing cores and a given real-time constraint, the proposed technique experiences upto 7% less SEUs compared to soft error-unaware optimization, while consuming only 3% more power. Rishad A. Shafik, Bashir M. Al-Hashimi, Krishnendu Chakrabarty |
DATE | 2 |
| 2010 | Scan based methodology for reliable state retention power gating designsabstractPower gating is an effective technique for reducing leakage power which involves powering off idle circuits through power switches, but those power-gated circuits which need to retain their states store their data in state retention registers. When power-gated circuits are switched from sleep to active mode, sudden rush of current has the potential of corrupting the stored data in the state retention registers which could be a reliability problem. This paper presents a methodology for improving the reliability of power-gated designs by protecting the integrity of state retention registers through state monitoring and correction. This is achieved by scan chain data encoding and decoding. The methodology is compatible with EDA tools design and power gating control flows. A detailed analysis of the proposed methodology's capability in detecting and correcting errors is given including the area overhead and energy consumption of the protection circuitry. The methodology is validate using FPGA and show that it is possible to correct all single errors with Hamming code and detect all multiple errors with CRC-16 code. To the best of our knowledge this is the first study in the area of reliable power gating designs through state monitoring and correction. Sheng Yang 0003, Bashir M. Al-Hashimi, David Flynn, S. Saqib Khursheed |
DATE | 2 |
| 2010 | Modeling the impact of process variation on resistive bridge defectsabstractRecent research has shown that tests generated without taking process variation into account may lead to loss of test quality. At present there is no efficient device-level modeling technique that models the effect of process variation on resistive bridges. This paper presents a fast and accurate technique to model the effect of process variation on resistive bridge defects. The proposed model is implemented in two stages: firstly, it employs an accurate transistor model (BSIM4) to calculate the critical resistance of a bridge; secondly, the effect of process variation is incorporated in this model by using three transistor parameters: gate length (L), threshold voltage (Vth) and effective mobility (μeff), where each follow Gaussian distribution. Experiments are conducted on a 65-nm gate library (for illustration purposes), and results show that on average the proposed modeling technique is more than 7 times faster and in the worst case, error in bridge critical resistance is 0.8% when compared with HSPICE. S. Saqib Khursheed, Shida Zhong, Robert C. Aitken, Bashir M. Al-Hashimi, Sandip Kundu |
ITC | 4 |
| 2010 | Design of Fixed-Point Processing Based Turbo Codes Using Extrinsic Information Transfer ChartsabstractThe operand-width specifications in fixed-point hardware implementations of turbo code decoders is an important design issue, since this governs the trade-off between the decoder's performance and its complexity, cost, area and energy consumption. The investigation of this issue would be extremely time-consuming in the conventional approach, which relies upon Monte-Carlo simulation based Bit Error Ratio (BER) analysis. In this paper, we propose a generic design method, which uses EXtrinsic Information Transfer (EXIT) chart analysis to simplify this design process. Our method is not only an order of magnitude faster than the conventional Monte-Carlo simulation based approach, but also offers deeper insights into why performance degradations are imposed by insufficient operand-width specifications. The benefits of our generic method are demonstrated in the context of a turbo decoder, allowing accurate specifications to be obtained and compared to those suggested by previous works. Liang Li 0013, Robert G. Maunder, Bashir M. Al-Hashimi, Lajos Hanzo |
VTC Fall | 3 |
| 2010 | Gate-Sizing-Based Single Vdd Test for Bridge Defects in Multivoltage DesignsabstractThe use of multiple voltage settings for dynamic power management is an effective design technique. Recent research has shown that testing for resistive bridging faults in such designs requires more than one voltage setting for 100% fault coverage; however, switching between several supply voltage settings has a detrimental impact on the overall cost of test. This paper proposes an effective gate sizing technique for reducing test cost of multi-Vdddesigns with bridge defects. Using synthesized ISCAS and ITC benchmarks and a parametric fault model, experimental results show that for all the circuits, the proposed technique achieves singleVddtest, without affecting the fault coverage of the original test. In addition, the proposed technique performs better in terms of timing, area, and power than the recently proposed test point insertion technique. This is the first reported work that achieves singleVddtest for resistive bridge defects, without compromising fault coverage in multi-Vdddesigns. S. Saqib Khursheed, Bashir M. Al-Hashimi, Krishnendu Chakrabarty, Peter Harrod |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Performability/Energy Tradeoff in Error-Control Schemes for On-Chip NetworksabstractHigh reliability against noise, high performance, and low energy consumption are key objectives in the design of on-chip networks. Recently some researchers have considered the impact of various error-control schemes on these objectives and on the tradeoff between them. In all these works performance and reliability are measured separately. However, we will argue in this paper that the use of error-control schemes in on-chip networks results indegradable systems, hence, performance and reliability must be measured jointly using a unified measure, i.e.,performability. Based on the traditional concept of performability, we provide a definition for the ¿Interconnect Performability¿. Analytical models are developed for interconnect performability and expected energy consumption. A detailed comparative analysis of the error-control schemes using the performability analytical models and SPICE simulations is provided taking into consideration voltage swing variations (used to reduce interconnect energy consumption) and variations in wire length. Furthermore, the impact of noise power and time constraint on the effectiveness of error-control schemes are analyzed. Alireza Ejlali, Bashir M. Al-Hashimi, Paul M. Rosinger, Seyed Ghassem Miremadi, Luca Benini |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Selective state retention design using symbolic simulationabstractAddressing both standby and active power is a major challenge in developing system-on-chip designs for battery-powered products. Powering off sections of logic or memories loses internal register and RAM states so designers have to weigh up the benefits and costs of implementing state retention on some or all of the power gated subsystems where state recovery has significant real-time or energy cost, compared to resetting the subsystem and re-acquiring state from scratch. Library IP and EDA tools can support state retention in hardware synthesized from standard RTL, but due to the silicon area costs there is strong interest in only retaining certain selective state for example the ldquoarchitectural staterdquo of a CPU to implement sleep modes. Currently there is no known rigourous technique for checking the integrity of selective state retention, and this is due to the complexity of checking that the correctness of the design is not compromised in any way. The complexity is exacerbated due to the interaction between the retained and the non-retained state, and exhaustive simulation rapidly becomes infeasible. This paper presents a case study based on symbolic simulation for assisting the designers to design and implement selective retention correctly. The main finding of our study is that the programmer visible state or the architectural state of the CPU needs to be implemented using retention registers whilst other micro-architectural enhancements such as pipeline registers, TLBs and caches can be implemented using normal registers without retention. This has a profound impact on power and area savings for chip design. By selectively retaining the state of the programmer's ldquoarchitecturalrdquo model and not the increasing proportion of extra state, one can incorporate energy-efficient sleep modes. To the best of our knowledge this is the first study in the area of rigourous design and implementation of selective state retention. Ashish Darbari, Bashir M. Al-Hashimi, David Flynn, John Biggs |
DATE | 2 |
| 2009 | Test cost reduction for multiple-voltage designs with bridge defects through Gate-SizingabstractMultiple-voltage is an effective dynamic power reduction design technique. Recent research has shown that testing for resistive bridging faults in such designs requires more than one voltage setting for 100% defect coverage; however switching between several supply voltage settings has a detrimental impact on the overall cost of test. This paper proposes an effective Gate Sizing technique for reducing test cost of multi-Vdd designs with bridge defects. Using synthesized ISCAS benchmarks and a parametric fault model, experimental results show that for all the circuits, the proposed technique achieves 100% defect coverage at a single Vdd setting; in addition it has a lower overhead than the recently proposed test point insertion technique in terms of timing, area and power. S. Saqib Khursheed, Bashir M. Al-Hashimi, Peter Harrod |
DATE | 2 |
| 2009 | Variation resilient adaptive controller for subthreshold circuitsabstractSubthreshold logic is showing good promise as a viable ultra-low-power circuit design technique for power-limited applications. For this design technique to gain widespread adoption, one of the most pressing concerns is how to improve the robustness of subthreshold logic to process and temperature variations. We propose a variation resilient adaptive controller for subthreshold circuits with the following novel features: new sensor based on time-to-digital converter for capturing the variations accurately as digital signatures, and an all-digital DC-DC converter incorporating the sensor capable of generating an operating operating Vddfrom 0 V to 1.2 V with a resolution of 18.75 mV, suitable for subthreshold circuit operation. The benefits of the proposed controller is reflected with energy improvement of up to 55% compared to when no controller is employed. The detailed implementation and validation of the proposed controller is discussed. Biswajit Mishra, Bashir M. Al-Hashimi, Mark Zwolinski |
DATE | 2 |
| 2009 | An automated design flow for vibration-based energy harvester systemsabstractThis paper proposes, for the first time, an automated energy harvester design flow which is based on a single HDL software platform that can be used to model, simulate, configure and optimise energy harvester systems. A demonstrator prototype incorporating an electromagnetic mechanical-vibration-based micro-generator and a limited number of library models has been developed and a design case study has been carried out. Experimental measurements have validated the simulation results which show that the outcome from the design flow can improve the energy harvesting efficiency by 75%. Leran Wang, Tom J. Kazmierski, Bashir M. Al-Hashimi, Stephen P. Beeby, Dibin Zhu |
DATE | 3 |
| 2009 | HSPICE implementation of a numerically efficient model of CNT transistor
Tom J. Kazmierski, Dafeng Zhou, Bashir M. Al-Hashimi |
FDL | 3 |
| 2009 | Energy-Aware Simulation for Wireless Sensor NetworksabstractEnergy-aware sensor nodes are usually tightly energy-constrained, execute energy-efficient algorithms, have the ability to interrogate and control the devices used for storing and consuming energy, and often feature one or more sources of energy harvesting. Due to the cost, time and expertise required to deploy a wireless sensor network (WSN), simulation is currently the most widely adopted evaluation method. Network simulation is well established for mobile ad hoc networks, using simulators such as the popular ns2. However, the differing characteristics and performance criteria of WSNs introduce additional simulation requirements, and this has resulted in a number of simulators and simulator extensions developed specifically for this purpose. This paper investigates the suitability of a number of state-of-the-art simulators for evaluating energy-aware WSNs, and subsequently proposes a novel structure for simulating energy-aware WSNs. The proposed structure provides diverse, flexible and extensible hardware and environment models, and integrates a structured architecture for embedded software to enhance the design of energy-aware sensor nodes. To illustrate an implementation of the structure, details of - and observations obtained using - an in-house simulator (WSNsim) are presented. Geoff V. Merrett, Neil M. White, Nick R. Harris, Bashir M. Al-Hashimi |
SECON | 4 |
| 2009 | Process Variation-Aware Test for Resistive BridgesabstractThis paper analyzes the behavior of resistive bridging faults under process variation and shows that process variation has a detrimental impact on test quality in the form of test escapes. To quantify this impact, a novel metric called test robustness is proposed and to mitigate test escapes, a new process variation-aware test generation method is presented. The method exploits the observation that logic faults that have high probability of occurrence and correspond to significant amounts of undetected bridge resistance have a high impact on test robustness and therefore should be targeted by test generation. Using synthesized International Symposium on Circuits and Systems benchmarks with realistic bridge locations, results show that for all the benchmarks, the method achieves better results (less test escapes) than tests generated without consideration of process variation. Urban Ingelsson, Bashir M. Al-Hashimi, S. Saqib Khursheed, Sudhakar M. Reddy, Peter Harrod |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Diagnosis of Multiple-Voltage Design With Bridge DefectabstractMultiple voltage is an effective dynamic power reduction design technique commonly used in low-power ICs. To the best of our knowledge, there is no reported work for diagnosing multiple-voltage enabled ICs, and the aim of this paper is to propose a method for diagnosing bridge defects in such ICs. By using synthesized ISCAS benchmarks, with realistic extracted bridges and a parametric fault model, this paper investigates the impact of varying supply voltage on the accuracy of diagnosis and demonstrates how the additional voltage settings can be leveraged to improve the diagnosis resolution through a novel multivoltage diagnosis algorithm. In addition, it also identifies the most useful voltage settings to reduce diagnosis cost by eliminating tests at certain voltage setting using the proposed multivoltage diagnosis approach, thereby achieving high diagnosis accuracy at reduced cost. S. Saqib Khursheed, Bashir M. Al-Hashimi, Sudhakar M. Reddy, Peter Harrod |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Transistor-level based defect tolerance for reliable nanoelectronicsabstractNanodevices based circuit design will be based on the acceptance that a high percentage of devices in the design will be defective. In this work, we investigate a defect tolerant technique that adds redundancy at the transistor level and provides built-in immunity to permanent defects (stuck-open, stuck-short and bridges). The proposed technique is based on replacing each transistor by quadded-transistor structure that guarantees defect tolerance of all single defects and a large number of multiple defects as validated by theoretical analysis and simulation. As demonstrated by extensive simulation results using ISCAS 85 and 89 benchmark circuits, the investigated technique achieves significantly higher defect tolerance than recently reported nanoelectronics defect-tolerant techniques (even with up to 4 to 5 times more transistor defect probability) and at reduced area overhead. Aiman H. El-Maleh, Bashir M. Al-Hashimi, Aissa Melouki |
AICCSA | 2 |
| 2008 | Variation Aware Analysis of Bridging Fault TestingabstractThis paper investigates the impact of process variation on test quality with regard to resistive bridging faults. The input logic threshold voltage and gate drives strength parameters are analyzed regarding their process variation induced influence on test quality. The impact of process variation on test quality is studied in terms of test escapes and measured by a robustness metric. It is shown that some bridges are sensitive to process variation in terms of logic behavior, but such variation does not necessarily compromise test quality if the test has high robustness. Experimental results of Monte-Carlo simulation based on recent process variation statistics are presented for ISCAS85 and -89 benchmark circuits, using a 45 nm gate library and realistic bridges. The results show that tests generated without consideration of process variation are inadequate in terms of test quality, particularly for small test sets. On the other hand, larger test sets detect more of the logic faults introduced by process variation and have higher test quality. Urban Ingelsson, Bashir M. Al-Hashimi, Peter Harrod |
ATS | 2 |
| 2008 | Efficient circuit-level modelling of ballistic CNT using piecewise non-linear approximation of mobile charge densityabstractThis paper presents a new carbon nanotube transistor (CNT) modelling technique which is based on an efficient numerical piece-wise non-linear approximation of the non-equilibrium mobile charge density. The technique facilitates the solution of the self-consistent voltage equation in a carbon nanotube such that the CNT drain-source current evaluation is accelerated by more than three orders of magnitude while maintaining high modelling accuracy. The model is currently limited to ballistic transport but can be extended to non-ballistic modes of transport when a suitable theory is developed while researchers study phenomena that sometimes prevent electrons in a carbon nanotube from going ballistic. Our results show that while the accuracy and speed of the proposed model vary with the number of piece-wise segments in the mobile charge approximation, it is possible to obtain a speed-up of more than 1000 times while maintaining the accuracy within less than 2% in terms of average RMS error compared with the state of the art theoretical reference CNT model implemented in FETToy. This numerical efficiency makes our model particularly suitable for implementation in circuit-level, eg. SPICE-like, simulators where large numbers of such devices may be used to build complex circuits. Tom J. Kazmierski, Dafeng Zhou, Bashir M. Al-Hashimi |
DATE | 3 |
| 2008 | Serialized Asynchronous Links for NoCabstractThis paper proposes an asynchronous serialized link for NoC that can achieve the same levels of performance in terms of flits per second as a synchronous link but with a reduced number of wires in the point to point switch links and reduced power consumption. This is achieved by employing serialization in the asynchronous domain as opposed to synchronous to facilitate the removal of global clocking on the serial links. Based on transistor level simulations using 0.12 μm foundry models it has been shown that it is possible to achieve the same level of performance as synchronous but with 75% reduction in wires and 65% reduction in power for a 300 MFlit/s link with 8 buffers with a switch clock speed of 300 MHz. Furthermore the paper presents the design requirements arising from interfacing switches of synchronous NoC and asynchronous serial links. Simon Ogg, Enrico Valli, Bashir M. Al-Hashimi, Alexandre Yakovlev, Crescenzo D'Alessandro, Luca Benini |
DATE | 3 |
| 2008 | Integrated approach to energy harvester mixed technology modelling and performance optimisationabstractThis paper presents an integrated approach to energy harvester modelling and performance optimisation where the complete mixed physical-domain energy harvester system (micro generator, voltage booster, storage element and load) can be modelled and optimised in a systematic manner using one simulation platform. We developed an accurate HDL model for the energy harvester and demonstrated its accuracy by validating it experimentally and comparing it with recently reported models. To address the performance loss due to the close mechanical-electrical interaction that takes place in energy harvesters, we proposed a holistic methodology to the energy harvester optimisation based on the HDL model. The effectiveness of employing such an approach has been demonstrated by showing that it is possible to improve vibration-based energy harvester efficiency (energy delivered to load/harvested energy) by 30% through optimising the micro-generator size and the voltage booster circuit components. Leran Wang, Tom J. Kazmierski, Bashir M. Al-Hashimi, Stephen P. Beeby, Russel N. Torah |
DATE | 3 |
| 2008 | Bridge Defect Diagnosis for Multiple-Voltage DesignabstractMultiple-voltage is an effective dynamic power reduction design technique, commonly used in low power ICs. To the best of our knowledge there is no reported work for diagnosing multiple-Vdd enabled ICs and the aim of this paper is to propose a method for diagnosing bridge defects in such ICs. Using synthesized ISCAS benchmarks, with realistic extracted bridges and parametric fault model; the paper investigates the impact of varying supply voltage on the accuracy of diagnosis and demonstrates how the additional voltage settings can be leveraged to improve the diagnosis resolution through a novel multi-Vdd diagnosis algorithm. S. Saqib Khursheed, Paul M. Rosinger, Bashir M. Al-Hashimi, Sudhakar M. Reddy, Peter Harrod |
ETS | 3 |
| 2008 | VHDL-AMS Implementation of a Numerical Ballistic CNT Model for Logic Circuit SimulationabstractThis paper introduces a novel numerical carbon nanotube transistor (CNT) modelling approach which brings in a flexible and efficient cubic spline non-linear approximation of the non-equilibrium mobile charge density. The spline algorithm creates a rapid and accurate solution of the numerical relationship between the charge density and the self-consistent voltage, which leads to the speed-up of deriving the current through the channel without losing much accuracy. This modelling method also allows the flexibility of choosing different cubic spline intervals which may affect the performance of the model, but it is still capable of obtaining an acceleration of more than a 100 times while maintaining the accuracy within less than 1.5% normalised RMS error compared with previous reported theoretical modelling approach. The model has been proved working as transistors in a logic inverter implemented using VHDL-AMS and simulated in SystemVision, which shows the availability of implementing a circuit-level simulators with our proposed model. Additionally, although this model is originally based on the ideal ballistic transport characteristics, it shows good flexibility that the extension with numbers of non-ballistic features are certainly acceptable. Dafeng Zhou, Tom J. Kazmierski, Bashir M. Al-Hashimi |
FDL | 3 |
| 2008 | Iterative Decoding for Redistributing Energy Consumption in Wireless Sensor NetworksabstractIn this paper, we propose a method for desirably redistributing a wireless sensor network's energy consumption from its sensor nodes (which may have scarce energy resources obtained through energy harvesting, for example) to its central node (which often has an abundant energy resource, such as the mains). At the cost of increasing the central node's decoding complexity, our method facilitates (1) a significant reduction in the number of times the sensor nodes are required to retransmit data owing to transmission errors and/or (2) a reduction of up to 3.99 dB in the sensor node's total transmit energy consumption. We show that our approach can reduce the overall energy consumption of transmitting sensor nodes by more than 20% in practice. Robert G. Maunder, Alex S. Weddell, Geoff V. Merrett, Bashir M. Al-Hashimi, Lajos Hanzo |
ICCCN | 4 |
| 2008 | A Structured Hardware/Software Architecture for Embedded Sensor NodesabstractOwing to the limited requirement for sensor processing in early networked sensor nodes, embedded software was generally built around the communication stack. Modern sensor nodes have evolved to contain significant on-board functionality in addition to communications, including sensor processing, energy management, actuation and locationing. The embedded software for this functionality, however, is often implemented in the application layer of the communications stack, resulting in an unstructured, top-heavy and complex stack. In this paper, we propose an embedded system architecture to formally specify multiple interfaces on a sensor node. This architecture differs from existing solutions by providing a sensor node with multiple stacks (each stack implements a separate node function), all linked by a shared application layer. This establishes a structured platform for the formal design, specification and implementation of modern sensor and wireless sensor nodes. We describe a practical prototype of an intelligent sensing, energy-aware, sensor node that has been developed using this architecture, implementing stacks for communications, sensing and energy management. The structure and operation of the intelligent sensing and energy management stacks are described in detail. The proposed architecture promotes structured and modular design, allowing for efficient code reuse and being suitable for future generations of sensor nodes featuring interchangeable components. Geoff V. Merrett, Alex S. Weddell, Nick R. Harris, Bashir M. Al-Hashimi, Neil M. White |
ICCCN | 4 |
| 2008 | An Empirical Energy Model for Supercapacitor Powered Wireless Sensor NodesabstractThe modeling of energy components in wireless sensor network (WSN) simulation is important for obtaining realistic lifetime predictions and ensuring the faithful operation of energy-aware algorithms. The use of supercapacitors as energy stores on WSN nodes is increasing, but their behavior differs from that of batteries. This paper proposes a model for a supercapacitor energy store based upon experimental results, and compares obtained simulation results to those using an 'ideal' energy store model. The proposed model also considers the variety and behavior of energy consumers, and finds that contrary to many existing models, the energy consumed depends on the store voltage (which varies considerably during supercapacitor discharge). Furthermore, energy models in a node's embedded firmware are shown to be paramount for providing energy-aware operation. Geoff V. Merrett, Alex S. Weddell, Adam P. Lewis, Nick R. Harris, Bashir M. Al-Hashimi, Neil M. White |
ICCCN | 5 |
| 2008 | A New Approach for Transient Fault Injection Using Symbolic SimulationabstractOne effective fault injection approach involves instrumenting the RTL in a controlled manner to incorporate fault injection, and evaluating the behaviour of the faulty RTL whilst running some benchmark programs. This approach relies on checking the effects of faults whilst the design is executing a specific binary image, and therefore the true impact of the fault is limited by the shadow of the program image. Another limitation of this approach is the use of extra hardware for fault injection which is not needed during the fault-free running of the design. The aim of this paper is to propose a new approach for transient fault injection based on symbolic simulation and model checking that circumvents the problems experienced due to application dependent fault injection and RTL modification. In this paper we present our approach and analyse the effect of transient faults on the fetch unit of a 32-bit multi-cycle RISC processor. Our approach can be applied generally to any faulty design, not necessarily a processor. Ashish Darbari, Bashir M. Al-Hashimi, Peter Harrod, Daryl Bradley |
IOLTS | 2 |
| 2008 | SystemC-Based Minimum Intrusive Fault Injection Technique with Improved Fault RepresentationabstractIn this paper, we propose a new SystemC-based fault injection technique that has improved fault representation in visible and on-the-fly data and signal registers. The technique is minimum intrusive since it only requires replacing the original data or signal types to fault injection enabler types. We compare the proposed simulation technique with recently reported SystemC-based techniques and show that our technique has fast simulation speed, better fault representation, while maintaining simplicity and minimum intrusion. We demonstrate fault injection capabilities in a behavioural SystemC description of MPEG-2 decoder using proposed technique and show that up to 98.9% fault representation within data and signal registers can be achieved. Rishad A. Shafik, Paul M. Rosinger, Bashir M. Al-Hashimi |
IOLTS | 3 |
| 2008 | Reduced Z-datapath Cordic RotatorabstractIn this article we propose a novel scheme based on virtually scaling-free COordinate Rotation DIgital Computer (CORDIC) algorithm to design a hardware efficient CORDIC rotator. For predicting rotation directions, less than 1/3rdof the elementary rotational stages require classical CORDIC iteration. The rest of the iteration directions could be computed in parallel and the corresponding z-datapath could be eliminated. A 16-bit implementation of the processor requires 0.23 mm2silicon area and consumes 967.8 μW power when synthesized in 0.18 μm technology. Koushik Maharatna, Karim El-Shabrawy, Bashir M. Al-Hashimi |
ISCAS | 3 |
| 2008 | SEU-Hardened Energy Recovery Pipelined Interconnects for On-Chip Networks
Alireza Ejlali, Bashir M. Al-Hashimi |
NOCS | 2 |
| 2008 | Thermal-Aware SoC Test Scheduling with Test Set Partitioning and Interleaving
Zhiyuan He 0002, Zebo Peng, Petru Eles, Paul M. Rosinger, Bashir M. Al-Hashimi |
J. Electron. Test. | 5 |
| 2008 | Bridging Fault Test Method With Adaptive Power Management AwarenessabstractA key design constraint of circuits used in hand-held devices is the power consumption, mainly due to battery-life limitations. Adaptive power management (APM) techniques aim at increasing the battery life of such devices by adjusting the supply voltage and operating frequency, and thus the power consumption, according to the workload. Testing for resistive bridging defects in APM-enabled designs raises a number of challenges due to their complex analog behavior. Testing at more than one supply voltage setting can be employed to improve defect coverage in such systems; however, switching between several supply voltage settings has a detrimental impact on the overall cost of test. This paper proposes a multi- automatic test generation method which delivers 100% resistive bridging defect coverage and also a way of reducing the number of supply voltage settings required during test through test point insertion. The proposed techniques have been experimentally validated using a number of benchmark circuits. S. Saqib Khursheed, Urban Ingelsson, Paul M. Rosinger, Bashir M. Al-Hashimi, Peter Harrod |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | Exploiting Power-Area Tradeoffs in Behavioural Synthesis through clock and operations throughput selectionabstractThis paper describes a new dynamic-power aware high level synthesis (HLS) data path approach that considers the close interrelation between clock choice and operations throughput selection whilst attempting to minimize area, power, or a combination thereof. It is shown that the proposed approach with its compound cost function and its novel clock and operations throughput selection algorithm, obtains solutions with lower power and area than using previous relevant work (Raghunathan, 1997), Moreover, different power-area tradeoffs can be explored due to the appropriate choice of clock period and operations throughput using our novel approach. M. A. Ochoa-Montiel, Bashir M. Al-Hashimi, Peter Kollig |
ASP-DAC | 2 |
| 2007 | Resistive Bridging Faults DFT with Adaptive Power Management AwarenessabstractA key design constraint of circuits used in handheld devices is the power consumption, due mainly to the limitations of battery life. The employment of adaptive power management (APM) methods optimizes the power consumption of such circuits. This paper describes an effective APM-aware DFT technique that consists of a Test Generation Suite, including fault list generation, test pattern generation and fault simulation. The test generation suite is capable of generating test patterns for multiple supply voltage (Vdd) settings to maximize coverage of resistive bridging faults; and a method to reduce the number of Vdd settings without compromising the fault coverage in order to reduce the cost of test. Preliminarily validations of the proposed DFT technique using a number of benchmark circuits demonstrate its effectiveness. Urban Ingelsson, Paul M. Rosinger, S. Saqib Khursheed, Bashir M. Al-Hashimi, Peter Harrod |
ATS | 4 |
| 2007 | Joint consideration of fault-tolerance, energy-efficiency and performance in on-chip networks
Alireza Ejlali, Bashir M. Al-Hashimi, Paul M. Rosinger, Seyed Ghassem Miremadi |
DATE | 2 |
| 2007 | Architecture Level Power-Performance Tradeoffs for Pipelined DesignsabstractThis paper presents a method to investigate power-performance tradeoffs in digital pipelined designs. The method is applied at the architectural level of the design. It is shown that addressing the tradeoffs at this level results in significant savings in power consumption without impacting the performance. The reduction in power is obtained through reducing the number of registers used in implementing the pipeline stages. The method has been validated by synthesizing a floating-point unit with different pipeline stages and power consumption of the designs were obtained using industry standard tools. It is shown that it is possible to obtain up to 18% reduction in power without affecting the clock period and with less area. Bashir M. Al-Hashimi |
ISCAS | 2 |
| 2007 | Reducing Interconnect Cost in NoC through Serialized Asynchronous LinksabstractThis work investigates the application of serialization as a means of reducing the number of wires in NoC combined with asynchronous links in order to simplify the clocking of the link. Throughput is reduced but savings in routing area and reduction in power could make this attractive Simon Ogg, Enrico Valli, Crescenzo D'Alessandro, Alexandre Yakovlev, Bashir M. Al-Hashimi, Luca Benini |
NOCS | 5 |
| 2007 | Workload-ahead-driven online energy minimization techniques for battery-powered embedded systems with time-constraintsabstractThis article proposes a new online voltage scaling (VS) technique for battery-powered embedded systems with real-time constraints. The VS technique takes into account the execution times and discharge currents of tasks to further reduce the battery charge consumption when compared to the recently reported slack forwarding technique [Ahmed and Chakrabarti 2004], while maintaining low online complexity of O(1). Furthermore, we investigate the impact of online rescheduling and remapping on the battery charge consumption for tasks with data dependency which has not been explicitly addressed in the literature and propose a novel rescheduling/remapping technique. Finally, we take leakage power into consideration and extend the proposed online techniques to include adaptive body biasing (ABB) which is used to reduce the leakage power. We demonstrate and compare the efficiency of the presented techniques using seven real-life benchmarks and numerous automatically generated examples. Marcus T. Schmitz, Bashir M. Al-Hashimi, Sudhakar M. Reddy |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2007 | Energy Optimization of Multiprocessor Systems on Chip by Voltage SelectionabstractDynamic voltage selection and adaptive body biasing have been shown to reduce dynamic and leakage power consumption effectively. In this paper, we optimally solve the combined supply voltage and body bias selection problem for multiprocessor systems with imposed time constraints, explicitly taking into account the transition overheads implied by changing voltage levels. Both energy and time overheads are considered. The voltage selection technique achieves energy efficiency by simultaneously scaling the supply and body bias voltages in the case of processors and buses with repeaters, while energy efficiency on fat wires is achieved through dynamic voltage swing scaling. We investigate the continuous voltage selection as well as its discrete counterpart, and we prove strong NP-hardness in the discrete case. Furthermore, the continuous voltage selection problem is solved using nonlinear programming with polynomial time complexity, while for the discrete problem, we use mixed integer linear programming and a polynomial time heuristic. We propose an approach that combines voltage selection and processor shutdown in order to optimize the total energy Alexandru Andrei, Petru Eles, Zebo Peng, Marcus T. Schmitz, Bashir M. Al-Hashimi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2006 | Cache size selection for performance, energy and reliability of time-constrained systemsabstractImproving performance, reducing energy consumption and enhancing reliability are three important objectives for embedded computing systems design. In this paper, we study the joint impact of cache size selection on these three objectives. For this purpose, we conduct extensive fault injection experiments on five benchmark examples using a cycle-accurate processor simulator. Performance and reliability are analyzed using the performability metric. Overall, our experiments demonstrate the importance of a careful cache size selection when designing energy-efficient and reliable systems. Furthermore, the experimental results show the existence of optimal or Pareto-optimal cache size selection to optimize the three design objectives Marcus T. Schmitz, Alireza Ejlali, Bashir M. Al-Hashimi, Sudhakar M. Reddy |
ASP-DAC | 4 |
| 2006 | Improving routing efficiency for network-on-chip through contention-aware input selectionabstractThe performance of network-on-chip (NoC) largely depends on the underlying routing techniques, which have two constituencies: output selection and input selection. Previous research on routing techniques for NoC has focused on the improvement of output selection. This paper investigates the impact of input selection, and presents a novel contention-aware input selection (CAIS) technique for NoC that improves the routing efficiency. When there are contentions of multiple input channels competing for the same output channel, CAIS decides which input channel obtains the access depending on the contention level of the upstream switches, which in turn removes possible network congestion. Simulation results with different synthetic and real-life traffic patterns show that, when combined with either deterministic or adaptive output selection, CAIS achieves significant better performance than the traditional first-come-first-served (FCFS) input selection, with low hardware overhead (<3%). Bashir M. Al-Hashimi, Marcus T. Schmitz |
ASP-DAC | 2 |
| 2006 | Minimizing test power in SRAM through reduction of pre-charge activityabstractIn this paper we analyze the test power of SRAM memories and demonstrate that the full functional pre-charge activity is not necessary during test mode because of the predictable addressing sequence. We exploit this observation in order to minimize power dissipation during test by eliminating the unnecessary power consumption associated with the pre-charge activity. This is achieved through a modified pre-charge control circuitry, exploiting the first degree of freedom of March tests, which allows choosing a specific addressing sequence. The efficiency of the proposed solution is validated through extensive Spice simulations Luigi Dilillo, Paul M. Rosinger, Bashir M. Al-Hashimi, Patrick Girard 0001 |
DATE | 3 |
| 2006 | Dynamic Voltage Scaling Aware Delay Fault TestingabstractThe application of Dynamic Voltage Scaling (DVS) to reduce energy consumption may have a detrimental impact on the quality of manufacturing tests employed to detect permanent faults. This paper analyses the influence of different voltage/frequency settings on fault detection within a DVS application. In particular, the effect of supply voltage on different types of delay faults is considered. This paper presents a study of these problems with simulation results. We have demonstrated that the test application time increases as we reduce the test voltage. We have also shown that for newer technologies we do not have to go to very low voltage levels for delay fault testing. We conclude that it is necessary to test at more than one operating voltage and that the lowest operating voltage does not necessarily give the best fault cover. Noohul Basheer Zain Ali, Mark Zwolinski, Bashir M. Al-Hashimi, Peter Harrod |
ETS | 3 |
| 2006 | On-Chip Time Measurement Architecture with Femtosecond Timing ResolutionabstractThis paper presents a new on-chip time measurement architecture which is based on the time-to-digital conversion (TDC) method that is capable of achieving a timing resolution of tens of femtoseconds without the use of external automatic test equipment (ATE). This is the highest temporal resolution that has been reported to-date and is achieved by the use of the homodyne technique. The proposed architecture has been designed using a 0.12mum CMOS process and simulation results based on foundry transistor models indicates that it is possible to achieve a timing resolution of 40 fs. The time measurement architecture is standalone and occupies a small silicon area, 150mum by 180mum, making it attractive for high resolution on-chip time measurement Matthew Collins, Bashir M. Al-Hashimi |
ETS | 2 |
| 2006 | Enhancing Delay Fault Coverage through Low Power Segmented ScanabstractReducing power dissipation during test has been an active area of academic and industrial research for the last few years and numerous low power DFT techniques and test generation procedures have been proposed. Segmented scan [17-20] has been shown to be an effective technique in addressing test power issues in industrial designs [18]. To achieve higher shipped product quality, tests for delay faults are becoming essential components of manufacturing test. This paper demonstrates, for the first time, that segmented scan facilitates increased delay fault coverage without degrading the reduction of the switching activity obtained by segmented scan. The increased transition delay fault coverage is achieved through careful selection of the capture cycle application. Experimental results on larger ISCAS-89 benchmarks show that using three segments, on average, fault coverage using launch off capture can be increased by about 5.4% while simultaneously reducing the peak switching activity caused by capture cycles by over 30%. Zhuo Zhang 0008, Sudhakar M. Reddy, Irith Pomeranz, Janusz Rajski, Bashir M. Al-Hashimi |
ETS | 5 |
| 2006 | New JETTA Editors, 2006
Bashir M. Al-Hashimi, Dimitris Gizopoulos, Manoj Sachdev, Adit D. Singh |
J. Electron. Test. | 1 |
| 2006 | Thermal-Safe Test Scheduling for Core-Based System-on-Chip Integrated CircuitsabstractOverheating has been acknowledged as a major problem during the testing of complex system-on-chip integrated circuits. Several power-constrained test-scheduling solutions have been recently proposed to tackle this problem during system integration. However, we show that these approaches cannot guarantee hot-spot-free test schedules because they do not take into account the nonuniform distribution of heat dissipation across the die and the physical adjacency of simultaneously active cores. This paper proposes a new test-scheduling approach that is able to produce short test schedules and guarantee thermal safety at the same time. Two thermal-safe test-scheduling algorithms are proposed. The first algorithm computes an exact (shortest) test schedule that is guaranteed to satisfy a given maximum temperature constraint. The second algorithm is a heuristic intended for complex systems with a large number of embedded cores, for which the exact thermal-safe test-scheduling algorithm may not be feasible. Based on a low-complexity test-session thermal-cost model, this algorithm produces near-optimal length test schedules with significantly less computational effort compared to the optimal algorithm Paul M. Rosinger, Bashir M. Al-Hashimi, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Dual Flow Nets: Modeling the control/data-flow relation in embedded systemsabstractThis paper addresses the interrelation between control and data flow in embedded system models through a new design representation, called Dual Flow Net (DFN). A modeling formalism with a very close-fitting control and data flow is achieved by this representation, as a consequence of enhancing its underlying Petri net structure. The work presented in this paper does not only tackle the modeling side in embedded systems design, but also the validation of embedded system models through formal methods. Various introductory examples illustrate the applicability of the DFN principles, whereas the capability of the model to with complex designs is demonstrated through the design and verification of a real-life Ethernet coprocessor. Mauricio Varea, Bashir M. Al-Hashimi, Luis Alejandro Cortés, Petru Eles, Zebo Peng |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2006 | Combined time and information redundancy for SEU-tolerance in energy-efficient real-time systemsabstractRecently, the tradeoff between energy consumption and fault-tolerance in real-time systems has been highlighted. These works have focused on dynamic voltage scaling (DVS) to reduce dynamic energy dissipation and on-time redundancy to achieve transient-fault tolerance. While the time redundancy technique exploits the available slack-time to increase the fault-tolerance by performing recovery executions, DVS exploits slack-time to save energy. Therefore, we believe there is a resource conflict between the time-redundancy technique and DVS. The first aim of this paper is to propose the use of information redundancy to solve this problem. We demonstrate through analytical and experimental studies that it is possible to achieve both higher transient fault-tolerance [tolerance to single event upsets (SEUs)] and less energy using a combination of information and time redundancy when compared with using time redundancy alone. The second aim of this paper is to analyze the interplay of transient-fault tolerance (SEU-tolerance) and adaptive body biasing (ABB) used to reduce static leakage energy, which has not been addressed in previous studies. We show that the same technique (i.e., the combination of time and information redundancy) is applicable to ABB-enabled systems and provides more advantages than time redundancy alone. Alireza Ejlali, Bashir M. Al-Hashimi, Marcus T. Schmitz, Paul M. Rosinger, Seyed Ghassem Miremadi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Quasi-Static Voltage Scaling for Energy Minimization with Time ConstraintsabstractSupply voltage scaling and adaptive body-biasing are important techniques that help to reduce the energy dissipation of embedded systems. This is achieved by dynamically adjusting the voltage and performance settings according to the application needs. In order to take full advantage of slack that arises from variations in the execution time, it is important to recalculate the voltage (performance) settings during run time, i.e., online. However voltage scaling (VS) is computationally expensive, and thus significantly hampers the possible energy savings. To overcome the online complexity, we propose a quasi-static voltage scaling scheme, with a constant online time complexity O(1). This allows us to increase the exploitable slack as well as to avoid the energy dissipated due to online recalculation of the voltage settings. We conduct several experiments that demonstrate the advantages of the proposed technique over the previously published voltage scaling approaches. Alexandru Andrei, Marcus T. Schmitz, Petru Eles, Zebo Peng, Bashir M. Al-Hashimi |
DATE | 5 |
| 2005 | Rapid Generation of Thermal-Safe Test SchedulesabstractOverheating has been acknowledged as a major issue in testing complex SoC. Several power constrained system-level DFT solutions (power constrained test scheduling) have recently been proposed to tackle this problem. However as is shown in this paper imposing a chip-level maximum power constraint does not necessarily avoid local overheating due to the nonuniform distribution of power across the chip. This paper proposes a new approach for dealing with overheating during test, by embedding thermal awareness into test scheduling. The proposed approach facilitates rapid generation of thermal-safe test schedules without requiring time-consuming thermal simulations. This is achieved by employing a low-complexity test session thermal model used to guide the test schedule generation algorithm. This approach reduces the chances of a design re-spin due to potential overheating during test. Paul M. Rosinger, Bashir M. Al-Hashimi, Krishnendu Chakrabarty |
DATE | 2 |
| 2005 | Power-Composition Profile Driven Co-Synthesis with Power Management Selection for Dynamic and Leakage Energy ReductionabstractRecent research has shown that the combination of dynamic voltage scaling (DVS) and adaptive body biasing (ABB) yields high energy reductions in embedded systems. Nevertheless, the implementation of DVS and ABB requires a significant system cost, making it less attractive for many small systems. In this paper we demonstrate that it is possible to reduce this system cost and to achieve comparable energy saving to that obtained using combined DVS and ABB scheme through a co-synthesis methodology which is aware of the tasks' power-composition profile (the ratio of the dynamic power to the leakage power). In particular, the presented methodology performs a power management selection at the architectural level, i. e., it decides upon which processing elements to be equipped with which power management scheme (DVS, ABB, or combined DVS and ABB) - with the aim to achieve high energy savings at a reduced implementation cost. The proposed technique maps, schedules, and voltage scales applications specified as task graphs with timing constraints. Detailed experiments including a real-life benchmark are conducted to demonstrate the effectiveness of the proposed methodology. Bashir M. Al-Hashimi, Marcus T. Schmitz, Petru Eles |
DSD | 2 |
| 2005 | A programmable time measurement architecture for embedded memory characterizationabstractThis paper describes a programmable time measurement architecture that facilitates memory characterization. We have created a standalone time measurement architecture that can measure rise time, fall time, pulse width and propagation delay time measurements without the need of additional circuitry as presented in M. J. Hsiaoet al. (2004) or circuit duplication based in T. Xia and J. C. Lo (2003). This is achieved by the use of time-to-digital conversion (TDC) based on the dual-slope principle. The key feature of the proposed architecture is programmability through the use of a novel programmable input stage. Furthermore, a current steering time-to-voltage converter (TVC) is used in order to improve the linearity and dynamic range as compared to recent designs. The proposed architecture has been designed using 0.18/spl mu/m CMOS process and results from simulations using foundry models suggest it is possible to achieve a timing resolution of 103ps. The measurement core size is 110/spl mu/m /spl times/ 75 /spl mu/m. Matthew Collins, Bashir M. Al-Hashimi, J. Neil Ross |
ETS | 2 |
| 2005 | Energy efficient SEU-tolerance in DVS-enabled real-time systems through information redundancyabstractConcerns about the reliability of real-time embedded systems that employ dynamic voltage scaling has recently been highlighted [1,2,3], focusing on transient-fault-tolerance techniques based on time-redundancy. In this paper we analyze the usage of information redundancy in DVS-enabled systems with the aim of improving both the system tolerance to transient faults as well as the energy consumption. We demonstrate through a case study that it is possible to achieve both higher fault-tolerance and less energy using a combination of information and time redundancy when compared with using time redundancy alone. This even holds despite the impact of the information redundancy hardware overhead and its associated switching activities Alireza Ejlali, Marcus T. Schmitz, Bashir M. Al-Hashimi, Seyed Ghassem Miremadi, Paul M. Rosinger |
ISLPED | 3 |
| 2005 | Cosynthesis of energy-efficient multimode embedded systems with consideration of mode-execution probabilitiesabstractWe present a novel co-design methodology for the synthesis of energy-efficient embedded systems. In particular, we concentrate on distributed embedded systems that accommodate several different applications within a single device, i.e., multimode embedded systems. Based on the key observation that operational modes are executed with different probabilities, that is, the system spends uneven amounts of time in the different modes, we develop a new co-design technique that exploits this property to significantly reduce energy dissipation. Energy and cost savings are achieved through a suitable synthesis process that yields better hardware-resource-sharing opportunities. We conduct several experiments, including a realistic smart phone example, that demonstrate the effectiveness of our approach. Reductions in power consumption of up to 64% are reported. Marcus T. Schmitz, Bashir M. Al-Hashimi, Petru Eles |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | Synchronization overhead in SOC compressed testabstractTest data compression is an enabling technology for low-cost test. Compression schemes however, require communication between the system under test and the automated test equipment. This communication, referred to in this paper as synchronization overhead, may hinder the effective deployment of this new test technology for core-based systems-on-chip. This paper analyzes the sources of synchronization overhead and discusses the different tradeoffs, such as area overhead, test time and automatic test equipment extensions. A novel scalable and programmable on-chip distribution architecture is proposed, which addresses the synchronization overhead problem and facilitates the use of low cost testers for manufacturing test. The design of the proposed architecture is introduced in a generic framework, and the implementation issues (including the test controller and test set preparation) have been considered for a particular case. Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Overhead-Conscious Voltage Selection for Dynamic and Leakage Energy Reduction of Time-Constrained SystemsabstractDynamic voltage scaling and adaptive body biasing have been shown to reduce dynamic and leakage power consumption effectively. In this paper, we optimally solve the combined supply voltage and body bias selection problem for multi-processor systems with imposed time constraints, explicitly taking into account the transition overheads implied by changing voltage levels. Both energy and time overheads are considered. We investigate the continuous voltage scaling as well as its discrete counterpart, and we prove NP-hardness in the discrete case. Furthermore, the continuous voltage scaling problem is formulated and solved using nonlinear programming with polynomial time complexity, while for the discrete problem we use mixed integer linear programming. Extensive experiments, conducted on several benchmarks and a real-life example, are used to validate the approaches. Alexandru Andrei, Marcus T. Schmitz, Petru Eles, Zebo Peng, Bashir M. Al-Hashimi |
DATE | 5 |
| 2004 | Minimization of Crosstalk Noise, Delay and Power Using a Modified Bus Invert TechniqueabstractPreviously reported bus encoding approaches reduce crosstalk delay but they ignore the effects of inductive coupling between the bus lines, i.e. crosstalk noise. Aiming to solve this issue, this paper presents a modified bus-invert technique which minimizes crosstalk noise, as well as delay and power, at the expense of a small area overhead. Matheos Lampropoulos, Bashir M. Al-Hashimi, Paul M. Rosinger |
DATE | 2 |
| 2004 | A compression-driven test access mechanism design approachabstractDriven by the industrial need for low-cost test methodologies, the academic community and the industry alike have put forth a number of efficient test data compression (TDC) methods. In addition, the need for core-based System-on-a-Chip (SoC) test led to considerable research in test access mechanism (TAM) design. While most previous work has considered TAM design and TDC independently, this work analyzes the interrelations between the two, outlining that a minimum test time solution obtained using TAM design will not necessarily correspond to a minimum test time solution when compression is applied. This is due to the dependency of some TDC methods on test bus width and care bit density, both of which are related to test time, and hence to TAM design. Therefore, this paper illustrates the importance of considering the characteristics of the compression method when performing TAM design, and it also shows how an existing TAM design method can be enhanced toward a compression-driven solution. Paul Theo Gonciari, Bashir M. Al-Hashimi |
ETS | 2 |
| 2004 | Simultaneous communication and processor voltage scaling for dynamic and leakage energy reduction in time-constrained systemsabstractWe propose a new technique for the combined voltage scaling of processors and communication links, taking into account dynamic as well as leakage power consumption. The voltage scaling technique achieves energy efficiency by simultaneously scaling the supply and body bias voltages in the case of processors and buses with repeaters, while energy efficiency on fat wires is achieved through dynamic voltage swing scaling. We also introduce a set of accurate communication models for the energy estimation of voltage scalable embedded systems. In particular, we demonstrate that voltage scaling of bus repeaters and dynamic adaption of the voltage swing on fat wires can significantly influence the system's energy consumption. Experimental results, conducted on numerous generated benchmarks and a real-life example, demonstrate that substantial energy savings can be achieved with the proposed techniques. Alexandru Andrei, Marcus T. Schmitz, Petru Eles, Zebo Peng, Bashir M. Al-Hashimi |
ICCAD | 5 |
| 2004 | Testability Trade-Offs for BIST Data Paths
Nicola Nicolici, Bashir M. Al-Hashimi |
J. Electron. Test. | 2 |
| 2004 | Scan architecture with mutually exclusive scan segment activation for shift- and capture-power reductionabstractPower dissipation during scan testing is becoming an important concern as design sizes and gate densities increase. While several approaches have been recently proposed for reducing power dissipation during the shift cycle (minimum transition don't care fill, special scan cells and scan chain partitioning), very little work has been carried out towards reducing the peak power during test response capture and the few existing approaches for reducing capture power rely on complex ATPG algorithms. This paper proposes a scan architecture with mutually exclusive scan segment activation which overcomes the shortcomings of previous approaches. The proposed architecture achieves both shift and capture power reduction with no impact on the performance of the design, and with minimal impact on area and testing time (typically 2-3%). An algorithmic procedure for assigning flip-flips to scan segments enables reuse of test patterns generated by standard ATPG tools. An implementation of the proposed method had been integrated into an automated design flow using commercial synthesis and simulation tools which was used on a wide range of benchmark designs. Reductions up to 57% in average power, and up to 44% and 34% in peak power dissipation during shift and capture cycles, respectively, were obtained when using two scan segments. Increasing the number of scan segments to six leads to reductions of 96% and 80% in average power and respectively maximum number of simultaneous transitions. Paul M. Rosinger, Bashir M. Al-Hashimi, Nicola Nicolici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Iterative schedule optimization for voltage scalable distributed embedded systemsabstractWe present an iterative schedule optimization for multirate system specifications, mapped onto heterogeneous distributed architectures containing dynamic voltage scalable processing elements (DVS-PEs). To achieve a high degree of energy reduction, we formulate a generalized DVS problem, taking into account the power variations among the executing tasks. An efficient heuristic is presented that identifies optimized supply voltages by not only "simply" exploiting slack time, but under the additional consideration of the power profiles. Thereby, this algorithm minimizes the energy dissipation of heterogeneous architectures, including power-managed processing elements, effectively. Further, we address the simultaneous schedule optimization toward timing behavior and DVS utilization by integrating the proposed DVS heuristic into a genetic list scheduling approach. We investigate and analyze the possible energy reduction at both steps of the co-synthesis (voltage scaling and scheduling), including the power variations effects. Extensive experiments indicate that the presented work produces solutions with high quality. Marcus T. Schmitz, Bashir M. Al-Hashimi, Petru Eles |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2003 | Test Data Compression: The System Integrator's Perspective
Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
DATE | 2 |
| 2003 | Versatile High-Level Synthesis of Self-Checking Datapaths Using an On-Line Testability Metric
Petros Oikonomakos, Mark Zwolinski, Bashir M. Al-Hashimi |
DATE | 3 |
| 2003 | A Co-Design Methodology for Energy-Efficient Multi-Mode Embedded Systems with Consideration of Mode Execution Probabilities
Marcus T. Schmitz, Bashir M. Al-Hashimi, Petru Eles |
DATE | 2 |
| 2003 | Scheduling and Mapping of Conditional Task Graphs for the Synthesis of Low Power Embedded Systems
Bashir M. Al-Hashimi, Petru Eles |
DATE | 2 |
| 2003 | A CAD methodology for switched current analog IP coresabstractAs system chips begin to absorb analog functionality there is increasing interest into what makes a good analog IP core. Research suggests that a solution could be the transistor-only switched current analog circuit technique, perfectly suited to modern digital processes. However, many designers are unfamiliar with this technology, favouring more conventional approaches not hindered by a lack of supporting CAD tools. To address this problem we have developed a CAD methodology, which allows rapid generation of switched current analog IP cores. This recently developed system has already shown good success with filter design and is being expanded to include other major analog functions. Reuben Wilcock, Bashir M. Al-Hashimi |
ETFA (1) | 2 |
| 2003 | Variable-length input Huffman coding for system-on-a-chip testabstractThis paper presents a new compression method for embedded core-based system-on-a-chip test. In addition to the new compression method, this paper analyzes the three test data compression environment (TDCE) parameters: compression ratio, area overhead, and test application time, and explains the impact of the factors which influence these three parameters. The proposed method is based on a new variable-length input Huffman coding scheme, which proves to be the key element that determines all the factors that influence the TDCE parameters. Extensive experimental comparisons show that, when compared with three previous approaches, which reduce some test data compression environment's parameters at the expense of the others, the proposed method is capable of improving on all the three TDCE parameters simultaneously. Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Addressing useless test data in core-based system-on-a-chip testabstractThis paper analyzes the test memory requirements for core-based systems-on-a-chips and identifies useless test data as one of the contributors to the total amount of test data. The useless test data comprises the padding bits necessary to compensate for the difference between the lengths of different chains in multiple scan chain designs. Although useless test data does not represent any relevant test information, it is often unavoidable, and leads to the tradeoff between the test bus width and the volume of test data in multiple scan chain-based cores. Ultimately, this tradeoff influences the test access mechanism design algorithms leading to solutions that have either short test time or low volume of test data. Therefore, in this paper, a novel test methodology is proposed which, by dividing the wrapper scan chains (WSCs) into two or more partitions, and by exploiting automated test equipment memory management features, reduces the amount of useless test data. Extensive experimental results using ISCAS'89 and ITC'02 benchmark circuits are provided to analyze the implications of the number of WSCs in the partition, and the number of partitions on the proposed methodology. Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2002 | Improving Compression Ratio, Area Overhead, and Test Application Time for System-on-a-Chip Test Data Compression/DecompressionabstractProposes a new test data compression/decompression method for systems-on-a-chip. The method is based on analyzing the factors that influence test parameters: compression ratio, area overhead and test application time. To improve compression ratio, the new method is based on a variable-length input Huffman coding (VIHC), which fully exploits the type and length of the patterns, as well as a novel mapping and reordering algorithm proposed in a pre-processing step. The new VIHC algorithm is combined with a novel parallel on-chip decoder that simultaneously leads to low test application time and low area overhead. It is shown that, unlike three previous approaches which reduce some test parameters at the expense of the others, the proposed method is capable of improving all the three parameters simultaneously. An experimental comparison on benchmark circuits validates the proposed method. Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
DATE | 2 |
| 2002 | Energy-Efficient Mapping and Scheduling for DVS Enabled Distributed Embedded SystemsabstractIn this paper, we present an efficient two-step iterative synthesis approach for distributed embedded systems containing dynamic voltage scalable processing elements (DVS-PEs), based on genetic algorithms. The approach partitions, schedules, and voltage scales multi-rate specifications given as task graphs with multiple deadlines. A distinguishing feature of the proposed synthesis is the utilisation of a generalised DVS method. In contrast to previous techniques, which "simply" exploit available slack time, this generalised technique additionally considers the PE power profile during a refined voltage selection to further increase the energy savings. Extensive experiments are conducted to demonstrate the efficiency of the proposed approach. We report up to 43.2% higher energy reductions compared to previous DVS scheduling approaches based on constructive techniques and total energy savings of up to 82.9% for mapping and scheduling optimised DVS systems. Marcus T. Schmitz, Bashir M. Al-Hashimi, Petru Eles |
DATE | 2 |
| 2002 | Low Power Mixed-Mode BIST Based on Mask Pattern Generation Using Dual LFSR Re-SeedingabstractLow power design techniques have been employed for more than two decades, however an emerging problem is satisfying the test power constraints for avoiding destructive test and improving the yield. Our research addresses this problem by proposing a new method which maintains the benefits of mixed-mode built-in self-test (BIST) (low test application time and high fault coverage), and reduces the excessive power dissipation associated with scan-based test. This is achieved by employing dual linear feedback shift register (LFSR) re-seeding and generating mask patterns to reduce the switching activity. Theoretical analysis and experimental results show that the proposed method consistently reduces the switching activity by 25% when compared to the traditional approaches, at the expense of a limited increase in storage requirements. Paul M. Rosinger, Bashir M. Al-Hashimi, Nicola Nicolici |
ICCD | 2 |
| 2002 | Integrated Test Data Decompression and Core Wrapper Design for Low-Cost System-on-a-Chip TestingabstractThis paper discusses an integrated solution for reducing the volume of test data for deterministic system-on-a-chip testing. The proposed solution is based on a new test data decompression architecture which exploits the features of a core wrapper design algorithm targeting the elimination of useless test data. The compressed test data can be transferred from the automatic test equipment to the on-chip decompression architecture using only one test pin, thus providing an efficient reduced pin count test methodology for multiple scan chains-based embedded cores. In addition to reducing the volume of test data, the proposed solution decreases the control overhead, test application time and power dissipation during scan. Further, it also requires lower on-chip area when compared to the testing scenarios which employ decompression architectures for every scan chain and it eliminates the synchronization overhead between the automatic test equipment and the system-on-a-chip. Moreover, the proposed solution is scalable and programmable and, since it can be considered as an add-on to a test access mechanism of a given width, it provides seamless integration with any design flow. Thus, the proposed integrated solution is an efficient low-cost test methodology for systems-on-a-chip. Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
ITC | 2 |
| 2002 | Useless Memory Allocation in System-on-a-Chip Test: Problems and SolutionsabstractUnlike the existing research direction that focuses on useful test data reduction, this paper analyzes the useless test data memory requirements for system-on-a-chip test. The proposed solution to minimize the useless test memory is based on a new test methodology which combines a novel core wrapper design algorithm with a new test vector deployment procedure stored in the automatic test equipment (ATE). To reduce memory requirements, the proposed core wrapper design finds the minimum number of wrapper scan chain partitions such that the useless memory allocation is minimized in each partition, which facilitates efficient usage of ATE capabilities. Further the new test vector deployment procedure provides a seamless integration with the ATE. When compared to the previously proposed core wrapper design algorithms, the proposed test methodology reduces the memory requirements up to 45%, without any penalties in test area overhead. Paul Theo Gonciari, Bashir M. Al-Hashimi, Nicola Nicolici |
VTS | 2 |
| 2002 | Multiple Scan Chains for Power Minimization during Test Application in Sequential CircuitsabstractThe paper presents a novel technique for power minimization during test application in sequential circuits using multiple scan chains. The technique is based on a new design for test architecture and a novel test application strategy which reduces spurious transitions in the circuit under test. To facilitate the reduction of spurious transitions, the proposed design for test architecture is based on classifying scan latches into compatible, incompatible and independent scan latches. Based on their classification, the scan latches are partitioned into multiple scan chains and a single extra test vector associated with each scan chain is computed. A new test application strategy which applies the extra test vector to primary inputs while shifting out test responses for each scan chain, minimizes power dissipation by eliminating the spurious transitions which occur in the combinational part of the circuit. The newly introduced multiple scan chain-based technique does not introduce performance degradation and minimizes clock tree power dissipation with minimal impact on both test area and test data overhead. Unlike previous approaches which are test set dependent, and hence are not able to handle large circuits due to the complexity of the design space, the paper shows that with low test area and test data overhead substantial savings in power dissipation during test application are achieved in very low computational time for both small and large test sets. For example, in the case of the benchmark circuit s15850, it takes <6009 in computational time and <1 percent in test area and test data overhead to achieve over 80 percent savings in power dissipation. Nicola Nicolici, Bashir M. Al-Hashimi |
IEEE Trans. Computers | 2 |
| 2002 | Power profile manipulation: a new approach for reducing test application time under power constraintsabstractThis paper proposes a power profile manipulation approach which merges two distinct research directions in low power testing: minimization of test power dissipation and test application time reduction under power constraints. It is shown how complementary techniques can be easily combined through this approach to significantly increase test concurrency under power constraints. This is achieved in two steps: in the first step power dissipation is considered a design objective and, consequently, it is minimized; results are further exploited in the second step, when power becomes a design constraint under which the test application time is reduced. A distinctive feature of the proposed power profile manipulation approach is that it can be included in, and consequently improve, any existing power constrained test scheduling algorithm. Extensive experimental results using benchmark circuits, considering test-per-clock, as well as test-per-scan schemes, show that by integrating the proposed power profile manipulation approach into any existing power constrained test scheduling algorithm, savings up to 41 % in test application time are achieved. Paul M. Rosinger, Bashir M. Al-Hashimi, Nicola Nicolici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Testability trade-offs for BIST RTL data paths: the case for three dimensional design spaceabstractPower dissipation during test application is an emerging problem due to yield and reliability concerns. This paper focuses on BIST for RTL data paths and discusses testability trade-offs in terms of test application time, BIST area overhead and power dissipation. Nicola Nicolici, Bashir M. Al-Hashimi |
DATE | 2 |
| 2001 | Dual transitions petri net based modelling technique for embedded systems specificationabstractThis paper presents a new modelling technique capable of modelling both control and data information using a single unified approach. This is achieved by modifying the classical Petri Net structure, allowing it to have two types of transitions and arcs. As a consequence, loops and conditional operations within complex specifications are easily identified. The system dynamic behaviour is modelled using a new marking scheme of the net consisting of a new element called value for data representation in addition to classical tokens used for control purpose. Structural definitions, behavioural rules and graphical representation of the new modelling technique are given. One potential application of the proposed modelling technique is the internal representation of embedded systems specification. Two examples are included illustrating the applicability and efficiency of the proposed modelling technique. Mauricio Varea, Bashir M. Al-Hashimi |
DATE | 2 |
| 2001 | Tackling test trade-offs for BIST RTL data paths: BIST area overhead, test application time and power dissipationabstractPower dissipation during test application is an emerging problem due to yield and reliability concerns. This paper focuses on BIST for RTL data paths and discusses testability trade-offs in terms of test application time, BIST area overhead and power dissipation. Using a complex validation flow and experimental data for over 30,000 testable data paths, it is shown how test application time decreases asymptotically when increasing power constraints. Further, it is experimentally demonstrated why power conscious test synthesis and test scheduling algorithms are required due to large variations in useless power dissipation as test application time decreases. Finally, while previous research has outlined that test application time decreases as BIST area overhead increases, this paper shows that in order to reach high quality solutions in terms of test application time and BIST area overhead under given power constraints, a three dimensional design space needs to be explored. Nicola Nicolici, Bashir M. Al-Hashimi |
ITC | 2 |
| 2000 | Scan Latch Partitioning into Multiple Scan Chains for Power Minimization in Full Scan Sequential CircuitsabstractPower dissipated during test application is substantially higher than power dissipated during functional operation which can decrease the reliability and lead to yield loss. This paper presents a new technique for power minimization during test application in full scan sequential circuits. The technique is based on classifying scan latches into compatible, incompatible and independent scan latches. Based on their classification, scan latches are partitioned into multiple scan chains. A new test application strategy which applies an extra test vector to primary inputs while shifting out test responses for each scan chain, minimizes power dissipation by eliminating the spurious transitions which occur in the combinational part of the circuit. Unlike previous approaches which are test vector and scan latch order dependent and hence are not able to handle large circuits due to the complexity of the design space, this paper shows that with low test area and test data overhead substantial savings in power dissipation during test application are achieved in a very low computational time. For example, in the case of benchmark circuit s15850 it takes <3600s in computational time and <1% in test area and test data overhead to achieve 80% savings in power dissipation. Nicola Nicolici, Bashir M. Al-Hashimi |
DATE | 2 |
| 2000 | Power conscious test synthesis and scheduling for BIST RTL data pathsabstractPrevious research has outlined that power dissipated during test application is substantially higher than during functional operation, which leads to loss of yield and decreases reliability. This paper shows for the first time how power is minimized in BIST RTL data paths by using power conscious test synthesis and test scheduling. According to the necessity for achieving the required test efficiency, power dissipation is classified into necessary and useless power dissipation. According to the occurrence during the testing process, power dissipation is classified into test application and shifting power dissipation. The effect of test synthesis and scheduling on power dissipation is analyzed and power minimization is achieved in two steps. Firstly, during the testable design space exploration only power conscious test synthesis moves are accepted leading to minimization of useless power dissipation. Secondly, module selection during power conscious test scheduling satisfies power constraints while reducing test application time. Experimental results using generic power models show savings up to 28% in test application power dissipation and up to 29% in shifting power dissipation. Nicola Nicolici, Bashir M. Al-Hashimi |
ITC | 2 |
| 2000 | BIST hardware synthesis for RTL data paths based on testcompatibility classesabstractA new built-in self-test (BIST) methodology for register transfer level (RTL) data paths is presented. The proposed BIST methodology takes advantage of the structural information of the RTL data path and reduces the test application time by grouping same-type modules into test compatibility classes (TCCs). During testing, compatible modules share a small number of test pattern generators at the same test time leading to significant reductions in BIST area overhead, performance degradation and test application time. Module output responses from each TCC are checked by comparators leading to substantial reduction in fault-escape probability. Only a single signature analysis register is required to compress the responses of each TCC which leads to high reductions in volume of output data and overall test application time (the sum of test application time and shifting time required to shift out test responses). This paper shows how the proposed TCC grouping methodology is a general case of the traditional BIST embedding methodology for RTL data paths with both uniform and variable bit width. A new BIST hardware synthesis algorithm employs efficient tabu search-based testable design space exploration which combines the accuracy of incremental test scheduling algorithms and the exploration speed of test scheduling algorithms based on fixed test resource allocation. To illustrate TCC grouping methodology efficiency, various benchmark and complex hypothetical data paths have been evaluated and significant improvements over the BIST embedding methodology are achieved. Nicola Nicolici, Bashir M. Al-Hashimi, Andrew D. Brown, Alan Christopher Williams |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Efficient BIST Hardware Insertion with Low Test Application Time for Synthesized Data PathsabstractIn this paper new and efficient BIST methodology and BIST hardware insertion algorithms are presented for RTL data paths obtained from high level synthesis. The methodology is based on concurrent testing of modules with identical physical information by sharing the test pattern generators in a partial intrusion BIST environment. Furthermore, to reduce the number of signature analysis registers and test application time the same type modules are grouped in test compatibility classes and n-input k-bit comparators are used to check the results. The test application time is computed using an incremental test scheduling approach. An existing test scheduling algorithm is modified to obtain an efficient trade-off between the algorithm complexity and testable design space exploration. A cost function based on both test application time and area overhead is defined and a tabu search-based heuristic capable of exploring the solution space in a very rapid time is presented. To reduce the computational time testable design space exploration is carried our in two phases: test application time reduction phase and BIST area reduction phase. Experimental results are included confirming the efficiency of the proposed methodology. Nicola Nicolici, Bashir M. Al-Hashimi |
DATE | 2 |
| 1999 | Reduction of Latency and Resource Usage in Bit-Level Pipelined Data Paths for FPGAsabstractPipelining of data path structures increases the throughput rate at the expense of enlarged resource usage and latency unless architectures optimised towards specific applications are used. This paper describes a novel methodology for the design of generic bit-level pipelined data paths that have the low resource usage and latency of specifically tailored architectures but still allow the flexible trade-off between speed and resource requirements inherent in generic circuits. This is achieved through the elimination of all skew and alignment flip-flops from the data path whilst still maintaining the original pipelining scheme, hence allowing more compact structures with decreased circuit delays. The resulting low latency is beneficial in the realisation of all recursive signal processing applications and the reduced resource usage enables particularly the efficient FPGA realisation of high performance signal processing functions. The design process is illustrated through the high level-based FPGA realisation of a 9th-order wave digital filter, demonstrating that high performance and efficient resource usage are possible. For example, the implementation of a digital filter with 10-bit signal word length and 6-bit coefficients using a Xilinx XC4013XL-1 device supports sample rates of 2.5MHz Peter Kollig, Bashir M. Al-Hashimi |
FPGA | 2 |
| 1998 | Correction to the Proof of Theorem 2 in "Parallel Signature Analysis Design with Bounds on Aliasing"
Nicola Nicolici, Bashir M. Al-Hashimi |
IEEE Trans. Computers | 2 |