Rishad A. Shafik

dblp:60/617 · also Rishad Ahmed Shafik · DBLP profile ↗
← Back
49ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0001-5444-537XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 6 first-author · 18 since 2021Software engineering, systems software and programming languages · 19 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ANIMATE: Automated Framework for Scalable Design of Tsetlin Machines Using 1-Safe Petri Nets
Alex Chan, Mohamed Tarraf, Rishad A. Shafik, Alexandre Yakovlev
PETRI NETS3
2026 AURORA - AUtomated 8T SRAM Wired-OR Logic Array for Boolean-Based Machine Learning
Komal Krishnamurthy, Marcos L. L. Sartori, Shengyu Duan, Alexandre Yakovlev, Rishad A. Shafik
DATE5
2026 Learning dynamics, pattern recognition capability and interpretability of the Tsetlin Machine
abstract
The inability to trace an AI’s reasoning process and understand why it makes each decision is known as the black box problem. This remains one of the major barriers to the trusted and widespread use of machine learning in many application domains. The paper explores pattern recognition performance and learning dynamics of the Tsetlin Machine – a new explainable logic-based machine-learning approach. Tsetlin Machine uses a collection of finite-state automata with a unique logic-based learning mechanism and provides a promising alternative to Artificial Neural Networks with several advantages, such as interpretability, low complexity, suitability for hardware implementation and high performance. This work investigates Tsetlin Machine’s mechanism for constructing conjunctive clauses from data and their interpretation for pattern recognition on several datasets. We demonstrate that during training the logical clauses learn persistent sub-patterns within the class. Each clause creates a class template by clustering a certain number of similar class samples, combining them through literal-wise logical conjunction (i.e., AND-ing). The number of class samples that each clause combines depends on Tsetlin Machine’s hyperparameters. The more class samples that are combined, the more general the clauses become. The paper aims at uncovering how Tsetlin Machine’s hyperparameters influence the balance between clause generalization and specialization and how this affects the accuracy of pattern recognition. It also studies the evolution of the machine’s internal state, its convergence and training completion.
Olga Tarasyuk, Anatoliy Gorbenko, Tousif Rahman, Lei Jiao 0001, Ole-Christoffer Granmo, Rishad A. Shafik, Alexandre Yakovlev
Pattern Recognit.6
2026 An All-Digital 8.6-nJ/Frame 65-nm Tsetlin Machine Image Classification Accelerator
abstract
We present an all-digital programmable machine learning accelerator chip for image classification, underpinning on the Tsetlin machine (TM) principles. The TM is an emerging machine learning algorithm founded on propositional logic, utilizing sub-pattern recognition expressions called clauses. The accelerator implements the coalesced TM version with convolution, and classifies booleanized images of$28\times 28$pixels with 10 categories. A configuration with 128 clauses is used in a highly parallel architecture. Fast clause evaluation is achieved by keeping all clause weights and Tsetlin automata (TA) action signals in registers. The chip is implemented in a 65 nm low-leakage CMOS technology, and occupies an active area of 2.7 mm2. At a clock frequency of 27.8 MHz, the accelerator achieves 60.3 k classifications per second, and consumes 8.6 nJ per classification. This demonstrates the energy-efficiency of the TM, which was the main motivation for developing this chip. The latency for classifying a single image is$25.4~\mu $s which includes system timing overhead. The accelerator achieves 97.42%, 84.54% and 82.55% test accuracies for the datasets MNIST, Fashion-MNIST and Kuzushiji-MNIST, respectively, matching the TM software models.
Svein Anders Tunheim, Yujin Zheng, Lei Jiao 0001, Rishad A. Shafik, Alexandre Yakovlev, Ole-Christoffer Granmo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Energy-Efficient Reconfigurable Skyrmion-Based Counter for Nanoscale Applications
abstract
Counters are essential building blocks in digital systems, widely used for tasks such as tracking events, generating timing signals, and performing arithmetic operations like counting, frequency division, and sequencing. As the demand for low-power, high-performance computing grows, particularly in applications like IoT and edge devices, energy-efficient counter designs become increasingly crucial. Skyrmions have recently emerged as promising candidates for future logic device design due to their low energy, non-volatility, and stability. Their integration into information processing systems, such as racetrack memory and logic circuits, highlights their potential to overcome the limitations of conventional CMOS technology. This work introduces a reconfigurable skyrmion-based 4-bit counter for high-speed nanoscale computing applications. This architecture offers low energy, reliable performance, and enhanced scalability, making it an attractive solution for next-generation digital circuits and emerging nanoscale applications such as machine learning. The simulation results demonstrate the counter’s ability to reduce energy consumption to approximately 1.27 aJ per transition, which is nearly 99% lower than traditional CMOS-based counters.
C. Kishore, Santhosh Sivasubramani, Sarwath Sara, Arabinda Haldar, Chandrasekhar Murapaka, Rishad A. Shafik, Amit Acharyya
ISCAS6
2025 HotReRAM: A Performance-Power-Thermal Simulation Framework for ReRAM-Based Caches
abstract
This article proposes a comprehensive thermal modeling and simulation framework, HotReRAM, for resistive RAM (ReRAM)-based caches that is verified against a memristor circuit-level model. The simulation is driven by power traces based on cache accesses for detailed temperature modeling over time. HotReRAM models power at a fine-grain level and generates temperature traces for different cache regions together with detailed analyses of thermal stability, retention time and write latency. Combining HotReRAM with gem5, a full-system simulator, and NVSim, a power simulator, for ReRAM enables temporal and spatial modeling of crucial ReRAM characteristics. This integration allows designers and architects to analyze various cache characteristics within a single cache bank and address thermal-induced issues when designing ReRAM caches. Our simulation results for an 8-MiB ReRAM cache show that the spatial thermal variance can be as high as 7 K for a single cache bank, whereas the temporal thermal variance is more than 40 K. Such temperature variances impact retention time with a standard deviation of 3.9–10.2 for a set of benchmark applications, where the write latency can increase by up to 14.5%.
Shounak Chakraborty 0001, Thanasin Bunnam, Jedsada Arunruerk, Sukarn Agarwal, Shengqi Yu, Rishad A. Shafik, Magnus Själander
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Dynamic Tsetlin Machine Accelerators for On-Chip Training Using FPGAs
abstract
The increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficiency and adaptability within the resource constraints of the nodes. Deploying and training Deep Neural Networks (DNNs)-based models at the edge, although accurate, posit significant challenges from the back-propagation algorithm’s complexity, bit precision trade-offs, and heterogeneity of DNN layers. This paper presents a Dynamic Tsetlin Machine (DTM) training accelerator as an alternative to DNN implementations. DTM utilizes logic-based on-chip inference with finite-state automata-driven learning within the same Field Programmable Gate Array (FPGA) package. Underpinned on the Vanilla and Coalesced Tsetlin Machine algorithms, the dynamic aspect of the accelerator design allows for a run-time reconfiguration targeting different datasets, model architectures, and model sizes without resynthesis. This makes the DTM suitable for targeting multivariate sensor-based edge tasks. Compared to DNNs, DTM trains with fewer multiply-accumulates, devoid of derivative computation. It is a data-centric ML algorithm that learns by aligning Tsetlin automata with input data to form logical propositions enabling efficient Look-up-Table (LUT) mapping and frugal Block RAM usage in FPGA training implementations. The proposed accelerator offers 2.54x more Giga operations per second per Watt (GOP/s per W) and uses 6x less power than the next-best comparable design.
Gang Mao, Tousif Rahman, Sidharth Maheshwari, Bob Pattison, Rishad A. Shafik, Alexandre Yakovlev
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 Tsetlin Machine-Based Image Classification FPGA Accelerator With On-Device Training
abstract
The Tsetlin Machine (TM) is a novel machine learning algorithm that uses Tsetlin automata (TAs) to define propositional logic expressions (clauses) for classification. This paper describes a field-programmable gate array (FPGA) accelerator for image classification based on the Convolutional Coalesced Tsetlin Machine. The accelerator classifies booleanized images of$28\times 28$pixels into 10 classes, and is configured with 128 clauses in a highly parallel architecture. To achieve fast clause evaluation and class prediction, the TA action signals and the clause weights per class are available from registers. Full on-device training is included, and the TAs are implemented with 34 Block RAM (BRAM) instances which operate in parallel. Each BRAM is addressed by the clause number and has a 72-bit word width that supports 8 TAs. The design is implemented in a Xilinx Zynq Ultrascale+ XCZU7 FPGA. Running at 50 MHz, the accelerator core achieves 134k image classifications per second, with an energy consumption per classification of$13.3~\mu $J. A single training epoch of 60k samples requires a processing time of 1.5 seconds. The accelerator obtains a test accuracy of 97.6% on MNIST, 84.1% on Fashion-MNIST and 82.8% on Kuzushiji-MNIST.
Svein Anders Tunheim, Lei Jiao 0001, Rishad A. Shafik, Alexandre Yakovlev, Ole-Christoffer Granmo
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Design of Event-Driven Tsetlin Machines Using Safe Petri Nets
Alex Chan, Adrian Wheeldon, Rishad A. Shafik, Alexandre Yakovlev
Petri Nets3
2024 MATADOR: Automated System-on-Chip Tsetlin Machine Design Generation for Edge Applications
abstract
System-on-Chip Field-Programmable Gate Arrays (SoC-FPGAs) offer significant throughput gains for machine learning (ML) edge inference applications via the design of co-processor accelerator systems. However, the design effort for training and translating ML models into SoC-FPGA solutions can be substantial and requires specialist knowledge aware trade-offs between model performance, power consumption, latency and resource utilization. Contrary to other ML algorithms, Tsetlin Machine (TM) performs classification by forming logic proposition between boolean actions from the Tsetlin Automata (the learning elements) and boolean input features. A trained TM model, usually, exhibits high sparsity and considerable overlapping of these logic propositions both within and among the classes. The model, thus, can be translated to RTL-level design using a miniscule number of AND and NOT gates. This paper presents MATADOR, an automated boolean-to-silicon tool with GUI interface capable of implementing optimized accelerator design of the TM model onto SoC-FPGA for inference at the edge. It offers automation of the full development pipeline: model training, system level design generation, design verification and deployment. It makes use of the logic sharing that ensues from propositional overlap and creates a compact design by effectively utilizing the TM model's sparsity. MATADOR accelerator designs are shown to be up to 13.4x faster, up to 7x more resource frugal and up to 2x more power efficient when compared to the state-of-the-art Quantized and Binary Deep Neural Network implementations.
Tousif Rahman, Gang Mao, Sidharth Maheshwari, Rishad A. Shafik, Alexandre Yakovlev
DATE4
2023 Stateful Energy Management for Multi-Source Energy Harvesting Transient Computing Systems
abstract
The intermittent and varying nature of energy harvesting (EH) entails dedicated energy management with large energy storage, which is a limiting factor for low-power/cost systems with small form factors. Transient computing allows system operations to be performed in the presence of power outages by saving the system state into a non-volatile memory (NVM), thereby reducing the size of this storage. These systems are often designed with a task-based strategy, which requires the storage to be sized for the most energy consuming task. That is, however, not ideal for most systems since their tasks have varying energy requirements, i.e., energy storage size and operating voltage. Hence, to overcome this issue, this paper proposes a novel energy management unit (EMU), tailored for multi-source EH transient systems, that allows selecting the storage size and operating voltage for the next task to be performed at run-time, thereby optimizing task-specific energy needs and startup times based on application requirements. For the first time in literature, we adopted a hybrid NVM+VM approach allowing our EMU to reliably and efficiently retain its internal state, i.e., stateful EMU, under even the most severe EH conditions. Extensive empirical evaluations validated the operation of the proposed stateful EMU at a small overhead (0.07mJ of energy to update the EMU state and a$\simeq 4\mu\mathrm{A}$of static current consumption of the EMU).
Sergey Mileiko, Oktay Cetinkaya, Rishad A. Shafik, Domenico Balsamo
DATE3
2023 Asynchronous Control for Tsetlin Machine with Binary Memristor-Transistor Array
abstract
Tsetlin machines (TMs) are a novel machine learning paradigm based on learning automata and Boolean logic inference, with better energy-efficiency and explainability than neural networks. This work exploits non-volatile ReRAM-transistor memory arrays to perform efficient in-memory TM computing. To accommodate the large timing variability of ReRAM devices and enhance energy efficiency and speed, the control path is implemented with quasi delay-insensitive (QDI) asynchronous circuits. The design of these circuits are derived and synthesized from their signal-transition graph specifications using the Workcraft tool. The resulting circuits offer high event-driven controllability and high variation tolerance for the mixed-signal ReRAM data path. Compared to state of the art TM hardware, the new TM design uses less than 5% of the power to achieve better than$4\times$the performance.
Omar Ghazal, Gang Mao, Jesse Ojukwu, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik
ISCAS7
2023 IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin Machines
abstract
In-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively.
Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik
ISLPED11
2023 A multi-step finite-state automaton for arbitrarily deterministic Tsetlin Machine learning
abstract
Abstract Due to the high arithmetic complexity and scalability challenges of deep learning, there is a critical need to shift research focus towards energy efficiency. Tsetlin Machines (TMs) are a recent approach to machine learning (ML) that has demonstrated significantly reduced energy compared to neural networks alike, while providing comparable accuracy on several benchmarks. However, TMs rely heavily on energy‐costly random number generation to stochastically guide a team of Tsetlin Automata (TA) in TM learning. In this paper, we propose a novel finite‐state learning automaton that can replace the TA in the TM, for increased determinism. The new automaton uses multi‐step deterministic state jumps to reinforce sub‐patterns, without resorting to randomization. A determinism parameter finely controls trading off the energy consumption of random number generation, against randomization for increased accuracy. Randomization is controlled by flipping a coin before every state jump, ignoring the state jump on tails. For example, makes every update random and makes the automaton completely deterministic. Both theoretically and empirically, we establish that the proposed automaton converges to the optimal action almost surely. Further, used together with the TM, only substantial degrees of determinism reduce accuracy. Energy‐wise, random number generation constitutes switching energy consumption of the TM, saving up to 11 mW power for larger datasets with high values. Our new learning automaton approach thus facilitates low‐energy ML.
Kuruge Darshana Abeyrathna, Ole-Christoffer Granmo, Rishad A. Shafik, Lei Jiao 0001, Adrian Wheeldon, Alexandre Yakovlev, Jie Lei 0007, Morten Goodwin
Expert Syst. J. Knowl. Eng.3
2023 Approximate digital-in analog-out multiplier with asymmetric nonvolatility and low energy consumption
abstract
Many modern compute-intensive applications require arithmetic results (usually multiplication) to be represented as analog signals. Using digital multipliers followed by digital-to-analog conversion (DAC) results in high energy and performance costs. This is because digital multipliers have costly carry propagation, and DAC circuits add associated conversion costs. Another concern, especially for arithmetic on the edge, is the need for nonvolatile operands in the face of power uncertainty. To deal with this, nonvolatile memory technologies have been combined with in-memory computing. This paper proposes a mixed-signal multiplier which directly generates an analog product based on two digital input operands. Fundamental to the design are transistor-memristor cells, organized in a crossbar structure. Using analog resistive partial product accumulation in the crossbar, the approximate multiplier eliminates the need for carry propagation and an explicit DAC. It also provides asymmetric nonvolatility making memristor writing a rare event, extending the application significance of the method. The design is shown to be functionally correct up to 4-bit, and achieves 8× to over 300× speedup, competitive peak-power and orders of magnitude energy reduction, compared with existing full-digital memristor-based multipliers and low-power multiplication DAC solutions.
Shengqi Yu, Fei Xia 0001, Rishad A. Shafik, Domenico Balsamo, Alexandre Yakovlev
Integr.3
2023 REDRESS: Generating Compressed Models for Edge Inference Using Tsetlin Machines
abstract
Inference at-the-edge using embedded machine learning models is associated with challenging trade-offs between resource metrics, such as energy and memory footprint, and the performance metrics, such as computation time and accuracy. In this work, we go beyond the conventional Neural Network based approaches to explore Tsetlin Machine (TM), an emerging machine learning algorithm, that uses learning automata to create propositional logic for classification. We use algorithm-hardware co-design to propose a novel methodology for training and inference of TM. The methodology, called REDRESS, comprises independent TM training and inference techniques to reduce the memory footprint of the resulting automata to target low and ultra-low power applications. The array of Tsetlin Automata (TA) holds learned information in the binary form as bits: {0,1}, called excludes and includes, respectively. REDRESS proposes a lossless TA compression method, called the include-encoding, that stores only the information associated with includes to achieve over 99% compression. This is enabled by a novel computationally minimal training procedure, called the Tsetlin Automata Re-profiling, to improve the accuracy and increase the sparsity of TA to reduce the number of includes, hence, the memory footprint. Finally, REDRESS includes an inherently bit-parallel inference algorithm that operates on the optimally trained TA in the compressed domain, that does not require decompression during runtime, to obtain high speedups when compared with the state-of-the-art Binary Neural Network (BNN) models. In this work, we demonstrate that using REDRESS approach, TM outperforms BNN models on all design metrics for five benchmark datasets viz. MNIST, CIFAR2, KWS6, Fashion-MNIST and Kuzushiji-MNIST. When implemented on an STM32F746G-DISCO microcontroller, REDRESS obtained speedups and energy savings ranging 5-5700× compared with different BNN models.
Sidharth Maheshwari, Tousif Rahman, Rishad A. Shafik, Alexandre Yakovlev, Ashur Rafiev, Lei Jiao 0001, Ole-Christoffer Granmo
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 A TEG-Based Non-Intrusive Ultrasonic System for Autonomous Water Flow Rate Measurement
abstract
Residential water meters accommodate various methods of power provisioning. Electromagnetic and ultrasonic meters, for example, often rely on a battery-like external power source, whereas mechanical meters harvest energy from water flow through an impeller. Although energy harvesting (EH) minimizes maintenance needs driven by battery depletion/replenishment, placing a physical element into the flow adversely affects water pressure. This intrusive EH/sensing technique is not user-friendly either since the meters with impellers need to be embedded into pipes by skilled personnel. Hence, this paper proposes a non-intrusive sensor system powered by thermoelectric generators (TEGs) forplug-and-playwater flow rate measurement. This system, equipped with a custom-made energy management unit (EMU), adopts ultrasonic sensors, a task-based computing scheme, and a LoRa module for autonomous sensing and reporting of the flow rate. After summarizing thermoelectricity and delta time-of-flight ($\Delta$ToF)-based ultrasonic sensing theory, we provide the system model and design details with a particular focus on the EMU. Then, we experimentally evaluate the system under varying conditions, demonstrating their impact on average sensing and transmission periods. The results unveil that our proposal can achieve high measurement precision ($\pm 1.4\%$), comparable to its intrusive and battery-powered counterparts, and thus has the potential of replacing the residential water meters.
Sergey Mileiko, Oktay Cetinkaya, Darren Mackie, Rishad A. Shafik, Domenico Balsamo
IEEE Trans. Sustain. Comput.4
2022 Runtime Energy Minimization of Distributed Many-Core Systems using Transfer Learning
abstract
The heterogeneity of computing resources continues to permeate into many-core systems making energy-efficiency a challenging objective. Existing rule-based and model-driven methods return sub-optimal energy-efficiency and limited scalability as system complexity increases to the domain of distributed systems. This is exacerbated further by dynamic variations of workloads and quality-of-service (QoS) demands. This work presents a QoS-aware runtime management method for energy minimization using a transfer learning (TL) driven exploration strategy. It enhances standard Q-learning to improve both learning speed and operational optimality (i.e., QoS and energy). The core to our approach is a multi-dimensional knowledge transfer across a task's state-action space. It accelerates the learning of dynamic voltage/frequency scaling (DVFS) control actions for tuning power/performance trade-offs. Firstly, the method identifies and transfers already learned policies between explored and behaviorally similar states referred to as Intra-Task Learning Transfer (ITLT). Secondly, if no similar “expert” states are available, it accelerates exploration at a local state's level through what's known as Intra-State Learning Transfer (ISLT). A comparative evaluation of the approach indicates faster and more balanced exploration. This is shown through energy savings ranging from 7.30% to 18.06%, and improved QoS from 10.43% to 14.3%, when compared to existing exploration strategies. This method is demonstrated under WordPress and TensorFlow workloads on a server cluster.
Dainius Jenkus, Fei Xia 0001, Rishad A. Shafik, Alexandre Yakovlev
DATE3
2022 Cyclostationary Random Number Sequences for the Tsetlin Machine
Svein Anders Tunheim, Rohan Kumar Yadav, Lei Jiao 0001, Rishad A. Shafik, Ole-Christoffer Granmo
IEA/AIE4
2022 Editable asynchronous control logic for SAR ADCs
abstract
This paper presents a novel design method for asynchronous control logic targeting successive approximation register (SAR) analog-to-digital converters (ADCs). This work is based on modeling the control logic for SAR ADCs using signal transition graphs (STGs). Different from conventional synchronous controllers, the proposed method results in asynchronous controllers driven by the causality of signals rather than relying on clocks to control the conversion process. Moreover, the proposed asynchronous control logic can be modularized through the handshake protocol, making it possible to build ADCs of arbitrary precision based on single-bit control units. This work results in a formal, model-based asynchronous design flow for SAR ADC control, which is shown to produce resulting circuits of similar speeds but great power efficiency improvements.
Fei Xia 0001, Gang Mao, Shengqi Yu, Rishad A. Shafik, Alexandre Yakovlev
ISCAS5
2022 Adaptive Intelligence for Batteryless Sensors Using Software-Accelerated Tsetlin Machines
abstract
Tsetlin Machine (TM) is a new machine learning algorithm that encodes propositional logic into learning automata---a set of logical expressions composed of boolean input features---to recognise patterns. The simplicity, efficiency, and accuracy of this logic-based algorithm encourage rethinking the application of traditional arithmetic-based neural networks (NNs) in intelligent sensors design. Indeed, TM is a promising candidate for embedding intelligence into tiny batteryless sensors with the potential to address two critical challenges: (1) computing under resource constraints and (2) demand for dynamic adaptation to the unpredictable nature of harvested energy. However, its structural model complexity manifests in two conflicting issues: large memory footprint and long latency. This paper addresses these shortcomings by proposing adaptive compression techniques exploiting the inherent redundancies observed in trained models. Through dynamically scaling the computational complexity based on available energy, our techniques significantly reduce the memory footprint and speed up the runtime execution. We evaluate our techniques against standard TMs and binarized neural networks (BNNs) for vision and acoustic workloads deployed on a TI MSP430 MCU operating under intermittent power supply conditions. We show that our techniques can achieve up to 99% compression of TM models and offer 13.5× latency and energy reductions when compared with the most efficient neural network configuration without compromising accuracy.
Abu Bakar, Tousif Rahman, Rishad A. Shafik, Fahim Kawsar, Alessandro Montanari
SenSys3
2022 An FPGA Based Energy-Efficient Read Mapper With Parallel Filtering and In-Situ Verification
abstract
In the assembly pipeline of Whole Genome Sequencing (WGS), read mapping is a widely used method to re-assemble the genome. It employs approximate string matching and dynamic programming-based algorithms on a large volume of data and associated structures, making it a computationally intensive process. Currently, the state-of-the-art data centers for genome sequencing incur substantial setup and energy costs for maintaining hardware, data storage and cooling systems. To enable low-cost genomics, we propose an energy-efficient architectural methodology for read mapping using a single system-on-chip (SoC) platform. The proposed methodology is based on the q-gram lemma and designed using a novel architecture for filtering and verification. The filtering algorithm is designed using a parallel sorted q-gram lemma based method for the first time, and it is complemented by an in-situ verification routine using parallel Myers bit-vector algorithm. We have implemented our design on the Zynq Ultrascale+ XCZU9EG MPSoC platform. It is then extensively validated using real genomic data to demonstrate up to 7.8× energy reduction and up to 13.3× less resource utilization when compared with the state-of-the-art software and hardware approaches.
Venkateshwarlu Y. Gudur, Sidharth Maheshwari, Amit Acharyya, Rishad A. Shafik
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 PLEDGER: Embedded Whole Genome Read Mapping using Algorithm-HW Co-design and Memory-aware Implementation
abstract
With over 6000 known genetic disorders, genomics is a key driver to transform the current generation of healthcare from reactive to personalized, predictive, preventive and participatory (P4) form. High throughput sequencing technologies produce large volumes of genomic data, making genome reassembly and analysis computationally expensive in terms of performance and energy. In this paper, we propose an algorithm-hardware co-design driven acceleration approach for enabling translational genomics. Core to our approach is a Pyopencl based tooL for gEnomic workloaDs tarGeting Embedded platforms (PLEDGER). PLEDGER is a scalable, portable and energy-efficient solution to genomics targeting low-cost embedded platforms. It is a read mapping tool to reassemble genome, which is a crucial prerequisite to genomics. Using bit-vectors and variable level optimisations, we propose a low-memory footprint, dynamic programming based filtration and verification kernel capable of accelerated parallel heterogeneous executions. We demonstrate, for the first time, mapping of real reads to whole human genome on a memory-restricted embedded platform using novel memory-aware preprocessed data structures. We compare the performance and accuracy of PLEDGER with state-of-the-art RazerS3, Hobbes3, CORAL and REPUTE on two systems: 1) Intel i7-8750H CPU + Nvidia GTX 1050 Ti, 2) Odroid N2 with 6 cores: 4xCortex-A73 + 2xCortex-A53 and Mali GPU. PLEDGER demonstrates persistent energy and accuracy advantages compared to state-of-the-art read mappers producing up to 11× speedups and 5.9× energy savings compared to state-of-the-art hardware resources.
Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Venkateshwarlu Y. Gudur, Amit Acharyya
DATE2
2021 Low-Latency Asynchronous Logic Design for Inference at the Edge
abstract
Modern internet of things (IoT) devices leverage machine learning inference using sensed data on-device rather than offloading them to the cloud. Commonly known as inference at-the-edge, this gives many benefits to the users, including personalization and security. However, such applications demand high energy efficiency and robustness. In this paper we propose a method for reduced area and power overhead of self-timed early-propagative asynchronous inference circuits, designed using the principles of learning automata. Due to natural resilience to timing as well as logic underpinning, the circuits are tolerant to variations in environment and supply voltage whilst enabling the lowest possible latency. Our method is exemplified through an inference datapath for a low power machine learning application. The circuit builds on the Tsetlin machine algorithm further enhancing its energy efficiency. Average latency of the proposed circuit is reduced by 10× compared with the synchronous implementation whilst maintaining similar area. Robustness of the proposed circuit is proven through post-synthesis simulation with 0.25 V to 1.2 V supply. Functional correctness is maintained and latency scales with gate delay as voltage is decreased.
Adrian Wheeldon, Alexandre Yakovlev, Rishad A. Shafik, Jordan Morris
DATE3
2021 Optimized Multi-Memristor Model based Low Energy and Resilient Current-Mode Multiplier Design
abstract
Multipliers are central to modern compute-intensive applications, such as signal processing and artificial intelligence (AI).However, the complex logic chain in conventional multipliers, particularly due to cascaded carry propagation circuits, contributes to high energy and performance costs.This paper proposes a novel current-mode multiplier design that reduces the carry propagation chain and improves the current amplification.Fundamental to this design is a one transistor multi-memristor (1TxM) cell architecture.In each cell, transistor can be switched ON/OFF to determine the cell selection, while the high/low resistive states of memristors determine the corresponding cell output current when selected.The memristor states as well as biasing configurations in each memristor are suitably optimized through a new memristor model.The number of memristors implementing this model in each cell is suitably determined depending on the cell significance to achieve the required amplification.Consequently, the design reduces the need to have current mirror circuits in each current path, while also ensuring high resilience in transitional bias voltages.Parallel cell currents are then directed to a common current accumulation path to generate the multiplier output without requiring any carry propagation chain.We carried out a wide range of experiments to extensively validate our multiplier design in Cadence Virtuoso analogue design environment for functional and parametric properties.The results show that the proposed multiplier reduces up to 85% latency and 99% energy cost when compared with the recently proposed approaches.
Shengqi Yu, Rishad A. Shafik, Thanasin Bunnam, Kaiyun Chen, Alexandre Yakovlev
DATE2
2021 Run-time Configurable Approximate Multiplier using Significance-Driven Logic Compression
abstract
Designing energy-efficient hardware continues to be challenging due to arithmetic complexities. The problem is further exacerbated in systems powered by energy harvesters as variable power levels can limit their computation capabilities. In this work, we propose a run-time configurable adaptive approximation method for multiplication that is capable of managing the energy and performance tradeoffs — ideally suited in these systems. Central to our approach is a Significance-Driven Logic Compression (SDLC) multiplier architecture that can dynamically adjust the level of approximation depending on the run-time power/accuracy constraints. The architecture can be configured to operate in the exact mode (no approximation) or in progressively higher approximation modes (i.e. 2 to 4-bit SDLC). Our method is implemented in both ASIC and FPGA. The implementation results indicate that our design has only a 2.3% silicon overhead, on top of what is required by a traditional exact multiplier. We evaluate the efficiency of the proposed design through a number of case studies. We show that our method achieves similar image fidelity as in the existing approximate methods, without a delay penalty. Further, the inclusion of the dynamic approximation techniques is justified by up to 62.6% energy savings when processing an image with a multiplier using 4-bit SDLC and 35% energy savings when using 2-bit SDLC. In addition, case study results show that the proposed approach incurs negligible loss in output quality with the worst PSNR of 30dB when using the 4-bit SDLC multiplier.
Ibrahim Haddadi, Issa Qiqieh, Rishad A. Shafik, Fei Xia 0001, Mohammed A. Noaman Al-Hayanni, Alexandre Yakovlev
ICCD3
2021 CORAL: Verification-Aware OpenCL Based Read Mapper for Heterogeneous Systems
abstract
Genomics has the potential to transform medicine from reactive to a personalized, predictive, preventive, and participatory (P4) form. Being a Big Data application with continuously increasing rate of data production, the computational costs of genomics have become a daunting challenge. Most modern computing systems are heterogeneous consisting of various combinations of computing resources, such as CPUs, GPUs, and FPGAs. They require platform-specific software and languages to program making their simultaneous operation challenging. Existing read mappers and analysis tools in the whole genome sequencing (WGS) pipeline do not scale for such heterogeneity. Additionally, the computational cost of mapping reads is high due to expensive dynamic programming based verification, where optimized implementations are already available. Thus, improvement in filtration techniques is needed to reduce verification overhead. To address the aforementioned limitations with regards to the mapping element of the WGS pipeline, we propose a Cross-platfOrm Read mApper using opencL (CORAL). CORAL is capable of executing on heterogeneous devices/platforms, simultaneously. It can reduce computational time by suitably distributing the workload without any additional programming effort. We showcase this on a quadcore Intel CPU along with two Nvidia GTX 590 GPUs, distributing the workload judiciously to achieve up to 2× speedup compared to when, only, the CPUs are used. To reduce the verification overhead, CORAL dynamically adapts k-mer length during filtration. We demonstrate competitive timings in comparison with other mappers using real and simulated reads. CORAL is available at: https://github.com/nclaes/CORAL.
Sidharth Maheshwari, Venkateshwarlu Y. Gudur, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 REPUTE: An OpenCL based Read Mapping Tool for Embedded Genomics
abstract
Genomics is transforming medicine from reactive to personalized, predictive, preventive and participatory (P4). The massive amount of data produced by genomics is a major challenge as it requires extensive computational capabilities, consuming large amounts of energy. A crucial prerequisite for computational genomics is genome assembly but the existing mapping tools used are predominantly software based, optimized for homogeneous high-performance systems. In this paper, we propose an OpenCL based REad maPper for heterogeneoUs sysTEms (REPUTE), which can use diverse and parallel compute and storage devices effectively. Core to this tool are dynamic programming based filtration and verification kernel to map the reads on multiple devices, concurrently. We show hardware/ software co-design and implementations of REPUTE across different platforms, and compare it with state-of-the-art mappers. We demonstrate the performance of mappers on two systems: 1) Intel CPU + 2×Nvidia GPUs; 2) HiKey970 embedded SoC with ARM Cortex-A73/A53 cores. The results show that REPUTE outperforms other read mappers in most cases producing up to 13× speedup with better or comparable accuracy. We also demonstrate that the embedded implementation can achieve up to 27× energy savings, enabling low-cost genomics.
Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya
DATE2
2020 Current-Mode Carry-Free Multiplier Design using a Memristor-Transistor Crossbar Architecture
abstract
Multipliers are a major energy and delay contributor in modern compute-intensive applications due to their complex logic architecture. As such, designing multipliers with reduced energy and faster speed has remained a thoroughgoing challenge. This paper presents a novel, carry-free multiplier, which is suitable for a new-generation of energy-constrained applications. The multiplier circuit consists of an array of memristor-transistor cells that can be selected (i.e., turned ON or OFF) using a combination of DC bias voltages based on the operand values. When a cell is selected it contributes to current in the array path, which is then amplified by current mirrors with variable transistor gate sizes. The different current paths are connected to a node for analogously accumulating the currents to produce the multiplier output directly. This removes the need for latency-sensitive carry propagation stages, typically seen in traditional multipliers. We conduct a number of experiments to validate the functional and parametric properties. Our experiments showed that proposed multiplier achieves 51.44% savings in energy at a similar accuracy when compared with recently proposed approaches.
Shengqi Yu, Ahmed Soltan, Rishad A. Shafik, Thanasin Bunnam, Fei Xia 0001, Domenico Balsamo, Alexandre Yakovlev
DATE3
2020 Explainability and Dependability Analysis of Learning Automata based AI Hardware
abstract
Explainability remains the holy grail in designing the next-generation pervasive artificial intelligence (AI) systems. Current neural network based AI design methods do not naturally lend themselves to reasoning for a decision making process from the input data. A primary reason for this is the overwhelming arithmetic complexity.Built on the foundations of propositional logic and game theory, the principles of learning automata are increasingly gaining momentum for AI hardware design. The lean logic based processing has been demonstrated with significant advantages of energy efficiency and performance. The hierarchical logic underpinning can also potentially provide opportunities for by-design explainable and dependable AI hardware. In this paper, we study explainability and dependability using reachability analysis in two simulation environments. Firstly, we use a behavioral SystemC model to analyze the different state transitions. Secondly, we carry out illustrative fault injection campaigns in a low-level SystemC environment to study how reachability is affected in the presence of hardware stuck-at 1 faults. Our analysis provides the first insights into explainable decision models and demonstrates dependability advantages of learning automata driven AI hardware design.
Rishad A. Shafik, Adrian Wheeldon, Alexandre Yakovlev
IOLTS1
2020 Accelerated Filtering and in situ Verification for Energy-Optimized Genome Read Mapping
abstract
Whole genome sequencing (WGS) includes sequencing and assembly pipelines to extract biological genomes for new advances in healthcare, agriculture and environmental research. It produces small random sections of the genome, called reads, and then re-assembled by mapping those reads to a reference genome. This process called read mapping produces a large volume of data, which are disparately processed by compute- and memory-intensive filtering and verification algorithms. As such, the problem of energy-frugal read mapping has remained an open challenge. In this paper, we propose an accelerated read mapping methodology with combined filtering and verification, implemented on an FPGA platform. Core to our methodology is an algorithm based on q-gram lemma for filtration with Myers bit-vector for verification in tandem. Through in situ verification, the proposed implementation optimizes resource utilization between filtration and verification and introduces parallel pipelines in computation and storage processes. Our experimental analysis shows that this methodology gives up to 8.7× energy efficiency when implemented on the Zynq Ultrascale+ FPGA platform, compared with the state-of-the-art software and hardware approaches.
Venkateshwarlu Y. Gudur, Sidharth Maheshwari, Rishad A. Shafik, Amit Acharyya
ISCAS3
2020 Dynamics of Time-Domain Power-Elastic Circuits for Pervasive Machine Learning
abstract
Time-domain data encoding, in the form of the duty cycle of a pulse width modulated (PWM) signal, has recently shown promising ways of building Machine Learning (ML) circuits. As the temporal signals approximately retain their “pseudo-analog” capacitive charging rates under voltage/ frequency variations, the circuits designed are inherently power elastic, offering the crucial leverage of energy autonomy for pervasive applications. This paper focuses on the analysis of dynamic parametric variations and their impact on the temporally encoded Machine Learning circuits. The aim is to investigate and suitably optimize these parameters for robustness, power elasticity and energy efficiency. Our study of dynamics includes how the selection of passive (R and C) components affects the dynamic range of operating frequency, which we term as “PWM carrier frequency”. We investigate how RC values define the performance and energy in terms of computation latency and energy per operation. Additionally, we demonstrates how the dynamic range of voltage and frequency variations affect functional and non-functional parameters of the PWM-based neural network solutions.
Sergey Mileiko, Thanasin Bunnam, Fei Xia 0001, Rishad A. Shafik, Alexandre Yakovlev
ISCAS4
2020 PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core Systems
abstract
Performance and energy efficiency considerations have shifted computing paradigms from single-core to many-core architectures. At the same time, traditional speedup models such as Amdahl's Law face challenges in the run-time reasoning for system performance and energy efficiency, because these models typically assume limited variations of the parallel fraction. Moreover, the parallel fraction, which varies dynamically in workloads, is generally unknown at run-time without application-level instrumentation. This article describes novel performance/energy trade-off models based on realistic architectural considerations, which describe the parallel fraction and speedup as functions of performance counter values available in modern processors, removing the need for application-level instrumentation. These are then used to develop a Parallelization-Aware Run-time Management (PARMA) approach. PARMA aims at controlling core allocations and operating voltage/frequency points for energy efficiency, according to the varying workload parallel fractions. The efficacy of our models and the PARMA approach is extensively validated using a number of PARSEC benchmark applications, involving two performance/energy trade-off metrics: energy-delay-product (EDP), typically used in high-performance applications and energy per instruction (EPI), suitable for energy-aware applications. Up to 48 and 68 percent improvements in EDP and EPI have been observed using the PARMA approach compared with parallelization-agnostic methods.
Mohammed A. Noaman Al-Hayanni, Ashur Rafiev, Fei Xia 0001, Rishad A. Shafik, Alexander B. Romanovsky, Alexandre Yakovlev
IEEE Trans. Computers4
2019 A Pulse Width Modulation based Power-elastic and Robust Mixed-signal Perceptron Design
abstract
Neural networks are exerting burgeoning influence in emerging artificial intelligence applications at the micro-edge, such as sensing systems and image processing. As many of these systems are typically self-powered, their circuits are expected to be resilient and efficient in the presence of continuous power variations caused by the harvesters. In this paper, we propose a novel mixed-signal (i.e. analogue/digital) approach of designing a power-elastic perceptron using the principle of pulse width modulation (PWM). Fundamental to the design are a number of parallel inverters that transcode the input-weight pairs based on the principle of PWM duty cycle. Since PWM-based inverters are typically agnostic to amplitude and frequency variations, the perceptron shows a high degree of power elasticity and robustness under these variations. We show extensive design analysis in Cadence Analog Design Environment tool using a 3 × 3 perceptron circuit as a case study to demonstrate the resilience in the presence of parameric variations.
Sergey Mileiko, Rishad A. Shafik, Alexandre Yakovlev, Jonathan Edwards
DATE2
2018 Real-Power Computing
abstract
The traditional hallmark in embedded systems is to minimize energy consumption considering hard or soft real-time deadlines. The basic principle is to transfigure the uncertainties of task execution times in the real world into energy saving opportunities. The energy saving is achieved by suitably controlling the reliable power supply at circuit or system-level with the aim of minimizing the slack times, while meeting the specified performance requirements. Computing paradigm for emerging ubiquitous systems, particularly for the energy-harvested ones, has clearly shifted from the traditional systems. The energy supply of these systems can vary temporally and spatially within a dynamic range, essentially making computation extremely challenging. Such a paradigm shift requires disruptive approaches to design computing systems that can provide continued functionality under unreliable supply power envelope and operate with autonomous survivability (i.e., the ability to automatically guarantee retention and/or completion of a given computation task). In this paper, we introduce Real-Power Computing, inspired by the above trends and tenets. We show how computation systems must be designed with power-proportionality to achieve sustained computation and survivability when operating at extreme power conditions. We present extensive analysis of the need for this new computing approach using definitions, where necessary, coupled with detailed taxonomies, empirical observations, a review of relevant research works and example scenarios using three case studies representing the proposed paradigm.
Rishad A. Shafik, Alexandre Yakovlev, Shidhartha Das
IEEE Trans. Computers1
2017 Machine learning for run-time energy optimisation in many-core systems
abstract
In recent years, the focus of computing has moved away from performance-centric serial computation to energy-efficient parallel computation. This necessitates run-time optimisation techniques to address the dynamic resource requirements of different applications on many-core architectures. In this paper, we report on intelligent run-time algorithms which have been experimentally validated for managing energy and application performance in many-core embedded system. The algorithms are underpinned by a cross-layer system approach where the hardware, system software and application layers work together to optimise the energy-performance trade-off. Algorithm development is motivated by the biological process of how a human brain (acting as an agent) interacts with the external environment (system) changing their respective states over time. This leads to a pay-off for the action taken, and the agent eventually learns to take the optimal/best decisions in future. In particular, our online approach uses a model-free reinforcement learning algorithm that suitably selects the appropriate voltage-frequency scaling based on workload prediction to meet the applications' performance requirements and achieve energy savings of up to 16% in comparison to state-of-the-art-techniques, when tested on four ARM A15 cores of an ODROID-XU3 platform.
Dwaipayan Biswas, Vibishna Balagopal, Rishad A. Shafik, Bashir M. Al-Hashimi, Geoff V. Merrett
DATE3
2017 Energy-efficient approximate multiplier design using bit significance-driven logic compression
abstract
Approximate arithmetic has recently emerged as a promising paradigm for many imprecision-tolerant applications. It can offer substantial reductions in circuit complexity, delay and energy consumption by relaxing accuracy requirements. In this paper, we propose a novel energy-efficient approximate multiplier design using a significance-driven logic compression (SDLC) approach. Fundamental to this approach is an algorithmic and configurable lossy compression of the partial product rows based on their progressive bit significance. This is followed by the commutative remapping of the resulting product terms to reduce the number of product rows. As such, the complexity of the multiplier in terms of logic cell counts and lengths of critical paths is drastically reduced. A number of multipliers with different bit-widths (4-bit to 128-bit) are designed in SystemVerilog and synthesized using Synopsys Design Compiler. Post-synthesis experiments showed that up to an order of magnitude energy savings, and reductions of 65% in critical delay and almost 45% in silicon area can be achieved for a 128-bit multiplier compared to an accurate equivalent. These gains are achieved with low accuracy losses estimated at less than 0.00071 mean relative error. Additionally, we demonstrate the energy-accuracy trade-offs for different degrees of compression, achieved through configurable logic clustering. In evaluating the effectiveness of our approach, a case study image processing application showed up to 68.3% energy reduction with negligible losses in image quality expressed as peak signal-to-noise ratio (PSNR).
Issa Qiqieh, Rishad A. Shafik, Ghaith Tarawneh, Danil Sokolov, Alexandre Yakovlev
DATE2
2016 Power-Aware Performance Adaptation of Concurrent Applications in Heterogeneous Many-Core Systems
abstract
Modern embedded systems execute multiple applications, both sequentially and concurrently. These applications are exercised on heterogeneous platforms generating varying power consumption and system workloads (CPU or memory intensive or both). As a result, determining the most energy-efficient system configuration (i.e. the number of parallel threads, their core allocations and operating frequencies) tailored for each kind of workload and application scenario is extremely challenging. In this paper, we propose a novel runtime optimization approach with the aim of achieving maximized power normalized performance considering dynamic variation of workload and application scenarios. Fundamental to this approach is a comprehensive study to investigate the tradeoffs between inter-application concurrency with performance and power consumption under different system configurations. Using real experimental measurements on an Odroid XU-3 heterogeneous platform with a number of PARSEC benchmark applications, we model power normalized performance (in terms of IPS/Watt) underpinning analytical power and performance models, derived through multivariate linear regression (MLR). Using these models, we show that with increasing number of concurrent CPU intensive applications show variable gains in IPS/Watt compared to the memory intensive applications in both sequential and concurrent application scenarios. Furthermore, we demonstrate that it is possible to continuously adapt system configuration through a low-cost and linear-complexity runtime algorithm, which can improve the IPS/Watt by up to 125% compared to the existing approach.
Ali Aalsaud, Rishad A. Shafik, Ashur Rafiev, Fei Xia 0001, Sheng Yang 0003, Alexandre Yakovlev
ISLPED2
2016 Learning Transfer-Based Adaptive Energy Minimization in Embedded Systems
abstract
Embedded systems execute applications with varying performance requirements. These applications exercise the hardware differently depending on the computation task, generating varying workloads with time. Energy minimization with such workload and performance variations within (intra) and across (inter) applications is particularly challenging. To address this challenge, we propose an online approach, capable of minimizing energy through adaptation to these variations. At the core of this approach is a reinforcement learning algorithm that suitably selects the appropriate voltage/frequency scaling (VFS) based on workload predictions to meet the applications' performance requirements. The adaptation is then facilitated and expedited through learning transfer, which uses the interaction between the application, runtime, and hardware layers to adjust the VFS. The proposed approach is implemented as a power governor in Linux and extensively validated on an ARM Cortex-A8 running different benchmark applications. We show that with intra- and inter-application variations, our proposed approach can effectively minimize energy consumption by up to 33% compared to the existing approaches. Scaling the approach to multicore systems, we also demonstrate that it can minimize energy by up to 18% with 2× reduction in the learning time when compared with an existing approach.
Rishad A. Shafik, Sheng Yang 0003, Anup Das 0001, Luis Alfonso Maeda-Nunez, Geoff V. Merrett, Bashir M. Al-Hashimi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Workload uncertainty characterization and adaptive frequency scaling for energy minimization of embedded systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi
DATE4
2015 Application-specific memory protection policies for energy-efficient reliable design
abstract
In this paper, we show that the vulnerability of memory components due to data retention in the presence of soft errors exhibit orders of magnitude variations with applications through extensive analysis of MiBench benchmarks. Underpinning such analysis, we propose a novel application-specific design flow for joint energy efficiency and reliability optimization. The energy efficiency is achieved through voltage/frequency scaling (VFS), while reliability is achieved through suitably choosing the appropriate protection policies (L1-Cache resizing and selective ECC) for hierarchical memory components. Fundamental to such joint optimization is a design analysis framework, which can analyze trade-off between memory protection policies considering the impact of VFS, and apply design optimization algorithm to provide with an energy-efficient design, while meeting a given reliability target. Using this framework the proposed design flow is validated through extensive number of application case studies based on ARMv7 processors modeled in GEM5. We show that the joint consideration of cache resizing and VFS can improve the L1-Cache reliability by up to 5x compared to VFS alone, while incurring <10% energy overhead. Additionally, using selective ECC for L2-Cache and DRAM, we show that energy consumption can be reduced by up to 40%.
Sheng Yang 0003, Rishad A. Shafik, S. Saqib Khursheed, David Flynn, Geoff V. Merrett, Bashir M. Al-Hashimi
RSP2
2015 A Low-Cost Unified Design Methodology for Secure Test and Intellectual Property Core Protection
abstract
On-chip security is an emerging challenge in the design of embedded systems with intellectual property (IP) cores. Traditionally this challenge is addressed using ad hoc design techniques with separate design objectives of secure design for testability (DfT), and IP core protection. However, in this paper, we will argue that such design approaches can incur high costs. Underpinning this argument, we propose a novel design methodology, called Secure TEst and IP core Protection (STEP), which aims to address the joint objective of IP core protection and secure testing. To ensure that this objective is achieved at a low cost, the STEP design methodology employs common key integrated hardware. This hardware is incorporated in the system through an automated design conversion technique, which can be easily merged into the electronic design automation (EDA) tool chain. We evaluate the effectiveness of our proposed design methodology considering various implementations of advanced encryption standard (AES) systems as case studies. We show that our proposed design methodology benefits from design automation with high security, and protection at the cost of low area, and power consumption overheads, when compared with traditional design methodologies.
Rishad A. Shafik, Jimson Mathew, Dhiraj K. Pradhan
IEEE Trans. Reliab.1
2014 Reinforcement Learning-Based Inter- and Intra-Application Thermal Optimization for Lifetime Improvement of Multicore Systems
abstract
The thermal profile of multicore systems vary both within an application's execution (intra) and also when the system switches from one application to another (inter). In this paper, we propose an adaptive thermal management approach to improve the lifetime reliability of multicore systems by considering both inter- and intra-application thermal variations. Fundamental to this approach is a reinforcement learning algorithm, which learns the relationship between the mapping of threads to cores, the frequency of a core and its temperature (sampled from on-board thermal sensors). Action is provided by overriding the operating system's mapping decisions using affinity masks and dynamically changing CPU frequency using in-kernel governors. Lifetime improvement is achieved by controlling not only the peak and average temperatures but also thermal cycling, which is an emerging wear-out concern in modern systems. The proposed approach is validated experimentally using an Intel quad-core platform executing a diverse set of multimedia benchmarks. Results demonstrate that the proposed approach minimizes average temperature, peak temperature and thermal cycling, improving the mean-time-to-failure (MTTF) by an average of 2x for intra-application and 3x for inter-application scenarios when compared to existing thermal management techniques. Furthermore, the dynamic and static energy consumption are also reduced by an average 10% and 11% respectively.
Anup Das 0001, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi, Akash Kumar 0001, Bharadwaj Veeravalli
DAC2
2014 A low power and robust carbon nanotube 6T SRAM design with metallic tolerance
abstract
Carbon nanotube field-effect transistor (CNTFET) is envisioned as a promising device to overcome the limitations of traditional CMOS based MOSFETs due to its favourable physical properties. This paper presents a novel six-transistor (6T) static random access memory (SRAM) bitcell design using CNTFETs. Extensive validations and comparative analyses are carried out with the proposed SRAM design using SPICE based simulations. We show that the proposed CNTFET based SRAM has a significantly better static noise margin (SNM) and write ability margin (WAM) compared to a CNTFET-based standard 6T bitcell, equivalent to isolated read-port 8T cell based on CNTFET, while consuming less dynamic power. We further demonstrate that it exhibits higher robustness under process, voltage and temperature (PVT) variations when compared with the traditional CMOS SRAM cell designs. Furthermore, metallic CNTs removal technique is used considering metallic tolerance to make the proposed SRAM design more reliable.
Luo Sun, Jimson Mathew, Rishad A. Shafik, Dhiraj K. Pradhan
DATE3
2013 A fast and Effective DFT for test and diagnosis of power switches in SoCs
abstract
Power switches are increasingly becoming dominant leakage power reduction technique for sub-100nm CMOS technologies. Hence, fast and effective DFT solution for test and diagnosis of power switches is much needed to facilitate faster identification of potential faults and their locations. In this paper, we present a novel, coarse-grain DFT solution enabling divide and conquer based test and diagnosis solution of power switches. The proposed solution benefits from exponential time savings compared to previously reported solutions. Our DFT solution requires only (2⌈log2m⌈ + 3) clock cycles in the worst case for test and diagnosis for m-segment power switches. These time savings are further substantiated by effective discharge circuit design, which eliminates the possibility of false test and hence significantly reducing the charge and discharge times. We validated the effectiveness of our proposed solution through SPICE simulations on a number of ISCAS benchmark circuits, synthesized using 90nm gate libraries.
Jimson Mathew, Rishad A. Shafik, Subhasis Bhattacharjee, Dhiraj K. Pradhan
DATE3
2013 Software Modification Aided Transient Error Tolerance for Embedded Systems
abstract
Commercial off-the-shelf (COTS) components are increasingly being employed in embedded systems due to their high performance at low cost. With emerging reliability requirements, design of these components using traditional hardware redundancy incur large overheads, time-demanding re-design and validation. To reduce the design time with shorter time-to-market requirements, software-only reliable design techniques can provide with an effective and low-cost alternative. This paper presents a novel, architecture-independent software modification tool, SMART (Software Modification Aided transient eRror Tolerance) for effective error detection and tolerance. To detect transient errors in processor data path, control flow and memory at reasonable system overheads, the tool incorporates selective and non-intrusive data duplication and dynamic signature comparison. Also, to mitigate the impact of the detected errors, it facilitates further software modification implementing software-based check-pointing. Due to automatic software based source-to-source modification tailored to a given reliability requirement, the tool requires no re-design effort, hardware- or compiler-level intervention. We evaluate the effectiveness of the tool using a Xentium processor based system as a case study of COTS based systems. Using various benchmark applications with single-event upset (SEUs) based error model, we show that up to 91% of the errors can be detected or masked with reasonable performance, energy and memory footprint overheads.
Rishad A. Shafik, Gerard K. Rauwerda, Jordy Potman, Kim Sunesen, Dhiraj K. Pradhan, Jimson Mathew, Ioannis Sourdis
DSD1
2012 STEP: a unified design methodology for secure test and IP core protection
abstract
Intellectual property (IP) core based embedded systems design is a pervasive practice in the semiconductor industry due to shorter time-to-market and tougher cost competitions. Protecting the design information in these IP cores and securing test from various attacks are two emerging challenges in today's embedded systems design. Recently reported techniques address these challenges considering secure test and IP core protection separately. However, for ensuring high security during IP core functionality and also during test, joint consideration of secure test and IP core protection is much needed. In this paper, we propose a novel and unified design methodology, called STEP (Secure TEst and IP core Protection), which addresses the joint objective of secure test and IP core protection. The aim of STEP design methodology is to achieve high security at low system cost using the same key integrated hardware during test and IP core functionality. We evaluate the effectiveness of STEP design methodology considering advanced encryption standard (AES) system as a case study. We show that proposed design methodology benefits from high security and test accuracy, requiring up to 9% higher area and 20% power overheads.
Pranav Yeolekar, Rishad A. Shafik, Jimson Mathew, Dhiraj K. Pradhan, Saraju P. Mohanty
ACM Great Lakes Symposium on VLSI2
2010 Soft error-aware design optimization of low power and time-constrained embedded systems
abstract
In this paper, we examine the impact of application task mapping on the reliability of MPSoC in the presence of single-event upsets (SEUs). We propose a novel soft error-aware design optimization using joint power minimization with voltage scaling and reliability improvement through application task mapping. The aim is to minimize the number of SEUs experienced by the MPSoC for a suitably identified voltage scaling of the system processing cores such that the power is reduced and the specified real-time constraint is met.We evaluate the effectiveness of the proposed optimization technique using an MPEG-2 decoder and random task graphs. We show that for an MPEG-2 decoder with four processing cores, our optimization technique produces a design that experiences 38% less SEUs than soft error-unaware design optimization for a soft error rate of 10-9, while consuMPEG-2 decoderming 9% less power and meeting a given real-time constraint. Furthermore, we investigate the impact of architecture allocation (varying the number of MPSoC cores) on the power consumption and SEUs experienced. We show that for an MPSoC with six processing cores and a given real-time constraint, the proposed technique experiences upto 7% less SEUs compared to soft error-unaware optimization, while consuming only 3% more power.
Rishad A. Shafik, Bashir M. Al-Hashimi, Krishnendu Chakrabarty
DATE1
2008 SystemC-Based Minimum Intrusive Fault Injection Technique with Improved Fault Representation
abstract
In this paper, we propose a new SystemC-based fault injection technique that has improved fault representation in visible and on-the-fly data and signal registers. The technique is minimum intrusive since it only requires replacing the original data or signal types to fault injection enabler types. We compare the proposed simulation technique with recently reported SystemC-based techniques and show that our technique has fast simulation speed, better fault representation, while maintaining simplicity and minimum intrusion. We demonstrate fault injection capabilities in a behavioural SystemC description of MPEG-2 decoder using proposed technique and show that up to 98.9% fault representation within data and signal registers can be achieved.
Rishad A. Shafik, Paul M. Rosinger, Bashir M. Al-Hashimi
IOLTS1