VLDB 2026 Research / reviewers in the wild / expert
Sangeet Saha
dblp:27/11467
· DBLP profile ↗
24ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0001-6119-4927ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 6 first-author · 15 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Scalability Challenges in LUT-Based Neural Networks via Pruning OptimisationsabstractModern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computational cost. To address this, Look-Up Table (LUT)-based matrix multiplications have emerged as a promising alternative for reducing the computational cost and time of the multiply-accumulate operations in a neural network. However, the LUT-based neural network still faces the scalability challenge due to the inherent limitations of LUT-based matrix multiplication. To mitigate these scalability limitations, this paper proposes a scalable and energy-efficient LUT-based approximate matrix multiplication unit (LUT-MU) constituting the basic component of the neural networks by integrating a pruning strategy on the MADDNESS algorithm, a LUT-based matrix multiplication methodology. With increasing problem size and precision demands in matrix multiplication, our proposed LUT-MU architecture effectively constrains resource expansion. The case study shows that deploying our LUT-MU in neural network architectures, including fully connected layers (MNIST) and ResNets (CIFAR-10, ImageNet)—on XCZU7EV and XCZU19EG FPGAs, produces up to 1.6× throughput improvement and 4.2× energy efficiency gains over mainstream CUDA-based network implementations, and 1.8× energy efficiency compared to leading quantised neural network implementations, with moderate impact on accuracy. Compared to original MADDNESS-based neural networks, our LUT-MU shows 1.3 to 2.6× resource savings based on various resolution configuration settings of MADDNESS. Xuqi Zhu, Huaizhi Zhang, Chandrajit Pal, Sangeet Saha, Klaus D. McDonald-Maier, Xiaojun Zhai |
IEEE Trans. Computers | 6 |
| 2025 | RENOWNED: A Real-Time Anomaly Detection and Mitigation Framework in Edge-Enabled IoVabstractThe rapid adoption of smart vehicles and their interconnection through the Internet of Vehicles (IoV) has increased the use of electronic control units (ECUs) in cars. These ECUs, while enabling advanced features, also present a larger target for cyberattacks, which can disrupt critical functions and jeopardize safety. The time-sensitive nature of automotive systems necessitates swift responses, making the protection of ECUs crucial. The imprecise computation (IC) task model can mitigate the risk of task completion failures by generating acceptable approximation results within deadlines when achieving absolute accuracy becomes difficult within fixed deadlines and energy budgets. This article introduces RENOWNED, a solution that ensures the normal functioning of these controller area networks (CAN) controlled ECUs even in the face of anomalies. It combines anomaly detection and mitigation through the HEALING module to maintain the desired performance. The anomaly detection module uses graph attention networks (GAT) to identify unusual processor behavior. If an anomaly is detected, the HEALING module takes over, reallocating tasks based on the available resources to guarantee that deadlines are met and energy constraints are not exceeded. Experiments have shown that RENOWNED delivers a Quality of Service (QoS) of 25% to 64% when system utilisation is varied in the range from 40% to 90%. It exhibits an excelling performance in detecting anomalies, achieving a 97.6% accuracy even when the magnitude mixed anomaly signals are very minute. Thus, our proposed RENOWNED offers a robust way to enhance the reliability and energy efficiency of safety-critical automotive applications prevalent in IoV. Chandrajit Pal, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
IEEE Internet Things J. | 2 |
| 2025 | PRECIOUS: Approximate Real-Time Computing in MLC-MRAM Based Heterogeneous CMPsabstractEnhancing quality of service (QoS) in approximate-computing (AC) based real-time systems, without violating power limits is becoming increasingly challenging due to contradictory constraints, i.e., power consumption and time criticality, as multicore computing platforms are becoming heterogeneous. To fulfill these constraints and optimise system QoS, AC tasks should be judiciously mapped on such platforms. However, prior approaches rarely considered the problem of AC task deployment on heterogeneous platforms. Moreover, the majority of prior approaches typically neglect the runtime architectural phenomena, which can be accounted for along with the approximation tolerance of the applications to enhance the QoS. We presentPRECIOUS, a novel hybrid offline-online approach that firstschedules AC real-timetasks on aheterogeneous multicorewith an objective to maximise QoS and determines the appropriate cluster for each task constrained by a system-wide power limit, deadline, and task-dependency. At runtime,PRECIOUSintroduces novel architectural techniques for the AC tasks, where tasks are executed on a heterogeneous platform equipped withmultilevel-cell (MLC)-MRAMbased last-level cache to improve energy efficiency and performance by prudentially leveraging storage density of MLC-MRAM while ameliorating associated high write latency and write energy. Our novel block management for the MLC-MRAM cache further improves performance of the system, which we exploit opportunistically to enhance system QoS, and turn off processor cores during the dynamically generated slacks.PRECIOUS-Offlineachieves up to 76% QoS for a specific task-set, surpassing prior art, whereasPRECIOUS-Onlineenhances QoS by 9.0% by reducing cache miss-rate by 19% on a 64-core heterogeneous system without incurring any energy overhead over a conventional MRAM based cache design. Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier |
IEEE Trans. Computers | 1 |
| 2025 | MESSI: Task Mapping and Scheduling Strategy for FPGA-based Heterogeneous Real-Time SystemsabstractContinuous demands for improved performance within constrained resource budgets are driving a move from homogeneous to heterogeneous processing platforms for the implementation of today’s Real-Time (RT) embedded systems. The applications executing on such systems are typically represented as a Precedence Task Graph (PTG), where a node represents a task or algorithm for one functionality and edges represent the complex interactions between multiple functionalities. Due to RT constraints, the task graph needs to be executed within a specified deadline. Although some existing studies have looked into solving this challenge, comprehensive studies that combine the theoretical features of RT task-graph mapping and scheduling with practical runtime architectural characteristics have mostly been ignored to date. Hence, in this article, we consider the challenge of scheduling an RT application modeled as a single PTG, with the objective of minimizing the overall execution time under Hardware (HW) resource and deadline constraints for heterogeneous Central Processing Unit (CPU) + Field Programmable Gate Array (FPGA) architectures. First, we introduce an optimal solution using Integer Linear Programming (ILP). However, this ILP-based optimal solution suffers from computational complexity and does not scale well even for moderately large problem sizes. Hence, we additionally propose heuristic algorithms for task mapping and scheduling. The efficiency of the proposed scheme, named MESSI, has been evaluated through experiments using PTG on a practical CPU+FPGA system regarding current technology restrictions. Our experiments demonstrate that performance gains of 55.6% and area usage reductions of 46.3% are possible compared to full Software (SW) and HW execution, respectively. Sallar Ahmadi-Pour, Sangeet Saha, Klaus D. McDonald-Maier, Rolf Drechsler |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | APPARENT: AI-Powered Platform Anomaly Detection in Edge ComputingabstractEmbedded systems serving as IoT nodes are often vulnerable to malicious and unknown runtime software that could compromise the system, steal sensitive data, and cause undesirable system behaviour. Commercially available embedded systems used in automation, medical equipment, and automotive industries, are especially exposed to this vulnerability since they lack the resources to incorporate conventional safety features and are challenging to mitigate through conventional approaches. We propose a novel system design coined as APPARENT which can identify program characteristics by monitoring and counting the maximum possible low-level hardware events from Hardware Performance Counters (HPCs) that occur during the program's execution and analyse the correlation among the counts of various monitored events. To further utilise these captured events as features we propose a self-supervised machine learning algorithm that combines a Graph Attention Network GAT and a Generative Topographic Mapping GTM to detect unusual program behaviour as anomalies to enhance the system security. Our proposed methodology takes advantage of attributes like program counter, cycles per instruction, and physical and virtual timers at various exception levels of the embedded processor to identify abnormal activity. APPARENT identifies unknown program behaviours not present in the training phase with an accuracy of over 98.46% on Autobench EEMBC benchmarks. Chandrajit Pal, Sangeet Saha, Xiaojun Zhai, Gareth Howells 0001, Klaus D. McDonald-Maier |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | MAFin: Maximizing Accuracy in FinFET based Approximated Real-Time ComputingabstractWe propose MAFin that exploits the unique temperature effect inversion (TEI) property of a FinFET based multicore platform, where processing speed increases with temperature, in the context of approximate real-time computing. In approximate real-time computing platforms, the execution of each task can be divided into two parts: (i) the mandatory part, execution of which provides a result of acceptable quality, followed by (ii) the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy (QoS) without violating deadlines. With an objective to maximize the QoS for a FinFET based multicore system, MAFin, our proposed real-time scheduler first derives a task-to-core allocation, while respecting system-wide constraints and prepares a schedule. During execution, MAFin further increases the achieved QoS, while balancing the performance and temperature on-the-fly by incorporating a prudential temperature cognizant frequency management mechanism and guarantees imposed constraints. Specifically, MAFin exploits the TEI property of FinFET based processors, where processor-speed is enhanced at the increased temperature, to reduce the execution time of the individual tasks. This reduced execution-time is then traded off either to enhance QoS by executing more from the tasks' optional parts or to improve energy efficiency by turning off the core. While surpassing prior art, MAFin achieves 70% QoS, which is further enhanced by 8.3% in online, with a maximum EDP gain of up to 12%, based on benchmark based evaluation on a 4-core based system. Shounak Chakraborty 0001, Sangeet Saha, Magnus Själander, Klaus D. McDonald-Maier |
DAC | 2 |
| 2024 | ARCTIC: Approximate Real-Time Computing in a Cache-Conscious Multicore EnvironmentabstractImproving result-accuracy in approximate computing (AC) based time-critical systems, without violating power constraints of the underlying circuitry, is gradually becoming challenging with the rapid progress in technology scaling. The execution span of each AC real-time tasks can be split into a couple of parts: (i) the mandatory part, execution of which offers a result of acceptable quality, followed by (ii) the optional part, which can be executed partially or completely to refine the initially obtained result in order to increase the result-accuracy, while respecting the time-constraint. In this article, we introduce a novel hybrid offline-online scheduling strategy, for AC real-time tasks. The goal of real-time scheduler of is to maximise the results-accuracy (QoS) of the task-set with opportunistic shedding of the optional part, while respecting system-wide constraints. During execution, retains exclusive copy of the private cache blocks only in the local caches in a multi-core system and no copies of these blocks are maintained at the other caches, and improves performance (i.e., reduces execution-time) by accumulating more live blocks on-chip. Combining offline scheduling with the online cache optimization improves both QoS and energy efficiency. While surpassing prior arts, our proposed strategy reduces the task-rejection-rate by up to 25%, whereas enhances QoS by 10%, with an average energy-delay-product gain of up to 9.1%, on an 8-core system. Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Magnus Själander, Klaus D. McDonald-Maier |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Bayesian Optimization for Efficient Heterogeneous MPSoC Based DNN Accelerator Runtime TuningabstractWith the explosive growth of Internet of Things (IoT) devices and applications, deploying Deep Neural Networks (DNNs) on resource-constrained embedded edge devices has become a popular research trend. Because such systems have limited resources, they need to rely on optimising resource utilisation to meet performance requirements. However, for scenarios where the DNN application and workloads are dynamically changing, the offline system optimisation technique cannot achieve optimal runtime performance in practical environments. Hence, in this PhD project, we propose a Bayesian Optimisation (BO)-based runtime tuning scheme for improving energy efficiency of heterogeneous MPSoC-based DNN accelerator in the context of DNN applications. By seeking suitable hardware configurations of the accelerator for dynamic DNN inference workloads ranging from 200 M to 600 M FLOPs (floating-point operations) at runtime, the recommended configuration can averagely save up to 15.33% energy consumption from a random configuration setting. Xuqi Zhu, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
FPL | 3 |
| 2023 | DELICIOUS: Deadline-Aware Approximate Computing in Cache-Conscious MulticoreabstractEnhancing result-accuracy in approximate computing (AC) based real-time systems, without violating power constraints of the underlying hardware, is a challenging problem. Execution of such AC real-time applications can be split into two parts: (i)the mandatory part, execution of which provides a result of acceptable quality, followed by (ii)the optional part, that can be executed partially or fully to refine the initially obtained result in order to increase the result-accuracy, without violating the time-constraint. This article introducesDELICIOUS, a novel hybrid offline-onlinescheduling strategyfor AC real-time dependent tasks. By employing an efficientheuristic algorithm,DELICIOUSfirst generates a schedule for a task-set with an objective to maximize the results-accuracy, while respecting system-wide constraints. During execution,DELICIOUSthen introduces aprudential cache resizingthat reduces temperature of the adjacent cores, by generating thermal buffers at the turned off cache ways.DELICIOUSfurther trades off this thermal benefits by enhancing the processing speed of the cores for a stipulated duration, calledV/F Spiking, without violating the power budget of the core, to shorten the execution length of the tasks. This reduced runtime is exploited either to enhance result-accuracy by dynamically adjusting the optional part, or to reduce temperature by enabling sleep mode at the cores. While surpassing the prior art,DELICIOUSoffers 80% result-accuracy with its scheduling strategy, which is further enhanced by 8.3% in online, while reducing runtime peak temperature by 5.8°C on average, as shown by benchmark based evaluation on a 4-core based multicore. Sangeet Saha, Shounak Chakraborty 0001, Sukarn Agarwal, Rahul Gangopadhyay, Magnus Själander, Klaus D. McDonald-Maier |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Task Mapping and Scheduling in FPGA-based Heterogeneous Real-time Systems: A RISC-V Case-StudyabstractHeterogeneous platforms, that integrate CPU and FPGA-based processing units, are emerging as a promising solution for accelerating various applications in the embedded system domain. However, in this context, so far, comprehensive studies that combine theoretical features of real-time task scheduling with practical runtime architectural characteristics have mostly been ignored. To fill this gap, in this paper we propose a real-time scheduling algorithm with the objective of minimizing the overall execution time under hardware resource constraints for heterogeneous CPU+FPGA architectures. In particular, we propose an Integer Linear Programming (ILP) based technique for task allocation and scheduling. We then show how to implement a given scheduling on a practical CPU+FPGA system regarding current technology restrictions and validate our methodology using a practical RISC-V case-study. Our experiments demonstrate that performance gains of 40 % and area usage reductions of 67 % are possible compared to a full software and hardware execution, respectively. Sallar Ahmadi-Pour, Sangeet Saha, Vladimir Herdt, Rolf Drechsler, Klaus D. McDonald-Maier |
DSD | 2 |
| 2022 | A Supervisory Control Approach for Scheduling Real-time Periodic Tasks on Dynamically Reconfigurable PlatformsabstractThe dynamic partial reconfiguration (DPR) feature offered by modern FPGAs provides the flexibility of adapting the underlying hardware according to the needs of a particular situation at runtime, in response to application requirements. In recent times, DPR along with drastically reduced reconfiguration overheads has allowed the possibility of scheduling multiple real-time applications on FPGA platforms. However, in order to effectively harness the computation capacity of an FPGA floor, efficient techniques which can schedule real-time applications over both space and time are required. It may be noted that safety-critical systems often require resource-optimal solutions to reduce size, weight, cost and power consumption of the system. However, the scheduling of real-time tasks on FPGAs in the presence of non-negligible reconfigurationlcontext-switching overheads requires careful exploration of the state space which often makes it prohibitively expensive to be applied on-line. Hence, off-line formal approaches are often preferred in the design of reconfiguration controllers (i.e., schedulers) that are correct-by-construction as well as optimal in terms of usage of resources. In this paper, we propose a formal scheduler synthesis framework that generates an optimal scheduler for a set of non-preemptive periodic real-time tasks executing on a FPGA platform. We show the practical viability of our proposed framework by synthesizing schedulers for real-world benchmark applications and implementing them on FPGAs. Cherinet Kejela, Rajesh Devaraj, Arnab Sarkar 0001, Sangeet Saha |
DSD | 4 |
| 2022 | SENAS: Security driven ENergy Aware Scheduler for Real Time Approximate Computing Tasks on Multi-Processor SystemsabstractPresent day real time approximate computing applications like image and video processing involves execution of a set of tasks before a certain amount of time or deadline. In addition to this, present day systems are associated with strict energy budget that cannot be changed post deployment. The tasks comprises of a mandatory and optional part. Completion of all mandatory portions of all tasks before deadline is much more important than result accuracy in such real time approximate computing applications. Based on the energy budget, the optional portions can be executed that determines the quality of service (QoS) of the system. In ideal scenario, sufficient energy budget is present that ensures completion of both mandatory and optional portions in a system with a pre-determined number of processors. However, if fault or malware attack occurs on one or more processors, then the system will cease to work and results may be fatal. In this work, we consider such a scenario where the processors may be faulty and stop functioning in post deployment phases or some malware may cause unexpected delays in processing or may cause unexpected power draining at runtime that will prevent the system from meeting its deadline. We propose a Security driven ENergy Aware Scheduler (SENAS) that works as a self aware agent. Initially, based on the available energy budget, SENAS determines which task is to be executed in which processor of a system. At runtime, SENAS constantly monitors the working of the processors and on detecting any anomaly in any of the processors, it reschedules its tasks at runtime by reducing execution of the optional portions of the tasks and ensuring completion before deadline with high QoS. Krishnendu Guha, Sangeet Saha, Klaus D. McDonald-Maier |
IOLTS | 2 |
| 2022 | Benchmark Tool for Detecting Anomalous Program Behaviour on Embedded DevicesabstractThis paper presents an open-source benchmark tool for anomaly detection in program behaviour, using program counter (PC) and instruction type information. It is introducing anomalies in artificial way, allowing for fine-grained evaluation with adjustable sliding window sizes and preprocessing configuration. The usage of the benchmark, including demonstrated data collection, does not require any additional hardware other than a standard computer. The benchmark uses the output of llvm-objdump program to focus on non-library code which allows for rapid evaluation of various detection methods with different configurations. The proposed tool extracts features derived from processor’s PC and instruction type information and then utilizes the features to identify abnormal behavior using 4 different anomaly detection algorithms. New detection methods can be easily incorporated into the benchmark, which provides a solid foundation for evaluating novel, previously unseen methods against methods we selected for our experiment. Michal Borowski, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
TrustCom | 2 |
| 2022 | ACCURATE: Accuracy Maximization for Real-Time Multicore Systems With Energy-Efficient Way-Sharing CachesabstractImproving result accuracy in approximate computing (AC)-based real-time applications without violating deadlines has recently become an active research domain. Execution time of AC real-time tasks can individually be separated into: execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve the result accuracy of the initial result within a given deadline. However, obtaining higher result accuracy at the cost of enhanced execution time may lead to deadline violation, along with higher energy usage. We present ACCURATE, a novel hybrid offline–online approximate real-time scheduling approach that first schedules AC-based tasks on multicore with an objective to maximize result accuracy and determines operational processing speeds for each task constrained by system-wide power limit, deadline, and task dependency. At runtime, by employing a way-sharing technique (WH_LLC) at the last level cache (LLC), ACCURATE improves performance, which is further leveraged, to enhance result accuracy by executing more from the optional part and to improve the energy efficiency of the cache by turning off a controlled number of cache ways. ACCURATE also exploits the slacks either to improve the result accuracy of the tasks or to enhance the energy efficiency of the underlying system, or both. ACCURATE achieves 85% QoS with 36% average reduction in cache leakage consumption with a 24% average gain in energy-delay product (EDP) for a 4-core-based chip multiprocessor (CMP) with 6.4% average improvement in performance. Sangeet Saha, Shounak Chakraborty 0001, Xiaojun Zhai, Shoaib Ehsan, Klaus D. McDonald-Maier |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | RASA: Reliability-Aware Scheduling Approach for FPGA-Based Resilient Embedded Systems in Extreme EnvironmentsabstractField-programmable gate arrays (FPGAs) offer the flexibility of general-purpose processors along with the performance efficiency of dedicated hardware that essentially renders it as a platform of choice for modern-day robotic systems for achieving real-time performance. Such robotic systems when deployed in harsh environments often get plagued by faults due to extreme conditions. Consequently, the real-time applications running on FPGA become susceptible to errors which call for a reliability-aware task scheduling approach, the focus of this article. We attempt to address this challenge using a hybrid offline-online approach. Given a set of periodic real-time tasks that require to be executed, the offline component generates a feasible preemptive schedule with specific preemption points. At runtime, these preemption events are utilized for fault detection. Upon detecting any faulty execution at such distinct points, the reliability-aware scheduling approach, RASA, orchestrates the recovery mechanism to remediate the scenario without jeopardizing the predefined schedule. Effectiveness of the proposed strategy has been verified through simulation-based experiments and we observed that the RASA is able to achieve 72% of task acceptance rate even under 70% of system workloads with high fault occurrence rates. Sangeet Saha, Xiaojun Zhai, Shoaib Ehsan, Shakaiba Majeed, Klaus D. McDonald-Maier |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | Energy-Aware Real-Time Tasks Processing for FPGA-Based Heterogeneous CloudabstractCloud computing is becoming a popular model of computing. Due to the increasing complexity of the cloud service request, it often exploits heterogeneous architecture. Moreover, some service requests (SRs)/tasks exhibit real-time features, which are required to be handled within a specified duration. Along with the stipulated temporal management, the strategy should also be energy efficient, as energy consumption in cloud computing is challenging. In this paper, we have proposed a strategy, called “Efficient Resource Allocation of Service Request” (ERASER) for energy efficient allocation and scheduling of periodic real-time SRs on cloud platform. The cloud platform is consists of Field Programmable Gate Arrays (FPGAs) as Processing Elements (PEs) along with the General Purpose Processors (GPP). We have further proposed, an SR migration technique to reduce the tasks rejection by serving maximum SRs. Simulation based experimental results demonstrate that the proposed methodology is capable to achieve upto 90 percent resource utilization with only 26 percent SR rejection rate over different experimental scenarios. Comparison results with other state-of-the-art techniques reveal that the proposed strategy outperforms the existing technique with 17 percent reduction in SR rejection rate and 21 percent reduction in energy consumption. Further, the simulation outcomes have been validated on real FPGA test-bed based on Xilinx Zynq SoC with standard benchmark tasks. Atanu Majumder, Sangeet Saha, Amlan Chakrabarti, Klaus D. McDonald-Maier |
IEEE Trans. Sustain. Comput. | 2 |
| 2021 | EnSuRe: Energy & Accuracy Aware Fault-tolerant Scheduling on Real-time Heterogeneous SystemsabstractThis paper proposes an energy efficient real-time scheduling strategy called EnSuRe, which (i) executes real-time tasks on low power consuming primary processors to enhance the system accuracy by maintaining the deadline and (ii) provides reliability against a fixed number of transient faults by selectively executing backup tasks on high power consuming backup processor. Simulation results reveal that EnSuRe consumes nearly 25% less energy, compared to existing techniques, while satisfying the fault tolerance requirements. EnSuRe is also able to achieve 75% system accuracy with 50% system utilisation. Further, the obtained simulation outcomes are validated on benchmark tasks via a fault injection framework on Xilinx ZYNQ APSoC heterogeneous dual core platform. Sangeet Saha, Adewale Adetomi, Xiaojun Zhai, Server Kasap, Shoaib Ehsan, Tughrul Arslan, Klaus D. McDonald-Maier |
IOLTS | 1 |
| 2021 | Design and Implementation of a RISC V Processor on FPGAabstractThe RISC-V ISA is becoming one of the leading instruction sets for the Internet-of-Things and System-on-Chip applications. Due to its strong security features and open-source nature, it is becoming a competitor to the popular ARM architecture. This paper describes the design of a light weight, open-source implementation of a RISCV processor using modern hardware design teclmiques, the implementation of the design onto a Field Programmable Gate Array (FPGA), and its testing. We wanted to create a RISC-V processor that is easy for beginners to learn from and lightweight enough to be implemented on even small FPGAs. While there are existing opensource implementations of RISC-V processors, none are intuitive enough for a beginner to follow. For this reason, in this paper we have minimised the use of conventions and components in modern processors that are not strictly necessary for a barebones implementation. For example, the processor does not include pipelining and uses a simple Harvard architecture. The barebones nature of the design allows for a lot of potential for upgradability. The implementation of each component, and the corresponding test benches, are written in concise and conventional System Verilog. The project produced a RISC-V processor with files for targeting Basys 3 Artix-7 FPGA. Performance was tested using the Dhyrstone benchmark and achieved a strong 2276 DMIPs/MHz, even outperforming the ARM Cortex-A9, while maintaining very low resource utilization on the FPGA. Ludovico Poli, Sangeet Saha, Xiaojun Zhai, Klaus D. McDonald-Maier |
MSN | 2 |
| 2021 | Prepare: Power-Aware Approximate Real-time Task Scheduling for Energy-Adaptive QoS MaximizationabstractAchieving high result-accuracy in approximate computing (AC) based real-time applications without violating power constraints of the underlying hardware is a challenging problem. Execution of such AC real-time tasks can be divided into the execution of the mandatory part to obtain a result of acceptable quality, followed by a partial/complete execution of the optional part to improve accuracy of the initially obtained result within the given time-limit. However, enhancing result-accuracy at the cost of increased execution length might lead to deadline violations with higher energy usage. We propose Prepare , a novel hybrid offline-online approximate real-time task-scheduling approach, that first schedules AC-based tasks and determines operational processing speeds for each individual task constrained by system-wide power limit, deadline, and task-dependency. At runtime, by employing fine-grained DVFS, the energy-adaptive processing speed governing mechanism of Prepare reduces processing speed during each last level cache miss induced stall and scales up the processing speed once the stall finishes to a higher value than the predetermined one. To ensure on-chip thermal safety, this higher processing speed is maintained only for a short time-span after each stall, however, this reduces execution times of the individual task and generates slacks. Prepare exploits the slacks either to enhance result-accuracy of the tasks, or to improve thermal and energy efficiency of the underlying hardware, or both. With a 70 - 80% workload, Prepare offers 75% result-accuracy with its constrained scheduling, which is enhanced by 5.3% for our benchmark based evaluation of the online energy-adaptive mechanism on a 4-core based homogeneous chip multi-processor, while meeting the deadline constraint. Overall, while maintaining runtime thermal safety, Prepare reduces peak temperature by up to 8.6 °C for our baseline system. Our empirical evaluation shows that constrained scheduling of Prepare outperforms a state-of-the-art scheduling policy, whereas our runtime energy-adaptive mechanism surpasses two current DVFS based thermal management techniques. Shounak Chakraborty 0001, Sangeet Saha, Magnus Själander, Klaus D. McDonald-Maier |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | A self-scrubbing scheme for embedded systems in radiation environmentsabstractAs one of the most important components in the embedded systems, the SRAM are sensitive to radiation effects. When the embedded systems working in the extreme radiation environments, the bit flips could occur frequently and decrease the reliability of the systems significantly. In this paper, the self-scrubbing RAM scheme is proposed for light wight embedded systems in the extreme radiation environments. In the scheme, both scrubbing and ECC are used to mitigate the large number of the errors in the RAMs. The separately scrubber is designed to scrub the RAM separately. Therefore it is can be able to operating the scrubbing, when the CPUs are busy. In addition, the scrubber is a portable modules and the hardware costs do not grow with the size of the available RAM. The results of the real world radiation experiments show that it can correct most errors in the RAM under neutron radiation where the errors rates in unhardened RAMs is approximately 1.2bit/(KB·h). The results of the 6 hours radiation experiments show that the error rates of in the conventional ECC RAM is approximately 4.3×10-4bit/(KB·h), while the self-scrubbing RAMs is less than 8.7×10-5bit/(KB·h). Yufan Lu, Xiaojun Zhai, Sangeet Saha, Shoaib Ehsan, Klaus D. McDonald-Maier |
IOLTS | 3 |
| 2020 | Proxy Circuits for Fault-Tolerant Primitive Interfacing in Reconfigurable Devices Targeting Extreme EnvironmentsabstractContinuous interface access to device-level primitives in reconfigurable devices in extreme environments is key to reliable operation. However, it is possible for a primitive's interface controller, which is static to be rendered non-operational by a permanent damage in the controller's circuitry. In order to mitigate this, this paper proposes the use of relocatable proxy circuits to provide remote interfacing capability to primitives from anywhere on a reconfigurable device. A demonstration with device register read controller shows that an improvement in fault-tolerance can be achieved. Adewale Adetomi, Sangeet Saha, Klaus D. McDonald-Maier, Tughrul Arslan |
ISCAS | 2 |
| 2020 | EAAM: Energy-aware application management strategy for FPGA-based IoT-Cloud environments
Atanu Majumder, Sangeet Saha, Amlan Chakrabarti |
J. Supercomput. | 2 |
| 2019 | Profi-Load: An FPGA-Based Solution for Generating Network Load in Profinet CommunicationabstractIndustrial automation has received a considerable attention in the last few years with the rise of Internet of Things (IoT). Specifically, industrial communication network technology such as Profinet has proved to be a major game changer for such automation. However, industrial automation devices often have to exhibit robustness to dynamically changing network conditions and thus, demand a rigorous testing environment to avoid any safety-critical failures. Hence, in this paper, we have proposed an FPGA-based novel framework called “Profi-Load” to generate Profinet traffic with specific intensities for a specified duration of time. The proposed Profi-Load intends to facilitate the performance testing of the industrial automated devices under various network conditions. By using the advantage of inherent hardware parallelism and re-configurable features of FPGA, Profi-Load is able to generate Profinet traffic efficiently. Moreover, it can be reconfigured on the fly as per the specific requirements. We have developed our proposed Profi-Load framework by employing the Xilinx-based “NetJury” device which belongs to Zynq-7000 FPGA family. A series of experiments have been conducted to evaluate the effectiveness of Profi-Load and it has been observed that Profi-Load is able to generate precise load at a constant rate for stringent timing requirements. Furthermore, a suitable Human Machine Interface (HMI) has also been developed for quick access to our framework. The HMI at the client side can directly communicate with the NetJury device and parameters such as, required load amount, number of packet(s) to be sent or desired time duration can be selected using the HMI. Ahmad Khaliq, Sangeet Saha, Bina Bhatt, Dongbing Gu, Klaus D. McDonald-Maier |
SMC | 2 |
| 2017 | Spatio-Temporal Scheduling of Preemptive Real-Time Tasks on Partially Reconfigurable SystemsabstractReconfigurable devices that promise to offer the twin benefits of flexibility as in general-purpose processors along with the efficiency of dedicated hardwares often provide a lucrative solution for many of today’s highly complex real-time embedded systems. However, online scheduling of dynamic hard real-time tasks on such systems with efficient resource utilization in terms of both space and time poses an enormously challenging problem. We attempt to solve this problem using a combined offline-online approach. The offline component generates and stores various optional feasible placement solutions for different sub-sets of tasks that may possibly be co-mapped together. Given a set of periodic preemptive real-time tasks that requires to be executed at runtime, the online scheduler first carries out an admission control procedure and then produces a schedule, which is guaranteed to meet all timing constraints provided it is spatially feasible to place designated subsets of these tasks at specified scheduling points within a future time interval. These feasibility checks are done and actual placement solutions are obtained through a low overhead search of the statically precomputed placement solutions. Based on this approach, we have proposed a periodic preemptive real-time scheduling methodology for runtime partially reconfigurable devices. Effectiveness of the proposed strategy has been verified through simulation based experiments and we observed that the strategy achieves high resource utilization with low task rejection rates over various simulation scenarios. Sangeet Saha, Arnab Sarkar 0001, Amlan Chakrabarti |
ACM Trans. Design Autom. Electr. Syst. | 1 |