Amit Kumar Singh 0002

dblp:10/7647-2 · DBLP profile ↗
← Back
76ranked-venue papers
15as first author
28since 2021 · last 2026
0000-0003-2056-0569ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 66 · 12 first-author · 25 since 2021Software engineering, systems software and programming languages · 12 · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Late Breaking Results: Adaptive Ensembles of Dynamic DNNs for Collaborative Edge Inference
abstract
Edge computing enables low-latency and privacy-preserving DNN inference, yet heterogeneous and dynamically changing device resources make it difficult to satisfy real-time constraints. In this paper, we present AdaEnsemble, an adaptive and collaborative ensemble inference framework that integrates Dynamic DNNs with deadline-aware scheduling. The system profiles accuracy and latency offline and selects both model widths and participating devices at runtime to maximize accuracy under a given deadline. Experiments on heterogeneous edge devices show that AdaEnsemble adapts effectively to different latency requirements and consistently outperforms the state-of-the-art.
Mingyu Hu, Amit Kumar Singh 0002, Jonathon S. Hare, Geoff V. Merrett
DATE2
2026 Adaptive Federated Learning Defense Against Byzantine Attacks and Concept Drift in IIoT
Assem Alhelou, Amit Kumar Singh 0002, Xiaohang Wang 0001
ISCAS2
2026 Online detection of hardware Trojan enabled packet tampering attack on network-on-chip: A Bayesian approach
Xiaohang Wang 0001, Ge Cao, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Liang Wang 0020, Fen Guo
Integr.5
2026 Device behavioural blueprint (DB2): A risk-aware framework for unique device behaviour profiling using microarchitectural variations
abstract
This paper introduces DB 2 , a risk-aware behavioural identity framework that derives device identity from CPU–RTC timing deviation and Performance Monitoring Unit (PMU) microarchitectural events, without relying on GPUs, radios, sensors, or dedicated hardware. The method captures oscillator-coupled timing variation and execution behaviour through a structured signal-processing pipeline, producing device-specific behavioural signatures that remain distinguishable across reboots, temperature variation, and core transitions. DB 2 structures identity assurance into three layers: closed-set identification, calibrated open-set rejection, and stability-aware risk scoring. Evaluation under a strict three-way split with reboot separation for training, calibration, and unseen testing yields a macro-F 1 of 0.957 on unseen reboots. The open-set layer rejects previously unseen devices with a mean true-positive rate of 0.990 at a calibrated event-level false-reject rate of approximately 0.08 under strict leave-one-device-out validation, with operating-point selection performed exclusively on the calibration split. A Dynamic-Aware Identification and Risk (DAIR) mechanism decomposes behavioural stability across temperature, reboot, and core factors to provide interpretable posture monitoring for enrolled devices. Under identity-claim manipulation via spoofing, Sybil, and relabelling scenarios involving cloning, targeted identities exhibit reduced identification consistency and elevated risk, while non-targeted devices remain stable under identical calibration settings. These results show that behavioural fingerprints can be derived from standard CPU, RTC, and PMU-accessible resources on edge devices, enabling device-identity and behavioural-assurance monitoring in IoT and edge environments without specialised hardware.
Muthupavithran Selvam, Safwana Haque, Amit Kumar Singh 0002, Zhan Cui, Muttukrishnan Rajarajan
J. Netw. Comput. Appl.3
2025 On Design Space Exploration of Cache System in Multi-Chiplet Systems
abstract
While multi-chiplet based many-core systems have emerged as a viable solution for heterogeneous integration and addressing manufacturing and technological challenges in the post-Moore’s Law era, their design and optimization remain highly complex and challenging. Among the various subsystems, the cache hierarchy has significant implications for overall system performance, yet its vast design space presents substantial optimization challenges. This complexity arises from factors such as the large number of chiplets in the system, the number of cores per chiplet, memory hierarchy variations, cache size variability, caching strategies, and inter-chiplet interconnection networks. Existing design space exploration methods, such as NN-Baton and IntLP, fail to optimize cache subsystem performance or thoroughly explore the design space. To address these limitations, we propose a novel design space exploration method for cache subsystem optimization. Our approach models cache miss rates and network latency as functions of cache hierarchy and inter-/intra-chiplet interconnection network parameters. We then define an optimization problem to minimize the concurrent average memory access time (C-AMAT) under cost and power consumption constraints. This problem is addressed using a bilevel optimization algorithm, which iteratively solves two independent subproblems: (1) cache subsystem optimization, and (2) inter-chiplet interconnection network optimization. Experimental results show that our method reduces the application execution time by 39.7% and 39.2%, on average, compared to architectures similar to AMD Zen 4 and Intel Sapphire Rapids, respectively, and by $\mathbf{2 5 . 9 1 \%}$ over IntLP. These results underscore the potential of the proposed method for optimizing cache subsystems in future multi-chiplet based many-core systems.
Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002
DAC4
2025 LEGOSim: A Unified Parallel Simulation Framework for Multi-chiplet Heterogeneous Integration
Tiantian Lin, Xiaohang Wang 0001, Ling Wang 0005, Zhulin Zheng, Yingtao Jiang, Amit Kumar Singh 0002, Jieming Yin, Sihai Qiu, Mingzhe Zhang 0005, Kui Ren 0001
MICRO7
2025 Detection and defence against thermal and timing covert channel attacks in multicore systems
abstract
As interest in multicore systems grows, so does the potential for information leakage through covert channel communication. Covert channel attacks pose severe risks because they can expose confidential information and data. Countering these attacks requires a deep understanding of various covert channel attack types and their characteristics. Thermal covert channel and covert timing channel attacks, which use temperature and timing, respectively to transfer information, are two dominant examples that can compromise sensitive data. In this paper, we propose a methodology for jointly detecting and mitigating these types of attacks, which has been lacking in the literature. Our experiments have demonstrated that the proposed countermeasures can increase the bit error rate (BER) for mitigation while maintaining comparable power consumption to that of the state-of-the-art.
Parisa Rahimi, Amit Kumar Singh 0002, Xiaohang Wang 0001, Seyedali Pourmoafi
J. Syst. Archit.2
2025 LUFT-CAN: A lightweight unsupervised learning based intrusion detection system with frequency-time analysis for vehicular CAN bus
Xiaohang Wang 0001, Li Lu 0008, Shuguo Zhuo, Yingtao Jiang, Amit Kumar Singh 0002, Kui Ren 0001, Mei Yang 0001, Kaiwei Wu
J. Syst. Archit.6
2025 On Task Mapping in Multi-chiplet Based Many-Core Systems to Optimize Inter- and Intra-chiplet Communications
abstract
Multi-chiplet system design, by integrating multiple chiplets/dielets within a single package, has emerged as a promising paradigm in the post-Moore era. This paper introduces a novel task mapping algorithm for multi-chiplet based many-core systems, addressing the unique challenges posed by intra- and inter-chiplet communications under power and thermal constraints. Traditional task mapping algorithms fail to account for the latency and bandwidth differences between these communications, leading to sub-optimal performance in multi-chiplet systems. Our proposed algorithm employs a two-step process: (1) task assignment to chiplets using binary linear programming, leveraging a totally unimodular constraint matrix, and (2) intra-chiplet mapping that minimizes communication latency while considering both thermal and power constraints. This method strategically positions tasks with extensive inter-chiplet communication near interface nodes and centralizes those with predominant intra-chiplet communication. Experimental results demonstrate that the proposed algorithm outperforms existing methods (DAR and IOA) with a 37.5% and 24.7% reduction in execution time, respectively. Communication latency is also reduced by up to 43.2% and 32.9%, compared to DAR and IOA. These findings affirm that the proposed task mapping algorithm aligns well with the characteristics of multi-chiplet based many-core systems, and thus improves optimal performance.
Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001
IEEE Trans. Computers4
2025 On Optimizing Inter- and Intra-Chiplet Interconnection Topologies for Robust Multi-Chiplet Systems
abstract
Inter- and intra-chiplet interconnection networks play a vital role in the operation of many core systems made of multiple chiplets. However, these networks are susceptible to faults caused by manufacturing defects and attacks resulting from the malicious insertion of hardware Trojans and backdoors. Unlike conventional fault-tolerant or countermeasure methods, this article focuses on optimizing network robustness to withstand both faults and attacks, while considering the constraints of chiplet area and power budget. To achieve this, this article first defines network robustness as a quantifiable measure based on various network parameters, after which an optimization problem is formulated to optimize the robustness of the network topology. To efficiently solve this problem, a reinforcement learning algorithm is proposed. Experimental results demonstrate that the proposed method is capable of generating inter- and intra-chiplet interconnection networks that are significantly more robust than existing topology generation methods. Specifically, the proposed method improves robustness over ButterDonut and Kite, respectively, by an average of 10.88% and 14.06% under random faults and by 9.37% and 7.81% under targeted attacks. These experimental results confirm that the proposed method is capable of generating robust inter- and intra-chiplet interconnection networks that can withstand both faults and attacks. By optimizing the network topology’s robustness, it provides a valuable contribution to the design and security of chiplet-based core systems.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Yingtao Jiang, Mei Yang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 On Improving the Performance of Intra- and Inter-chiplet Interconnection Networks in Multi-chiplet Systems for Accelerating FHE Encrypted Neural Network Applications
abstract
Fully Homomorphic Encryption (FHE) is regarded as a promising way to protect data privacy with encrypted computation. Due to high computation overhead, hardware based FHE accelerators were proposed to speed up FHE applications. To support complicated FHE-encrypted neural network applications, multi-chiplet based FHE accelerators were further proposed for scaling up system size, whereas one of the challenges is designing efficient intra- and inter-chiplet interconnection networks to accelerate data transfer. Conventional regular topologies like mesh or Kite either lead to high inter-chiplet transmission latency or excessive power consumption as these topologies assume uniform bandwidth or radix for nodes/links, ignoring the highly irregular distribution of inter-chiplet communication volumes. On the other hand, the problem of generating customized intra- and inter-chiplet interconnection networks has high complexity and previous network-on-chip topology generation works cannot efficiently improve the performance of intra- and inter-chiplet interconnection networks. In this article, the intra- and inter-chiplet interconnection optimization problem is defined, aiming to minimize the execution time of FHE applications under cost and power constraints. To efficiently solve this problem, we propose a bilevel optimization algorithm, which decomposes the problem into three sub-problems: (1) FHE parameters selection, (2) task-to-core mapping, and (3) intra-/inter-chiplet interconnection network topology generation. These sub-problems are then solved iteratively. Experimental results demonstrate that our proposed method reduces execution time by 51.66%, 43.16%, 39.44%, 43.34%, and 27.70% compared with REED and four multi-chiplet based FHE accelerators with mesh, Kite, Butterfly, and Florets as inter-chiplet interconnection networks. Therefore, the proposed method can effectively accelerate FHE applications on large-scale multi-chiplet systems.
Zewei Lai, Jinhui Ye, Xiaohang Wang 0001, Zheang Fu, Amit Kumar Singh 0002, Yingtao Jiang, Kui Ren 0001, Mei Yang 0001, Sihai Qiu, Mingzhe Zhang 0005
ACM Trans. Embed. Comput. Syst.5
2024 Special Session: Emerging Architecture Design, Control, and Security Challenges in Software Defined Vehicles
abstract
Software Defined Vehicles (SDVs) represent a paradigm shift in the automotive industry, where vehicles are increasingly controlled and managed through software, while relying less on mechanical and hardware components. While this allows considerable flexibility in the introduction of new “smart” features and fast tracks innovations in multiple domains, it also creates new challenges and opportunities in architecture design, control, and security. By adopting modular architectures, adaptive control strategies, and robust security measures, SDVs can pave the way for a safer and more efficient future of transportation. In this paper, we cover perspectives from both, industry and academia, in this area. They provide embedded systems researchers an overview of recent developments and emerging challenges in SDV from the perspective of architecture design, control, and security. The emerging challenges also set the foundations for future research in this domain.
Aya El-Fatyany, Xiaohang Wang 0001, Parasara Sridhar Duggirala, Samarjit Chakraborty, Sudeep Pasricha, Amit Kumar Singh 0002
CODES+ISSS6
2024 Fluid Dynamic DNNs for Reliable and Adaptive Distributed Inference on Edge Devices
abstract
Distributed inference is a popular approach for efficient DNN inference at the edge. However, traditional Static and Dynamic DNNs are not distribution-friendly, causing system reliability and adaptability issues. In this paper, we introduce Fluid Dynamic DNNs (Fluid DyDNNs), tailored for distributed inference. Distinct from Static and Dynamic DNNs, Fluid DyDNNs utilize a novel nested incremental training algorithm to enable independent and combined operation of its sub-networks, enhancing system reliability and adaptability. Evaluation on embedded Arm CPUs with a DNN model and the MNIST dataset, shows that in scenarios of single device failure, Fluid Dy DNNs ensure continued inference, whereas Static and Dynamic DNNs fail. When devices are fully operational, Fluid DyDNNs can operate in either a High-Accuracy mode and achieve comparable accuracy with Static DNNs, or in a High-Throughput mode and achieve 2.5x and 2x throughput compared with Static and Dynamic DNNs, respectively.
Lei Xun, Mingyu Hu, Hengrui Zhao, Amit Kumar Singh 0002, Jonathon S. Hare, Geoff V. Merrett
DATE4
2023 Content- and Lighting-Aware Adaptive Brightness Scaling for Improved Mobile User Experience
abstract
For an improved user experience, the display sub-system is expected to provide superior resolution and optimal brightness despite its impact on battery life. Existing brightness scaling approaches set the display brightness statically or adaptively in response to predefined events such as low-battery or ambient light of the environment, which are independent of the displayed content. Approaches that consider the displayed content are either limited to video content or do not account for the user's expected battery life, thereby failing to maximise the user experience. This paper proposes Content- and ambient Lighting-aware Adaptive Brightness Scaling in mobile devices that maximises user experience while meeting battery life expectations. The approach employs a content- and ambient lighting-aware profiler that learns and classifies each sample into predefined clusters at runtime by leveraging insights on user perceptions of content and ambient luminance variations. We maximise user experience through adaptive scaling of the display's brightness using an energy prediction model that determines appropriate brightness levels while meeting expected battery life. The evaluation of the proposed approach on a commercial smartphone improves Quality of Experience (QoE) by up to 24.5 % compared to state-of-art.
Samuel Isuwa, David Amos, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett
DATE3
2023 Mobility-aware fog computing in dynamic networks with mobile nodes: A survey
abstract
Fog computing is an evolving paradigm that addresses the latency-oriented performance and spatio-temporal issues of the cloud services by providing an extension to the cloud computing and storage services in the vicinity of the service requester. In dynamic networks, where both the mobile fog nodes and the end users exhibit time-varying characteristics, including dynamic network topology changes, there is a need of mobility-aware fog computing, which is very challenging due to various dynamisms, and yet systematically uncovered. This paper presents a comprehensive survey on the fog computing compliant with the OpenFog (IEEE 1934) standardised concept, where the mobility of fog nodes constitutes an integral part. A review of the state-of-the-art research in fog computing implemented with mobile nodes is conducted. The review includes the identification of several models of fog computing concept established on the principles of opportunistic networking, social communities, temporal networks, and vehicular ad-hoc networks. Relevant to these models, the contributing research studies are critically examined to provide an insight into the open issues and future research directions in mobile fog computing research.
Krzysztof Ostrowski, Krzysztof Malecki, Piotr Dziurzanski, Amit Kumar Singh 0002
J. Netw. Comput. Appl.4
2023 Maximising mobile user experience through self-adaptive content- and ambient-aware display brightness scaling
abstract
Display subsystems have become the predominant user interface on mobile devices, serving as both input and output interfaces. For a better quality of user experience (QoE), the display subsystem is expected to provide appropriate resolution and brightness despite its impact on battery life. Existing display brightness approaches either consider content- and ambient-light in isolation or do not account for the user’s expected battery life, thereby failing to maximise the QoE. This paper proposes aCADS, a self-Adaptive Content- and Ambient-aware Display brightness Scaling in mobile devices that maximises QoE while meeting battery life expectations. The approach employs a content- and ambient lighting-aware profiler that learns and classifies each sample into predefined clusters at runtime by leveraging insights on user perceptions of content and ambient luminances variations. We maximise QoE through adaptive scaling of the display’s brightness using an energy model that determines appropriate brightness levels while meeting expected battery life. The evaluation on a commercial smartphone shows that aCADS improves QoE by up to 32.5 % compared to state-of-the-art.
Samuel Isuwa, David Amos, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett
J. Syst. Archit.3
2023 Detection of Thermal Covert Channel Attacks Based on Classification of Components of the Thermal Signal Features
abstract
In response to growing security challenges facing many-core systems imposed by thermal covert channel (TCC) attacks, a number of threshold-based detection methods have been proposed. In this paper, we show that these threshold-based detection methods are inadequate to detect TCCs that harness advanced signaling and specific modulation techniques. Since the frequency representation of a TCC signal is found to have multiple side lobes, this important feature shall be explored to enhance the TCC detection capability. To this end, we present a pattern-classification-based TCC detection method using an artificial neural network that is trained with a large volume of spectrum traces of TCC signals. After proper training, this classifier is applied at runtime to infer TCCs, should they exist. The proposed detection method is able to achieve a detection accuracy of 99%, even in the presence of the stealthiest TCCs ever discovered. Because of its low runtime overhead ($< 0.187\%$) and low energy overhead ($< 0.072\%$), this proposed detection method can be indispensable in fighting against TCC attacks in many-core systems. With such a high accuracy in detecting TCCs, powerful countermeasures, like the ones based on dynamic voltage and frequency scaling (DVFS), can be rightfully applied to neutralize any malicious core participating in a TCC attack.
Xiaohang Wang 0001, Hengli Huang, Ruolin Chen, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Computers5
2023 Modeling and Analysis of Thermal Covert Channel Attacks in Many-core Systems
abstract
In a many-core chip, thermal flux and thermal correlation among the cores can be explored to create a thermal covert channel (TCC). In this paper, we provide an analytical model to quickly determine the key TCC performance metrics, in terms of bit error rate (BER), signal to noise ratio (SNR), and channel capacity, without going through lengthy computer simulation and/or physical experiments that are normally needed in current TCC performance studies. According to our model, the TCC’s BER is proportional to the square root of the transmission frequency, which can be explored quantitatively to boost the TCC’s transmission efficiency by letting the TCC’s thermal signal be transmitted at a higher frequency. In addition, our proposed model also links the jamming noise and application of Dynamic Voltage Frequency Scaling (DVFS) to TCC’s BER performance, a feature that can be explored to design/optimize the countermeasures against the TCC attacks. The TCC performance predicted by the proposed theoretical model is found in a good agreement with that obtained from computer simulations, with an average error lower than 7%.
Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Computers4
2022 On Evaluation of On-chip Thermal Covert Channel Attacks
abstract
Thermal covert channel (TCC) attacks have been a serious security concern to the use of many-core chips. Severity of these attacks is directly linked to the TCC’s transmission rate and its BER (bit error rate) performance, both of which are impacted by the transmission characteristics of thermal signals and adopted encoding, modulation, and multiplexing schemes. This paper examines, compares, and analyzes various TCCs built upon different combinations of encoding, modulation, and multiplexing. In particular, our study shows that TCC using non-return-to-zero (NRZ) line coding and frequency shift keying (FSK) modulation achieves the highest throughput of 120 bps and BER of below 10%.
Jiachen Wang 0011, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
CASES4
2022 Performance Optimization of Many-Core Systems by Exploiting Task Migration and Dark Core Allocation
abstract
As an effective scheme often adopted for performance tuning in many-core processors, task migration provides an opportunity for “hot” tasks to be migrated to run on a “cool” core that has a lower temperature. When a task needs to migrate from one processor core to another, the migration can embark on numerous modes defined by the migration paths undertaken and/or the destinations of the migration. Selecting the right migration mode that a task shall follow has always been difficult, and it can be more challenging with the existence of dark cores that can be called back to service (reactivated), which ushers in additional task migration modes. Previous works have demonstrated that dark cores can be placed near the active cores to reduce power density so that the active cores can run at higher voltage/frequency levels for higher performance. However, the existing task migration schemes neither consider the impact of dark cores on each application's performance, nor exploit performance trade-off under different migration modes. Unlike the existing task migration schemes, in this article, a runtime task migration algorithm that simultaneously takes both migration modes and dark cores into consideration is proposed, and it essentially has two major steps. In the first step, for a specific migration mode that is tied to an application whose tasks need to be migrated, the number of dark cores is determined so that the overall performance is maximized. The second step is to find an appropriate core region and its location for each application to optimize the communication latency and computation performance; during this step, focus is placed on reducing the fragmentation of the free core regions resulting from the task migration. Experimental results have confirmed that our approach achieves over 50 percent reduction in total response time when compared to recently proposed thermal-aware runtime task migration approachess.
Shengyan Wen, Xiaohang Wang 0001, Amit Kumar Singh 0002, Yingtao Jiang, Mei Yang 0001
IEEE Trans. Computers3
2022 Detection of and Countermeasure Against Thermal Covert Channel in Many-Core Systems
abstract
The thermal covert channels (TCCs) in many-core systems can cause detrimental data breaches. In this article, we present a three-step scheme to detect and fight against such TCC attacks. Specifically, in the detection step, each core calculates the spectrum of its own CPU workload traces that are collected over a few fixed time intervals, and then it applies a frequency scanning method to detect if there exists any TCC attack. In the next positioning step, the logical cores running the transmitter threads are located. In the last step, the physical CPU cores suspiciously engaging in a TCC attack have to undertake dynamic voltage frequency scaling (DVFS) such that any possible TCC trace will be essentially wiped out. Our experiments have confirmed that on average 97% of the TCC attacks can be detected, and with the proposed defense, the packet error rate (PER) of a TCC attack can soar to more than 70%, literally shutting down the attack in practical terms. The performance penalty caused by the inclusion of the proposed DVFS countermeasures is found to be only 3% for an$8\times 8$many-core system.
Hengli Huang, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Combating Stealthy Thermal Covert Channel Attack With Its Thermal Signal Transmitted in Direct Sequence Spread Spectrum
abstract
Many-core systems are susceptible to attacks launched by thermal covert channel (TCC) attacks. Detection of TCC attacks often relies on the use of threshold-based approaches or variants, and a countermeasure to thwart the channel can be applied only after an attack is deemed to be present. In this article, we describe a direct sequence spread spectrum (DSSS)-based TCC, where its thermal data are modulated by a pseudo-random bit sequence. Unfortunately, such DSSS-based TCC has an extremely low signal strength that the signal is nearly indistinguishable from the noise and thus cannot be detected by any existing threshold-based detection methods. To combat this stealthy TCC, we propose a novel detection scheme that lets the received signal pass through a differential filter where irrelevant frequency components occupied mainly by the noise gets eliminated and the filtered signal is next compared against a threshold for successful detection. Experimental results show that the DSSS-based TCC can effectively survive detection by the existing detection methods with its BER as low as 4%. In contrast, with the proposed detection and countermeasure applied, the detection accuracy jumps to 89%, and the BER of the DSSS-based TCC soars to 50%, which indicates that the TCC is practically shut down.
Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Secured Data Transmission Over Insecure Networks-on-Chip by Modulating Inter-Packet Delays
abstract
As the network-on-chip (NoC) integrated into an SoC design can come from an untrusted third party, there is a growing risk that data integrity and security get compromised when supposedly sensitive data flows through such an untrusted NoC. We thus introduce a new method that can ensure secure and secret data transmission over such an untrusted NoC. Essentially, the proposed scheme relies on encoding binary data as delays between packets travelling across the source and destination pair. The maximum data transmission rate of this inter-packet-delay (IPD)-based communication channel can be determined from the analytical model developed in this article. To further improve the undetectability and robustness of the proposed data transmission scheme, a new block coding method and communication protocol are also proposed. Experimental results show that the proposed IPD-based method can achieve a packet error rate (PER) of as low as 0.3% and an effective throughput of$\boldsymbol {2.3\times 10^{5}}$b/s, outperforming the methods of thermal covert channel, cache covert channel, and circuit-based encryption and, thus, is suitable for secure data transmission in unsecure systems.
Jiaen Xu, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Chongyan Gu, Letian Huang, Mei Yang 0001, Shunbin Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 QUAREM: Maximising QoE Through Adaptive Resource Management in Mobile MPSoC Platforms
abstract
Heterogeneous multi-processor system-on-chip (MPSoC) smartphones are required to offer increasing performance and user quality-of-experience (QoE) , despite comparatively slow advances in battery technology. Approaches to balance instantaneous power consumption, performance and QoE have been reported, but little research has considered how to perform longer-term budgeting of resources across a complete battery discharge cycle. Approaches that have considered this are oblivious to the daily variability in the user’s desired charging time-of-day (plug-in time), resulting in a failure to meet the user’s battery life expectations, or else an unnecessarily over-constrained QoE. This paper proposes QUAREM, an adaptive resource management approach in mobile MPSoC platforms that maximises QoE while meeting battery life expectations. The proposed approach utilises a model that learns and then predicts the dynamics of the energy usage pattern and plug-in times. Unlike state-of-the-art approaches, we maximise the QoE through the adaptive balancing of the battery life and the quality of service (QoS) for the duration of the battery discharge. Our model achieves a good degree of accuracy with a mean absolute percentage error of 3.47% and 2.48% for the energy demand and plug-in times, respectively. Experimental evaluation on an off-the-shelf commercial smartphone shows that QUAREM achieves the expected battery life of the user within 20–25% energy demand variation with little or no QoE degradation.
Samuel Isuwa, Somdip Dey, Andre P. Ortega, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett
ACM Trans. Embed. Comput. Syst.4
2021 I2UTS: An IoT based Intelligent Urban Traffic System
abstract
Growing population and migration to cities have given birth to multiple urban issues. Traffic congestion is one of the most prominent ones with severe side effects like fuel wastage, loss of lives, and slow productivity. The traditional traffic control system deploys programming logic control (PLC) which uses round-robin scheduling algorithm. However, few recent works have proposed IoT-based framework which requires the deployment of a series of sensors. In this paper, we propose an IoT-based framework that uses the existing network of CCTV cameras at the junction. An edge device is used to estimate the traffic density and detect emergency vehicles using YOLO v3 -Efficient Net. These two parameters are used as an input to a novel traffic control algorithm. The performance of the proposed framework has been evaluated by analyzing its properties using the UA-DETRAC dataset. The proposed framework achieves 68.10% vehicle detection accuracy.
Vejey Pradeep Suresh Achari, Zeba Khanam, Amit Kumar Singh 0002, Anish Jindal, Alok Prakash, Neeraj Kumar 0001
HPSR3
2021 An enhanced planned obsolescence attack by aging networks-on-chip
Yinyuan Zhao, Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
J. Syst. Archit.5
2021 Longevity Framework: Leveraging Online Integrated Aging-Aware Hierarchical Mapping and VF-Selection for Lifetime Reliability Optimization in Manycore Processors
abstract
Rapid device aging in the nano era threatens system lifetime reliability, posing a major intrinsic threat to system functionality. Traditional techniques to overcome the aging-induced device slowdown, such as guardbanding are static and incur performance, power, and area penalties. In a manycore processor, the system-level design abstraction offers dynamic opportunities through the control of task-to-core mappings and per-core operation frequency towards more balanced core aging profile across the chip, optimizing the system lifetime reliability while meeting the application performance requirements. This article presents Longevity Framework (LF) that leverages online integrated aging-aware hierarchical mapping and voltage frequency (VF)-selection for lifetime reliability optimization in manycore processors. The mapping exploration is hierarchical to achieve scalability. The VF-selection builds on the trade-offs involved between power, performance, and aging as the VF is scaled while leveraging the per-core DVFS capabilities. The methodology takes the chip-wide process variation into account. Extensive experimentation, comparing the proposed approach with two state-of-the-art methods, for 64-core and 256-core systems running applications from PARSEC and SPLASH-2 benchmark suites, show an improvement of up to 3.2 years in the system lifetime reliability and 4× improvement in the average core health.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001
IEEE Trans. Computers3
2021 On Performance Optimization and Quality Control for Approximate-Communication-Enabled Networks-on-Chip
abstract
For many applications showing error forgiveness, approximate computing is a new design paradigm that trades application output accuracy for mitigating computation/communication effort, which results in performance/energy benefit. Since networks-on-chip (NoCs) are one of the major contributors to system performance and power consumption, the underlying communication is approximated to achieve time/energy improvement. However, performing approximation blindly causes unacceptable quality loss. In this article, first, an optimization problem to maximize NoC performance is formulated with the constraint of application quality requirement, and the application quality loss is studied. Second, a congestion-aware quality control method is proposed to improve system performance by aggressively dropping network data, which is based on flow prediction and a lightweight heuristic. In the experiments, two recent approximation methods for NoCs are augmented with our proposed control method to compare with their original ones. Experimental results show that our proposed method can speed up execution by as much as 29.42% over the two state-of-the-art works.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Liang Wang 0020, Terrence S. T. Mak
IEEE Trans. Computers4
2020 Temporal Motionless Analysis of Video using CNN in MPSoC
abstract
This paper proposes a novel human-inspired methodology called IRON-MAN (Integrated RatiONal prediction and Motionless ANalysis of videos) on mobile multi-processor systems-on-chips (MPSoCs). The methodology integrates analysis of the previous image frames of the video to represent the analysis of the current frame in order to perform Temporal Motionless Analysis of the Video (TMAV). This is the first work on TMAV using Convolutional Neural Network (CNN) for scene prediction in MPSoCs. Experimental results show that our methodology outperforms state-of-the-art. We also introduce a metric named, Energy Consumption per Training Image (ECTI) to assess the suitability of using a CNN model in mobile MPSoCs with a focus on energy consumption of the device.
Somdip Dey, Amit Kumar Singh 0002, Dilip K. Prasad, Klaus D. McDonald-Maier
ASAP2
2020 On Countermeasures Against the Thermal Covert Channel Attacks Targeting Many-core Systems
abstract
Although it has been demonstrated in multiple studies that serious data leaks could occur to many-core systems thanks to the existence of the thermal covert channels (TCC), little has been done to produce effective countermeasures that are necessary to fight against such TCC attacks. In this paper, we propose a three-step countermeasure to address this critical defense issue. Specifically, the countermeasure includes detection based on signal frequency scanning, positioning affected cores, and blocking based on Dynamic Voltage Frequency Scaling (DVFS) technique. Our experiments have confirmed that on average 98% of the TCC attacks can be detected, and with the proposed defense, the bit error rate of a TCC attack can soar to 92%, literally shutting down the attack in practical terms. The performance penalty caused by the inclusion of the proposed countermeasures is only 3% for an 8×8 system.
Hengli Huang, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Letian Huang
DAC4
2020 User Interaction Aware Reinforcement Learning for Power and Thermal Efficiency of CPU-GPU Mobile MPSoCs
abstract
Mobile user’s usage behaviour changes throughout the day and the desirable Quality of Service (QoS) could thus change for each session. In this paper, we propose a QoS aware agent to monitor mobile user’s usage behaviour to find the target frame rate, which satisfies the desired user’s QoS, and applies reinforcement learning based DVFS on a CPU-GPU MPSoC to satisfy the frame rate requirement. Experimental study on a real Exynos hardware platform shows that our proposed agent is able to achieve a maximum of 50% power saving and 29% reduction in peak temperature compared to stock Android’s power saving scheme. It also outperforms the existing state-of-the-art power and thermal management scheme by 41% and 19%, respectively.
Somdip Dey, Amit Kumar Singh 0002, Xiaohang Wang 0001, Klaus D. McDonald-Maier
DATE2
2020 On hardware-trojan-assisted power budgeting system attack targeting many core systems
Xiaohang Wang 0001, Yingtao Jiang, Liang Wang 0020, Mei Yang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
J. Syst. Archit.6
2020 Collaborative Adaptation for Energy-Efficient Heterogeneous Mobile SoCs
abstract
Heterogeneous Mobile System-on-Chips (SoCs) containing CPU and GPU cores are becoming prevalent in embedded computing, and they need to execute applications concurrently. However, existing run-time management approaches do not perform adaptive mapping and thread-partitioning of applications while exploiting both CPU and GPU cores at the same time. In this paper, we propose an adaptive mapping and thread-partitioning approach for energy-efficient execution of concurrent OpenCL applications on both CPU and GPU cores while satisfying performance requirements. To start execution of concurrent applications, the approach makes mapping (number of cores and operating frequencies) and partitioning (distribution of threads between CPU and GPU) decisions to satisfy performance requirements for each application. The mapping and partitioning decisions are made by having a collaboration between the CPU and GPU cores' processing capabilities such that balanced execution can be performed. During execution, adaptation is triggered when new application(s) arrive, or an executing one finishes, that frees cores. The adaptation process identifies a new mapping and thread-partitioning in a similar collaborative manner for remaining applications provided it leads to an improvement in energy efficiency. The proposed approach is experimentally validated on the Odroid-XU3 hardware platform with varying set of applications. Results show an average energy saving of 37%, compared to existing approaches while satisfying the performance requirements.
Amit Kumar Singh 0002, Basireddy Karunakar Reddy, Alok Prakash, Geoff V. Merrett, Bashir M. Al-Hashimi
IEEE Trans. Computers1
2020 AdaMD: Adaptive Mapping and DVFS for Energy-Efficient Heterogeneous Multicores
abstract
Modern heterogeneous multicore systems, containing various types of cores, are increasingly dealing with concurrent execution of dynamic application workloads. Moreover, the performance constraints of each application vary, and applications enter/exit the system at any time. Existing approaches are not efficient in such dynamic scenarios, especially if applications are unknown, as they require extensive offline application analysis and do not consider the runtime execution scenarios (application arrival/completion, and workload and performance variations) for runtime management. To address this, we present AdaMD, an adaptive mapping and dynamic voltage and frequency scaling (DVFS) approach for improving energy consumption and performance. The key feature of the proposed approach is the elimination of dependency on offline profiled results while making runtime decisions. This is achieved through a performance prediction model having a maximum error of 7.9% lower than the previously reported model and a mapping approach that allocates processing cores to applications while respecting performance constraints. Furthermore, AdaMD adapts to runtime execution scenarios efficiently by monitoring the application status, and performance/workload variations to adjust the previous DVFS settings and thread-to-core mappings. The proposed approach is experimentally validated on the Odroid-XU3, with various combinations of diverse multithreaded applications from PARSEC and SPLASH benchmarks. Results show energy savings of up to 28% compared to the recently proposed approach while meeting performance constraints.
Basireddy Karunakar Reddy, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Combating Enhanced Thermal Covert Channel in Multi-/Many-Core Systems With Channel-Aware Jamming
abstract
As a means to thwart thermal covert channel attack in a multi-/many-core system, a strong heat noise whose frequency band coincides with that occupied by the thermal covert channel is injected to jam the channel. However, this undiscriminating channel jamming-based countermeasure will fail if a thermal covert channel is allowed to change its transmission frequency dynamically in response to the jamming. To combat this enhanced thermal covert channel, a more advanced countermeasure is needed and thus proposed that checks the frequency spectrum and tracks any possible covert channel. Only after a channel is detected to be susceptible, a thermal noise with this channel frequency is then emitted to jam the covert channel. The communication protocols and frequency changing scheme pertaining to this enhanced thermal covert channel are described in this article. The experimental results confirm that, when the proposed countermeasure is applied, the enhanced thermal covert channel, much more resilient to jamming, suffers from an extremely high packet error rate (PER), which makes any meaningful data leakage practically impossible. As the proposed countermeasure method is poised to contain dangerous thermal covert channel attacks with an anti-jamming capability, it lends itself well to secure multi-/many-core systems.
Jiachen Wang 0011, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Letian Huang, Mei Yang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Energy Minimization for Multicore Platforms Through DVFS and VR Phase Scaling With Comprehensive Convex Model
abstract
Energy management is a critical challenge in multicore processors due to continuous technology scaling. Previous methods have mostly focused on the energy minimization of the processor cores. However, energy overhead of the off-chip voltage regulator (VR) has recently shown to be a nontrivial part of the total energy consumption and has been previously overlooked. In this paper, we propose an overall energy optimization method for the system that minimizes both per-core energy consumption and VR energy consumption using dynamic voltage frequency scaling and VR phase scaling by solving a comprehensive convex model. In order to improve the accuracy of the task latency model, a new task model considering both computation and memory access of the task is also developed. Furthermore, for better scalability and lower online overhead, we decompose our proposed convex method into two stages: 1) an offline stage and 2) an online stage. During the offline stage, we explore the convex model by assuming different numbers of active phases of the VR, various workload pressures and workload characteristics to collect the optimal frequency assignments under different scenarios. During the online stage, the specific frequency assignment for cores and optimal active phase number of the VR are selected and applied based on the actual workload pressure and its characteristics running on the cores. Experiments on real benchmarks show that when compared with the state-of-the-art approaches, which are oblivious to VR overheads and exploit slack time to achieve energy minimization, our method can achieve a significant energy saving of up to 22.4% with negligible online overhead.
Zuomin Zhu, Wei Zhang 0012, Vivek Chaturvedi, Amit Kumar Singh 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 LifeGuard: A Reinforcement Learning-Based Task Mapping Strategy for Performance-Centric Aging Management
abstract
Device scaling to subdeca nanometer has pushed device aging as a primary design concern. In manycore systems, inevitable process variation further adds to delay degradation and, coupled with the scalability issues in manycores, makes aging management, while meeting performance demands, a complex problem. LifeGuard is a performance-centric reinforcement learning-based task mapping strategy that leverages the different impact of applications on aging for improving system health. Experimental results, comparing LifeGuard with two state-of-the-art aging optimizing techniques, on a 256-core system, showed that LifeGuard led to improved health for, respectively, 57% and 74% of the cores, and also an enhanced aggregate core frequency.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001
DAC3
2019 TEEM: Online Thermal- and Energy-Efficiency Management on CPU-GPU MPSoCs
abstract
Heterogeneous Multiprocessor System-on-Chip (MPSoC) are progressively becoming predominant in most modern mobile devices. These devices are required to perform processing of applications within thermal, energy and performance constraints. However, most stock power and thermal management mechanisms either neglect some of these constraints or rely on frequency scaling to achieve energy-efficiency and temperature reduction on the device. Although this inefficient technique can reduce temporal thermal gradient, but at the same time hurts the performance of the executing task. In this paper, we propose a thermal and energy management mechanism which achieves reduction in thermal gradient as well as energy-efficiency through resource mapping and thread-partitioning of applications with online optimization in heterogeneous MPSoCs. The efficacy of the proposed approach is experimentally appraised using different applications from Polybench benchmark suite on Odroid-XU4 developmental platform. Results show 28% performance improvement, 28.32% energy saving and reduced thermal variance of over 76% when compared to the existing approaches. Additionally, the method is able to free more than 90% in memory storage on the MPSoC, which would have been previously utilized to store several task-to-thread mapping configurations.
Samuel Isuwa, Somdip Dey, Amit Kumar Singh 0002, Klaus D. McDonald-Maier
DATE3
2019 ACDC: An Accuracy- and Congestion-aware Dynamic Traffic Control Method for Networks-on-Chip
abstract
Many applications exhibit error forgiving features. For these applications, approximate computing provides the opportunity of accelerating the execution time or reducing power consumption, by mitigating computation effort to get an approximate result. Among the components on a chip, network-on-chip (NoC) contributes a large portion to system power and performance. In this paper, we exploit the opportunity of aggressively reducing network congestion and latency by selectively dropping data. Essentially, the importance of the dropped data is measured based on a quality model. An optimization problem is formulated to minimize the network congestion with constraint of the result quality. A lightweight online algorithm is proposed to solve this problem. Experiments show that on average, our proposed method can reduce the execution time by as much as 12.87% and energy consumption by 12.42% under strict quality requirement, speed up execution by 19.59% and reduce energy consumption by 21.20% under relaxed requirement, compared to a recent work on approximate computing approach for NoCs.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Terrence S. T. Mak
DATE4
2019 Towards Scalable Lifetime Reliability Management for Dark Silicon Manycore Systems
abstract
Aggressive technology scaling enabled very high integration density. Unfortunately, it also led to issues such as process variation, increased power density and consequently rising chip temperature resulting in accelerated device aging and poor lifetime reliability of different components in a manycore system. Moreover, thermal and power limitations let only a fraction of the chip function at full speed; the rest is the dark silicon. Most of the lifetime reliability enhancement solutions for the multi-/manycore systems in the literature are heuristic-based, while some use standard compute-intensive methods to solve the optimization problem making them not scale well with the manycore size. The heuristic-based solutions are formulated to search through the design space of a fine granularity making it huge, limiting their scalability. Also, these approaches do not account for the impact of different applications' execution behavior on the aging of the underlying cores, and their performance requirement distribution across the cores to their advantage. In this paper, we present our resource management strategies towards building scalable lifetime reliability enhancement solutions for dark silicon manycore systems. The first technique, Hierarchical Mapping approach (HiMap), maps a periodic workload employing a block-based hierarchical method that leverages dark cores for thermal mitigation. The second approach, LifeGuard, uses reinforcement learning to learn the applications' aging behavior, and is aware of the performance requirement pattern onto the core frequencies. It maps randomly arriving requests and is scalable to the number of applications and the size of a manycore.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001
IOLTS3
2019 On Runtime Communication and Thermal-Aware Application Mapping and Defragmentation in 3D NoC Systems
abstract
Many-core systems connected by 3D Networks-on-Chip (NoC) are emerging as a promising computation engine for systems like cloud computing servers, big data systems, etc. Mapping applications at runtime to 3D NoCs is the key to maintain high throughput of the overall chip under a thermal/power constraint. However, the goals of optimizing both the communication latency and chip peak temperature are contradicting due to several reasons. First, exploiting the vertical TSV links can accelerate communications, while low peak temperature prefers that the tasks to be mapped closer to the heat sink, instead of using the vertical links. Second, mapping tasks in close proximity can reduce communication latency, but at the cost of poor heat dissipation. To address these issues, in this paper, we propose an efficient runtime mapping algorithm to reduce both communication latency and overall application running time under thermal constraint. In essence, this algorithm first selects a 3D cuboid core region of a specific shape for each incoming application by setting the region's number of occupied vertical layers and its distance to the heat sink, in order to optimize its communication performance and peak temperature. Next, the exact locations of the core regions in the chip are determined, followed by a task-to-core mapping. A defragmentation algorithm is also proposed to keep free core regions contiguous. The experimental results have confirmed that, compared to two recently proposed runtime mapping algorithms, our proposed approach can reduce the total running time by up to 48% and communication cost by up to 44%, with a low runtime overhead.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
IEEE Trans. Parallel Distributed Syst.3
2019 Predictive Thermal Management for Energy-Efficient Execution of Concurrent Applications on Heterogeneous Multicores
abstract
Current multicore platforms contain different types of cores, organized in clusters (e.g., ARM's big.LITTLE). These platforms deal with concurrently executing applications, having varying workload profiles and performance requirements. Runtime management is imperative for adapting to such performance requirements and workload variabilities and to increase energy and temperature efficiency. Temperature has also become a critical parameter since it affects reliability, power consumption, and performance and, hence, must be managed. This paper proposes an accurate temperature prediction scheme coupled with a runtime energy management approach to proactively avoid exceeding temperature thresholds while maintaining performance targets. Experiments show up to 20% energy savings while maintaining high-temperature averages and peaks below the threshold. Compared with state-of-the-art temperature predictors, this paper predicts 35% faster and reduces the mean absolute error from 3.25 to 1.15 °C for the evaluated applications' scenarios.
Eduardo Wächter, Cedric de Bellefroid, Basireddy Karunakar Reddy, Amit Kumar Singh 0002, Bashir M. Al-Hashimi, Geoff V. Merrett
IEEE Trans. Very Large Scale Integr. Syst.4
2018 HiMap: A hierarchical mapping approach for enhancing lifetime reliability of dark silicon manycore systems
abstract
Technology scaling into the nano-scale CMOS regime has resulted in increased leakage and roadblock on voltage scaling, which has led to several issues like high power density and elevated on-chip temperature. This consequently aggravates device aging, compromising lifetime reliability of the manycore systems. This paper proposes HiMap, a dynamic hierarchical mapping approach to maximize lifetime reliability of manycore systems while satisfying performance, power, and thermal constraints. HiMap is process variation- and aging-aware. It comprises of two levels: (1) it identifies a region of cores suitable for mapping, and (2) it maps threads in the region and intersperses dark cores for thermal mitigation while considering the current health of the cores. Both the levels strive to reduce aging variance across the chip. We evaluated HiMap for 64-core and 256-core systems. Results demonstrate an improved system lifetime reliability by up to 2 years at the end of 3.25 years of use, as compared to the state-of-the-art.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, R. Rohith, Siew-Kei Lam, Muhammad Shafique 0001
DATE3
2018 Online concurrent workload classification for multi-core energy management
abstract
Modern embedded multi-core processors are organized as clusters of cores, where all cores in each cluster operate at a common Voltage-frequency (V-f). Such processors often need to execute applications concurrently, exhibiting varying and mixed workloads (e.g. compute- and memory-intensive) depending on the instruction mix and resource sharing. Runtime adaptation is key to achieving energy savings without trading-off application performance with such workload variabilities. In this paper, we propose an online energy management technique that performs concurrent workload classification using the metric Memory Reads Per Instruction (MRPI) and pro-actively selects an appropriate V-fsetting through workload prediction. Subsequently, it monitors the workload prediction error and performance loss, quantified by Instructions Per Second (IPS) at runtime and adjusts the chosen V-fto compensate. We validate the proposed technique on an Odroid-XU3 with various combinations of benchmark applications. Results show an improvement in energy efficiency of up to 69% compared to existing approaches.
Basireddy Karunakar Reddy, Geoff V. Merrett, Bashir M. Al-Hashimi, Amit Kumar Singh 0002
DATE4
2018 Exploiting Dark Cores for Performance Optimization via Patterning for Many-core Chips in the Dark Silicon Era
abstract
All the cores of a many-core chip cannot be active at the same time, due to reasons like low CPU utilization in server systems and limited power budget in dark silicon era. These free cores (referred to as bubbles) can be placed near active cores for heat dissipation so that the active cores can run at a higher frequency level, boosting the performance of active cores and applications. In the literature, this approach is referred as static patterning. Patterning for performance boost has the following challenges. First, communication distance increases when a bubble is inserted between two communicating tasks, leading to performance degradation. Second, budgeting too many bubbles as cooler to running applications leads to insufficient cores for future applications. In addition, task-migration-based dynamic patterning can further improve the performance of the system. In this paper, a static and a dynamic patterning approaches are proposed to budget free cores to each application so as to optimize the throughput of the whole system. Essentially, the proposed static patterning algorithm determines the number and locations of bubbles to optimize the performance and waiting time of each application, followed by tasks of each application being mapped to a core region. The dynamic patterning algorithm first selects the best pattern, the bubble number and the core region shape for each application that results in maximal performance, followed by choosing the location for each application's core region. Experiments show that our approach achieves 50% higher throughput when compared to state-of-the-art thermal-aware runtime task mapping approaches.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Shengyan Wen
NOCS2
2018 Effectiveness of HT-assisted sinkhole and blackhole denial of service attacks targeting mesh networks-on-chip
Xiaohang Wang 0001, Yingtao Jiang, Mei Yang 0001, Terrence S. T. Mak, Amit Kumar Singh 0002
J. Syst. Archit.6
2018 Bubble Budgeting: Throughput Optimization for Dynamic Workloads by Exploiting Dark Cores in Many Core Systems
abstract
All the cores of a many-core chip cannot be active at the same time, due to reasons like low CPU utilization in server systems and limited power budget in dark silicon era. These free cores (referred to as bubbles) can be placed near active cores for heat dissipation so that the active cores can run at a higher frequency level, boosting the performance of applications that run on active cores. Budgeting inactive cores (bubbles) to applications to boost performance has the following three challenges. First, the number of bubbles varies due to open workloads. Second, communication distance increases when a bubble is inserted between two communicating tasks (a task is a thread or process of a parallel application), leading to performance degradation. Third, budgeting too many bubbles as coolers to running applications leads to insufficient cores for future applications. In order to address these challenges, in this paper, a bubble budgeting scheme is proposed to budget free cores to each application so as to optimize the throughput of the whole system. Throughput of the system depends on the execution time of each application and the waiting time incurred for newly arrived applications. Essentially, the proposed algorithm determines the number and locations of bubbles to optimize the performance and waiting time of each application, followed by tasks of each application being mapped to a core region. A Rollout algorithm is used to budget power to the cores as the last step. Experiments show that our approach achieves 50 percent higher throughput when compared to state-of-the-art thermal-aware runtime task mapping approaches. The runtime overhead of the proposed algorithm is in the order of 1M cycles, making it an efficient runtime task management method for large-scale many-core systems.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
IEEE Trans. Computers2
2017 Two-stage thermal-aware scheduling of task graphs on 3D multi-cores exploiting application and architecture characteristics
abstract
In this paper, we propose a two-stage thermal-aware task scheduling policy which exploits the application and system architecture characteristics to decouple the mapping of task-graphs for the performance and peak temperature optimization into two stages. At the first stage, the algorithm collects the best mapping of task-graphs exploiting the application and architecture characteristics to minimize the makespan of the task-graphs. At the second stage, a light-weight online algorithm comprised of efficient thermal rank and combined power models is performed to map the task nodes to the real cores for temperature minimization while maintaining the best possible performance achieved in the first stage. Compared to the previous approaches which perform the performance and temperature optimization together, our method can reduce the online mapping algorithm complexity and improve its efficiency. Experiments on real benchmarks show that an average of 6.3°C peak temperature reduction and 6.8% performance improvement can be achieved compared to other existing methods.
Zuomin Zhu, Vivek Chaturvedi, Amit Kumar Singh 0002, Wei Zhang 0012, Yingnan Cui
ASP-DAC3
2017 On Runtime Communication- and Thermal-aware Application Mapping in 3D NoC
abstract
Many-core systems connected by 3D Network-on-Chips (NoC) are emerging as a promising computation engine for systems like cloud computing servers, big data systems, etc. Mapping applications at runtime to 3D NoCs is the key to maintain high throughput of the overall chip under a thermal/power constraint. However, the goals of optimizing both the communication latency and chip peak temperature are contradicting due to several reasons. Firstly, exploiting the vertical TSV links can accelerate communications, while low peak temperature prefers that the tasks to be mapped closer to the heat sink, instead of using the vertical links. Secondly, mapping tasks in close proximity can reduce communication latency, but at the cost of poor heat dissipation. To address these issues, in this paper, we propose an efficient runtime mapping algorithm to reduce both communication latency and overall application running time under thermal constraint. In essence, this algorithm first selects a 3D cuboid core region of a specific shape for each incoming application by setting the region's number of occupied vertical layers and its distance to the heat sink, in order to optimize its communication performance and peak temperature. Next, the exact locations of the core regions in the chip are determined, followed by a task-to-core mapping. The experimental results have confirmed that, compared to two recently proposed runtime mapping algorithms, our proposed approach can reduce the total running time by up to 48% and communication cost by up to 44%, with a low runtime overhead.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
NOCS3
2017 Energy-Efficient Run-Time Mapping and Thread Partitioning of Concurrent OpenCL Applications on CPU-GPU MPSoCs
abstract
Heterogeneous Multi-Processor Systems-on-Chips (MPSoCs) containing CPU and GPU cores are typically required to execute applications concurrently. However, as will be shown in this paper, existing approaches are not well suited for concurrent applications as they are developed either by considering only a single application or they do not exploit both CPU and GPU cores at the same time. In this paper, we propose an energy-efficient run-time mapping and thread partitioning approach for executing concurrent OpenCL applications on both GPU and GPU cores while satisfying performance requirements. Depending upon the performance requirements, for each concurrently executing application, the mapping process finds the appropriate number of CPU cores and operating frequencies of CPU and GPU cores, and the partitioning process identifies an efficient partitioning of the applications’ threads between CPU and GPU cores. We validate the proposed approach experimentally on the Odroid-XU3 hardware platform with various mixes of applications from the Polybench benchmark suite. Additionally, a case-study is performed with a real-world application SLAMBench. Results show an average energy saving of 32% compared to existing approaches while still satisfying the performance requirements.
Amit Kumar Singh 0002, Alok Prakash, Basireddy Karunakar Reddy, Geoff V. Merrett, Bashir M. Al-Hashimi
ACM Trans. Embed. Comput. Syst.1
2016 Energy-Aware Resource Allocation in Multi-mode Automotive Applications with Hard Real-Time Constraints
abstract
This paper presents an energy aware resource allocation approach that benefits from modal nature of hard real-time systems under consideration. The modal nature of considered applications made it possible to decrease the number of active cores consuming high power in certain modes or to switch into core states with lower power consumption, which lead to considerable energy savings while still not violating any of timing constraints. For the considered automotive use case, the number of required cores has been decreased by up to 75% in a particular mode and relatively low amount of data is to be migrated during the mode change. The trade-off between the amount of data to be migrated and energy dissipation in the subsequent state is also analysed.
Piotr Dziurzanski, Amit Kumar Singh 0002, Leandro Soares Indrusiak
ISORC2
2016 Value and Energy Aware Adaptive Resource Allocation of Soft Real-Time Jobs on Many-Core HPC Data Centers
abstract
Modern high performance computing (HPC) data centers consume huge energy to operate them. Therefore, appropriate measures are required to reduce their energy consumption. Existing efforts for such measures focus on consolidation and dynamic voltage and frequency scaling (DVFS). However, most of them do not perform adaptive resource allocation for the executing dependent tasks (or jobs) in order to optimize both value and energy. The value is achieved by completing the execution of a job and it depends on the completion time. A high value is achieved if the job is completed before its deadline, otherwise a lower value. In this paper, we propose an adaptive resource allocation approach that uses design-time profiling results of jobs for efficient allocation and adaptation in order to optimize both value and energy while executing dependent tasks. The profiling results for each job are obtained by exploiting efficient allocation combined with identification of voltage/frequency levels of used system cores and used in adapting to different number of cores based on the monitored execution progress of the job and available cores. Experiments show that the proposed approach enhances the overall value by about 10% when compare to existing approaches while showing reduction in energy consumption and percentage of rejected jobs leading to zero value.
Amit Kumar Singh 0002, Piotr Dziurzanski, Leandro Soares Indrusiak
ISORC1
2016 Bubble budgeting: throughput optimization for dynamic workloads by exploiting dark cores in many core systems
abstract
All the cores of a many-core chip cannot be active at the same time, due to reasons like low CPU utilization in server systems and limited power budget in dark silicon era. These free cores (referred to as bubbles) can be placed near active cores for heat dissipation so that the active cores can run at a higher frequency level, boosting the performance of active cores and applications. Budgeting inactive cores (bubbles) to workloads to boost performance has the following three challenges. First, the number of bubbles varies due to dynamic workloads. Second, communication distance increases when a bubble is inserted between two communicating tasks, leading to performance degradation. Third, budgeting too many bubbles as cooler to running applications leads to insufficient cores for future applications. In order to address these challenges, in this paper, a bubble budgeting scheme is proposed to budget free cores to each application so as to optimize the throughput of the whole system, including the execution time of each application and the waiting time incurred for newly arrived applications. Essentially, the proposed algorithm determines the number and locations of bubbles to optimize the performance and waiting time of each application, followed by tasks of each application being mapped to a core region. Experiments show that our approach achieves 50% higher throughput when compared to state-of-the-art thermal-aware runtime task mapping approaches.
Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
NOCS2
2016 Resource and Throughput Aware Execution Trace Analysis for Efficient Run-Time Mapping on MPSoCs
abstract
There have been several efforts on run-time mapping of applications on multiprocessor-systems-on-chip. These traditional efforts perform either on-the-fly processing or use design-time analyzed results. However, on-the-fly processing often leads to low-quality mappings, and design-time analysis becomes computationally costly for large-size problems and require huge storage for large number of applications. In this paper, we present a novel run-time mapping approach, where identification of an efficient mapping for a use-case is done by the online execution trace analysis of the active applications. The trace analysis facilitates for fast identification of the mapping while optimizing for the system resource usage and throughput of the active applications, leading to reduced energy consumption as well. By rapidly identifying the efficient mapping at run-time, the proposed approach overcomes the mappings' exploration time bottleneck for large-size problems and their storage overhead problem when compared to the traditional approaches. Our experiments show that on average the exploration time to identify the mapping is reduced $14 {\times }$ when compared to state-of-the-art approaches and storage overhead is reduced by 92%. Additionally, energy and resource savings are achieved along with identification of high-quality mapping.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Defragmentation for Efficient Runtime Resource Management in NoC-Based Many-Core Systems
abstract
Efficient runtime resource allocation is critical to the overall performance and energy consumption of many-core systems. A region of free cores is allocated for each newly launched application. The cores are deallocated when the corresponding applications finish execution. The frequent allocations and deallocations of the cores might leave free cores scattered (not forming a contiguous region). This situation is referred to as fragmentation. Fragmentation could cause the inefficient mapping of the incoming applications, i.e., long communication distance between communicating cores. This further leads to poor performance and high energy consumption. In this paper, we propose a runtime defragmentation scheme that collects and reallocates the scattered cores in close proximity. We first define a fragmentation metric that is able to evaluate the scatteredness level of the free cores. Based on this, the proposed algorithm is executed to bring the scattered free cores together when the fragmentation metric is over a certain predefined threshold. In this way, the contiguous free core region is formed to facilitate the efficient mapping of the incoming applications. Moreover, the proposed algorithm also aims to minimize the negative impact on the performance of existing applications. Experimental results show that the proposed defragmentation scheme reduces the overall execution time and the energy consumption by 42% and 41%, respectively, when it is augmented to existing runtime mapping algorithms. Moreover, a negligible overhead, accounting for only less than 2.6% of the overall execution time, is required for the proposed defragmentation process. The proposed defragmentation scheme is an effective resource management enhancement to existing runtime mapping algorithms for many-core systems.
Jim Ng, Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Analysis and Mapping for Thermal and Energy Efficiency of 3-D Video Processing on 3-D Multicore Processors
abstract
Three-dimensional video processing has high computation requirements and multicore processors realized in 3-D integrated circuits (ICs) provide promising high performance computing platforms. However, the conventional approaches to accelerate the computations involved in 3-D video processing do not exploit the high performance potential of 3-D ICs. In this paper, we propose an application-driven methodology that performs efficient mapping of 3-D video applications' components on 3-D multicores to achieve high performance (throughput). The methodology involves an extensive application analysis to exploit the spatial and temporal correlation available in 3-D neighborhood. Afterward, it leverages the correlation and thermal properties of different 3-D views to perform an efficient mapping of 3-D video processing on cores available at different layers of 3-D IC. The goal is to optimize energy consumption and peak temperature while meeting the throughput requirement. Experiments show 76% reduction in communication energy along with reduction in peak temperature when compared with approaches exploiting architecture characteristics only.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Value and Energy Optimizing Dynamic Resource Allocation in Many-Core HPC Systems
abstract
The conventional approaches to reduce the energy consumption of high performance computing (HPC) data centers focus on consolidation and dynamic voltage and frequency scaling (DVFS). Most of these approaches consider independent tasks (or jobs) and do not jointly optimize for energy and value. In this paper, we propose DVFS-aware profiling and non-profiling based approaches that use design-time profiling results and perform all the computations at run-time, respectively. The profiling based approach is suitable for the scenarios when the jobs or their structure is known at design-time, otherwise, the non-profiling based approach is more suitable. Both the approaches consider jobs containing dependent tasks and exploit efficient allocation combined with identification of voltage/frequency levels of used system cores to jointly optimize value and energy. Experiments show that the proposed approaches reduce energy consumption by 15% when compared to existing approaches while achieving significant amount of value and reducing percentage of rejected jobs leading to zero value.
Amit Kumar Singh 0002, Piotr Dziurzanski, Leandro Soares Indrusiak
CloudCom1
2015 Exploiting loop-array dependencies to accelerate the design space exploration with high level synthesis
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001, Khin Mi Mi Aung
DATE2
2015 DeFrag: Defragmentation for Efficient Runtime Resource Allocation in NoC-Based Many-core Systems
abstract
Efficient runtime resource allocation is critical to the overall performance and energy consumption of many-core systems. However, due to the applications' unknown arrival and departure time under dynamic workloads, the runtime system resource management is challenging. The frequent allocations and deal locations of the applications might leave on-chip free cores scattered due to the lack of design-time knowledge of their finishing time. This situation is referred to as fragmentation. In order to optimize the performance and energy consumption of the system in such situations, in this paper, we propose a runtime defragmentation approach that collects and reshapes the scattered cores in close proximity. We also propose a fragmentation metric which is able to evaluate the scatteredness of the free cores. Based on this, the proposed algorithm will be executed to bring the scattered free cores together when the metric is over a certain predefined threshold. In this way, the contiguous free core region is formed to facilitate efficient mapping of the incoming applications. Moreover, the proposed algorithm is also aware of the existing applications and minimizes their performance impact. Experimental results demonstrated that the proposed defragmentation approach reduces the overall execution time and energy consumption by 42% and 41%, respectively when compared to some of the existing approaches. Moreover, a negligible overhead, accounting for only less than 2.6% of the overall execution time, is required for the defragmentation process.
Jim Ng, Xiaohang Wang 0001, Amit Kumar Singh 0002, Terrence S. T. Mak
PDP3
2015 Execution Trace-Driven Energy-Reliability Optimization for Multimedia MPSoCs
abstract
Multiprocessor systems-on-chip (MPSoCs) are becoming a popular design choice in current and future technology nodes to accommodate the heterogeneous computing demand of a multitude of applications enabled on these platform. Streaming multimedia and other communication-centric applications constitute a significant fraction of the application space of these devices. The mapping of an application on an MPSoC is an NP-hard problem. This has attracted researchers to solve this problem both as stand-alone (best-effort) and in conjunction with other optimization objectives, such as energy and reliability. Most existing studies on energy-reliability joint optimization are static—that is, design time based. These techniques fail to capture runtime variability such as resource unavailability and dynamism associated with application behaviors, which are typical of multimedia applications. The few studies that consider dynamic mapping of applications do not consider throughput degradation, which directly impacts user satisfaction. This article proposes a runtime technique to analyze the execution trace of an application modeled as Synchronous Data Flow Graphs (SDFGs) to determine its mapping on a multiprocessor system with heterogeneous processing units for different fault scenarios. Further, communication energy is minimized for each of these mappings while satisfying the throughput constraint. Experiments conducted with synthetic and real SDFGs demonstrate that the proposed technique achieves significant improvement with respect to the state-of-the-art approaches in terms of throughput and storage overhead with less than 20% energy overhead.
Anup Das 0001, Amit Kumar Singh 0002, Akash Kumar 0001
ACM Trans. Reconfigurable Technol. Syst.2
2014 Design Space Exploration to Accelerate Nelder-Mead Algorithm Using FPGA
abstract
Nelder-Mead algorithm (NMA) is the best-known algorithm for multidimensional optimization without involving derivative computations. Due to the simplicity in implementation and the fast convergent property of NMA, it is widely used in the fields of statistics, engineering, physics and medical sciences. In practice, when objective function is complicated, the optimization procedure requires a lot of computation efforts, leading to a time-consuming process. This work introduces a NMA solver engine fully implemented on FPGA hardware and performs design space exploration to provide various solutions suitable for FPGA device.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001, Khin Mi Mi Aung
FCCM2
2014 Leakage and performance aware resource management for 2D dynamically reconfigurable FPGA architectures
abstract
The variety of applications for field programmable gate arrays (FPGAs) is continuously growing, thus it is important to address power consumption issues during the operation. As technological node shrinks, leakage power becomes increasingly critical in overall power consumption of FPGA. The technique of configuration pre-fetching (loads configurations as soon as possible) adopted to achieve high performance is one of the major reasons of leakage waste since regions containing reconfiguration information cannot be powered down in between the time gap of reconfiguration and execution. In this work, we present a heuristic approach to minimize the leakage power consumption for two-dimensional reconfigurable FPGA architectures. The heuristic scheduler is based on list scheduling and exploits dynamic priority for sorting the tasks into schedule order and a cost function for cell allocation. Farthest placement scheme is adopted for anti-fragmentation purpose. The cost function provides control to compromise between leakage dissipation and schedule length.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001
FPL3
2014 A multi-stage leakage aware resource management technique for reconfigurable architectures
abstract
Shrinking size of transistors has enabled us to integrate more and more logic elements into FPGA chips leading to higher computing power. However, it also brings serious concern to the leakage power dissipation of the FPGA devices. One of the major reasons for leakage power dissipation in FPGA is the utilization of prefetching technique to minimize the reconfiguration overhead (delay) in Partially Reconfigurable (PR) FPGAs. This technique creates delays between the reconfiguration and execution parts of a task, which may lead up to 44% leakage power of FPGA since the SRAM-cells containing reconfiguration information cannot be powered down. In this work, a resource management approach containing scheduling, placement and post-placement stages has been proposed to address the aforementioned issue. In scheduling stage, a leakage-aware cost function is derived to cope with the leakage power. The placement stage uses a cost function that allows designers to decide a trade-off between performance and leakage-saving. The post-placement stage employs a heuristic approach and shows further improvements. Experiments show that our approach can achieve large leakage savings for both synthetic and real life applications with acceptable extended deadline. Furthermore, different variants of the proposed approach can reduce leakage power by 40-65% when compared to a performance-driven approach and by 15-43% when compared to state-of-the-art works.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI2
2014 Thermal-aware task scheduling for peak temperature minimization under periodic constraint for 3D-MPSoCs
abstract
3D-MPSoC offer great performance and scalability benefits. However, due to strong vertical thermal correlation and increased power density, thermal challenges in 3D-MPSoC are critical. In this paper, we propose a novel thermal aware task scheduling technique that combine intelligent task mapping with DVFS to minimize the peak temperature of the system. Particularly, our approach leverages on the fundamental thermal characteristics of 3D architecture when mapping tasks to processing cores and employing DVFS at design time followed by a simple thermal optimization step at run time. Our experiments validate the efficiency of our approach in peak temperature minimization up to 14°C compared to other existing methods.
Vivek Chaturvedi, Amit Kumar Singh 0002, Wei Zhang 0012, Thambipillai Srikanthan
RSP2
2014 A multi-stage thermal management strategy for 3D multicores
abstract
3D integration technology has the potential to enhance IC performance, improve functionality and lessen wiring of ICs. However, it poses several challenges, where the key challenge is heat generation from internal active layers due to power dissipation. To mitigate this challenge, thermal aware design has become a necessity. Towards thermal aware design, this paper proposes a two stage design technique. In the first stage, a temperature-power thermal model is created to calculate power dissipated by an IC at an input temperature. The proposed model calculates power dissipated by 2D and 3D ICs with an average error of 0.37% and 25% respectively. Power calculation helps in process variation, validation of power models and minimization of temperature gradients. In the second stage, thermal aware mapping is performed for the ICs. For thermal aware mapping, three mapping algorithms are proposed to account for different resource (processor) availability scenarios. Each algorithm utilizes temperature-power thermal model (from the first design stage) to map applications to processing elements in a 3D IC. The proposed two stage design technique performs faster temperature to power calculations than existing techniques. It provides a simplified approach to mapping compared to existing techniques by utilizing power dissipated by processing elements to map applications.
Dipika Suresh, Amit Kumar Singh 0002, Akash Kumar 0001
RSP2
2013 Energy optimization by exploiting execution slacks in streaming applications on multiprocessor systems
abstract
Dynamic voltage and frequency scaling (DVFS) offers great potential for optimizing the energy efficiency of Multiprocessor Systems-on-Chip (MPSoCs). The conventional approaches for processor voltage and frequency adjustment are not suitable for streaming multimedia applications due to the cyclic nature of dependencies in the executing tasks which can potentially violate the throughput constraints. In this paper, we propose a methodology that applies DVFS for such cyclic dependent tasks. The methodology involves an off-line analysis that assumes worst-case execution times of tasks to identify the executions that can be slowed down and an on-line analysis to utilize the slacks arising from tasks that finish their execution before the worst-case execution times. Thus, the methodology minimizes energy consumption during both off-line and on-line analysis while satisfying the throughput constraints. Experiments based on models of real-life streaming multimedia applications show that the proposed methodology reduces the overall energy consumption by 43% when compared to existing approaches.
Amit Kumar Singh 0002, Anup Das 0001, Akash Kumar 0001
DAC1
2013 Mapping on multi/many-core systems: survey of current and emerging trends
abstract
The reliance on multi/many-core systems to satisfy the high performance requirement of complex embedded software applications is increasing. This necessitates the need to realize efficient mapping methodologies for such complex computing platforms. This paper provides an extensive survey and categorization of state-of-the-art mapping methodologies and highlights the emerging trends for multi/many-core systems. The methodologies aim at optimizing system's resource usage, performance, power consumption, temperature distribution and reliability for varying application models. The methodologies perform design-time and run-time optimization for static and dynamic workload scenarios, respectively. These optimizations are necessary to fulfill the end-user demands. Comparison of the methodologies based on their optimization aim has been provided. The trend followed by the methodologies and open research challenges have also been discussed.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
DAC1
2013 Incorporating Energy and Throughput Awareness in Design Space Exploration and Run-Time Mapping for Heterogeneous MPSoCs
abstract
The advancement in process technology has enabled integration of different types of processing cores into a single chip towards creating heterogeneous Multiprocessor Systems-on-Chip (MPSoCs). While providing high level of computation power to support complex applications, these modern systems also introduce novel challenges for system designers, like managing a huge number of mappings (application tasks to processing cores allocations) that increases exponentially with the number of cores and their types. This paper presents a mapping approach that computes multiple energy-throughput trade-off points (mappings) at design-time and uses one of these points at run-time based on desired throughput and current resource availability while optimizing for the overall energy consumption. While significantly reducing the complexity of the design space exploration (DSE) to compute mappings at design-time, the proposed strategy still evaluates mappings for all the resource combinations of the platform, providing efficient mapping solutions for all the scenarios of system architecture at run-time. Moreover, the proposed approach performs energy-aware mapping at run-time while utilizing the DSE results. Experimental results show that proposed strategy achieves better energy-throughput trade-off points, covers all the resource combinations and reduces energy consumption up to 24.93% at design-time and additionally 17.8% at run-time when compared to state-of-the-art techniques.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001, Khin Mi Mi Aung
DSD2
2013 RAPIDITAS: RAPId Design-Space-Exploration Incorporating Trace-Based Analysis and Simulation
abstract
Simulation-based Design Space Exploration (DSE) to evaluate all possible mappings for a given application and Multiprocessor-System-on-Chip (MPSoC) platform is computationally costly for large problems. Even using efficient exploration methodologies to evaluate the mappings cannot overcome the evaluation time bottleneck. This paper presents a novel DSE methodology that analyzes the execution trace to prune the vast design space. Simulations are employed only on the pruned design points (mappings), hence reducing the number of simulations. The methodology performs iterative exploration and provides premier mappings requiring different number of processors, which can be used at run-time subject to desired performance and available platform processors. We evaluate our methodology by using models of real-life multimedia applications and demonstrate that the DSE time is reduced by 72% while generating high quality mappings.
Amit Kumar Singh 0002, Anup Das 0001, Akash Kumar 0001
DSD1
2013 CADSE: communication aware design space exploration for efficient run-time MPSoC management
Amit Kumar Singh 0002, Akash Kumar 0001, Jigang Wu, Thambipillai Srikanthan
Frontiers Comput. Sci.1
2012 Accelerating throughput-aware runtime mapping for heterogeneous MPSoCs
abstract
Modern embedded systems need to support multiple time-constrained multimedia applications that often employ multiprocessor-systems-on-chip (MPSoCs). Such systems need to be optimized for resource usage and energy consumption. It is well understood that a design-time approach cannot provide timing guarantees for all the applications due to its inability to cater for dynamism in applications. However, a runtime approach consumes large computation requirements at runtime and hence may not lend well to constrained-aware mapping. In this article, we present a hybrid approach for efficient mapping of applications in such systems. For each application to be supported in the system, the approach performs extensive design-space exploration (DSE) at design time to derive multiple design points representing throughput and energy consumption at different resource combinations. One of these points is selected at runtime efficiently, depending upon the desired throughput while optimizing for energy consumption and resource usage. While most of the existing DSE strategies consider a fixed multiprocessor platform architecture, our DSE considers a generic architecture, making DSE results applicable to any target platform. All the compute-intensive analysis is performed during DSE, which leaves for minimum computation at runtime. The approach is capable of handling dynamism in applications by considering their runtime aspects and providing timing guarantees. The presented approach is used to carry out a DSE case study for models of real-life multimedia applications: H.263 decoder, H.263 encoder, MPEG-4 decoder, JPEG decoder, sample rate converter, and MP3 decoder. At runtime, the design points are used to map the applications on a heterogeneous MPSoC. Experimental results reveal that the proposed approach provides faster DSE, better design points, and efficient runtime mapping when compared to other approaches. In particular, we show that DSE is faster by 83% and runtime mapping is accelerated by 93% for some cases. Further, we study the scalability of the approach by considering applications with large numbers of tasks.
Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan
ACM Trans. Design Autom. Electr. Syst.1
2011 A hybrid strategy for mapping multiple throughput-constrained applications on MPSoCs
abstract
Modern embedded systems are based on Multiprocessor-Systems-on-Chip (MPSoCs) to meet the strict timing deadlines of multiple applications. MPSoC resources must be utilized efficiently by mapping the applications in throughput-aware manner in order to meet throughput constraints for each of them. A design-time methodology is applicable only to predefined set of applications with static behavior, which is incapable of handling dynamism in applications. On the other hand, a run-time approach can cater to the dynamism but cannot provide timing guarantees for all the applications due to large computation requirements at run-time. This paper presents a hybrid flow which performs compute intensive analysis at design-time to derive multiple resource-throughput trade-off points and selects one of these at run-time subject to available resources and desired throughput. Experimental results show that the design-time analysis is faster by 39%, provides better trade-off points and the run-time mapping is speeded up by 93% when compared to state-of-the-art techniques.
Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan
CASES1
2010 Mapping real-life applications on run-time reconfigurable NoC-based MPSoC on FPGA
abstract
Multiprocessor systems-on-chip (MPSoC) are required to fulfill the performance demand of modern real-life embedded applications. These MPSoCs are employing Network-on-Chip (NoC) for reasons of efficiency and scalability. Additionally, these systems need to support run-time reconfiguration of their components to cater to dynamically changing demands of the system. Designing and programming such systems for real-life applications prove to be a major challenge. This paper demonstrates the designing of reconfigurable NoC-based MPSoC and programming it for real-life applications. The NoC is reconfigured at run-time to support different combinations of multiple applications at different times. The platform is verified with a case study executing the parallelized C-codes of a simple producer-consumer and JPEG decoder applications on a NoC-based MPSoC on a Xilinx FPGA. Based on our investigations to map the applications on a 3 × 3 platform, we show that the NoC reconfiguration overhead is kept at a minimum and the platform utilizes 85% of the total available slices of Virtex-5 FPGA. Moreover, we show that the proposed approach is highly scalable when targeting for large number of applications.
Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan, Yajun Ha
FPT1
2010 Communication-aware heuristics for run-time task mapping on NoC-based MPSoC platforms
Amit Kumar Singh 0002, Thambipillai Srikanthan, Akash Kumar 0001, Jigang Wu
J. Syst. Archit.1
2009 Mapping Algorithms for NoC-Based Heterogeneous MPSoC Platforms
abstract
Mapping of applications onto multiprocessor system-on-chip (MPSoC) can be realized either at design-time or run-time. At any time the number of tasks executing in MPSoC platform can exceed the available resources, requiring efficient run-time mapping techniques to meet the real-time constraints of the applications. This paper presents two run-time mapping heuristics for mapping the tasks of an application in close proximity so as to minimize the communication overhead. In particular, the communication overhead between two adjacent hardware tasks is eliminated by mapping them onto the same reconfigurable processing node. We show that the proposed approach is capable of alleviating network-on-chip (NoC) congestion bottlenecks to optimize the overall performance. Based on our investigations to map the tasks of applications' at run-time onto an 8×8 NoC-based heterogeneous MPSoC, our mapping heuristics are capable of reducing total execution time and average channel load of applications when compared to state-of-the-art runtime mapping heuristics.
Amit Kumar Singh 0002, Jigang Wu, Alok Prakash, Thambipillai Srikanthan
DSD1
2009 Rapid design exploration framework for application-aware customization of soft core processors
abstract
Off-the-shelf soft core processors are becoming increasingly popular in embedded systems design today as they provide for application specific customization, in particular through instruction subsetting. However, choosing the right processor configuration remains a challenge as the search space becomes prohibitively large when the configurable options increase. In this paper we propose a framework to rapidly explore the processor configuration design space for a given application. Unlike existing approaches that require time-consuming synthesis process, the proposed method relies only on a single-pass output of the LLVM compiler infrastructure. Experimental results based on widely used benchmarks show that the proposed framework can reliably predict the actual performance and area trends of various configurable options.
Alok Prakash, Siew-Kei Lam, Amit Kumar Singh 0002, Thambipillai Srikanthan
FPL3