EDBT 2026 Demo / reviewers in the wild / expert
Pasi Liljeberg
dblp:22/660
· DBLP profile ↗
125ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-9392-3589ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 84 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 6 since 2021Software engineering, systems software and programming languages · 12 · 3 since 2021Computer networks · 4 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring PPG-Guided Knowledge Distillation for Contactless Respiration Estimation
Trishna Saikia, Anup Kumar Gupta 0001, Puneet Gupta 0002, Pasi Liljeberg |
FG | 4 |
| 2025 | HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge PlatformsabstractEdge inference techniques partition and distribute Deep Neural Network (DNN) inference tasks among multiple edge nodes for low latency inference, without considering the core-level heterogeneity of edge nodes. Further, default DNN inference frameworks also do not fully utilize the resources of heterogeneous edge nodes, resulting in higher inference latency. In this work, we propose a hierarchical DNN partitioning strategy (HiDP) for distributed inference on heterogeneous edge nodes. Our strategy hierarchically partitions DNN workloads at both global and local levels by considering the core-level heterogeneity of edge nodes. We evaluated our proposed HiDP strategy against relevant distributed inference techniques over widely used DNN models on commercial edge devices. On average our strategy achieved 38% lower latency, 46% lower energy, and 56% higher throughput in comparison with other relevant approaches, Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
DATE | 4 |
| 2025 | Invited Paper: Mindful AI for Pervasive Health and Wellbeing (PHW)abstractEmerging AI-driven pervasive health and wellbeing (PHW) services (e.g., personalized health assistants and mobile health applications) face critical challenges in handling noisy/intermittent sensory data, integrating cross-modal insights, and stringent energy and compute constraints. We present Mindful AI, a cognitive-inspired framework designed to enable adaptive, resilient, and efficient PHW services in real-world conditions. Our dual-mode intelligence—Automatic (System 1) and Reflective (System 2)—selectively directs system attention toward the most relevant sensing and compute contexts, unifying bottom-up stimuli (driven by input quality, inference demands and model confidence, and resource availability) with top-down insights (reflecting user demands, system goals/constraints, and contextual information). Our framework distills and orchestrates insights across sensing, communication, and computation through hybrid attention toward bottom-up and top-down insights that support cross-layer sense-compute co-optimization to achieve resilient, low-latency, and energy-efficient PHW services. We evaluate our approach on multi-tier device-edge-cloud platforms, using real-world case studies in pain assessment, stress monitoring, and human activity recognition to demonstrate adaptation to real-world uncertainties (e.g., sensor degradation, context drift, network variability), while maintaining strict QoS, accuracy, and latency guarantees. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
ICCAD | 3 |
| 2025 | Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge PlatformsabstractCompound AI (cAI) systems chain multiple AI models to solve complex problems. cAI systems are typically composed of deep neural networks (DNNs), transformers, and large language models (LLMs), exhibiting a high degree of computational diversity and dynamic workload variation. Deploying cAI services on mobile edge platforms poses a significant challenge in scheduling concurrent DNN-transformer inference tasks, which arrive dynamically in an unknown sequence. Existing mobile edge AI inference strategies manage multi-DNN or transformer-only workloads, relying on design-time profiling, and cannot handle concurrent inference of DNNs and transformers required by cAI systems. In this work, we address the challenge of scheduling cAI systems on heterogeneous mobile edge platforms. We present Twill, a run-time framework to handle concurrent inference requests of cAI workloads through task affinity-aware cluster mapping and migration, priority-aware task freezing/unfreezing, and Dynamic Voltage/Frequency Scaling (DVFS), while minimizing inference latency within power budgets. We implement and deploy our Twill framework on the Nvidia Jetson Orin NX platform. We evaluate Twill against state-of-the-art edge AI inference techniques over contemporary DNNs and LLMs, reducing inference latency by 54% on average, while honoring power budgets. Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
ICCAD | 4 |
| 2025 | Exploiting Approximation for Run-time Resource Management of Embedded HMPsabstractRun-time resource management (RTM) of multi-programmed workloads on heterogeneous multi-core platforms is challenging due to (i) fixed power budget of the device, (ii) variable performance requirements of the workloads, and (iii) unknown arrival of the applications. Existing RTM solutions lack power-performance coordination, resulting in performance degradation during power actuation or power violations during performance provisioning. Exploiting inherent error-resilience of the applications can address the performance loss incurred in power actuation, by combining run-time approximation with traditional power knobs (including Dynamic Voltage/Frequency Scaling, Task Migration, Degree of Parallelism, and CPU Quota ). In this work, we present an accuracy-aware resource management framework that jointly actuates run-time approximation and traditional power knobs for efficient power-performance management of multi-programmed and multi-threaded workloads running on heterogeneous mobile platforms. Our strategy configures the accuracy of the applications at run-time to exploit accuracy-performance trade-offs, by considering system-wide power-performance dynamics. We use heuristic estimation models to jointly enforce accuracy configuration and traditional power knobs settings at run-time. We evaluated our framework on real-world embedded mobile platforms, including Odroid XU3 and Asus Tinker Edge R boards to demonstrate the efficiency of our proposed approach across multiple workload scenarios. Our approach achieved 25% lower performance violations against the state-of-the-art run-time resource management policies at the cost of 2.2% accuracy loss across six applications. Zain Taufique, Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Cristiana Bolchini, Nikil Dutt, Pasi Liljeberg |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2024 | Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge PlatformsabstractDNN inference can be accelerated by distributing the workload among a cluster of collaborative edge nodes. Heterogeneity among edge devices and accuracy-performance trade-offs of DNN models present a complex exploration space while catering to the inference performance requirements. In this work, we propose adaptive workload distribution for DNN inference, jointly considering node-level heterogeneity of edge devices, and application-specific accuracy and performance requirements. Our proposed approach combinatorially optimizes heterogeneity-aware workload partitioning and dynamic accuracy configuration of DNN models to ensure performance and accuracy guarantees. We tested our approach on an edge cluster of Odroid XU4, Raspberry Pi4, and Jetson Nano boards and achieved an average gain of 41.52% in performance and 5.2% in output accuracy as compared to state-of-the-art workload distribution strategies. Zain Taufique, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
ASPDAC | 3 |
| 2024 | ECG Unveiled: Analysis of Client Re-Identification Risks in Real-World ECG DatasetsabstractWhile ECG data is crucial for diagnosing and monitoring heart conditions, it also contains unique biometric information that poses significant privacy risks. Existing ECG re-identification studies rely on exhaustive analysis of numerous deep learning features, confining to ad-hoc explainability towards clinicians decision making. In this work, we delve into explainability of ECG re-identification risks using transparent machine learning models. We use SHapley Additive exPlanations (SHAP) analysis to identify and explain the key features contributing to re-identification risks. We conduct an empirical analysis of identity re-identification risks using ECG data from five diverse real-world datasets, encompassing 223 participants. By employing transparent machine learning models, we reveal the diversity among different ECG features in contributing towards re-identification of individuals with an accuracy of 0.76 for gender, 0.67 for age group, and 0.82 for participant ID re-identification. Our approach provides valuable insights for clinical experts and guides the development of effective privacy-preserving mechanisms. Further, our findings emphasize the necessity for robust privacy measures in real-world health applications and offer detailed, actionable insights for enhancing data anonymization techniques. Anil Kanduri, Seyed Amir Hossein Aqajari, Salar Jafarlou, Sanaz R. Mousavi, Pasi Liljeberg, Shaista Malik, Amir-Mohammad Rahmani |
BSN | 6 |
| 2024 | Attention-Based Explainable AI for Wearable Multivariate Data: A Case Study on Affect Status PredictionabstractWearable technology enables ubiquitous health monitoring where multivariate physiological and behavioral data can be captured over time. Such multivariate time series (MTS) data in healthcare applications needs technique to interpret the analysis results. However, existing deep learning models for MTS data analysis often lack interpretability, and current explainable AI (xAI) techniques fail to capture the temporal and inter-variable complexities inherent in MTS. This hinders the trust and integration of these AI-based systems in clinical decision-making. In this paper, we propose an attention-based xAI method to classify and interpret MTS data collected from wearable devices. Our approach leverages self-attention mechanisms and graph attention layers (GAT) to capture both temporal and inter-variable dependencies, providing interpretability at both the temporal and modality levels. We evaluate our method using a longitudinal affect status monitoring. The dataset was collected from 21 college students via wearable devices over one year. We train separate models for positive (PA) and negative affect (NA) prediction, and compare their performance with a Transformer-based method. Our method achieves robust classification performance, with 78.62% accuracy for PA and 76.30% for NA, while offering transparent explanations of its decisions. These findings highlight the potential of our xAI method for reliable and interpretable MTS classification in healthcare applications. Zhongqi Yang, Iman Azimi, Amir-Mohammad Rahmani, Pasi Liljeberg |
BSN | 5 |
| 2024 | Work-in-Progress: Context and Noise Aware Resilience for Autonomous Driving ApplicationsabstractAutonomous Vehicles (AVs) often use noise prone sensory data from cameras and LiDAR for perception. In specific noisy scenarios, different object detection models exhibit non-intuitive and varying degrees of resilience, necessitating adaptive model selection. In this work, we develop a context and noise aware framework for run-time adaptive configuration of objection models for high accuracy and low latency inference. We combine driving scene context and input data noise to prioritize among input modalities, followed by selection and configuration of most resilient object detection model appropriate for the context. Our evaluation for 2D object detection on nuScenes dataset provided average 1.83x speedup in latency compared to baseline while preserving average prediction confidence. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
CODES+ISSS | 3 |
| 2024 | SEAL: Sensing Efficient Active Learning on Wearables through Context-awarenessabstractIn this paper, we introduce SEAL, a co-optimization framework designed to enhance both sensing and querying strategies in wearable devices for mHealth applications. Employing Reinforcement Learning (RL), SEAL strategically utilizes user contextual information and the machine learning model's confidence levels to make efficient decisions. This innovative approach is particularly significant in addressing the challenge of battery drain due to continuous physiological signal sensing, such as Photoplethysmography (PPG). Our framework demonstrates its effectiveness in a stress monitoring application, achieving a substantial reduction of 76% in the volume of PPG signals collected, while only experiencing a minor 6% decrease in user-labeled data quality. This balance showcases SEAL's potential in optimizing data collection in a way that is considerate of both device constraints and data integrity. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
DATE | 4 |
| 2024 | Tango: Low Latency Multi-DNN Inference on Heterogeneous Edge PlatformsabstractThere is an increasing demand to run DNN applications on edge platforms for low-latency inference. Executing multi-DNN workloads with diverse compute and latency requirements on resource-constrained heterogeneous edge platforms poses a significant scheduling challenge. In this work, we present Tango framework for orchestrating multi-DNN inference on heterogeneous edge platforms. Our approach uses a Proximal Policy-based Reinforcement Learning agent to jointly optimize cluster selection, accuracy configuration, and frequency scaling to minimize inference latency with a tolerable accuracy loss. We implemented the proposed Tango framework as a portable middleware and deployed it on real hardware of the Jetson TX edge platform. Our evaluation against relevant multi-DNN scheduling strategies demonstrates 61 % lower latency and 48.4 % lower energy consumption at a maximum accuracy loss of 1.59 %. Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
ICCD | 4 |
| 2024 | EA^2: Energy Efficient Adaptive Active Learning for Smart WearablesabstractMobile Health (mHealth) applications rely on supervised Machine Learning (ML) algorithms, requiring end-user-labeled data for the training phase. The gold standard for obtaining such labeled data is by sending queries to users and gathering responses for the corresponding label, which was conventionally done through triggering questions sent at random. Active Learning (AL) methods use intelligent query-sending policies by incorporating users' contextual information to maximize the response rate and informativeness of the collected labeled data. However, wearable devices' substantial battery drainage associated with the sensing of physiological signals underscores the need for developing an efficient sensing policy in addition to a query-sending policy. In this work, we present a co-optimization framework for both sensing and querying strategies within wearable devices, leveraging contextual information and ML model's prediction confidence. We designed a Reinforcement Learning (RL) agent to quantify different contextual parameters combined with model confidence to determine sensing and querying decisions. Our evaluation of an exemplar stress monitoring application showed a 76% reduction in sensing and data transmission energy consumption, with only a 6% drop in user-labeled data. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
ISLPED | 4 |
| 2024 | Personalized and adaptive neural networks for pain detection from multi-modal physiological features
Mingzhe Jiang, Riitta Rosio, Sanna Salanterä, Amir-Mohammad Rahmani, Pasi Liljeberg, Daniel Santos da Silva, Victor Hugo C. de Albuquerque |
Expert Syst. Appl. | 5 |
| 2023 | End-to-End PPG Processing Pipeline for Wearables: From Quality Assessment and Motion Artifacts Removal to HR/HRV Feature ExtractionabstractThe rapid development of wearable technology has enabled remote photoplethysmography (PPG)-based health monitoring in everyday settings, offering real-time and continuous monitoring of cardiovascular parameters, such as heart rate (HR) and heart rate variability (HRV). However, PPG signals collected in daily life are prone to artifacts and noise, posing challenges to HR and HRV extraction. The existing HR and HRV extraction methods cannot effectively handle noisy PPG signals and ensure accurate results. Additionally, current Python packages were primarily designed for analyzing "clean" PPG signals, limiting their performance in handling artifacts and noise and resulting in unreliable HR and HRV measurements. In this paper, we propose a robust end-to-end PPG processing pipeline to reliably extract HR and HRV from PPG signals collected in free-living settings. The pipeline comprises three machine learning-based PPG analysis methods: signal quality assessment, reconstruction of noisy signal, and systolic peak detection. We assess the proposed PPG pipeline using a dataset including PPG and Electrocardiogram (ECG) signals recorded from 46 individuals by smartwatches. Our evaluation demonstrates the proposed pipeline’s superior performance compared to two established benchmark methods in terms of correlation and mean absolute error with ECG as the reference. We also provide the Python implementation of our pipeline for the research community to facilitate integration into their solutions. Mohammad Feli, Kianoosh Kazemi, Iman Azimi, Amir-Mohammad Rahmani, Pasi Liljeberg |
BIBM | 6 |
| 2023 | Quantifying Movement Behavior of Chronic Low Back Pain Patients in Virtual RealityabstractChronic low back pain (CLBP) is a globally common musculoskeletal problem. Measuring the sensation of pain and the effect of a treatment has always been a challenge for healthcare. Here, we study how the movement data, collected while using a virtual reality (VR) program, could be used as an objective measurement in patients with CLBP. A specific data collection method based on VR was developed and used with CLBP patients and healthy volunteers. We demonstrate that the movement data in VR can be used to classify individuals in these two groups with a high accuracy by using logistic regression. The most discriminative features are the duration of the movements and the total variation of movement velocity. Furthermore, we show that hidden Markov models can divide movement data into meaningful segments, which creates possibilities for defining even more detailed features, with potential to improve accuracy, when larger datasets become available in the future. Tommi Gröhn, Sammeli Liikkanen, Teppo Huttunen, Mika Mäkinen, Pasi Liljeberg, Pekka Marttinen |
ACM Trans. Comput. Heal. | 5 |
| 2023 | A Deep Learning-based PPG Quality Assessment Approach for Heart Rate and Heart Rate VariabilityabstractPhotoplethysmography (PPG) is a non-invasive optical method to acquire various vital signs, including heart rate (HR) and heart rate variability (HRV). The PPG method is highly susceptible to motion artifacts and environmental noise. Unfortunately, such artifacts are inevitable in ubiquitous health monitoring, as the users are involved in various activities in their daily routines. Such low-quality PPG signals negatively impact the accuracy of the extracted health parameters, leading to inaccurate decision-making. PPG-based health monitoring necessitates a quality assessment approach to determine the signal quality according to the accuracy of the health parameters. Different studies have thus far introduced PPG signal quality assessment methods, exploiting various indicators and machine learning algorithms. These methods differentiate reliable and unreliable signals, considering morphological features of the PPG signal and focusing on the cardiac cycles. Therefore, they can be utilized for HR detection applications. However, they do not apply to HRV, as only having an acceptable shape is insufficient, and other signal factors may also affect the accuracy. In this article, we propose a deep learning–based PPG quality assessment method for HR and various HRV parameters. We employ one customized one-dimensional (1D) and three 2D Convolutional Neural Networks (CNN) to train models for each parameter. Reliability of each of these parameters will be evaluated against the corresponding electrocardiogram signal, using 210 hours of data collected from a home-based health monitoring application. Our results show that the proposed 1D CNN method outperforms the other 2D CNN approaches. Our 1D CNN model obtains the accuracy of 95.63%, 96.71%, 91.42%, 94.01%, and 94.81% for the HR, average of normal to normal interbeat (NN) intervals, root mean square of successive NN interval differences, standard deviation of NN intervals, and ratio of absolute power in low frequency to absolute power in high frequency ratios, respectively. Moreover, we compare the performance of our proposed method with state-of-the-art algorithms. We compare our best models for HR-HRV health parameters with six different state-of-the-art PPG signal quality assessment methods. Our results indicate that the proposed method performs better than the other methods. We also provide the open source model implemented in Python for the community to be integrated into their solutions. Emad Kasaeyan Naeini, Fatemeh Sarhaddi, Iman Azimi, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Trans. Comput. Heal. | 4 |
| 2022 | AMSER: Adaptive Multimodal Sensing for Energy Efficient and Resilient eHealth SystemsabstracteHealth systems deliver critical digital healthcare and wellness services for users by continuously monitoring physiological and contextual data. eHealth applications use multi-modal machine learning kernels to analyze data from different sensor modalities and automate decision-making. Noisy inputs and motion artifacts during sensory data acquisition affect the i) prediction accuracy and resilience of eHealth services and ii) energy efficiency in processing garbage data. Monitoring raw sensory inputs to identify and drop data and features from noisy modalities can improve prediction accuracy and energy efficiency. We propose a closed-loop monitoring and control framework for multi-modal eHealth applications, AMSER, that can mitigate garbage-in garbage-out by i) monitoring input modalities, ii) analyzing raw input to selectively drop noisy data and features, and iii) choosing appropriate machine learning models that fit the configured data and feature vector - to improve prediction accuracy and energy efficiency. We evaluate our AMSER approach using multi-modal eHealth applications of pain assessment and stress monitoring over different levels and types of noisy components incurred via different sensor modalities. Our approach achieves up to 22% improvement in prediction accuracy and 5.6× energy consumption reduction in the sensing phase against the state-of-the-art multi-modal monitoring application. Emad Kasaeyan Naeini, Sina Shahhosseini, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
DATE | 4 |
| 2022 | Exploring computation offloading in IoT systemsabstractInternet of Things (IoT) paradigm raises challenges for devising efficient strategies that offload applications to the fog or the cloud layer while ensuring the optimal response time for a service. Traditional computation offloading policies assume the response time is only dominated by the execution time. However, the response time is a function of many factors including contextual parameters and application characteristics that can change over time. For the computation offloading problem, the majority of existing literature presents efficient solutions considering a limited number of parameters (e.g., computation capacity and network bandwidth) neglecting the effect of the application characteristics and dataflow configuration. In this paper, we explore the impact of the computation offloading on total application response time in three-layer IoT systems considering more realistic parameters, e.g., application characteristics, system complexity, communication cost, and dataflow configuration. This paper also highlights the impact of a new application characteristic parameter defined as Output–Input Data Generation (OIDG) ratio and dataflow configuration on the system behavior. In addition, we present a proof-of-concept end-to-end dynamic computation offloading technique, implemented in a real hardware setup, that observes the aforementioned parameters to perform real-time decision-making. Sina Shahhosseini, Arman Anzanpour, Iman Azimi, Sina Labbaf, Dongjoo Seo, Sung-Soo Lim, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
Inf. Syst. | 7 |
| 2022 | Confidence-Enhanced Early Warning Score Based on Fuzzy LogicabstractAbstract Cardiovascular diseases are one of the world’s major causes of loss of life. The vital signs of a patient can indicate this up to 24 hours before such an incident happens. Healthcare professionals use Early Warning Score (EWS) as a common tool in healthcare facilities to indicate the health status of a patient. However, the chance of survival of an outpatient could be increased if a mobile EWS system would monitor them during their daily activities to be able to alert in case of danger. Because of limited healthcare professional supervision of this health condition assessment, a mobile EWS system needs to have an acceptable level of reliability - even if errors occur in the monitoring setup such as noisy signals and detached sensors. In earlier works, a data reliability validation technique has been presented that gives information about the trustfulness of the calculated EWS. In this paper, we propose an EWS system enhanced with the self-aware property confidence, which is based on fuzzy logic. In our experiments, we demonstrate that - under adverse monitoring circumstances (such as noisy signals, detached sensors, and non-nominal monitoring conditions) - our proposed Self-Aware Early Warning Score (SA-EWS) system provides a more reliable EWS than an EWS system without self-aware properties. Maximilian Götzinger, Arman Anzanpour, Iman Azimi, Nima Taherinejad, Axel Jantsch, Amir-Mohammad Rahmani, Pasi Liljeberg |
Mob. Networks Appl. | 7 |
| 2022 | Concurrent Application Bias Scheduling for Energy Efficiency of Heterogeneous Multi-Core PlatformsabstractMinimizing energy consumption of concurrent applications on heterogeneous multi-core platforms is challenging given the diversity in energy-performance profiles of both the applications and hardware. Adaptive learning techniques made the exhaustive Pareto-optimal space exploration practically feasible to identify an energy efficient configuration. Existing approaches consider a single application's characteristic for optimizing energy consumption. However, an optimal configuration for a given application in isolation may not be optimal when other applications are run concurrently. Approaches that consider concurrent application scenarios overlook the weight of total energy consumption per application, restricting them from prioritizing among applications. We address this limitation by considering the mutual effect of concurrent applications on system wide energy consumption to adapt resource configuration at run-time. We characterize each application's power-performance profile as a weighted bias through off-line profiling. We infer this model combined with an on-line predictive strategy to make resource allocation decisions for minimizing energy consumption while honoring performance requirements. The proposed strategy is implemented as a user-space process and evaluated on a heterogeneous hardware platform of Odroid XU3 over the Rodinia benchmark suite. Experimental results show up to 61 percent of energy saving compared to the standard baseline of Linux governors and up to 27 percent of energy gain compared to state-of-the-art adaptive learning-based resource management techniques. Elham Shamsa, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani |
IEEE Trans. Computers | 3 |
| 2021 | Energy-Performance Co-Management of Mixed-Sensitivity Workloads on Heterogeneous Multi-core SystemsabstractSatisfying performance of complex workload scenarios with respect to energy consumption on Heterogeneous Multi-core Platforms (HMPs) is challenging when considering i) the increasing variety of applications, and ii) the large space of resource management configurations. Existing run-time resource management approaches use online and offline learning to handle such complexity. However, they focus on one type of application, neglecting concurrent execution of mixed sensitivity workloads. In this work, we propose an energy-performance co-management method which prioritizes mixed type of applications at run-time, and searches in the configuration space to find the optimal configuration for each application which satisfies the performance requirements while saving energy. We evaluate our approach on a real Odroid XU3 platform over mixed-sensitivity embedded workloads. Experimental results show our approach provides 54% lower performance violation with 50% higher energy saving compared to the existing approaches. Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg |
ASP-DAC | 4 |
| 2021 | UBAR: User- and Battery-aware Resource Management for SmartphonesabstractSmartphone users require high Battery Cycle Life (BCL) and high Quality of Experience (QoE) during their usage. These two objectives can be conflicting based on the user preference at run-time. Finding the best trade-off between QoE and BCL requires an intelligent resource management approach that considers and learns user preference at run-time. Current approaches focus on one of these two objectives and neglect the other, limiting their efficiency in meeting users’ needs. In this article, we present UBAR, User- and Battery-aware Resource management, which considers dynamic workload, user preference, and user plug-in/out pattern at run-time to provide a suitable trade-off between BCL and QoE. UBAR personalizes this trade-off by learning the user’s habits and using that to satisfy QoE, while considering battery temperature and State of Charge (SOC) pattern to maximize BCL. The evaluation results show that UBAR achieves 10% to 40% improvement compared to the existing state-of-the-art approaches. Elham Shamsa, Alma Pröbstl, Nima Taherinejad, Anil Kanduri, Samarjit Chakraborty, Amir-Mohammad Rahmani, Pasi Liljeberg |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2020 | Context-Aware Sensing via Dynamic Programming for Edge-Assisted Wearable SystemsabstractHealthcare applications supported by the Internet of Things enable personalized monitoring of a patient in everyday settings. Such applications often consist of battery-powered sensors coupled to smart gateways at the edge layer. Smart gateways offer several local computing and storage services (e.g., data aggregation, compression, local decision making), and also provide an opportunity for implementing local closed-loop optimization of different parameters of the sensor layer, particularly energy consumption. To implement efficient optimization methods, information regarding the context and state of patients need to be considered to find opportunities to adjust energy to demanded accuracy. Edge-assisted optimization can manage energy consumption of the sensor layer but may also adversely affect the quality of sensed data, which could compromise the reliable detection of health deterioration risk factors. In this article, we propose two approaches: myopic and Markov decision processes (MDPs)—to consider both energy constraints and risk factor requirements for achieving a twofold goal: energy savings while satisfying accuracy requirements of abnormality detection in a patient’s vital signs. Vital signs, including heart rate, respiration rate, and oxygen saturation, are extracted from a photoplethysmogram signal and errors of extracted features are compared to a ground truth that is modeled as a Gaussian distribution. We control the sensor’s sensing energy to minimize the power consumption while meeting a desired level of satisfactory detection performance. We present experimental results on realistic case studies using a reconfigurable photoplethysmogram sensor in an IoT system, and show that compared to nonadaptive methods, myopic reduces an average of 16.9% in sensing energy consumption with the maximum probability of abnormality misdetection on the order of 0.17 in a 24-hour health monitoring system. In addition, over 4 weeks of monitoring, we demonstrate that our MDP policy can extend the battery life on average of more than 2x while fulfilling the same average probability of misdetection compared to the myopic method. We illustrate results comparing myopic , MDP, and nonadaptive methods to monitor 14 subjects over 1 month. Delaram Amiri, Arman Anzanpour, Iman Azimi, Marco Levorato, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Trans. Comput. Heal. | 5 |
| 2020 | Edge-Assisted Control for Healthcare Internet of Things: A Case Study on PPG-Based Early Warning ScoreabstractRecent advances in pervasive Internet of Things technologies and edge computing have opened new avenues for development of ubiquitous health monitoring applications. Delivering an acceptable level of usability and accuracy for these healthcare Internet of Things applications requires optimization of both system-driven and data-driven aspects, which are typically done in a disjoint manner. Although decoupled optimization of these processes yields local optima at each level, synergistic coupling of the system and data levels can lead to a holistic solution opening new opportunities for optimization. In this article, we present an edge-assisted resource manager that dynamically controls the fidelity and duration of sensing w.r.t. changes in the patient’s activity and health state, thus fine-tuning the trade-off between energy efficiency and measurement accuracy. The cornerstone of our proposed solution is an intelligent low-latency real-time controller implemented at the edge layer that detects abnormalities in the patient’s condition and accordingly adjusts the sensing parameters of a reconfigurable wireless sensor node. We assess the efficiency of our proposed system via a case study of the photoplethysmography-based medical early warning score system. Our experiments on a real full hardware-software early warning score system reveal up to 49% power savings while maintaining the accuracy of the sensory data. Arman Anzanpour, Delaram Amiri, Iman Azimi, Marco Levorato, Nikil Dutt, Pasi Liljeberg, Amir-Mohammad Rahmani |
ACM Trans. Internet Things | 6 |
| 2019 | Goal-Driven Autonomy for Efficient On-chip Resource Management: Transforming Objectives to GoalsabstractRun-time resource allocation of heterogeneous multi-core systems is challenging with varying workloads and limited power and energy budgets. User interaction within these systems changes the performance requirements, often conflicting with concurrent applications' objective and system constraints. Current resource allocation approaches focus on optimizing fixed objective, ignoring the variation in system and applications' objective at run-time. For an efficient resource allocation, the system has to operate autonomously by formulating a hierarchy of goals. We present goal-driven autonomy (GDA) for on-chip resource allocation decisions, which allows systems to generate and prioritize goals in response to the workload and system dynamic variation. We implemented a proof-of-concept resource management framework that integrates the proposed goal management control to meet power, performance and user requirements simultaneously. Experimental results on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) show the effectiveness of GDA in efficient resource allocation in comparison with existing fixed objective policies. Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DATE | 4 |
| 2019 | Dynamic Computation Migration at the Edge: Is There an Optimal Choice?abstractIn the era of Fog computing where one can decide to compute certain time-critical tasks at the edge of the network, designers often encounter a question whether the sensor layer provides the optimal response time for a service, or the Fog layer, or their combination. In this context, minimizing the total response time using computation migration is a communication-computation co-optimization problem as the response time does not depend only on the computational capacity of each side. In this paper, we aim at investigating this question and addressing it in certain situations. We formulate this question as a static or dynamic computation migration problem depending on whether certain communication and computation characteristics of the underlying system is known at design-time or not. We first propose a static approach to find the optimal computation migration strategy using models known at design-time. We then make a more realistic assumption that several sources of variation can affect the system's response latency (e.g., the change in computation time, bandwidth, transmission channel reliability, etc.), and propose a dynamic computation migration approach which can adaptively identify the latency optimal computation layer at runtime. We evaluate our solution using a case-study of artificial neural network based arrhythmia classification using a simulation environment as well as a real test-bed. Sina Shahhosseini, Iman Azimi, Arman Anzanpour, Axel Jantsch, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Great Lakes Symposium on VLSI | 5 |
| 2019 | Missing data resilient decision-making for healthcare IoT through personalization: A case study on maternal healthabstractRemote health monitoring is an effective method to enable tracking of at-risk patients outside of conventional clinical settings, providing early-detection of diseases and preventive care as well as diminishing healthcare costs. Internet-of-Things (IoT) technology facilitates developments of such monitoring systems although significant challenges need to be addressed in the real-world trials. Missing data is a prevalent issue in these systems, as data acquisition may be interrupted from time to time in long-term monitoring scenarios. This issue causes inconsistent and incomplete data and subsequently could lead to failure in decision making. Analysis of missing data has been tackled in several studies. However, these techniques are inadequate for real-time health monitoring as they neglect the variability of the missing data. This issue is significant when the vital signs are being missed since they depend on different factors such as physical activities and surrounding environment. Therefore, a holistic approach to customize missing data in real-time health monitoring systems is required, considering a wide range of parameters while minimizing the bias of estimates. In this paper, we propose a personalized missing data resilient decision-making approach to deliver health decisions 24/7 despite missing values. The approach leverages various data resources in IoT-based systems to impute missing values and provide an acceptable result. We validate our approach via a real human subject trial on maternity health, in which 20 pregnant women were remotely monitored for 7 months. In this setup, a real-time health application is considered, where maternal health status is estimated utilizing maternal heart rate. The accuracy of the proposed approach is evaluated, in comparison to existing methods. The proposed approach results in more accurate estimates especially when the missing window is large. Iman Azimi, Tapio Pahikkala, Amir-Mohammad Rahmani, Hannakaisa Niela-Vilén, Anna Axelin, Pasi Liljeberg |
Future Gener. Comput. Syst. | 6 |
| 2019 | Energy efficient fog-assisted IoT system for monitoring diabetic patients with cardiovascular disease
Tuan Nguyen Gia, Imed Ben Dhaou, Mai Ali, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen |
Future Gener. Comput. Syst. | 6 |
| 2019 | Energy-Aware VM Consolidation in Cloud Data Centers Using Utilization Prediction ModelabstractVirtual Machine (VM) consolidation provides a promising approach to save energy and improve resource utilization in data centers. Many heuristic algorithms have been proposed to tackle the VM consolidation as a vector bin-packing problem. However, the existing algorithms have focused mostly on the number of active Physical Machines (PMs) minimization according to their current resource requirements and neglected the future resource demands. Therefore, they generate unnecessary VM migrations and increase the rate of Service Level Agreement (SLA) violations in data centers. To address this problem, we propose a VM consolidation approach that takes into account both the current and future utilization of resources. Our approach uses a regression-based model to approximate the future CPU and memory utilization of VMs and PMs. We investigate the effectiveness of virtual and physical resource utilization prediction in VM consolidation performance using Google cluster and PlanetLab real workload traces. The experimental results show, our approach provides substantial improvement over other heuristic and meta-heuristic algorithms in reducing the energy consumption, the number of VM migrations and the number of SLA violations. Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Nguyen Trung Hieu, Hannu Tenhunen |
IEEE Trans. Cloud Comput. | 3 |
| 2018 | Approximation-aware coordinated power/performance management for heterogeneous multi-coresabstractRun-time resource management of heterogeneous multi-core systems is challenging due to i) dynamic workloads, that often result in ii) conflicting knob actuation decisions, which potentially iii) compromise on performance for thermal safety. We present a runtime resource management strategy for performance guarantees under power constraints using functionally approximate kernels that exploit accuracy-performance trade-offs within error resilient applications. Our controller integrates approximation with power knobs - DVFS, CPU quota, task migration - in coordinated manner to make performance-aware decisions on power management under variable workloads. Experimental results on Odroid XU3 show the effectiveness of this strategy in meeting performance requirements without power violations compared to existing solutions. Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Cristiana Bolchini, Nikil Dutt |
DAC | 4 |
| 2018 | Trends in On-chip Dynamic Resource ManagementabstractThe Complexity of emerging multi/many-core architectures and diversity of modern workloads demands coordinated dynamic resource management methods. We introduce a classification for these methods capturing the utilized resources and metrics. In this work, we use this classification to survey the key efforts in dynamic resource management. We first cover heuristic and optimization methods used to manage resources such as power, energy, temperature, Quality-of-Service (QoS) and reliability of the system. We then identify some of the machine learning based methods used in tuning architectural parameters in computer systems. In many cases, resource managers need to enforce design constraints during runtime with a certain level of guarantee. Hence, we also study the trend in deploying formal control theoretic approaches in order to achieve efficient and robust dynamic resource management. Kasra Moazzemi, Anil Kanduri, David Juhasz, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DSD | 6 |
| 2018 | Edge-Assisted Sensor Control in Healthcare IoTabstractThe Internet of Things is a key enabler of mobile health-care applications. However, the inherent constraints of mobile devices, such as limited availability of energy, can impair their ability to produce accurate data and, in turn, degrade the output of algorithms processing them in real-time to evaluate the patient's state. This paper presents an edge-assisted framework, where models and control generated by an edge server inform the sensing parameters of mobile sensors. The objective is to maximize the probability that anomalies in the collected signals are detected over extensive periods of time under battery-imposed constraints. Although the proposed concept is general, the control framework is made specific to a use-case where vital signs -heart rate, respiration rate and oxygen saturation- are extracted from a Photoplethysmogram (PPG) signal to detect anomalies in real-time. Experimental results show a 16.9% reduction in sensing energy consumption in comparison to a constant energy consumption with the maximum misdetection probability of 0.17 in a 24-hour health monitoring system. Delaram Amiri, Arman Anzanpour, Iman Azimi, Marco Levorato, Amir-Mohammad Rahmani, Pasi Liljeberg, Nikil Dutt |
GLOBECOM | 6 |
| 2018 | Approximation for Run-time Power ManagementabstractPerformance and energy efficiency of multi-core and many-core systems are restricted by increasing power densities and/or limited energy resources. Maximizing performance while minimizing power and energy consumption becomes challenging with emerging workloads. Approximate computing is an alternative solution that offers the required performance and energy gains, leveraging inherent error resilience of specific application domains. Dynamic power management using approximation as another knob can maximize performance and energy efficiency within fixed power budgets. Disciplined tuning of approximation along with other traditional power knobs requires efficient runtime resource management techniques. We present our strategy for using approximation as another knob for tuning the performance loss incurred in power actuation in many-core systems, which is also portable for heterogeneous multi-core systems. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg |
ISCAS | 4 |
| 2018 | Exploiting smart e-Health gateways at the edge of healthcare Internet-of-Things: A fog computing approach
Amir-Mohammad Rahmani, Tuan Nguyen Gia, Behailu Negash, Arman Anzanpour, Iman Azimi, Mingzhe Jiang, Pasi Liljeberg |
Future Gener. Comput. Syst. | 7 |
| 2018 | adBoost: Thermal Aware Performance Boosting Through Dark Silicon PatterningabstractIncreasing power densities of many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. Efficient management of critical interlinked parameters - power, performance and temperature can improve resource utilization and mitigate dark silicon. In this paper, we present a run-time resource management system for thermal aware performance boosting using a dark silicon aware run-time application mapping strategy. The mapping policy patterns inactive cores among active cores for relatively lower and even distribution of operating temperatures. This provides enough thermal headroom for boosting the frequency of active cores upon performance surges and allows sustained boosting periods, improving the performance further. We design a controller for thermal aware performance boosting that decides on efficient allocation utilization of power budget and thermal headroom obtained from patterning. Our strategy yields up to 37 percent better throughput, 29 percent lower waiting time and up to 2 x longer boosting periods, in comparison with other state-of-the-art run-time mapping policies. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Muhammad Shafique 0001, Axel Jantsch, Pasi Liljeberg |
IEEE Trans. Computers | 6 |
| 2018 | IoT-Based Remote Pain Monitoring System: From Device to Cloud PlatformabstractFacial expressions are among behavioral signs of pain that can be employed as an entry point to develop an automatic human pain assessment tool. Such a tool can be an alternative to the self-report method and particularly serve patients who are unable to self-report like patients in the intensive care unit and minors. In this paper, a wearable device with a biosensing facial mask is proposed to monitor pain intensity of a patient by utilizing facial surface electromyogram (sEMG). The wearable device works as a wireless sensor node and is integrated into an Internet of Things (IoT) system for remote pain monitoring. In the sensor node, up to eight channels of sEMG can be each sampled at 1000 Hz, to cover its full frequency range, and transmitted to the cloud server via the gateway in real time. In addition, both low energy consumption and wearing comfort are considered throughout the wearable device design for long-term monitoring. To remotely illustrate real-time pain data to caregivers, a mobile web application is developed for real-time streaming of high-volume sEMG data, digital signal processing, interpreting, and visualization. The cloud platform in the system acts as a bridge between the sensor node and web browser, managing wireless communication between the server and the web application. In summary, this study proposes a scalable IoT system for real-time biopotential monitoring and a wearable solution for automatic pain assessment via facial expressions. Geng Yang 0003, Mingzhe Jiang, Wei Ouyang 0001, Guangchao Ji, Haibo Xie, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
IEEE J. Biomed. Health Informatics | 7 |
| 2017 | Ultra-short-term analysis of heart rate variability for real-time acute pain monitoring with wearable electronicsabstractIn medical care, it is essential to assess and manage acute painful conditions adequately. Heart rate variability (HRV) analysis is based on the acquisition of electrocardiogram (ECG), which is available from both patient monitor and wearable device. As HRV analysis can reflect autonomic nervous system activity which is unconsciously regulated, HRV analysis in ultra-short-term is getting attention in indicating the reaction due to acute pain. Different HRV features in different window lengths are involved in pain monitoring studies as a signal index or part of a multi-parameter model. In this work, seven HRV features and median heart rate (HR) in ultra-short-term are evaluated for their competence in indicating experimental acute pain. Also, the choice of time window length in HRV analysis and its relation with pain detection are discussed. The results of the normalized HRV analysis from healthy volunteers show that the changes of lnRMSSD, pNN20 and median HR associated with the intensity of experimental electrical pain; and in the tests with experimental thermal pain, lnLF and ln(LF/HF) changed along with pain intensity. The fusion of the HRV features could tell pain from no pain. With either experimental pain stimulation, optimal time window length was observed around or larger than 40 seconds with better correlation analysis result and HRV feature fusion performance. Mingzhe Jiang, Riitta Mieronkoski, Amir-Mohammad Rahmani, Nora Hagelberg, Sanna Salanterä, Pasi Liljeberg |
BIBM | 6 |
| 2017 | From threads to events: Adapting a lightweight middleware for Contiki OSabstractInteroperability is one of the key requirements in the Internet of Things considering the diverse platforms, communication standards and specifications available today. Inherent resource constraints in the majority of IoT devices makes it very difficult to use existing solutions for interoperability, thus demanding new approaches. This paper presents the process of adapting a lightweight interoperability middleware for IoT, LISA, from RIOT to Contiki OS and evaluates memory and power overheads. The middleware follows a service oriented architecture and classifies devices according to available resources to assign different roles, such as Application, Service and Manager Nodes. These roles live in different tiers in a generic IoT architecture, where the Manager nodes are located in the intermediate Fog layer. To adapt to an event based kernel of Contiki, the middleware defines and handles a set of events that are used to communicate with the user application. A network of nodes is simulated to show the architecture promoted by the middleware and the results are presented. Uzair A. Noman, Behailu Negash, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
CCNC | 4 |
| 2017 | Smart energy efficient gateway for Internet of mobile thingsabstractInternet of Things (IoT) is a fast developing vision in which physical quantities are digitized, processed and analyzed. Internet of Mobile Things (IoMT) as one of new domains of IoT, due to mobility, requires a more demanding and rigorous solution in many aspects, especially in terms of energy efficiency. We propose a solution consisting of energy efficient and fast hardware platform for building IoMT Fog layer facilities. Experimental results are presented to prove superiority of the proposed hardware in several aspects to popular general purpose platforms. Igor Tcarenko, Yuxiang Huan, David Juhasz, Amir-Mohammad Rahmani, Zhuo Zou, Tomi Westerlund, Pasi Liljeberg, Lirong Zheng 0001, Hannu Tenhunen |
CCNC | 7 |
| 2017 | Self-awareness in remote health monitoring systems using wearable electronicsabstractIn healthcare, effective monitoring of patients plays a key role in detecting health deterioration early enough. Many signs of deterioration exist as early as 24 hours prior having a serious impact on the health of a person. As hospitalization times have to be minimized, in-home or remote early warning systems can fill the gap by allowing in-home care while having the potentially problematic conditions and their signs under surveillance and control. This work presents a remote monitoring and diagnostic system that provides a holistic perspective of patients and their health conditions. We discuss how the concept of self-awareness can be used in various parts of the system such as information collection through wearable sensors, confidence assessment of the sensory data, the knowledge base of the patient's health situation, and automation of reasoning about the health situation. Our approach to self-awareness provides (i) situation awareness to consider the impact of variations such as sleeping, walking, running, and resting, (ii) system personalization by reflecting parameters such as age, body mass index, and gender, and (iii) the attention property of self-awareness to improve the energy efficiency and dependability of the system via adjusting the priorities of the sensory data collection. We evaluate the proposed method using a full system demonstration. Arman Anzanpour, Iman Azimi, Maximilian Götzinger, Amir-Mohammad Rahmani, Nima Taherinejad, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DATE | 6 |
| 2017 | Autonomous Patient/Home Health Monitoring Powered by Energy HarvestingabstractThis paper presents the design of an autonomous smart patient/home health monitoring system. Both patient physiological parameters as well as room conditions are being monitored continuously to insure patient safety. The sensors are connected on an IoT regime, where the collected data is wirelessly transferred to a nearby gateway which performs preliminary data analysis, commonly referred to as fog computing, to make sure emergency personnel and healthcare providers are notified in case patient being monitored is at risk. To achieve power autonomy three energy harvesting sources are proposed, namely, solar, RF and thermal. The design of RF energy harvesting system is demonstrated, where novel multiband antenna is fabricated as well as an efficient RF- DC rectifier achieving maximum efficiency of 84%. Finally, the sensor node is tested with different type of sensors and settings while being solely powered by a Photovoltaic (PV) solar cell. Mai Ali, Tuan Nguyen Gia, Abd-Elhamid M. Taha, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen |
GLOBECOM | 6 |
| 2017 | Low-latency hardware architecture for cipher-based message authentication codeabstractCipher-based message authentication code, CMAC, is a NIST approved standard for checking message integrity and authentication. This work presents a low-latency AES architecture for CMAC. The architecture uses intensive parallel processing per round and takes advantage of the BRAM present in modern FPGA. Experimental results show that for typical IoT application, the proposed architecture has a latency of 10 clock cycles, consumes 1355 slices, 2 BRAMs and achieves a throughput of 3.8Gbps. Imed Ben Dhaou, Tuan Nguyen Gia, Pasi Liljeberg, Hannu Tenhunen |
ISCAS | 3 |
| 2017 | Low-cost fog-assisted health-care IoT system with energy-efficient sensor nodesabstractA better lifestyle starts with a healthy heart. Unfortunately, millions of people around the world are either directly affected by heart diseases such as coronary artery disease and heart muscle disease (Cardiomyopathy), or are indirectly having heart-related problems like heart attack and/or heart rate irregularity. Monitoring and analyzing these heart conditions in some cases could save a life if proper actions are taken accordingly. A widely used method to monitor these heart conditions is to use ECG or electrocardiography. However, devices used for ECG are costly, energy inefficient, bulky, and mostly limited to the ambulatory environment. With the advancement and higher affordability of Internet of Things (IoT), it is possible to establish better health-care by providing real-time monitoring and analysis of ECG. In this paper, we present a low-cost health monitoring system that provides continuous remote monitoring of ECG together with automatic analysis and notification. The system consists of energy-efficient sensor nodes and a fog layer altogether taking advantage of IoT. The sensor nodes collect and wirelessly transmit ECG, respiration rate, and body temperature to a smart gateway which can be accessed by appropriate care-givers. In addition, the system can represent the collected data in useful ways, perform automatic decision making and provide many advanced services such as real-time notifications for immediate attention. Tuan Nguyen Gia, Mingzhe Jiang, Victor K. Sarker, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen |
IWCMC | 6 |
| 2017 | Special issue on energy efficient multi-core and many-core systems, Part II
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum |
J. Parallel Distributed Comput. | 2 |
| 2017 | Performance/Reliability-Aware Resource Management for Many-Cores in Dark Silicon EraabstractAggressive technology scaling has enabled the fabrication of many-core architectures while triggering challenges such as limited power budget and increased reliability issues, like aging phenomena. Dynamic power management and runtime mapping strategies can be utilized in such systems to achieve optimal performance while satisfying power constraints. However, lifetime reliability is generally neglected. We propose a novel lifetime reliability/performance-aware resource co-management approach for many-core architectures in the dark silicon era. The approach is based on a two-layered architecture, composed of a long-term runtime reliability controller and a short-term runtime mapping and resource management unit. The former evaluates the cores' aging status w.r.t. a target reference specified by the designer, and performs recovery actions on highly stressed cores by means of power capping. The aging status is utilized in runtime application mapping to maximize system performance while fulfilling reliability requirements and honoring the power budget. Experimental evaluation demonstrates the effectiveness of the proposed strategy, which outperforms most recent state-of-the-art contributions. M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
IEEE Trans. Computers | 4 |
| 2017 | HiCH: Hierarchical Fog-Assisted Computing Architecture for Healthcare IoTabstractThe Internet of Things (IoT) paradigm holds significant promises for remote health monitoring systems. Due to their life- or mission-critical nature, these systems need to provide a high level of availability and accuracy. On the one hand, centralized cloud-based IoT systems lack reliability, punctuality and availability (e.g., in case of slow or unreliable Internet connection), and on the other hand, fully outsourcing data analytics to the edge of the network can result in diminished level of accuracy and adaptability due to the limited computational capacity in edge nodes. In this paper, we tackle these issues by proposing a hierarchical computing architecture, HiCH, for IoT-based health monitoring systems. The core components of the proposed system are 1) a novel computing architecture suitable for hierarchical partitioning and execution of machine learning based data analytics, 2) a closed-loop management technique capable of autonomous system adjustments with respect to patient’s condition. HiCH benefits from the features offered by both fog and cloud computing and introduces a tailored management methodology for healthcare IoT systems. We demonstrate the efficacy of HiCH via a comprehensive performance assessment and evaluation on a continuous remote health monitoring case study focusing on arrhythmia detection for patients suffering from CardioVascular Diseases (CVDs). Iman Azimi, Arman Anzanpour, Amir-Mohammad Rahmani, Tapio Pahikkala, Marco Levorato, Pasi Liljeberg, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2017 | Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient ApplicationsabstractPower capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Reliability-Aware Runtime Power Management for Many-Core Systems in the Dark Silicon EraabstractPower management of networked many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates considering network characteristics at runtime to achieve better performance while honoring the peak power upper bound. On the other hand, power management has a direct effect on chip temperature, which is the main driver of the aging effects. Therefore, alongside performance fulfillment, the controlling mechanism must also consider the current cores' reliability in its actuator manipulation to enhance the overall system lifetime in the long term. In this paper, we propose a multiobjective dynamic power management technique that uses current power consumption and other network characteristics including the reliability of the cores as the feedback while utilizing fine-grained voltage and frequency scaling and per-core power gating as the actuators. In addition, disturbance rejecter and reliability balancer are designed to help the controller to better smooth power consumption in the short term and reliability in the long term, respectively. Simulations of dynamic workloads and mixed criticality application profiles show that our method not only is effective in honoring the power budget while considerably boosting the system throughput, but also increases the overall system lifetime by minimizing aging effects by means of power consumption balancing. Amir-Mohammad Rahmani, M. H. Haghbayan, Antonio Miele, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Self-Adaptive Resource Management System in IaaS CloudsabstractResource management in cloud infrastructures is one of the most challenging problems due to the heterogeneity of resources, variability of the workload and scale of data centers. Efficient management of physical and virtual resources can be achieved considering performance requirements of hosted applications and infrastructure costs. In this paper, we present a self-adaptive resource management system based on a hierarchical multi-agent based architecture. The system uses novel adaptive utilization threshold mechanism and benefits from reinforcement learning technique to dynamically adjust CPU and memory thresholds for each Physical Machine (PM). It periodically runs a Virtual Machine (VM) placement optimization algorithm to keep the total resource utilization of each PM within given thresholds for improving Service Level Agreement (SLA) compliance. More-over, the algorithm consolidates VMs into the minimum number of active PMs in order to reduce the energy consumption. Experimental results on real workload traces show that our recourse management system can provide substantial improvement over other approaches in terms of performance requirements, energy consumption and the number of VM migrations. Fahimeh Farahnakian, Rami Bahsoon, Pasi Liljeberg, Tapio Pahikkala |
CLOUD | 3 |
| 2016 | A lifetime-aware runtime mapping approach for many-core systems in the dark silicon era
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
DATE | 4 |
| 2016 | Approximation knob: power capping meets energy efficiencyabstractPower Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen |
ICCAD | 4 |
| 2016 | Special issue on energy efficient multi-core and many-core systems, Part I
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum |
J. Parallel Distributed Comput. | 2 |
| 2016 | A Power-Aware Approach for Online Test Scheduling in Many-Core ArchitecturesabstractAggressive technology scaling triggers novel challenges to the design of multi-/many-core systems, such as limited power budget and increased reliability issues. Today's many-core systems employ dynamic power management and runtime mapping strategies trying to offer optimal performance while fulfilling power constraints. On the other hand, due to the reliability challenges, online testing techniques are becoming a necessity in current and near future technologies. However, state-of-the-art techniques are not aware of the other power/performance requirements. This paper proposes a power-aware non-intrusive online testing approach for many-core systems. The approach schedules software based self-test routines on the various cores during their idle periods, while honoring the power budget and limiting delays in the workload execution. A test criticality metric, based on a device aging model, is used to select cores to be tested at a time. Moreover, power and reliability issues related to the testing at different voltage and frequency levels are also handled. Extensive experimental results reveal that the proposed approach can i) efficiently test the cores within the available power budget causing a negligible performance penalty, ii) adapt the test frequency to the current cores' aging status, and iii) cover available voltage and frequency levels during the testing. M. H. Haghbayan, Amir-Mohammad Rahmani, Antonio Miele, Mohammad Fattah, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen |
IEEE Trans. Computers | 6 |
| 2015 | Utilization Prediction Aware VM Consolidation Approach for Green Cloud ComputingabstractDynamic Virtual Machine (VM) consolidation is one of the most promising solutions to reduce energy consumption and improve resource utilization in data centers. Since VM consolidation problem is strictly NP-hard, many heuristic algorithms have been proposed to tackle the problem. However, most of the existing works deal only with minimizing the number of hosts based on their current resource utilization and these works do not explore the future resource requirements. Therefore, unnecessary VM migrations are generated and the rate of Service Level Agreement (SLA) violations are increased in data centers. To address this problem, our VM consolidation method which is formulated as a bin-packing problem considers both the current and future utilization of resources. The future utilization of resources is accurately predicted using a k-nearest neighbor regression based model. In this paper, we investigate the effectiveness of VM and host resource utilization predictions in the VM consolidation task using real workload traces. The experimental results show that our approach provides substantial improvement over other heuristic algorithms in reducing energy consumption, number of VM migrations and number of SLA violations. Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
CLOUD | 3 |
| 2015 | Smart e-Health Gateway: Bringing intelligence to Internet-of-Things based ubiquitous healthcare systemsabstractThere have been significant advances in the field of Internet of Things (IoT) recently. At the same time there exists an ever-growing demand for ubiquitous healthcare systems to improve human health and well-being. In most of IoT-based patient monitoring systems, especially at smart homes or hospitals, there exists a bridging point (i.e., gateway) between a sensor network and the Internet which often just performs basic functions such as translating between the protocols used in the Internet and sensor networks. These gateways have beneficial knowledge and constructive control over both the sensor network and the data to be transmitted through the Internet. In this paper, we exploit the strategic position of such gateways to offer several higher-level services such as local storage, real-time local data processing, embedded data mining, etc., proposing thus a Smart e-Health Gateway. By taking responsibility for handling some burdens of the sensor network and a remote healthcare center, a Smart e-Health Gateway can cope with many challenges in ubiquitous healthcare systems such as energy efficiency, scalability, and reliability issues. A successful implementation of Smart e-Health Gateways enables massive deployment of ubiquitous health monitoring systems especially in clinical environments. We also present a case study of a Smart e-Health Gateway called UTGATE where some of the discussed higher-level features have been implemented. Our proof-of-concept design demonstrates an IoT-based health monitoring system with enhanced overall system energy efficiency, performance, interoperability, security, and reliability. Amir-Mohammad Rahmani, Nanda Kumar Thanigaivelan, Tuan Nguyen Gia, Jose David Granados Vergara, Behailu Negash, Pasi Liljeberg, Hannu Tenhunen |
CCNC | 6 |
| 2015 | Power-aware online testing of manycore systems in the dark silicon era
M. H. Haghbayan, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Zainalabedin Navabi, Hannu Tenhunen |
DATE | 4 |
| 2015 | Dark silicon aware runtime mapping for many-core systems: A patterning approachabstractLimitation on power budget in many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. In such systems, an efficient run-time application mapping approach can considerably enhance resource utilization and mitigate the dark silicon phenomenon. In this paper, we propose a dark silicon aware runtime application mapping approach that patterns active cores alongside the inactive cores in order to evenly distribute power density across the chip. This approach leverages dark silicon to balance the temperature of active cores to provide higher power budget and better resource utilization, within a safe peak operating temperature. In contrast with exhaustive search based mapping approach, our agile heuristic approach has a negligible runtime overhead. Our patterning strategy yields a surplus power budget of up to 17% along with an improved throughput of up to 21% in comparison with other state-of-the-art run-time mapping strategies, while the surplus budget is as high as 40% compared to worst case scenarios. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
ICCD | 4 |
| 2015 | Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approachabstractPower management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies. Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen |
ISLPED | 5 |
| 2015 | A Low-Overhead, Fully-Distributed, Guaranteed-Delivery Routing Algorithm for Faulty Network-on-ChipsabstractThis paper introduces a new, practical routing algorithm, Maze-routing, to tolerate faults in network-on-chips. The algorithm is the first to provide all of the following properties at the same time: 1) fully-distributed with no centralized component, 2) guaranteed delivery (it guarantees to deliver packets when a path exists between nodes, or otherwise indicate that destination is unreachable, while being deadlock and livelock free), 3) low area cost, 4) low reconfiguration overhead upon a fault. To achieve all these properties, we propose Maze-routing, a new variant of face routing in on-chip networks and make use of deflections in routing. Our evaluations show that Maze-routing has 16X less area overhead than other algorithms that provide guaranteed delivery. Our Maze-routing algorithm is also high performance: for example, when up to 5 links are broken, it provides 50% higher saturation throughput compared to the state-of-the-art. Mohammad Fattah, Antti Airola, Rachata Ausavarungnirun, Nima Mirzaei, Pasi Liljeberg, Juha Plosila, Siamak Mohammadi, Tapio Pahikkala, Onur Mutlu, Hannu Tenhunen |
NOCS | 5 |
| 2015 | MapPro: Proactive Runtime Mapping for Dynamic Workloads by Quantifying Ripple Effect of Applications on Networks-on-ChipabstractIncreasing dynamic workloads running on NoC-based many-core systems necessitates efficient runtime mapping strategies. With an unpredictable nature of application profiles, selecting a rational region to map an incoming application is an NP-hard problem in view of minimizing congestion and maximizing performance. In this paper, we propose a proactive region selection strategy which prioritizes nodes that offer lower congestion and dispersion. Our proposed strategy, MapPro, quantitatively represents the propagated impact of spatial availability and dispersion on the network with every new mapped application. This allows us to identify a suitable region to accommodate an incoming application that results in minimal congestion and dispersion. We cluster the network into squares of different radii to suit applications of different sizes and proactively select a suitable square for a new application, eliminating the overhead caused with typical reactive mapping approaches. We evaluated our proposed strategy over different traffic patterns and observed gains of up to 41% in energy efficiency, 28% in congestion and 21% dispersion when compared to the state-of-the-art region selection methods. M. H. Haghbayan, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
NOCS | 4 |
| 2015 | PDNOC: Partially diagonal network-on-chip for high efficiency multicore systemsabstractSummary With the constantly increasing of number of cores in multicore processors, more emphasis should be paid to the on‐chip interconnect. Performance and power consumption of an on‐chip interconnect are directly affected by the network topology. Researchers have proposed various topologies to optimize these metrics. The efficiency can also be optimized by proper mapping of applications. Therefore in this paper, we propose a novel partially diagonal network‐on‐chip (PDNOC) design that takes advantage of both heterogeneous network topology and congestion‐aware application mapping. We analyse the partially diagonal network in terms of interconnect structure, area usage, power consumption, routing algorithm and implementation complexity. The key insight that enables the PDNOC is that most communication patterns in real‐world applications are hot‐spot and bursty. We implement a full system simulation environment using SPLASH‐2 benchmarks. Performance metrics of standard mesh, concentrated mesh, full diagonal mesh and four types of the proposed PDNOC are measured in terms of network latency, application execution time and energy delay product. Evaluation results show that on average, the proposed PDNOC designs provide up to 36% improvement in execution time over concentrated mesh, and 3.6× better energy delay product over fully connected diagonal network. PDNOC design with two adjacent PD networks is a better candidate for higher efficiency, while four PD networks provide better performance. Copyright © 2014 John Wiley & Sons, Ltd. Thomas Canhao Xu, Ville Leppänen, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Using Ant Colony System to Consolidate VMs for Green Cloud ComputingabstractHigh energy consumption of cloud data centers is a matter of great concern. Dynamic consolidation of Virtual Machines (VMs) presents a significant opportunity to save energy in data centers. A VM consolidation approach uses live migration of VMs so that some of the under-loaded Physical Machines (PMs) can be switched-off or put into a low-power mode. On the other hand, achieving the desired level of Quality of Service (QoS) between cloud providers and their users is critical. Therefore, the main challenge is to reduce energy consumption of data centers while satisfying QoS requirements. In this paper, we present a distributed system architecture to perform dynamic VM consolidation to reduce energy consumption of cloud data centers while maintaining the desired QoS. Since the VM consolidation problem is strictly NP-hard, we use an online optimization metaheuristic algorithm called Ant Colony System (ACS). The proposed ACS-based VM Consolidation (ACS-VMC) approach finds a near-optimal solution based on a specified objective function. Experimental results on real workload traces show that ACS-VMC reduces energy consumption while maintaining the required performance levels in a cloud data center. It outperforms existing VM consolidation approaches in terms of energy consumption, number of VM migrations, and QoS requirements concerning performance. Fahimeh Farahnakian, Adnan Ashraf, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Ivan Porres, Hannu Tenhunen |
IEEE Trans. Serv. Comput. | 4 |
| 2014 | Energy-Aware Dynamic VM Consolidation in Cloud Data Centers Using Ant Colony SystemabstractAs the scale of a cloud data center becomes larger and larger, the energy consumption of the data center also grows rapidly. Dynamic consolidation of Virtual Machines (VMs) presents a significant opportunity to save energy by turning off unused Physical Machines (PMs) in data centers. In this paper, we present a distributed controller to perform dynamic VM consolidation to improve the resource utilizations of PMs and to reduce their energy consumption. Moreover, we use the ant colony system to find a near-optimal VM placement solution based on the specified objective function. Experimental results on the real workload traces from more than a thousand PlanetLab VMs show that the proposed approach reduces energy consumption and maintains required performance levels in a large-scale data center. Fahimeh Farahnakian, Adnan Ashraf, Pasi Liljeberg, Tapio Pahikkala, Juha Plosila, Ivan Porres, Hannu Tenhunen |
IEEE CLOUD | 3 |
| 2014 | Hierarchical Agent-Based Architecture for Resource Management in Cloud Data CentersabstractIn order to resource management in a large-scale data center, we present a hierarchical agent-based architecture. In this architecture, multi agents cooperate together to minimize the number of active physical machines according to the current resource requirements. We proposed a local agent in each physical machine (PM) to determine the PM's status and a global agent to optimizes VM placement based on PM's status. Experimental results show the proposed architecture can minimize energy consumption while maintaining an acceptable QoS. Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila |
IEEE CLOUD | 3 |
| 2014 | Adjustable contiguity of run-time task allocation in networked many-core systemsabstractIn this paper, we propose a run-time mapping algorithm, CASqA, for networked many-core systems. In this algorithm, the level of contiguousness of the allocated processors (α) can be adjusted in a fine-grained fashion. A strictly contiguous allocation (α = 0) decreases the latency and power dissipation of the network and improves the applications execution time. However, it limits the achievable throughput and increases the turnaround time of the applications. As a result, recent works consider non-contiguous allocation (α = 1) to improve the throughput traded off against applications execution time and network metrics. In contradiction, our experiments show that a higher throughput (by 3%) with improved network performance can be achieved when using intermediate α values. More precisely, up to 35% drop in the network costs can be gained by adjusting the level of contiguity compared to non-contiguous cases, while the achieved throughput is kept constant. Moreover, CASqA provides at least 32% energy saving in the network compared to other works. Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
ASP-DAC | 2 |
| 2014 | Hierarchical VM Management Architecture for Cloud Data CentersabstractEfficient energy use has become a critical issue for designing and managing of cloud data centers. Virtualization is a key technology for reducing energy cost and improving resource utilization in data centers. One of the challenges faced by virtualized data centers is to decide how to pack VMs on the least number of physical machines. This paper presents a VM management framework which is based on a multi-agent system to minimize energy consumption and Service Level Agreement (SLA) violations. The proposed agents are arranged in a three level hierarchical structure to perform VM assignment, VM placement and VM consolidation in a data center efficiently. Experimental results demonstrate that the framework achieves high quality solution in spite of its simplicity and scalability. Fahimeh Farahnakian, Pasi Liljeberg, Tapio Pahikkala, Juha Plosila, Hannu Tenhunen |
CloudCom | 2 |
| 2014 | SHiFA: System-Level Hierarchy in Run-Time Fault-Aware Management of Many-Core SystemsabstractA system-level approach to fault-aware resource management of many-core systems is proposed. The proposed approach, called SHiFA, is able to tolerate run-time faults at system level without any hardware overhead. In contrast to the existing system-level methods, network resources are also considered to be potentially faulty. Accordingly, applications are mapped onto healthy nodes of the system at run-time such that their interaction will not require the use of faulty elements. By utilizing the simple routing approach, results show 100% utilizability of PEs and 99.41% of successful mapping when up to 8 links are broken. SHiFA design is based on distributed operating systems, such that it is kept scalable for future many-core systems. A significant improvement in scalability properties is observed compared to the state-of-the-art distributed approaches. Mohammad Fattah, Maurizio Palesi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DAC | 3 |
| 2014 | Online testing of many-core systems in the Dark Silicon eraabstractAs the dark silicon era is about to embrace, it is not anymore possible to attain commensurate performance benefits by increasing the number of transistors due to thermal design power. Dark Silicon issue stresses that a fraction of silicon chip being able to switch in full frequency is dropping and designers will soon face the growing underutilization inherent in future technologies. On the other hand, by reducing the transistor size, susceptibility to internal defects drastically increases and large ranges of defects such as aging or transient faults will be shown up more frequently. In this paper, we propose an online test scheduling algorithm using software based self-test for dark silicon era to test dark cores while considering thermal design power of the system. As the dark area of the system is dynamic and reshapes at a runtime, the tested cores can be used by other applications in the near future. Empirical results show the effectiveness of the proposed algorithm in terms of applicability and fault coverage with a negligible negative impact on the system throughput. M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DDECS | 3 |
| 2014 | Dark silicon aware power management for manycore systems under dynamic workloadsabstractDark Silicon denotes the phenomenon that, due to thermal and power constraints, the fraction of transistors that can operate at full frequency is decreasing with each technology generation. We propose a PID (Proportional Integral Derivative) controller based dynamic power management method that considers an upper bound on power consumption (called the Thermal Design Power (TDP)). To avoid violation of the TDP constraint for manycore systems running highly dynamic workloads, it provides fine-grained DVFS (Dynamic Voltage and Frequency Scaling) including near-threshold operation. In addition, the method distinguishes applications with hard Real-Time, soft Real-Time and no Real-Time constraints and treats them with appropriate priorities. In simulations with dynamic workloads mixed-critical application profiles, we show that the method is effective in honoring the TDP bound and it can boost system throughput by over 43% compared to a naive TDP scheduling policy. M. H. Haghbayan, Amir-Mohammad Rahmani, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen |
ICCD | 4 |
| 2014 | Energy-Efficient Virtual Machines Consolidation in Cloud Data Centers Using Reinforcement LearningabstractDynamic consolidation techniques optimize resource utilization and reduce energy consumption in Cloud data centers. They should consider the variability of the workload to decide when idle or underutilized hosts switch to sleep mode in order to minimize energy consumption. In this paper, we propose a Reinforcement Learning-based Dynamic Consolidation method (RL-DC) to minimize the number of active hosts according to the current resources requirement. The RL-DC utilizes an agent to learn the optimal policy for determining the host power mode by using a popular reinforcement learning method. The agent learns from past knowledge to decide when a host should be switched to the sleep or active mode and improves itself as the workload changes. Therefore, RL-DC does not require any prior information about workload and it dynamically adapts to the environment to achieve online energy and performance management. Experimental results on the real workload traces from more than a thousand PlanetLab virtual machines show that RL-DC minimizes energy consumption and maintains required performance levels. Fahimeh Farahnakian, Pasi Liljeberg, Juha Plosila |
PDP | 2 |
| 2014 | Mixed-Criticality Run-Time Task Mapping for NoC-Based Many-Core SystemsabstractContiguous processor allocation improves both the network and the application performance, by decreasing the congestion probability among communication of different applications. Consequently, the average, standard deviation and worst-case latency of the network is decreased significantly. This makes the contiguous allocation a good solution for time-critical applications with bounded deadlines. On the other hand, non-contiguous allocation will increase the system throughput significantly. Isolated nodes are utilized and more applications can finish their job in a time unit. However, this will lead to poor network metrics, unsuitable for real-time applications. In this work, we combine these two approaches in order to manage workloads with mixed-critical characteristics. Real-time applications are mapped contiguously, while non-critical applications are allowed to get dispersed over the available system nodes. Results show over 50% improvement in worst-case latency and 100 times improvement in deadline misses. Mohammad Fattah, Amir-Mohammad Rahmani, Thomas Canhao Xu, Anil Kanduri, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 5 |
| 2014 | Multi Rectangle Modeling Approach for Application Mapping on a Many-Core SystemabstractThe importance of first node selection in run-time resource management is shown in our previous work, SHiC. It is desired in SHiC to find the optimum node in an agile manner. Accordingly, the current mapping picture of the system is simplified to SHiC by modeling each application as a rectangle of occupied nodes. However, the algorithm performance can be influenced significantly with the rectangle model of each application. In this work, we introduce a precise description of our new accurate rectangle modeling algorithm. Moreover, we show that it is not sufficient to always model an application with only one rectangle, as dispersion of the allocated nodes is an irrepressible phenomenon. Accordingly, our algorithm enables to model a mapped application with several rectangles by tuning the model accuracy against the algorithm complexity. However, the algorithm is not in the critical path of the applications executions. Our results shows up to 5% reduction in power dissipation of the network. Igor Tcarenko, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2014 | Path-Based Partitioning Methods for 3D Networks-on-Chip with Minimal Adaptive RoutingabstractCombining the benefits of 3D ICs and Networks-on-Chip (NoCs) schemes provides a significant performance gain in Chip Multiprocessors (CMPs) architectures. As multicast communication is commonly used in cache coherence protocols for CMPs and in various parallel applications, the performance of these systems can be significantly improved if multicast operations are supported at the hardware level. In this paper, we present several partitioning methods for the path-based multicast approach in 3D mesh-based NoCs, each with different levels of efficiency. In addition, we develop novel analytical models for unicast and multicast traffic to explore the efficiency of each approach. In order to distribute the unicast and multicast traffic more efficiently over the network, we propose the Minimal and Adaptive Routing (MAR) algorithm for the presented partitioning methods. The analytical and experimental results show that an advantageous method named Recursive Partitioning (RP) outperforms the other approaches. RP recursively partitions the network until all partitions contain a comparable number of switches and thus the multicast traffic is equally distributed among several subsets and the network latency is considerably decreased. The simulation results reveal that the RP method can achieve performance improvement across all workloads while performance can be further improved by utilizing the MAR algorithm. Nineteen percent average and 42 percent maximum latency reduction are obtained on SPLASH-2 and PARSEC benchmarks running on a 64-core CMP. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, José Flich, Hannu Tenhunen |
IEEE Trans. Computers | 3 |
| 2014 | High-Performance and Fault-Tolerant 3D NoC-Bus Hybrid Architecture Using ARB-NET-Based Adaptive Monitoring PlatformabstractThe emerging three-dimensional integrated circuits (3D-ICs) achieve greater device integration and enhanced system performance at lower cost and reduced area footprint, thereby offering higher order of connectivity and greater design choices and possibilities. To exploit the intrinsic capability of reduced communication distances in 3D-ICs, three-dimensional NoC-bus hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored possibility to implement an efficient system-wide monitoring network. In this paper, an efficient three-dimensional NoC architecture is proposed which is optimized for system performance, power consumption, and reliability. The mechanism benefits from a congestion-aware and bus failure-tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we have integrated a low-cost monitoring platform on top of the three-dimensional NoC-Bus Hybrid mesh architecture that can be efficiently used for various system management purposes such as traffic monitoring, fault tolerance, and thermal management. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware and interlayer fault-tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of information generated within bus arbiters. Compared to recently proposed stacked mesh three-dimensional NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power, performance, and reliability improvements with a negligible hardware overhead. Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
IEEE Trans. Computers | 4 |
| 2014 | Adaptive load balancing in learning-based approaches for many-core embedded systems
Fahimeh Farahnakian, Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
J. Supercomput. | 4 |
| 2014 | Special section on advances in methods for adaptive multicore systems
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Supercomput. | 2 |
| 2013 | Smart hill climbing for agile dynamic mapping in many-core systemsabstractStochastic hill climbing algorithm is adapted to rapidly find the appropriate start node in the application mapping of network-based many-core systems. Due to highly dynamic and unpredictable workload of such systems, an agile run-time task allocation scheme is required. The scheme is desired to map the tasks of an incoming application at run-time onto an optimum contiguous area of the available nodes. Contiguous and un-fragmented area mapping is to settle the communicating tasks in close proximity. Hence, the power dissipation, the congestion between different applications, and the latency of the system will be significantly reduced. To find an optimum region, we first propose an approximate model that quickly estimates the available area around a given node. Then the stochastic hill climbing algorithm is used as a search heuristic to find a node that has the required number of available nodes around it. Presented agile climber takes the steps using an adapted version of hill climbing algorithm named Smart Hill Climbing, SHiC, which takes the runtime status of the system into account. Finally, the application mapping is performed starting from the selected first node. Experiments show significant gain in the mapping contiguousness which results in better network latency and power dissipation, compared to state-of-the-art works. Mohammad Fattah, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
DAC | 3 |
| 2013 | Enhanced fault-tolerant Network-on-Chip architecture using hierarchical agentsabstractThe reliability is a vital aspect in the design of Network-on-Chip (NoC) based systems because a fault in communication medium may cause an overall system failure. On the other hand, performance degradation is an inescapable consequence of fault-tolerant architectures. In this paper, we propose a fault-tolerant NoC architecture that attains higher performance by using low cost agents in a hierarchical manner. These agents which are distributed all over the network, collect, process, and distribute different fault information. Moreover, we propose an enhanced fault-tolerant and congestion-aware routing method that exploits the classified fault information related to the permanent faults that might occur inside the links, network interfaces and different parts of the routers. The experimental results reveal that the proposed NoC architecture imposes small area and power overheads. Mojtaba Valinataj, Pasi Liljeberg, Juha Plosila |
DDECS | 2 |
| 2013 | DyXYZ: Fully Adaptive Routing Algorithm for 3D NoCsabstractTraditional methods in 3D NoCs simply use a deterministic routing algorithm to deliver packets from a source to a destination node. However, deterministic methods are unable to distribute the traffic load over the network, which results in degrading the performance. In this paper, we present a fully adaptive routing algorithm for 3D NoCs, named DyXYZ. In DyXYZ, the congestion information at the input buffer of the neighboring routers is used as congestion metric to select among the output channels. This algorithm is proven to be deadlock free by using 4, 4, and 2 virtual channels along the X, Y, and Z dimensions, respectively. Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen |
PDP | 5 |
| 2013 | Enhancing Performance of 3D Interconnection Networks using Efficient Multicast Communication ProtocolabstractThree-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. They also provide greater design flexibility by allowing heterogeneous integration. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement efficient multicast routings for 3D networks-on-chip. In this paper, we propose a novel multicast partitioning and routing strategy for the 3D NoC-Bus Hybrid mesh architectures to enhance the overall system performance and reduce the power consumption. The proposed architecture exploits the beneficial attribute of a single-hop (bus-based) interlayer communication of the 3D stacked mesh architecture to provide high-performance hardware multicast support. To this end, a customized partitioning method and an efficient routing algorithm are presented to reduce the average hop count and latency of the network. Compared to the recently proposed 3D NoC architectures being capable of supporting hardware multicasting, our extensive simulations with different traffic profiles reveal that our architecture using the proposed multicast routing strategy can help achieve significant performance improvements. Sanaz Rahimi Moosavi, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2013 | Cluster-based topologies for 3D Networks-on-Chip using advanced inter-layer bus architecture
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Comput. Syst. Sci. | 3 |
| 2013 | Developing a power-efficient and low-cost 3D NoC using smart GALS-based vertical channels
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Comput. Syst. Sci. | 2 |
| 2013 | A systematic reordering mechanism for on-chip networks using efficient congestion-aware method
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Syst. Archit. | 3 |
| 2013 | Special issue on network-based many-core embedded systems
Masoud Daneshtalab, Pasi Liljeberg, Mehdi Modarressi, Leandro Soares Indrusiak |
J. Syst. Archit. | 2 |
| 2013 | Design space exploration of thermal-aware many-core systems
Kameswar Rao Vaddina, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila |
J. Syst. Archit. | 4 |
| 2013 | Optimal placement of vertical connections in 3D Network-on-Chip
Thomas Canhao Xu, Gert Schley, Pasi Liljeberg, Martin Radetzki, Juha Plosila, Hannu Tenhunen |
J. Syst. Archit. | 3 |
| 2012 | ARB-NET: A novel adaptive monitoring platform for stacked mesh 3D NoC architecturesabstractThe emerging three-dimensional integrated circuits (3D ICs) offer a promising solution to mitigate the barriers of interconnect scaling in modern systems. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is presented taking advantage from viable information generated within bus arbiters. Our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help achieving significant power and performance improvements compared to recently proposed stacked mesh 3D NoCs. Amir-Mohammad Rahmani, Khalid Latif 0002, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
ASP-DAC | 4 |
| 2012 | A Cluster-Based Core Protection Technique for Networks-on-ChipabstractPartial Virtual channel Sharing (PVS) architecture has been proposed to enhance the performance of Networks-on-Chip (NoC) based systems. In this paper, a cluster based processing core protection technique for NoC systems using PVS approach is presented. In case of network level faults, the processing core of faulty node can use any other router in the cluster for transmission or reception of data packets with proposed architecture. Simulation results show significant reduction in average packet latency at the expense of negligible area overhead. Khalid Latif 0002, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen, Tiberiu Seceleanu |
COMPSAC | 3 |
| 2012 | CATRA- congestion aware trapezoid-based routing algorithm for on-chip networksabstractCongestion occurs frequently in Networks-on-Chip when the packets demands exceed the capacity of network resources. Congestion-aware routing algorithms can greatly improve the network performance by balancing the traffic load in adaptive routing. Commonly, these algorithms either rely on purely local congestion information or take into account the congestion conditions of several nodes even though their statuses might be out-dated for the source node, because of dynamically changing congestion conditions. In this paper, we propose a method to utilize both local and non-local network information to determine the optimal path to forward a packet. The non-local information is gathered from the nodes that not only are more likely to be chosen as intermediate nodes in the routing path but also provide up-to-date information to a given node. Moreover, to collect and deliver the non-local information, a distributed propagation system is presented. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DATE | 3 |
| 2012 | Power and Thermal Analysis of Stacked Mesh 3D NoC Using AdaptiveXYZ Routing AlgorithmabstractThree-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement an integrated low-cost system-wide monitoring network. In this paper, a generic monitoring and management platform called ARB-NET is presented. Based on the ARB-NET monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is provided which takes advantage from viable information generated within the monitoring network. In addition, we address both the power and thermal issues of a stacked mesh 3D network on chips using AdaptiveXYZ routing. To this end, a thermal model of a 3D stacked NoC system in a modern flip-chip package is developed. Thermal and power analysis are performed in order to investigate the impact of the proposed adaptive routing from the power and thermal perspectives. Our experiments with a videoconference encoder as a real application show significant power, performance and peak temperature improvements compared to a typical stacked mesh 3D NoC. Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DSD | 3 |
| 2012 | NoC-AXI interface for FPGA-based MPSoC platformsabstractStreaming applications are a keystone in several emerging multimedia services like DVB-IPTV, VoD and on-line gaming. Due to the high computing requirements and real-time constraints inherent to this kind of applications multi-processor system-on-chip (MPSoCs) have been proposed as a solution. In addition, the FPGA technology has become popular among systems-on-chip (SoCs) designers due to its low development cost and short time to market. Here we present a FPGA-based MPSoC platform for streaming applications where the important component of this platform is the AXI interface. Marco Ramírez 0001, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg |
FPL | 4 |
| 2012 | CoNA: Dynamic application mapping for congestion reduction in many-core systemsabstractIncreasing the number of processors in a single chip toward network-based many-core systems requires a run-time task allocation algorithm. We propose an efficient mapping algorithm that assigns communicating tasks of incoming applications onto resources of a many-core system utilizing Network-on-Chip paradigm. In our contiguous neighborhood allocation (CoNA) algorithm, we target at the reduction of both internal and external congestion due to detrimental impact of congestion on the network performance. We approach the goal by keeping the mapped region contiguous and placing the communicating tasks in a close neighborhood. A completely synthesizable simulation environment where none of the system objects are assumed to be ideal is provided. Experiments show at least 40% gain in different mapping cost functions, as well as 16% reduction in average network latency compared to existing algorithms. Mohammad Fattah, Marco Ramírez 0001, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila |
ICCD | 4 |
| 2012 | HARAQ: Congestion-Aware Learning Model for Highly Adaptive Routing Algorithm in On-Chip NetworksabstractThe occurrence of congestion in on-chip networks can severely degrade the performance due to increased message latency. In mesh topology, minimal methods can propagate messages over two directions at each switch. When shortest paths are congested, sending more messages through them can deteriorate the congestion condition considerably. In this paper, we present an adaptive routing algorithm for on-chip networks that provide a wide range of alternative paths between each pair of source and destination switches. Initially, the algorithm determines all permitted turns in the network including 180-degree turns on a single channel without creating cycles. The implementation of the algorithm provides the best usage of all allowable turns to route messages more adaptively in the network. On top of that, for selecting a less congested path, an optimized and scalable learning method is utilized. The learning method is based on local and global congestion information and can estimate the latency from each output channel to the destination region. Masoumeh Ebrahimi, Masoud Daneshtalab, Fahimeh Farahnakian, Juha Plosila, Pasi Liljeberg, Maurizio Palesi, Hannu Tenhunen |
NOCS | 5 |
| 2012 | Generic Monitoring and Management Infrastructure for 3D NoC-Bus Hybrid ArchitecturesabstractThree-dimensional integrated circuits (3D ICs) achieve enhanced system integration and improved performance at lower cost and reduced area footprint. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed which provides performance, power consumption, and area benefits. Besides its various advantages, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes such as traffic monitoring, thermal management and fault tolerance. The proposed generic monitoring and management infrastructure called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring and management platform, a fully congestion-aware and inter-layer fault tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of viable information generated using bus arbiter network. In addition, we propose a thermal monitoring and management strategy on top of our ARB-NET infrastructure. Compared to recently proposed stacked mesh 3D NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power and performance improvements while preserving the system reliability with negligible hardware overhead. Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 4 |
| 2012 | LEAR - A Low-Weight and Highly Adaptive Routing Method for Distributing Congestions in On-chip NetworksabstractCongestion-aware routing algorithms can improve network throughput by avoiding packets to be routed through congested areas. In this paper, we propose a minimal/non-minimal routing algorithm to alleviate congestion in the network by making use of all available paths between sources and destinations. The simplicity of the proposed algorithm provides a cost and power efficient solution for Networks-on-Chip while the high degree of adaptive ness, achieved by using an additional virtual channel along the Y dimension, leads to an increased performance. In this method, different restrictions are imposed on the use of each virtual channel, so that the prohibited turns in one virtual channel are permitted in the other one. By fully exploiting of the eligible turns in the network, a large number of output channels can be provided by the proposed method. Based on this method, a packet is routed along the non-minimal path when the neighboring routers in the minimal path are congested. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2012 | An Efficient Hybridization Scheme for Stacked Mesh 3D NoC ArchitectureabstractThree-dimensional (3D) integration is a viable design paradigm to overcome the existing interconnect bottleneck in integrated systems and enhance system power/performance characteristics. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, stacked mesh 3D NoC architecture was proposed. However, this architecture suffers from naive and straightforward hybridization between NoC and bus media. In this paper, an efficient hybridization scheme is presented to enhance system performance, power consumption, and area of stacked mesh 3D NoC architectures. By utilizing a routing rule called LastZ the proposed hybridization scheme offers many advantages investigated in detail to emphasize the significant achievements. Our extensive simulations with synthetic and real benchmarks, including an integrated videoconference application show that compared to a typical 3D NoC-Bus Hybrid Mesh architecture, our hybridization scheme achieves significant power, performance, and area improvements. Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 2 |
| 2012 | Memory-Efficient On-Chip Network With Adaptive InterfacesabstractTo achieve higher memory bandwidth in network-based multiprocessor architectures, multiple dynamic random access memories can be accessed simultaneously. In such architectures, not only resource utilization and latency are the critical issues but also a reordering mechanism is required to deliver the response transactions of concurrent memory accesses in-order. In this paper, we present a memory-efficient on-chip network architecture to cope with these issues efficiently. Each node of the network is equipped with a novel network interface (NI) to deal with out-of-order delivery, and a priority-based router to decrease the network latency. The proposed NI exploits a streamlined reordering mechanism to handle the in-order delivery and utilizes the advance extensible interface transaction-based protocol to maintain compatibility with existing intellectual property cores. To improve the memory utilization and reduce the memory latency, an optimized memory controller is integrated in the presented NI. Experimental results with synthetic test cases demonstrate that the proposed on-chip network architecture provides significant improvements in average network latency (16%), average memory access latency (19%), and average memory utilization (22%). Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Enhancing Performance of NoC-Based Architectures Using Heuristic Virtual-Channel Sharing ApproachabstractThis paper presents a novel virtual-channel (VC) sharing technique for NoC architecture. The proposed architecture improves the utilization of resources to enhance the performance with minimal overheads. A heuristic approach towards the proper VC sharing strategy is proposed, which is performed by an adaptive algorithm that configures the VC sharing based on link load parameters. Architectural design to realize the adaptive VC sharing in generic router is elaborated. The technique can be applied to any NoC architecture, including 3-D NoCs. Extensive quantitative experiments with synthetic and real benchmarks, including an integrated video conference application, demonstrate considerable improvement in area and power efficiency compared to existing VC-based 2D/3D NoC architectures. Khalid Latif 0002, Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Tiberiu Seceleanu, Pasi Liljeberg, Hannu Tenhunen |
COMPSAC | 5 |
| 2011 | Evaluating Sustainability, Environmental Assessment and Toxic Emissions during Manufacturing Process of RFID Based SystemsabstractThe present state of the art research in the direction of embedded systems demonstrate that analysis of life-cycle, sustainability and environmental assessment have not been a core focus for researchers. To maximize a researcher's contribution in formulating environmentally friendly products, devising green manufacturing processes and services, there is a strong need to enhance life-cycle awareness and sustainability understandings among embedded systems researchers, so that the next generation of engineers will be able to realize the goal of a sustainable life-cycle. In this work an attempt has been made to investigate and evaluate the life-cycle management and environmental assessment in fabricating processes of the RFID based systems. We have chosen a general life cycle assessment approach which involves the collection and evaluation of quantitative data on the inputs and outputs of materials and energy associated with the RFID based systems. Based on the developed generic models, we have obtained the results in terms of environmental emissions for a production of paper substrate printed RFID antennas. We also make an attempt to raise several sustainability issues and quantify the toxic emissions during the manufacturing process. Rajeev Kumar Kanth, Pasi Liljeberg, Hannu Tenhunen, Qiansu Wan, Yasar Amin, Botao Shao, Qiang Chen 0014, Lirong Zheng 0001 |
DASC | 2 |
| 2011 | Optimal number and placement of Through Silicon Vias in 3D Network-on-ChipabstractIn this paper, we analyze the performance impact of different number of Through Silicon Vias (TSVs) in 3D Network-on-Chip (NoC). The adoption of a 3D NoC design depends on the performance and manufacturing cost of the chip. Therefore, a study of the placement of the TSV, that connects different layers of a 3D chip, is crucial. A 64-core 3D NoC is modeled based on state-of-the-art 2D chips. We discuss the number of TSVs required for a 3D NoC. Different placements of layer-layer connections are explored. We present benchmark results using a cycle accurate full system simulator based on realistic workloads. Experiments show that under different workloads, the average network latencies in two configurations (full and quarter connection) are reduced by 14.78% and 7.38% respectively, compared with the one-eighth connection design. The improvement of performance is a trade-off of manufacturing cost. Our analysis and experiment results provide a guideline for selecting optimal number of TSVs in 3D NoCs. Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen |
DDECS | 2 |
| 2011 | Enhancing Performance Sustainability of Fault Tolerant Routing Algorithms in NoC-Based ArchitecturesabstractReliability of embedded systems and devices is becoming a challenge with technology scaling. To deal with the reliability issues, fault tolerant solutions are needed. The design paradigm for future System-on-Chip (SoC) implementation is Network-on-Chip (NoC). Fault tolerance in NoC can be achieved at many abstraction levels. Many fault tolerant architectures and routing algorithms have already been proposed for NoC but the utilization of resources, affected indirectly by faults is yet to be addressed. In this paper, we propose a NoC architecture, which sustains the overall system performance by utilizing resources, which cannot be used by other architectures under faults. An approach towards a proper virtual-channel (VC) sharing strategy is proposed, based on communication bandwidth requirements. The technique can be applied to any NoC architecture, including 3-D NoCs. Extensive quantitative experiments with synthetic benchmarks, including uniform, transpose and negative exponential distribution (NED), demonstrate considerable improvement in terms of performance sustainability under faulty conditions compared to existing VC-based NoC architectures. Khalid Latif 0002, Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Tiberiu Seceleanu, Pasi Liljeberg, Hannu Tenhunen |
DSD | 5 |
| 2011 | LastZ: An Ultra Optimized 3D Networks-on-Chip Architectureabstract3D IC technology enables NoC architectures to offer greater device integration and shorter interlayer interconnects. The primary 3D NoC architectures such as Symmetric 3D Mesh NoC could not exploit the beneficial feature of a negligible inter-layer distance in 3D chips. To cope with this, 3D NoC-Bus Hybrid architecture was proposed which is a hybrid between packet-switched network and a bus. This architecture is feasible providing both performance and area benefits, while still suffering from naive and straightforward hybridization between NoC and bus media. In this paper, an ultra optimized hybridization scheme is proposed to enhance system performance, power consumption, area and thermal issues of 3D NoC-Bus Hybrid Mesh. The scheme benefits from a rule called LastZ which enables ultra optimization of the inter-layer communication architecture. In addition, we present a wrapper to preserve the backward compatibility of the proposed architecture for connecting with the existing network interfaces. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. Our extensive simulations demonstrate significant area, power, and performance improvements compared to a typical 3D NoC-Bus Hybrid Mesh architecture. Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DSD | 2 |
| 2011 | Thermal Analysis of Job Allocation and Scheduling Schemes for 3D Stacked NoC'sabstractThree-dimensional technology offers greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In this work, a 3D thermal model of a stacked network-on-chip system is developed and thermal analysis is performed in order to analyze different job allocation and scheduling schemes using finite element simulations. The steady-state heat transfer analysis on the 3D stacked structure has been performed. We have analyzed the effect of variation of die power consumption, with and without hotspots, on peak temperatures in different layers of the stack. The optimal die placement solution is also provided based on the maximum temperature attained by the individual silicon dies. Kameswar Rao Vaddina, Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila |
DSD | 4 |
| 2011 | A Minimal Average Accessing Time Scheduler for Multicore Processors
Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen |
ICA3PP (2) | 2 |
| 2011 | Exploring partitioning methods for 3D Networks-on-Chip utilizing adaptive routing modelabstractThree-Dimensional (3D) integration is a solution to the interconnect bottleneck in Two-Dimensional (2D) MultiProcessor System on Chip (MPSoC). 3D IC design improves performance and decreases power consumption by replacing long horizontal interconnects with shorter vertical ones. As the multicast communication is utilized commonly in various parallel applications, the performance can be significantly improved by supporting of multicast operations at the hardware level. In this paper, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method named Recursive Partitioning (RP) in which the network is recursively partitioned until all partitions contain comparable number of nodes. By this approach, the multicast traffic is distributed among several subsets and the network latency is considerably decreased. We also present Minimal Adaptive Routing (MAR) algorithm for the unicast and multicast traffic in 3D-mesh Networks-on-Chip (NoCs). The idea behind the MAR algorithm is utilizing the Hamiltonian path to provide a set of alternative paths. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 3 |
| 2011 | Congestion aware, fault tolerant, and thermally efficient inter-layer communication scheme for hybrid NoC-bus 3D architecturesabstractThree-dimensional IC technology offers greater device integration and shorter interlayer interconnects. In order to take advantage of these attributes, 3D stacked mesh architecture was proposed which is a hybrid between packet-switched network and a bus. Stacked mesh is a feasible architecture which provides both performance and area benefits, while suffering from inefficient intermediate buffers. In this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we hybridize the proposed adaptive routing with available algorithms to mitigate the thermal issues by herding most of the switching activities closer to the heat sink. Our extensive simulations with synthetic and real benchmarks, including the one with an integrated videoconference application, demonstrate significant power, performance, and peak temperature improvements compared to a typical stacked mesh 3D NoC. Amir-Mohammad Rahmani, Pasi Liljeberg, Khalid Latif 0002, Juha Plosila, Kameswar Rao Vaddina, Hannu Tenhunen |
NOCS | 2 |
| 2011 | A Stacked Mesh 3D NoC Architecture Enabling Congestion-Aware and Reliable Inter-layer CommunicationabstractIn this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. Stacked mesh is a feasible architecture which takes advantage of the short inter-layer wiring delays, while suffering from inefficient intermediate buffers. To cope with this, an inter-layer communication mechanism is developed to enhance the buffer utilization, load balancing, and system fault-tolerance. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm for vertical communication. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. In addition, a video conference encoder has been used as a real application for system analysis. Our extensive experiments show significant power and performance improvements compared to a typical stacked mesh 3D NoC. Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2011 | Agent-based on-chip network using efficient selection methodabstractCongestion in on-chip networks may cause many drawbacks in multiprocessor systems including throughput reduction, increase in latency, and additional power consumption. Furthermore, conventional congestion control methods, employed for on-chip networks, cannot efficiently collect congestion information and distribute them over the on-chip network. In this paper, we present a novel structure for on-chip networks, named Agent-based Network-on-Chip (ANoC), to diagnose the congested areas. In addition to the presented structure, an efficient Congestion-Aware Selection (CAS) method is proposed to reduce overall network latency. CAS is capable of selecting an appropriate output channel to route packets along a less congested path. 29% average and 35% maximum latency reduction are achieved on SPLASH-2 and PARSEC benchmarks running on a 36-core Chip Multi-Processor. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
VLSI-SoC | 3 |
| 2011 | A generic adaptive path-based routing method for MPSoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen |
J. Syst. Archit. | 4 |
| 2010 | Partitioning methods for unicast/multicast traffic in 3D NoC architectureabstractAs the scale of integration grows, the interconnection problem becomes one of the major design considerations of Multi Processor System on Chip (MPSoC). In recent years, many researchers have conducted studies on 3D IC designs stacking multiple layers on top of each other. In order to decrease the transmission delay of unicast/multicast messages in a network based multicore system, the network is divided into several partitions. In this paper, we first introduce a novel idea of balanced partitioning that allows the network to be partitioned effectively. Then, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method based on the idea of balanced partitioning to provide a high degree of parallelism with a considerable reduction of packet delay in unicast/multicast traffic. Simulations are provided to evaluate and compare the performance of proposed methods. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
DDECS | 3 |
| 2010 | Developing reconfigurable FIFOs to optimize power/performance of Voltage/Frequency Island-based networks-on-chipabstractNetwork-on-chip architectures partitioned into several Voltage/Frequency Islands (VFIs) have been proposed to alleviate problems related to integration, excessive energy consumption and clock distribution. The architecture is composed of synchronous switches that communicate with each other using bi-synchronous FIFOs. However, these FIFOs are not needed if adjacent switches belong to the same clock domain. In this paper, a Reconfigurable Synchronous/Bi-Synchronous (RSBS) FIFO is proposed which can operate in either synchronous or bi-synchronous mode. The FIFO is scalable and synthesizable in synchronous standard cells and also a technique for mesochronous adaptation has been recommended. In addition, some techniques are suggested to show how the FIFO could be utilized in a VFI-based NoC. Our results reveal that compared to a non-reconfigurable system architecture, the RSBS FIFOs help to achieve up to 15% savings in average power consumption of NoC switches and 29% improvement in total average packet latency in the case of MPEG-4 encoder application. Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DDECS | 2 |
| 2010 | A fault-tolerant and congestion-aware routing algorithm for Networks-on-ChipabstractThis paper presents a fault-tolerant routing algorithm for mesh-based Networks-on-Chip (NoC) with faulty links. It is a distributed, adaptive and congestion-aware routing algorithm where only two virtual channels are used for both adaptiveness and fault-tolerance. The proposed routing method has a multilevel fault-tolerance capability and therefore it is capable to tolerate more faulty links in more complicated faulty situations with additional hardware costs. The network performance, fault-tolerance capability and hardware overhead are evaluated through appropriate simulations. The experimental results show that the overall reliability of a Network-on-Chip is significantly enhanced against multiple link failures or partially faulty routers with only a small hardware overhead. Mojtaba Valinataj, Siamak Mohammadi, Juha Plosila, Pasi Liljeberg |
DDECS | 4 |
| 2010 | Power-aware NoC router using central forecasting-based dynamic virtual channel allocationabstractIn this paper, we propose a high performance central dynamic virtual channel allocation mechanism for on-chip routers. This central management unit devotes each input port a number of virtual channels (VC) among a shared VC bank based on a traffic forecasting technique. The forecasting technique exploits the link and VC utilizations in predicting the traffic. Based on the predicted traffic, for each input port, the number of active virtual channels may be increased, decreased, or kept unchanged. The clock-gating power management technique is used to activate/deactivate the VCs. Simulation results using uniform and Negative Exponential Distribution (NED) traffic profiles show that a considerable power savings in the virtual channels and overall router power consumption may be achieved especially in low traffic loads. The area overhead of the technique is negligible. Amir-Mohammad Rahmani, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
ISCAS | 3 |
| 2010 | A Low-Latency and Memory-Efficient On-chip NetworkabstractUsing multiple SDRAMs in MPSoCs and NoCs to increase memory parallelism is very common nowadays. In-order delivery, resource utilization, and latency are the most critical issues in such architectures. In this paper, we present a novel network interface architecture to cope with these issues efficiently. The proposed network interface exploits a resourceful reordering mechanism to handle the in-order delivery and to increase the resource utilization. A brilliant memory controller is efficiently integrated into this network interface to improve the memory utilization and reduce both memory and network latencies. In addition, to bring compatibility with existing IP cores the proposed network interface utilizes AXI transaction based protocol. Experimental results with synthetic test cases demonstrate that the proposed architecture gives significant improvements in average network latency (12%), average memory access latency (19%), and average memory utilization (22%). Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 3 |
| 2010 | A High-Performance Network Interface Architecture for NoCs Using Reorder Buffer SharingabstractIncreasing memory parallelism in MPSoCs to provide higher memory bandwidth is achieved by accessing multiple memories simultaneously. Inasmuch as the response transactions of concurrent memory accesses must be in-order, a reordering mechanism is required. To our knowledge the resource utilization of conventional reordering mechanisms is low. In this paper, we present a novel network interface architecture for on-chip networks to increase the resource utilization and to improve overall performance. Also, based on the proposed architecture, a hybrid network interface is presented to integrate both memory and processor in a tile. The proposed architecture exploits AXI transaction based protocol to be compatible with existing IP cores. Experimental results with synthetic test cases demonstrate that the proposed architecture outperforms the conventional architecture in terms of latency. Also, the cost of the presented architecture is evaluated with UMC 0.09 ¿ m technology. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 3 |
| 2010 | HAMUM - A Novel Routing Protocol for Unicast and Multicast Traffic in MPSoCsabstractMany parallel applications in MPSoCs take advantage of multicast communication. Several multicast schemes such as path-based, tree-based, and unicast-based have been proposed in interconnection networks. Path-based multicast scheme has been proven to be more efficient than the other schemes in on-chip interconnection network. A new adaptive routing model based on Hamiltonian path for both the multicast and unicast traffics, called Hamiltonian Adaptive Multicast and Unicast Model (HAMUM), is presented. Results obtained in both multicast and mixed traffic models show that the proposed adaptive algorithm for multicast aspect has lower latency and power dissipation compared to previously proposed path-based multicasting algorithms with less than 0.5% hardware overhead. Additionally, for the unicast aspect the proposed adaptive model outperforms the other unicast turn models. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
PDP | 3 |
| 2010 | Self-Adaptive System for Addressing Permanent Errors in On-Chip InterconnectsabstractWe present a self-contained adaptive system for detecting and bypassing permanent errors in on-chip interconnects. The proposed system reroutes data on erroneous links to a set of spare wires without interrupting the data flow. To detect permanent errors at runtime, a novel in-line test (ILT) method using spare wires and a test pattern generator is proposed. In addition, an improved syndrome storing-based detection (SSD) method is presented and compared to the ILT method. Each detection method (ILT and SSD) is integrated individually into the noninterrupting adaptive system, and a case study is performed to compare them with Hamming and Bose-Chaudhuri-Hocquenghem (BCH) code implementations. In the presence of permanent errors, the probability of correct transmission in the proposed systems is improved by up to 140% over the standalone Hamming code. Furthermore, our methods achieve up to 38% area, 64% energy, and 61% latency improvements over the BCH implementation at comparable error performance. Teijo Lehtonen, David Wolpert 0001, Pasi Liljeberg, Juha Plosila, Paul Ampadu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Self-timed thermal sensing and monitoring of multicore systemsabstractAs the number of cores increases thermal challenges increase, thereby degrading the performance and reliability of the system. We approach this challenge with a self-timed thermal monitoring method which is based on the use of thermal sensors. Since leakage currents are sensitive to temperature and increase with scaling, we propose the use of a leakage current based thermal sensing for monitoring purposes. In this work we have implemented a novel thermal sensing circuit in 65 nm CMOS technology, which converts analog temperature information into digital form. We have also proposed a novel thermal sensing and monitoring interconnection network structure based on self-timed signaling, comprising of an encoder/transmitter and decoder/ receiver. We have performed power supply noise, additive noise on sensor input signal and dynamic power supply voltage variation analysis on the thermal sensing circuit and show that it is robust enough under different operating temperatures. Kameswar Rao Vaddina, Ethiopia Nigussie, Pasi Liljeberg, Juha Plosila |
DDECS | 3 |
| 2009 | An Adaptive Unicast/Multicast Routing Algorithm for MPSoCsabstractSeveral parallel applications in MPSoCs take advantage of multicast communication. Path-based multicast scheme has been proven to be more efficient than the others multicast schemes in on-chip interconnection network. We present a new adaptive path based model for both the multicast and unicast wormhole routing protocols. The proposed model under mixed traffic models has lower latency than the previous path-based methods with negligible hardware overhead. Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
DSD | 3 |
| 2009 | Architectural Exploration of Per-Core DVFS for Energy-Constrained On-Chip NetworksabstractA feasible and scalable per-core DVFS architecture for on-chip network is presented. The supplies are dynamically adjusted at a very fine granularity based on the local traffic status. The adoption of multiple voltage supply networks and power selecting transistors provides the architecture with scalability and feasibility superior to existing similar techniques. With high-level simulation using 65 nm power model obtained from widely-acknowledged tools, the effectiveness of the technique is demonstrated with quantitative analysis of energy overhead and latency penalty. Under various traffic patterns, the average flit energy is reduced considerably, ranging from 45% to 60%, with moderately increased but stable transmission latency. Alexander Wei Yin, Liang Guang, Ethiopia Nigussie, Pasi Liljeberg, Jouni Isoaho, Hannu Tenhunen |
DSD | 4 |
| 2009 | Explorations of Honeycomb Topologies for Network-on-ChipabstractRectangular mesh and torus are the mostly used topologies in network-on-chip (NoC) based systems. In this paper, we quantitatively illustrate that the honeycomb topology is an advantageous design alternative in terms of network cost which is one of the most important parameters that reflects both network performance and implementation cost. Comparing with the rectangular mesh and torus, honeycomb mesh and torus topologies lead to 40% decrease of the network cost. Then we explore the NoC related topological properties of both honeycomb mesh and torus topologies. By transforming the honeycomb topologies into rectangular brick shapes, we demonstrate that the honeycomb topologies are feasible to be implemented with rectangular devices. We also propose a 3D honeycomb topology since 3D IC has become an emerging and promising technique. Another contribution of this paper is the proposal of deadlock free routing algorithms. Based on either the concept of turn model or the logical network, deadlock free routing for all the discussed honeycomb topologies can be achieved. Alexander Wei Yin, Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen |
NPC | 3 |
| 2007 | Fault Tolerance Analysis of NoC ArchitecturesabstractThe paper presents an approach for analyzing and improving fault tolerance aspects in NoC architectures. This is a necessary step to be taken in order to implement reliable systems in future nanoscale technologies. Several NoC architectures and the router structures as well as the network interface needed for them are presented and compared for their fault tolerance, area and performance. The results indicate that a network structure built from simple 3-port routers provides better fault tolerance than a structure based on more complex multiport routers, and that the area overhead can be kept moderate Teijo Lehtonen, Pasi Liljeberg, Juha Plosila |
ISCAS | 2 |
| 2005 | Modelling and Refinement of an On-Chip Communication Architecture
Juha Plosila, Pasi Liljeberg, Jouni Isoaho |
ICFEM | 2 |
| 2004 | Self-timed communication platform for implementing high-performance systems-on-chip
Pasi Liljeberg, Juha Plosila, Jouni Isoaho |
Integr. | 1 |
| 2003 | Self-Timed Approach for Reducing On-Chip Switching Noise
Johanna Tuominen, Pasi Liljeberg, Jouni Isoaho |
VLSI-SOC | 2 |