EDBT 2026 Demo / reviewers in the wild / expert
Amir-Mohammad Rahmani
dblp:41/923 · also Amir M. Rahmani
· DBLP profile ↗
112ranked-venue papers
19as first author
35since 2021 · last 2026
0000-0003-0725-1155ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 15 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 17 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 2 since 2021Computer networks · 5 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Introduction to the Special Issue on Large Language Models, Conversational Systems, and Generative AI in Health - Part 2
Manas Gaur, Amir-Mohammad Rahmani, Sharath Chandra Guntuku, Xiaofan Jiang 0001, Tristan Naumann |
ACM Trans. Comput. Heal. | 3 |
| 2025 | AGILE: A Multi-Task Contrastive Learning Framework with Adversarial Gradient Iterative Learning for Bio-Signal Anonymization
Tamonash Bhattacharyya, Farshad Firouzi, Amir-Mohammad Rahmani, Sanaz R. Mousavi, Krishnendu Chakrabarty |
BSN | 3 |
| 2025 | Domain-Specific Constitutional AI: Enhancing Safety in LLM-Powered Mental Health Chatbots
Chenhan Lyu, Yutong Song, Amir-Mohammad Rahmani |
BSN | 4 |
| 2025 | MedCoT-RAG: Causal Chain-of-Thought RAG for Medical Question AnsweringabstractLarge language models (LLMs) have shown promise in medical question answering but often struggle with hallucinations and shallow reasoning, particularly in tasks requiring nuanced clinical understanding. Retrieval-augmented generation (RAG) offers a practical and privacy-preserving way to enhance LLMs with external medical knowledge. However, most existing approaches rely on surface-level semantic retrieval and lack the structured reasoning needed for clinical decision support. We introduce MedCoT-RAG, a domain-specific framework that combines causal-aware document retrieval with structured chain-of-thought prompting tailored to medical workflows. This design enables models to retrieve evidence aligned with diagnostic logic and generate step-by-step causal reasoning reflective of realworld clinical practice. Experiments on three diverse medical QA benchmarks show that MedCoT-RAG outperforms strong baselines by up to$10.3\%$over vanilla RAG and$6.4\%$over advanced domain-adapted methods, improving accuracy, interpretability, and consistency in complex medical tasks. Elahe Khatibi, Amir-Mohammad Rahmani |
BSN | 3 |
| 2025 | FoodAgent: A Multi-Modal Mixture of Experts Reasoning Agent for Divide-and-Conquer Food Nutrition EstimationabstractEstimating nutrition from food images remains a challenging task, particularly for complex, multi-component dishes. While computer vision methods are effective at recognizing food elements, they typically treat entire meals as monolithic inputs, lacking the ability to decompose visual scenes into individual components. Large language models (LLMs), in contrast, offer strong identification and qualitative reasoning capabilities but struggle with quantitative estimation, especially for assessing volume and mass of individual elements. In this work, we propose FoodAgent, a multi-modal Mixture-of-Experts (MoE) reasoning framework that improves nutrition estimation through a divide-and-conquer strategy. By decomposing dishes into distinct food components, FoodAgent dynamically routes each element to one of three specialized expert modules: (1) monocular volume estimation for nutritionally important and visually clear elements, (2) Retrieval-Augmented Generation (RAG) for important but not clear elements, and (3) direct LLM inference for minor components. This conditional expert selection aligns estimation strategies with the visual and semantic characteristics of each food element, significantly reducing cumulative errors. Experiments show that our element-wise, MoE-driven approach outperforms holistic methods, especially in real-world dietary scenarios involving diverse and complex meals. Yutong Song, Chenhan Lyu, Amir-Mohammad Rahmani |
BSN | 5 |
| 2025 | Invited Paper: Mindful AI for Pervasive Health and Wellbeing (PHW)abstractEmerging AI-driven pervasive health and wellbeing (PHW) services (e.g., personalized health assistants and mobile health applications) face critical challenges in handling noisy/intermittent sensory data, integrating cross-modal insights, and stringent energy and compute constraints. We present Mindful AI, a cognitive-inspired framework designed to enable adaptive, resilient, and efficient PHW services in real-world conditions. Our dual-mode intelligence—Automatic (System 1) and Reflective (System 2)—selectively directs system attention toward the most relevant sensing and compute contexts, unifying bottom-up stimuli (driven by input quality, inference demands and model confidence, and resource availability) with top-down insights (reflecting user demands, system goals/constraints, and contextual information). Our framework distills and orchestrates insights across sensing, communication, and computation through hybrid attention toward bottom-up and top-down insights that support cross-layer sense-compute co-optimization to achieve resilient, low-latency, and energy-efficient PHW services. We evaluate our approach on multi-tier device-edge-cloud platforms, using real-world case studies in pain assessment, stress monitoring, and human activity recognition to demonstrate adaptation to real-world uncertainties (e.g., sensor degradation, context drift, network variability), while maintaining strict QoS, accuracy, and latency guarantees. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
ICCAD | 4 |
| 2025 | Introduction to the Special Issue on Large Language Models, Conversational Systems, and Generative AI in Health - Part 1abstractDialogue systems are designed to offer human users social support or functional services through natural language interactions. Traditional conversation research has put significant emphasis on a system’s response-ability, including its capacity to understand dialogue context and generate appropriate responses. However, the key element of proactive behavior—a crucial aspect of intelligent conversations—is often overlooked in these studies. Proactivity empowers conversational agents to lead conversations towards achieving pre-defined targets or fulfilling specific goals on the system side. Proactive dialogue systems are equipped with advanced techniques to handle complex tasks, requiring strategic and motivational interactions, thus representing a significant step towards artificial general intelligence. Motivated by the necessity and challenges of building proactive dialogue systems, we provide a comprehensive review of various prominent problems and advanced designs for implementing proactivity into different types of dialogue systems, including open-domain dialogues, task-oriented dialogues, and information-seeking dialogues. We also discuss real-world challenges that require further research attention to meet application needs in the future, such as proactivity in dialogue systems that are based on large language models, proactivity in hybrid dialogues, evaluation protocols and ethical considerations for proactive dialogue systems. By providing a quick access and overall picture of the proactive dialogue systems domain, we aim to inspire new research directions and stimulate further advancements towards achieving the next level of conversational AI capabilities, paving the way for more dynamic and intelligent interactions within various application domains. Manas Gaur, Amir-Mohammad Rahmani, Sharath Chandra Guntuku, Xiaofan Jiang 0001, Tristan Naumann |
ACM Trans. Comput. Heal. | 3 |
| 2025 | Exploiting Approximation for Run-time Resource Management of Embedded HMPsabstractRun-time resource management (RTM) of multi-programmed workloads on heterogeneous multi-core platforms is challenging due to (i) fixed power budget of the device, (ii) variable performance requirements of the workloads, and (iii) unknown arrival of the applications. Existing RTM solutions lack power-performance coordination, resulting in performance degradation during power actuation or power violations during performance provisioning. Exploiting inherent error-resilience of the applications can address the performance loss incurred in power actuation, by combining run-time approximation with traditional power knobs (including Dynamic Voltage/Frequency Scaling, Task Migration, Degree of Parallelism, and CPU Quota ). In this work, we present an accuracy-aware resource management framework that jointly actuates run-time approximation and traditional power knobs for efficient power-performance management of multi-programmed and multi-threaded workloads running on heterogeneous mobile platforms. Our strategy configures the accuracy of the applications at run-time to exploit accuracy-performance trade-offs, by considering system-wide power-performance dynamics. We use heuristic estimation models to jointly enforce accuracy configuration and traditional power knobs settings at run-time. We evaluated our framework on real-world embedded mobile platforms, including Odroid XU3 and Asus Tinker Edge R boards to demonstrate the efficiency of our proposed approach across multiple workload scenarios. Our approach achieved 25% lower performance violations against the state-of-the-art run-time resource management policies at the cost of 2.2% accuracy loss across six applications. Zain Taufique, Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Cristiana Bolchini, Nikil Dutt, Pasi Liljeberg |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Personalized Causal Graph Reasoning for LLMs: An Implementation for Dietary RecommendationsabstractLarge Language Models (LLMs) excel at general-purpose reasoning by leveraging broad commonsense knowledge, but they remain limited in tasks requiring personalized reasoning over multifactorial personal data. This limitation constrains their applicability in domains such as healthcare, where decisions must adapt to individual contexts. We introduce Personalized Causal Graph Reasoning, a framework that enables LLMs to reason over individual-specific causal graphs constructed from longitudinal data. Each graph encodes how user-specific factors influence targeted outcomes. In response to a query, the LLM traverses the graph to identify relevant causal pathways, rank them by estimated impact, simulate potential outcomes, and generate tailored responses. We implement this framework in the context of nutrient-oriented dietary recommendations, where variability in metabolic responses demands personalized reasoning. Using counterfactual evaluation, we assess the effectiveness of LLM-generated food suggestions for glucose control. Our method reduces postprandial glucose iAUC across three time windows compared to prior approaches. Additional LLM-as-a-judge evaluations further confirm improvements in personalization quality. Zhongqi Yang, Amir-Mohammad Rahmani |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Graph-Augmented LLMs for Personalized Health Insights: A Case Study in Sleep AnalysisabstractHealth monitoring systems have revolutionized mod-ern healthcare by enabling the continuous capture of physio-logical and behavioral data, essential for preventive measures and early intervention. Integrating this data with Large Lan-guage Models (LLMs) shows promise in delivering interactive health advice, but traditional methods like Retrieval-Augmented Generation (RAG) and fine-tuning often fail to fully utilize the complex, multi-dimensional data from wearable devices. These approaches typically provide limited actionable and personalized health insights due to their inadequate capacity to dynamically integrate and interpret diverse health data streams. This pa-per introduces a graph-augmented LLM framework designed to enhance the personalization and clarity of health insights. Utilizing a hierarchical graph structure, the framework captures inter and intra-patient relationships, enriching LLM prompts with feature importance scores from a Random Forest Model. The effectiveness of this approach is demonstrated through a sleep analysis case study involving 20 college students during the COVID-19 lockdown, highlighting the potential of our model to generate actionable and personalized health insights efficiently. We leverage another LLM to evaluate the insights for relevance, comprehensiveness, actionability, and personalization. Our findings show that augmenting prompts with our framework yields significant improvements in all four criteria, eliciting well-crafted, thoughtful responses tailored to a specific patient. Ajan Subramanian, Zhongqi Yang, Iman Azimi, Amir-Mohammad Rahmani |
BSN | 4 |
| 2024 | ECG Unveiled: Analysis of Client Re-Identification Risks in Real-World ECG DatasetsabstractWhile ECG data is crucial for diagnosing and monitoring heart conditions, it also contains unique biometric information that poses significant privacy risks. Existing ECG re-identification studies rely on exhaustive analysis of numerous deep learning features, confining to ad-hoc explainability towards clinicians decision making. In this work, we delve into explainability of ECG re-identification risks using transparent machine learning models. We use SHapley Additive exPlanations (SHAP) analysis to identify and explain the key features contributing to re-identification risks. We conduct an empirical analysis of identity re-identification risks using ECG data from five diverse real-world datasets, encompassing 223 participants. By employing transparent machine learning models, we reveal the diversity among different ECG features in contributing towards re-identification of individuals with an accuracy of 0.76 for gender, 0.67 for age group, and 0.82 for participant ID re-identification. Our approach provides valuable insights for clinical experts and guides the development of effective privacy-preserving mechanisms. Further, our findings emphasize the necessity for robust privacy measures in real-world health applications and offer detailed, actionable insights for enhancing data anonymization techniques. Anil Kanduri, Seyed Amir Hossein Aqajari, Salar Jafarlou, Sanaz R. Mousavi, Pasi Liljeberg, Shaista Malik, Amir-Mohammad Rahmani |
BSN | 8 |
| 2024 | Attention-Based Explainable AI for Wearable Multivariate Data: A Case Study on Affect Status PredictionabstractWearable technology enables ubiquitous health monitoring where multivariate physiological and behavioral data can be captured over time. Such multivariate time series (MTS) data in healthcare applications needs technique to interpret the analysis results. However, existing deep learning models for MTS data analysis often lack interpretability, and current explainable AI (xAI) techniques fail to capture the temporal and inter-variable complexities inherent in MTS. This hinders the trust and integration of these AI-based systems in clinical decision-making. In this paper, we propose an attention-based xAI method to classify and interpret MTS data collected from wearable devices. Our approach leverages self-attention mechanisms and graph attention layers (GAT) to capture both temporal and inter-variable dependencies, providing interpretability at both the temporal and modality levels. We evaluate our method using a longitudinal affect status monitoring. The dataset was collected from 21 college students via wearable devices over one year. We train separate models for positive (PA) and negative affect (NA) prediction, and compare their performance with a Transformer-based method. Our method achieves robust classification performance, with 78.62% accuracy for PA and 76.30% for NA, while offering transparent explanations of its decisions. These findings highlight the potential of our xAI method for reliable and interpretable MTS classification in healthcare applications. Zhongqi Yang, Iman Azimi, Amir-Mohammad Rahmani, Pasi Liljeberg |
BSN | 4 |
| 2024 | Work-in-Progress: Context and Noise Aware Resilience for Autonomous Driving ApplicationsabstractAutonomous Vehicles (AVs) often use noise prone sensory data from cameras and LiDAR for perception. In specific noisy scenarios, different object detection models exhibit non-intuitive and varying degrees of resilience, necessitating adaptive model selection. In this work, we develop a context and noise aware framework for run-time adaptive configuration of objection models for high accuracy and low latency inference. We combine driving scene context and input data noise to prioritize among input modalities, followed by selection and configuration of most resilient object detection model appropriate for the context. Our evaluation for 2D object detection on nuScenes dataset provided average 1.83x speedup in latency compared to baseline while preserving average prediction confidence. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
CODES+ISSS | 4 |
| 2024 | SEAL: Sensing Efficient Active Learning on Wearables through Context-awarenessabstractIn this paper, we introduce SEAL, a co-optimization framework designed to enhance both sensing and querying strategies in wearable devices for mHealth applications. Employing Reinforcement Learning (RL), SEAL strategically utilizes user contextual information and the machine learning model's confidence levels to make efficient decisions. This innovative approach is particularly significant in addressing the challenge of battery drain due to continuous physiological signal sensing, such as Photoplethysmography (PPG). Our framework demonstrates its effectiveness in a stress monitoring application, achieving a substantial reduction of 76% in the volume of PPG signals collected, while only experiencing a minor 6% decrease in user-labeled data quality. This balance showcases SEAL's potential in optimizing data collection in a way that is considerate of both device constraints and data integrity. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
DATE | 5 |
| 2024 | EA^2: Energy Efficient Adaptive Active Learning for Smart WearablesabstractMobile Health (mHealth) applications rely on supervised Machine Learning (ML) algorithms, requiring end-user-labeled data for the training phase. The gold standard for obtaining such labeled data is by sending queries to users and gathering responses for the corresponding label, which was conventionally done through triggering questions sent at random. Active Learning (AL) methods use intelligent query-sending policies by incorporating users' contextual information to maximize the response rate and informativeness of the collected labeled data. However, wearable devices' substantial battery drainage associated with the sensing of physiological signals underscores the need for developing an efficient sensing policy in addition to a query-sending policy. In this work, we present a co-optimization framework for both sensing and querying strategies within wearable devices, leveraging contextual information and ML model's prediction confidence. We designed a Reinforcement Learning (RL) agent to quantify different contextual parameters combined with model confidence to determine sensing and querying decisions. Our evaluation of an exemplar stress monitoring application showed a 76% reduction in sensing and data transmission energy consumption, with only a 6% drop in user-labeled data. Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
ISLPED | 5 |
| 2024 | Personalized and adaptive neural networks for pain detection from multi-modal physiological features
Mingzhe Jiang, Riitta Rosio, Sanna Salanterä, Amir-Mohammad Rahmani, Pasi Liljeberg, Daniel Santos da Silva, Victor Hugo C. de Albuquerque |
Expert Syst. Appl. | 4 |
| 2024 | PERFECT: Personalized Exercise Recommendation Framework and architECTureabstractBackground : The health benefits of regular physical activity (PA) are well-established and widely acknowledged. Through the integration of wearable trackers, the Internet of Things (IoT)—a network of interconnected devices capable of collecting and exchanging data—coupled with mobile health (mHealth), which refers to the use of mobile devices to support medical and public health practices, it is now feasible to systematically gather and present individual exercise behaviors. This advanced approach enables the precise correlation of users’ physiological data and daily activities with their specific fitness needs, offering a personalized pathway to improving health outcomes. Objective : This study aims to enhance PA levels among individuals by developing a personalized exercise recommendation system. Utilizing reinforcement learning, the system proposes tailored exercise plans based on biomarkers and the user’s specific context. Methods : In this study, we developed applications for smartphones and smartwatches designed to gather, monitor, and recommend exercise routines through the application of a contextual multi-arm bandit algorithm. To evaluate the efficacy of this mHealth exercise regimen, we enlisted the participation of twenty female college students. Results : The outcomes of our investigation revealed a significant enhancement in the average daily duration of exercise (P \({\lt}\) . 001). Participants expressed high levels of satisfaction with both the walking program and the recommendation system, achieving average ratings of 4.31 (SD \(=\) 0.60) and 3.69 (SD \(=\) 0.95), respectively, on a 5-point scale. Furthermore, the average scores for participants’ confidence in safely performing the recommended walking exercises, as well as their perception of the study’s effectiveness in meeting their PA needs, were both above 4, indicating a positive reception and confidence in the program’s design and implementation. Conclusions : The evolution of the IoT and wearable technology has marked the beginning of a new era for mHealth systems, particularly in the personalization of health interventions. Such advancements enable the precise personalization of PA recommendations, potentially enhancing user engagement and performance outcomes. This paper introduces a novel exercise recommendation system that utilizes reinforcement learning to personalize walking exercises based on the user’s biomarkers and context, aiming to improve the user’s aerobic capacity significantly. Milad Asgari Mehrabadi, Elahe Khatibi, Tamara Jimah, Sina Labbaf, Holly Borg, Laura Narvaez, Pamela Pimentel, Arlene Turner, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Trans. Comput. Heal. | 11 |
| 2023 | End-to-End PPG Processing Pipeline for Wearables: From Quality Assessment and Motion Artifacts Removal to HR/HRV Feature ExtractionabstractThe rapid development of wearable technology has enabled remote photoplethysmography (PPG)-based health monitoring in everyday settings, offering real-time and continuous monitoring of cardiovascular parameters, such as heart rate (HR) and heart rate variability (HRV). However, PPG signals collected in daily life are prone to artifacts and noise, posing challenges to HR and HRV extraction. The existing HR and HRV extraction methods cannot effectively handle noisy PPG signals and ensure accurate results. Additionally, current Python packages were primarily designed for analyzing "clean" PPG signals, limiting their performance in handling artifacts and noise and resulting in unreliable HR and HRV measurements. In this paper, we propose a robust end-to-end PPG processing pipeline to reliably extract HR and HRV from PPG signals collected in free-living settings. The pipeline comprises three machine learning-based PPG analysis methods: signal quality assessment, reconstruction of noisy signal, and systolic peak detection. We assess the proposed PPG pipeline using a dataset including PPG and Electrocardiogram (ECG) signals recorded from 46 individuals by smartwatches. Our evaluation demonstrates the proposed pipeline’s superior performance compared to two established benchmark methods in terms of correlation and mean absolute error with ECG as the reference. We also provide the Python implementation of our pipeline for the research community to facilitate integration into their solutions. Mohammad Feli, Kianoosh Kazemi, Iman Azimi, Amir-Mohammad Rahmani, Pasi Liljeberg |
BIBM | 5 |
| 2023 | Impact of COVID-19 Pandemic on Sleep Including HRV and Physical Activity as Mediators: A Causal ML ApproachabstractSleep quality is crucial to both mental and physical well-being. The COVID-19 pandemic, which has notably affected the population’s health worldwide, has been shown to deteriorate people’s sleep quality. Numerous studies have been conducted to evaluate the impact of the COVID-19 pandemic on sleep efficiency, investigating their relationships using correlation-based methods. These methods merely rely on learning spurious correlation rather than the causal relations among variables. Furthermore, they fail to pinpoint potential sources of bias and mediators and envision counterfactual scenarios, leading to a poor estimation. In this paper, we develop a Causal Machine Learning method, which encompasses causal discovery and causal inference components, to extract the causal relations between the COVID-19 pandemic (treatment variable) and sleep quality (outcome) and estimate the causal treatment effect, respectively. We conducted a wearable-based health monitoring study to collect data, including sleep quality, physical activity, and Heart Rate Variability (HRV) from college students before and after the COVID-19 lockdown in March 2020. Our causal discovery component generates a causal graph and pinpoints mediators in the causal model. We incorporate the strongly contributing mediators (i.e., HRV and physical activity) into our causal inference component to estimate the robust, accurate, and explainable causal effect of the pandemic on sleep quality. Finally, we validate our estimation via three refutation analysis techniques. Our experimental results indicate that the pandemic exacerbates college students’ sleep scores by 8%. Our validation results show significant p-values confirming our estimation. Elahe Khatibi, Mahyar Abbasian, Iman Azimi, Sina Labbaf, Mohammad Feli, Jessica L. Borelli, Nikil Dutt, Amir-Mohammad Rahmani |
BSN | 8 |
| 2023 | Loneliness Forecasting Using Multi-modal Wearable and Mobile Sensing in Everyday SettingsabstractThe adverse effects of loneliness on both physical and mental well-being are profound. Although previous research has utilized mobile sensing techniques to detect mental health issues, few studies have utilized state-of-the-art wearable devices to forecast loneliness and comprehend the physiological manifestations of loneliness and its predictive nature. The primary objective of this study is to examine the feasibility of forecasting loneliness by employing wearable devices, such as smart rings and watches, to monitor early physiological indicators of loneliness. Furthermore, smartphones are employed to capture initial behavioral signs of loneliness. To accomplish this, we employed personalized machine learning techniques, leveraging a comprehensive dataset comprising physiological and behavioral information obtained during our study involving the monitoring of college students. Through the development of personalized models, we achieved a notable accuracy of 0.82 and an F-1 score of 0.82 in forecasting loneliness levels seven days in advance. Additionally, the application of Shapley values facilitated model explainability. The wealth of data provided by this study, coupled with the forecasting methodology employed, possesses the potential to augment interventions and facilitate the early identification of loneliness within populations at risk. Zhongqi Yang, Iman Azimi, Salar Jafarlou, Sina Labbaf, Jessica L. Borelli, Nikil Dutt, Amir-Mohammad Rahmani |
BSN | 7 |
| 2023 | Towards Deep Personal Lifestyle Models Using Multimodal N-of-1 Data
Nitish Nagesh, Iman Azimi, Tom Andriola, Amir-Mohammad Rahmani, Ramesh Jain 0001 |
MMM (1) | 4 |
| 2023 | A Deep Learning-based PPG Quality Assessment Approach for Heart Rate and Heart Rate VariabilityabstractPhotoplethysmography (PPG) is a non-invasive optical method to acquire various vital signs, including heart rate (HR) and heart rate variability (HRV). The PPG method is highly susceptible to motion artifacts and environmental noise. Unfortunately, such artifacts are inevitable in ubiquitous health monitoring, as the users are involved in various activities in their daily routines. Such low-quality PPG signals negatively impact the accuracy of the extracted health parameters, leading to inaccurate decision-making. PPG-based health monitoring necessitates a quality assessment approach to determine the signal quality according to the accuracy of the health parameters. Different studies have thus far introduced PPG signal quality assessment methods, exploiting various indicators and machine learning algorithms. These methods differentiate reliable and unreliable signals, considering morphological features of the PPG signal and focusing on the cardiac cycles. Therefore, they can be utilized for HR detection applications. However, they do not apply to HRV, as only having an acceptable shape is insufficient, and other signal factors may also affect the accuracy. In this article, we propose a deep learning–based PPG quality assessment method for HR and various HRV parameters. We employ one customized one-dimensional (1D) and three 2D Convolutional Neural Networks (CNN) to train models for each parameter. Reliability of each of these parameters will be evaluated against the corresponding electrocardiogram signal, using 210 hours of data collected from a home-based health monitoring application. Our results show that the proposed 1D CNN method outperforms the other 2D CNN approaches. Our 1D CNN model obtains the accuracy of 95.63%, 96.71%, 91.42%, 94.01%, and 94.81% for the HR, average of normal to normal interbeat (NN) intervals, root mean square of successive NN interval differences, standard deviation of NN intervals, and ratio of absolute power in low frequency to absolute power in high frequency ratios, respectively. Moreover, we compare the performance of our proposed method with state-of-the-art algorithms. We compare our best models for HR-HRV health parameters with six different state-of-the-art PPG signal quality assessment methods. Our results indicate that the proposed method performs better than the other methods. We also provide the open source model implemented in Python for the community to be integrated into their solutions. Emad Kasaeyan Naeini, Fatemeh Sarhaddi, Iman Azimi, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Trans. Comput. Heal. | 6 |
| 2023 | An Accurate Non-accelerometer-based PPG Motion Artifact Removal Technique using CycleGANabstractA photoplethysmography (PPG) is an uncomplicated and inexpensive optical technique widely used in the healthcare domain to extract valuable health-related information, e.g., heart rate variability, blood pressure, and respiration rate. PPG signals can easily be collected continuously and remotely using portable wearable devices. However, these measuring devices are vulnerable to motion artifacts caused by daily life activities. The most common ways to eliminate motion artifacts use extra accelerometer sensors, which suffer from two limitations: (i) high power consumption, and (ii) the need to integrate an accelerometer sensor in a wearable device (which is not required in certain wearables). This paper proposes a low-power non-accelerometer-based PPG motion artifacts removal method outperforming the accuracy of the existing methods. We use Cycle Generative Adversarial Network to reconstruct clean PPG signals from noisy PPG signals. Our novel machine-learning-based technique achieves 9.5 times improvement in motion artifact removal compared to the state-of-the-art without using extra sensors such as an accelerometer, which leads to 45% improvement in energy efficiency. Amir Hosein Afandizadeh Zargari, Seyed Amir Hossein Aqajari, Hadi Khodabandeh, Amir-Mohammad Rahmani, Fadi J. Kurdahi |
ACM Trans. Comput. Heal. | 4 |
| 2022 | AMSER: Adaptive Multimodal Sensing for Energy Efficient and Resilient eHealth SystemsabstracteHealth systems deliver critical digital healthcare and wellness services for users by continuously monitoring physiological and contextual data. eHealth applications use multi-modal machine learning kernels to analyze data from different sensor modalities and automate decision-making. Noisy inputs and motion artifacts during sensory data acquisition affect the i) prediction accuracy and resilience of eHealth services and ii) energy efficiency in processing garbage data. Monitoring raw sensory inputs to identify and drop data and features from noisy modalities can improve prediction accuracy and energy efficiency. We propose a closed-loop monitoring and control framework for multi-modal eHealth applications, AMSER, that can mitigate garbage-in garbage-out by i) monitoring input modalities, ii) analyzing raw input to selectively drop noisy data and features, and iii) choosing appropriate machine learning models that fit the configured data and feature vector - to improve prediction accuracy and energy efficiency. We evaluate our AMSER approach using multi-modal eHealth applications of pain assessment and stress monitoring over different levels and types of noisy components incurred via different sensor modalities. Our approach achieves up to 22% improvement in prediction accuracy and 5.6× energy consumption reduction in the sensing phase against the state-of-the-art multi-modal monitoring application. Emad Kasaeyan Naeini, Sina Shahhosseini, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt |
DATE | 5 |
| 2022 | Flexible and Personalized Learning for Wearable Health Applications using HyperDimensional ComputingabstractHealth and wellness applications increasingly rely on machine learning techniques to learn end-user physiological and behavioral patterns in everyday settings, posing two key challenges: inability to perform on-device online learning for resource-constrained wearables, and learning algorithms that support privacy-preserving personalization. We exploit a Hyperdimensional computing (HDC) solution for wearable devices that offers flexibility, high efficiency, and performance while enabling on-device personalization and privacy protection. We evaluate the efficacy of our approach using three case studies and show that our system improves performance of training by up to 35.8x compared with the state-of-the-art while offering a comparable accuracy. Sina Shahhosseini, Yang Ni 0001, Emad Kasaeyan Naeini, Mohsen Imani, Amir-Mohammad Rahmani, Nikil Dutt |
ACM Great Lakes Symposium on VLSI | 5 |
| 2022 | Hyperdimensional Hybrid Learning on End-Edge-Cloud NetworksabstractIn this paper, we present Hyperdimensional Hybrid Learning (HDHL), which combines model-free and model-based Reinforcement Learning, to effectively reduce the computational cost and environment interaction for optimizing an intelligent cloud service. We first show that Hyperdimensional Q-Learning (QHD), the state-of-the-art Hyperdimensional Computing value-based Reinforcement Learning algorithm, is computationally faster than the Deep Q-Network (DQN) for this task. In addition, we demonstrate how HDHL reduces the number of environment interactions by 4.8× to learn the near optimal configuration. Our evaluation shows that HDHL is computationally more efficient than both Q-Learning algorithms, with the total time being reduced by 21.0× compared to DQN and 16.5× compared to QHD. Mariam Issa, Sina Shahhosseini, Yang Ni 0001, Danny Abraham, Amir-Mohammad Rahmani, Nikil Dutt, Mohsen Imani |
ICCD | 6 |
| 2022 | Exploring computation offloading in IoT systemsabstractInternet of Things (IoT) paradigm raises challenges for devising efficient strategies that offload applications to the fog or the cloud layer while ensuring the optimal response time for a service. Traditional computation offloading policies assume the response time is only dominated by the execution time. However, the response time is a function of many factors including contextual parameters and application characteristics that can change over time. For the computation offloading problem, the majority of existing literature presents efficient solutions considering a limited number of parameters (e.g., computation capacity and network bandwidth) neglecting the effect of the application characteristics and dataflow configuration. In this paper, we explore the impact of the computation offloading on total application response time in three-layer IoT systems considering more realistic parameters, e.g., application characteristics, system complexity, communication cost, and dataflow configuration. This paper also highlights the impact of a new application characteristic parameter defined as Output–Input Data Generation (OIDG) ratio and dataflow configuration on the system behavior. In addition, we present a proof-of-concept end-to-end dynamic computation offloading technique, implemented in a real hardware setup, that observes the aforementioned parameters to perform real-time decision-making. Sina Shahhosseini, Arman Anzanpour, Iman Azimi, Sina Labbaf, Dongjoo Seo, Sung-Soo Lim, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
Inf. Syst. | 9 |
| 2022 | Confidence-Enhanced Early Warning Score Based on Fuzzy LogicabstractAbstract Cardiovascular diseases are one of the world’s major causes of loss of life. The vital signs of a patient can indicate this up to 24 hours before such an incident happens. Healthcare professionals use Early Warning Score (EWS) as a common tool in healthcare facilities to indicate the health status of a patient. However, the chance of survival of an outpatient could be increased if a mobile EWS system would monitor them during their daily activities to be able to alert in case of danger. Because of limited healthcare professional supervision of this health condition assessment, a mobile EWS system needs to have an acceptable level of reliability - even if errors occur in the monitoring setup such as noisy signals and detached sensors. In earlier works, a data reliability validation technique has been presented that gives information about the trustfulness of the calculated EWS. In this paper, we propose an EWS system enhanced with the self-aware property confidence, which is based on fuzzy logic. In our experiments, we demonstrate that - under adverse monitoring circumstances (such as noisy signals, detached sensors, and non-nominal monitoring conditions) - our proposed Self-Aware Early Warning Score (SA-EWS) system provides a more reliable EWS than an EWS system without self-aware properties. Maximilian Götzinger, Arman Anzanpour, Iman Azimi, Nima Taherinejad, Axel Jantsch, Amir-Mohammad Rahmani, Pasi Liljeberg |
Mob. Networks Appl. | 6 |
| 2022 | Mobile Health Technology: From Daily Care and Pandemics to their Energy Consumption and Environmental Impact
Nima Taherinejad, Paolo Perego, Amir-Mohammad Rahmani |
Mob. Networks Appl. | 3 |
| 2022 | Concurrent Application Bias Scheduling for Energy Efficiency of Heterogeneous Multi-Core PlatformsabstractMinimizing energy consumption of concurrent applications on heterogeneous multi-core platforms is challenging given the diversity in energy-performance profiles of both the applications and hardware. Adaptive learning techniques made the exhaustive Pareto-optimal space exploration practically feasible to identify an energy efficient configuration. Existing approaches consider a single application's characteristic for optimizing energy consumption. However, an optimal configuration for a given application in isolation may not be optimal when other applications are run concurrently. Approaches that consider concurrent application scenarios overlook the weight of total energy consumption per application, restricting them from prioritizing among applications. We address this limitation by considering the mutual effect of concurrent applications on system wide energy consumption to adapt resource configuration at run-time. We characterize each application's power-performance profile as a weighted bias through off-line profiling. We infer this model combined with an on-line predictive strategy to make resource allocation decisions for minimizing energy consumption while honoring performance requirements. The proposed strategy is implemented as a user-space process and evaluated on a heterogeneous hardware platform of Odroid XU3 over the Rodinia benchmark suite. Experimental results show up to 61 percent of energy saving compared to the standard baseline of Linux governors and up to 27 percent of energy gain compared to state-of-the-art adaptive learning-based resource management techniques. Elham Shamsa, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani |
IEEE Trans. Computers | 4 |
| 2022 | Online Learning for Orchestration of Inference in Multi-user End-edge-cloud NetworksabstractDeep-learning-based intelligent services have become prevalent in cyber-physical applications, including smart cities and health-care. Deploying deep-learning-based intelligence near the end-user enhances privacy protection, responsiveness, and reliability. Resource-constrained end-devices must be carefully managed to meet the latency and energy requirements of computationally intensive deep learning services. Collaborative end-edge-cloud computing for deep learning provides a range of performance and efficiency that can address application requirements through computation offloading. The decision to offload computation is a communication-computation co-optimization problem that varies with both system parameters (e.g., network condition) and workload characteristics (e.g., inputs). However, deep learning model optimization provides another source of tradeoff between latency and model accuracy. An end-to-end decision-making solution that considers such computation-communication problem is required to synergistically find the optimal offloading policy and model for deep learning services. To this end, we propose a reinforcement-learning-based computation offloading solution that learns optimal offloading policy considering deep learning model selection techniques to minimize response time while providing sufficient accuracy. We demonstrate the effectiveness of our solution for edge devices in an end-edge-cloud system and evaluate with a real-setup implementation using multiple AWS and ARM core configurations. Our solution provides 35% speedup in the average response time compared to the state-of-the-art with less than 0.9% accuracy reduction, demonstrating the promise of our online learning framework for orchestrating DL inference in end-edge-cloud systems. Sina Shahhosseini, Dongjoo Seo, Anil Kanduri, Sung-Soo Lim, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2021 | Energy-Performance Co-Management of Mixed-Sensitivity Workloads on Heterogeneous Multi-core SystemsabstractSatisfying performance of complex workload scenarios with respect to energy consumption on Heterogeneous Multi-core Platforms (HMPs) is challenging when considering i) the increasing variety of applications, and ii) the large space of resource management configurations. Existing run-time resource management approaches use online and offline learning to handle such complexity. However, they focus on one type of application, neglecting concurrent execution of mixed sensitivity workloads. In this work, we propose an energy-performance co-management method which prioritizes mixed type of applications at run-time, and searches in the configuration space to find the optimal configuration for each application which satisfies the performance requirements while saving energy. We evaluate our approach on a real Odroid XU3 platform over mixed-sensitivity embedded workloads. Experimental results show our approach provides 54% lower performance violation with 50% higher energy saving compared to the existing approaches. Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg |
ASP-DAC | 3 |
| 2021 | Configurable DSI partitioned approximate multiplier
Fahimeh Hajizadeh, Mohammadreza Binesh Marvasti, Seyyed Amir Asghari, Mostafa Abbas Mollaei, Amir-Mohammad Rahmani |
Future Gener. Comput. Syst. | 5 |
| 2021 | SEAMS: Self-Optimizing Runtime Manager for Approximate Memory HierarchiesabstractMemory approximation techniques are commonly limited in scope, targeting individual levels of the memory hierarchy. Existing approximation techniques for a full memory hierarchy determine optimal configurations at design-time provided a goal and application. Such policies are rigid: they cannot adapt to unknown workloads and must be redesigned for different memory configurations and technologies. We propose SEAMS: the first self-optimizing runtime manager for coordinating configurable approximation knobs across all levels of the memory hierarchy. SEAMS continuously updates and optimizes its approximation management policy throughout runtime for diverse workloads. SEAMS optimizes the approximate memory configuration to minimize energy consumption without compromising the quality threshold specified by application developers. SEAMS can (1) learn a policy at runtime to manage variable application quality of service ( QoS ) constraints, (2) automatically optimize for a target metric within those constraints, and (3) coordinate runtime decisions for interdependent knobs and subsystems. We demonstrate SEAMS’ ability to efficiently provide functions (1)–(3) on a RISC-V Linux platform with approximate memory segments in the on-chip cache and main memory. We demonstrate SEAMS’ ability to save up to 37% energy in the memory subsystem without any design-time overhead. We show SEAMS’ ability to reduce QoS violations by 75% with < 5% additional energy. Biswadip Maity, Bryan Donyanavard, Anmol Surhonne, Amir-Mohammad Rahmani, Andreas Herkersdorf, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | UBAR: User- and Battery-aware Resource Management for SmartphonesabstractSmartphone users require high Battery Cycle Life (BCL) and high Quality of Experience (QoE) during their usage. These two objectives can be conflicting based on the user preference at run-time. Finding the best trade-off between QoE and BCL requires an intelligent resource management approach that considers and learns user preference at run-time. Current approaches focus on one of these two objectives and neglect the other, limiting their efficiency in meeting users’ needs. In this article, we present UBAR, User- and Battery-aware Resource management, which considers dynamic workload, user preference, and user plug-in/out pattern at run-time to provide a suitable trade-off between BCL and QoE. UBAR personalizes this trade-off by learning the user’s habits and using that to satisfy QoE, while considering battery temperature and State of Charge (SOC) pattern to maximize BCL. The evaluation results show that UBAR achieves 10% to 40% improvement compared to the existing state-of-the-art approaches. Elham Shamsa, Alma Pröbstl, Nima Taherinejad, Anil Kanduri, Samarjit Chakraborty, Amir-Mohammad Rahmani, Pasi Liljeberg |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2020 | Emergent Control of MPSoC Operation by a Hierarchical Supervisor / Reinforcement Learning ApproachabstractMPSoCs increasingly depend on adaptive resource management strategies at runtime for efficient utilization of resources when executing complex application workloads. In particular, conflicting demands for adequate computation performance and power-/energy-efficiency constraints make desired application goals hard to achieve. We present a hierarchical, cross-layer hardware/software resource manager capable of adapting to changing workloads and system dynamics with zero initial knowledge. The manager uses rule-based reinforcement learning classifier tables (LCTs) with an archive-based backup policy as leaf controllers. The LCTs directly manipulate and enforce MPSoC building block operation parameters in order to explore and optimize potentially conflicting system requirements (e.g., meeting a performance target while staying within the power constraint). A supervisor translates system requirements and application goals into per-LCT objective functions (e.g., core instructions-per-second (IPS). Thus, the supervisor manages the possibly emergent behavior of the low-level LCT controllers in response to 1) switching between operation strategies (e.g., maximize performance vs. minimize power; and 2) changing application requirements. This hierarchical manager leverages the dual benefits of a software supervisor (enabling flexibility), together with hardware learners (allowing quick and efficient optimization). Experiments on an FPGA prototype confirmed the ability of our approach to identify optimized MPSoC operation parameters at runtime while strictly obeying given power constraints. Florian Maurer 0003, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt, Andreas Herkersdorf |
DATE | 3 |
| 2020 | Context-Aware Sensing via Dynamic Programming for Edge-Assisted Wearable SystemsabstractHealthcare applications supported by the Internet of Things enable personalized monitoring of a patient in everyday settings. Such applications often consist of battery-powered sensors coupled to smart gateways at the edge layer. Smart gateways offer several local computing and storage services (e.g., data aggregation, compression, local decision making), and also provide an opportunity for implementing local closed-loop optimization of different parameters of the sensor layer, particularly energy consumption. To implement efficient optimization methods, information regarding the context and state of patients need to be considered to find opportunities to adjust energy to demanded accuracy. Edge-assisted optimization can manage energy consumption of the sensor layer but may also adversely affect the quality of sensed data, which could compromise the reliable detection of health deterioration risk factors. In this article, we propose two approaches: myopic and Markov decision processes (MDPs)—to consider both energy constraints and risk factor requirements for achieving a twofold goal: energy savings while satisfying accuracy requirements of abnormality detection in a patient’s vital signs. Vital signs, including heart rate, respiration rate, and oxygen saturation, are extracted from a photoplethysmogram signal and errors of extracted features are compared to a ground truth that is modeled as a Gaussian distribution. We control the sensor’s sensing energy to minimize the power consumption while meeting a desired level of satisfactory detection performance. We present experimental results on realistic case studies using a reconfigurable photoplethysmogram sensor in an IoT system, and show that compared to nonadaptive methods, myopic reduces an average of 16.9% in sensing energy consumption with the maximum probability of abnormality misdetection on the order of 0.17 in a 24-hour health monitoring system. In addition, over 4 weeks of monitoring, we demonstrate that our MDP policy can extend the battery life on average of more than 2x while fulfilling the same average probability of misdetection compared to the myopic method. We illustrate results comparing myopic , MDP, and nonadaptive methods to monitor 14 subjects over 1 month. Delaram Amiri, Arman Anzanpour, Iman Azimi, Marco Levorato, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Trans. Comput. Heal. | 7 |
| 2020 | CAST: Content-Aware STT-MRAM Cache Write Management for Different Levels of ApproximationabstractSpin transfer torque magnetic RAM (STT-MRAM) technology is one of the most promising alternative for static RAM (SRAM) for implementing on-chip memories. Compared with SRAMs, STT-MRAMs benefit from higher density and near-zero leakage power, nonetheless they impose high energy consumption for reliable write operations. However, in many applications, absolute data integrity is not required; thus, acting on the current applied in the write operations may represent a novel knob for disciplined approximate computing to obtain energy saving with a minimal quality loss in applications' outputs. This article proposes CAST, a hardware/software approach to adjust the energy/quality of write operations in STT-MRAM caches in multicore systems based on the content of requested write operations. CAST utilizes fine-grained cache-line-level actuation knobs with different levels of quality for individual write operations. This unique feature of STT-MRAMs allows to avoid interapplication actuation interference suffered by SRAMs, and makes the approach particularly suitable for systems running multiple applications with mixed accuracy sensitivity. Moreover, CAST exploits another peculiarity of STT-MRAMs represented by the asymmetry and transition-dependency of the write error rate, to further tune in a fine-grained manner the write current to achieve an additional energy saving, even in full-accurate applications. Our evaluations on workloads of full-approximate, mixed-criticality, and full-accurate applications demonstrate up to 57%, 34%, and 21% energy savings over a baseline STT-MRAM cache, respectively, with an acceptable quality of the generated outputs. Amir Mahdi Hosseini Monazzah, Amir-Mohammad Rahmani, Antonio Miele, Nikil Dutt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Edge-Assisted Control for Healthcare Internet of Things: A Case Study on PPG-Based Early Warning ScoreabstractRecent advances in pervasive Internet of Things technologies and edge computing have opened new avenues for development of ubiquitous health monitoring applications. Delivering an acceptable level of usability and accuracy for these healthcare Internet of Things applications requires optimization of both system-driven and data-driven aspects, which are typically done in a disjoint manner. Although decoupled optimization of these processes yields local optima at each level, synergistic coupling of the system and data levels can lead to a holistic solution opening new opportunities for optimization. In this article, we present an edge-assisted resource manager that dynamically controls the fidelity and duration of sensing w.r.t. changes in the patient’s activity and health state, thus fine-tuning the trade-off between energy efficiency and measurement accuracy. The cornerstone of our proposed solution is an intelligent low-latency real-time controller implemented at the edge layer that detects abnormalities in the patient’s condition and accordingly adjusts the sensing parameters of a reconfigurable wireless sensor node. We assess the efficiency of our proposed system via a case study of the photoplethysmography-based medical early warning score system. Our experiments on a real full hardware-software early warning score system reveal up to 49% power savings while maintaining the accuracy of the sensory data. Arman Anzanpour, Delaram Amiri, Iman Azimi, Marco Levorato, Nikil Dutt, Pasi Liljeberg, Amir-Mohammad Rahmani |
ACM Trans. Internet Things | 7 |
| 2020 | Homecare Robotic Systems for Healthcare 4.0: Visions and Enabling TechnologiesabstractPowered by the technologies that have originated from manufacturing, the fourth revolution of healthcare technologies is happening (Healthcare 4.0). As an example of such revolution, new generation homecare robotic systems (HRS) based on the cyber-physical systems (CPS) with higher speed and more intelligent execution are emerging. In this article, the new visions and features of the CPS-based HRS are proposed. The latest progress in related enabling technologies is reviewed, including artificial intelligence, sensing fundamentals, materials and machines, cloud computing and communication, as well as motion capture and mapping. Finally, the future perspectives of the CPS-based HRS and the technical challenges faced in each technical area are discussed. Geng Yang 0003, Zhibo Pang, M. Jamal Deen, Mianxiong Dong, Yuan-Ting Zhang, Nigel H. Lovell, Amir-Mohammad Rahmani |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Guest Editorial Enabling Technologies in Health Engineering and Informatics for the New Revolution of Healthcare 4.0abstractThe eleven papers presented in this special issue provide a snapshot of the latest advances in the field of enabling technologies in health engineering and health informatics for the new revolution of Healthcare 4.0, hoping to further enable, drive and accelerate the research, development, and application of key technologies into healthcare systems. Geng Yang 0003, Zhibo Pang, Amir-Mohammad Rahmani, Mianxiong Dong, Yuan-Ting Zhang, M. Jamal Deen, Nigel H. Lovell |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | A novel method for victim block selection for NAND flash-based solid state drives based on scoring
Asal Khanbadr, Mohammadreza Binesh Marvasti, Seyyed Amir Asghari, Sohrab Khanbadr, Amir-Mohammad Rahmani |
J. Supercomput. | 5 |
| 2019 | Goal-Driven Autonomy for Efficient On-chip Resource Management: Transforming Objectives to GoalsabstractRun-time resource allocation of heterogeneous multi-core systems is challenging with varying workloads and limited power and energy budgets. User interaction within these systems changes the performance requirements, often conflicting with concurrent applications' objective and system constraints. Current resource allocation approaches focus on optimizing fixed objective, ignoring the variation in system and applications' objective at run-time. For an efficient resource allocation, the system has to operate autonomously by formulating a hierarchy of goals. We present goal-driven autonomy (GDA) for on-chip resource allocation decisions, which allows systems to generate and prioritize goals in response to the workload and system dynamic variation. We implemented a proof-of-concept resource management framework that integrates the proposed goal management control to meet power, performance and user requirements simultaneously. Experimental results on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) show the effectiveness of GDA in efficient resource allocation in comparison with existing fixed objective policies. Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DATE | 3 |
| 2019 | Dynamic Computation Migration at the Edge: Is There an Optimal Choice?abstractIn the era of Fog computing where one can decide to compute certain time-critical tasks at the edge of the network, designers often encounter a question whether the sensor layer provides the optimal response time for a service, or the Fog layer, or their combination. In this context, minimizing the total response time using computation migration is a communication-computation co-optimization problem as the response time does not depend only on the computational capacity of each side. In this paper, we aim at investigating this question and addressing it in certain situations. We formulate this question as a static or dynamic computation migration problem depending on whether certain communication and computation characteristics of the underlying system is known at design-time or not. We first propose a static approach to find the optimal computation migration strategy using models known at design-time. We then make a more realistic assumption that several sources of variation can affect the system's response latency (e.g., the change in computation time, bandwidth, transmission channel reliability, etc.), and propose a dynamic computation migration approach which can adaptively identify the latency optimal computation layer at runtime. We evaluate our solution using a case-study of artificial neural network based arrhythmia classification using a simulation environment as well as a real test-bed. Sina Shahhosseini, Iman Azimi, Arman Anzanpour, Axel Jantsch, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani |
ACM Great Lakes Symposium on VLSI | 7 |
| 2019 | SOSA: Self-Optimizing Learning with Self-Adaptive Control for Hierarchical System-on-Chip ManagementabstractResource management strategies for many-core systems dictate the sharing of resources among applications such as power, processing cores, and memory bandwidth in order to achieve system goals. System goals require consideration of both system constraints (e.g., power envelope) and user demands (e.g., response time, energy-efficiency). Existing approaches use heuristics, control theory, and machine learning for resource management. They all depend on static system models, requiring a priori knowledge of system dynamics, and are therefore too rigid to adapt to emerging workloads or changing system dynamics. Bryan Donyanavard, Tiago Rogério Mück, Amir-Mohammad Rahmani, Nikil Dutt, Armin Sadighi, Florian Maurer 0003, Andreas Herkersdorf |
MICRO | 3 |
| 2019 | Missing data resilient decision-making for healthcare IoT through personalization: A case study on maternal healthabstractRemote health monitoring is an effective method to enable tracking of at-risk patients outside of conventional clinical settings, providing early-detection of diseases and preventive care as well as diminishing healthcare costs. Internet-of-Things (IoT) technology facilitates developments of such monitoring systems although significant challenges need to be addressed in the real-world trials. Missing data is a prevalent issue in these systems, as data acquisition may be interrupted from time to time in long-term monitoring scenarios. This issue causes inconsistent and incomplete data and subsequently could lead to failure in decision making. Analysis of missing data has been tackled in several studies. However, these techniques are inadequate for real-time health monitoring as they neglect the variability of the missing data. This issue is significant when the vital signs are being missed since they depend on different factors such as physical activities and surrounding environment. Therefore, a holistic approach to customize missing data in real-time health monitoring systems is required, considering a wide range of parameters while minimizing the bias of estimates. In this paper, we propose a personalized missing data resilient decision-making approach to deliver health decisions 24/7 despite missing values. The approach leverages various data resources in IoT-based systems to impute missing values and provide an acceptable result. We validate our approach via a real human subject trial on maternity health, in which 20 pregnant women were remotely monitored for 7 months. In this setup, a real-time health application is considered, where maternal health status is estimated utilizing maternal heart rate. The accuracy of the proposed approach is evaluated, in comparison to existing methods. The proposed approach results in more accurate estimates especially when the missing window is large. Iman Azimi, Tapio Pahikkala, Amir-Mohammad Rahmani, Hannakaisa Niela-Vilén, Anna Axelin, Pasi Liljeberg |
Future Gener. Comput. Syst. | 3 |
| 2019 | Energy efficient fog-assisted IoT system for monitoring diabetic patients with cardiovascular disease
Tuan Nguyen Gia, Imed Ben Dhaou, Mai Ali, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen |
Future Gener. Comput. Syst. | 4 |
| 2019 | Hierarchical adaptive Multi-objective resource management for many-core systems
Andre L. M. Martins, Alzemiro Henrique Lucas da Silva, Amir-Mohammad Rahmani, Nikil Dutt, Fernando Gehm Moraes |
J. Syst. Archit. | 3 |
| 2019 | HESSLE-FREE: <u>He</u>terogeneou<u>s</u> <u>S</u>ystems <u>Le</u>veraging <u>F</u>uzzy Control for <u>R</u>untim<u>e</u> Resourc<u>e</u> ManagementabstractAs computing platforms increasingly embrace heterogeneity, runtime resource managers need to efficiently, dynamically, and robustly manage shared resources (e.g., cores, power budgets, memory bandwidth). To address the complexities in heterogeneous systems, state-of-the-art techniques that use heuristics or machine learning have been proposed. On the other hand, conventional control theory can be used for formal guarantees, but may face unmanageable complexity for modeling system dynamics of complex heterogeneous systems. We address this challenge through HESSLE-FREE (Heterogeneous Systems Leveraging Fuzzy Control for Runtime Resource Management): an approach leveraging fuzzy control theory that combines the strengths of classical control theory together with heuristics to form a light-weight, agile, and efficient runtime resource manager for heterogeneous systems. We demonstrate the efficacy of HESSLE-FREE executing on a NVIDIA Jetson TX2 platform (containing a heterogeneous multi-processor with a GPU) to show that HESSLE-FREE: 1) provides opportunity for optimization in the controller and stability analysis to enhance the confidence in the reliability of the system; 2) coordinates heterogeneous compute units to achieve desired objectives (e.g., QoS, optimal power references, FPS) efficiently and with lower complexity , and 3) eases the burden of system specification. Kasra Moazzemi, Biswadip Maity, Saehanseul Yi, Amir-Mohammad Rahmani, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | SPECTR: Formal Supervisory Control and Coordination for Many-core Systems Resource ManagementabstractResource management strategies for many-core systems need to enable sharing of resources such as power, processing cores, and memory bandwidth while coordinating the priority and significance of system- and application-level objectives at runtime in a scalable and robust manner. State-of-the-art approaches use heuristics or machine learning for resource management, but unfortunately lack formalism in providing robustness against unexpected corner cases. While recent efforts deploy classical control-theoretic approaches with some guarantees and formalism, they lack scalability and autonomy to meet changing runtime goals. We present SPECTR, a new resource management approach for many-core systems that leverages formal supervisory control theory (SCT) to combine the strengths of classical control theory with state-of-the-art heuristic approaches to efficiently meet changing runtime goals. SPECTR is a scalable and robust control architecture and a systematic design flow for hierarchical control of many-core systems. SPECTR leverages SCT techniques such as gain scheduling to allow autonomy for individual controllers. It facilitates automatic synthesis of the high-level supervisory controller and its property verification. We implement SPECTR on an Exynos platform containing ARM»s big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that SPECTR»s use of SCT is key to managing multiple interacting resources (e.g., chip power and processing cores) in the presence of competing objectives (e.g., satisfying QoS vs. power capping). The principles of SPECTR are easily applicable to any resource type and objective as long as the management problem can be modeled using dynamical systems theory (e.g., difference equations), discrete-event dynamic systems, or fuzzy dynamics. Amir-Mohammad Rahmani, Bryan Donyanavard, Tiago Rogério Mück, Kasra Moazzemi, Axel Jantsch, Onur Mutlu, Nikil Dutt |
ASPLOS | 1 |
| 2018 | Approximation-aware coordinated power/performance management for heterogeneous multi-coresabstractRun-time resource management of heterogeneous multi-core systems is challenging due to i) dynamic workloads, that often result in ii) conflicting knob actuation decisions, which potentially iii) compromise on performance for thermal safety. We present a runtime resource management strategy for performance guarantees under power constraints using functionally approximate kernels that exploit accuracy-performance trade-offs within error resilient applications. Our controller integrates approximation with power knobs - DVFS, CPU quota, task migration - in coordinated manner to make performance-aware decisions on power management under variable workloads. Experimental results on Odroid XU3 show the effectiveness of this strategy in meeting performance requirements without power violations compared to existing solutions. Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Cristiana Bolchini, Nikil Dutt |
DAC | 3 |
| 2018 | Gain scheduled control for nonlinear power management in CMPsabstractDynamic voltage and frequency scaling (DVFS) is a well-established technique for power management of thermal-or energy-sensitive chip multiprocessors (CMPs). In this context, linear control theoretic solutions have been successfully implemented to control the voltage-frequency knobs. However, modern CMPs with a large range of operating frequencies and multiple voltage levels display nonlinear behavior in the relationship between frequency and power. State-of-the-art linear controllers therefore under-optimize DVFS operation. We propose a Gain Scheduled Controller (GSC) for nonlinear runtime power management of CMPs that simplifies the controller implementation of systems with varying dynamic properties by utilizing an adaptive control theoretic approach in conjunction with static linear controllers. Our design improves the accuracy of the controller over a static linear controller with minimal overhead. We implement our approach on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that the system's response to changes in target power is improved by 2x while operating up to 12% more efficiently for tracking accuracy. Bryan Donyanavard, Amir-Mohammad Rahmani, Tiago Rogério Mück, Kasra Moazemmi, Nikil Dutt |
DATE | 2 |
| 2018 | Design methodologies for enabling self-awareness in autonomous systemsabstractThis paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications. Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi |
DATE | 7 |
| 2018 | Trends in On-chip Dynamic Resource ManagementabstractThe Complexity of emerging multi/many-core architectures and diversity of modern workloads demands coordinated dynamic resource management methods. We introduce a classification for these methods capturing the utilized resources and metrics. In this work, we use this classification to survey the key efforts in dynamic resource management. We first cover heuristic and optimization methods used to manage resources such as power, energy, temperature, Quality-of-Service (QoS) and reliability of the system. We then identify some of the machine learning based methods used in tuning architectural parameters in computer systems. In many cases, resource managers need to enforce design constraints during runtime with a certain level of guarantee. Hence, we also study the trend in deploying formal control theoretic approaches in order to achieve efficient and robust dynamic resource management. Kasra Moazzemi, Anil Kanduri, David Juhasz, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DSD | 5 |
| 2018 | Edge-Assisted Sensor Control in Healthcare IoTabstractThe Internet of Things is a key enabler of mobile health-care applications. However, the inherent constraints of mobile devices, such as limited availability of energy, can impair their ability to produce accurate data and, in turn, degrade the output of algorithms processing them in real-time to evaluate the patient's state. This paper presents an edge-assisted framework, where models and control generated by an edge server inform the sensing parameters of mobile sensors. The objective is to maximize the probability that anomalies in the collected signals are detected over extensive periods of time under battery-imposed constraints. Although the proposed concept is general, the control framework is made specific to a use-case where vital signs -heart rate, respiration rate and oxygen saturation- are extracted from a Photoplethysmogram (PPG) signal to detect anomalies in real-time. Experimental results show a 16.9% reduction in sensing energy consumption in comparison to a constant energy consumption with the maximum misdetection probability of 0.17 in a 24-hour health monitoring system. Delaram Amiri, Arman Anzanpour, Iman Azimi, Marco Levorato, Amir-Mohammad Rahmani, Pasi Liljeberg, Nikil Dutt |
GLOBECOM | 5 |
| 2018 | Approximation for Run-time Power ManagementabstractPerformance and energy efficiency of multi-core and many-core systems are restricted by increasing power densities and/or limited energy resources. Maximizing performance while minimizing power and energy consumption becomes challenging with emerging workloads. Approximate computing is an alternative solution that offers the required performance and energy gains, leveraging inherent error resilience of specific application domains. Dynamic power management using approximation as another knob can maximize performance and energy efficiency within fixed power budgets. Disciplined tuning of approximation along with other traditional power knobs requires efficient runtime resource management techniques. We present our strategy for using approximation as another knob for tuning the performance loss incurred in power actuation in many-core systems, which is also portable for heterogeneous multi-core systems. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg |
ISCAS | 3 |
| 2018 | Enhancing transient fault tolerance in embedded systems through an OS task level redundancy approach
Seyyed Amir Asghari, Mohammadreza Binesh Marvasti, Amir-Mohammad Rahmani |
Future Gener. Comput. Syst. | 3 |
| 2018 | Internet-of-Things and big data for smarter healthcare: From device to architecture, applications and analytics
Farshad Firouzi, Amir-Mohammad Rahmani, Kunal Mankodiya, Mustafa Badaroglu, Geoff V. Merrett, Bahareh J. Farahani |
Future Gener. Comput. Syst. | 2 |
| 2018 | Exploiting smart e-Health gateways at the edge of healthcare Internet-of-Things: A fog computing approach
Amir-Mohammad Rahmani, Tuan Nguyen Gia, Behailu Negash, Arman Anzanpour, Iman Azimi, Mingzhe Jiang, Pasi Liljeberg |
Future Gener. Comput. Syst. | 1 |
| 2018 | Platform-Centric Self-Awareness as a Key Enabler for Controlling Changes in CPSabstractFuture cyber-physical systems will host a large number of coexisting distributed applications on hardware platforms with thousands to millions of networked components communicating over open networks. These applications and networks are subject to continuous change. The current separation of design process and operation in the field will be superseded by a life-long design process of adaptation, infield integration, and update. Continuous change and evolution, application interference, environment dynamics and uncertainty lead to complex effects which must be controlled to serve a growing set of platform and application needs. Self-adaptation based on self-awareness and self-configuration has been proposed as a basis for such a continuous in-field process. Research is needed to develop automated in-field design methods and tools with the required safety, availability, and security guarantees. The paper shows two complementary use cases of self-awareness in architectures, methods, and tools for cyber-physical systems. The first use case focuses on safety and availability guarantees in self-aware vehicle platforms. It combines contracting mechanisms, tool based self-analysis and self-configuration. A software architecture and a runtime environment executing these tools and mechanisms autonomously are presented including aspects of self-protection against failures and security threats. The second use case addresses variability and long term evolution in networked MPSoC integrating hardware and software mechanisms of surveillance, monitoring, and continuous adaptation. The approach resembles the logistics and operation principles of manufacturing plants which gave rise to the metaphoric term of an Information Processing Factory that relies on incremental changes and feedback control. Both use cases are investigated by larger research groups. Despite their different approaches, both use cases face similar design and design automation challenges which will be summarized in the end. We will argue that seemingly unrelated research challenges, such as in machine learning and security, could also profit from the methods and superior modeling capabilities of self-aware systems. Mischa Möstl, Johannes Schlatow, Rolf Ernst, Nikil Dutt, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Fadi J. Kurdahi, Thomas Wild, Armin Sadighi, Andreas Herkersdorf |
Proc. IEEE | 6 |
| 2018 | adBoost: Thermal Aware Performance Boosting Through Dark Silicon PatterningabstractIncreasing power densities of many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. Efficient management of critical interlinked parameters - power, performance and temperature can improve resource utilization and mitigate dark silicon. In this paper, we present a run-time resource management system for thermal aware performance boosting using a dark silicon aware run-time application mapping strategy. The mapping policy patterns inactive cores among active cores for relatively lower and even distribution of operating temperatures. This provides enough thermal headroom for boosting the frequency of active cores upon performance surges and allows sustained boosting periods, improving the performance further. We design a controller for thermal aware performance boosting that decides on efficient allocation utilization of power budget and thermal headroom obtained from patterning. Our strategy yields up to 37 percent better throughput, 29 percent lower waiting time and up to 2 x longer boosting periods, in comparison with other state-of-the-art run-time mapping policies. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Muhammad Shafique 0001, Axel Jantsch, Pasi Liljeberg |
IEEE Trans. Computers | 3 |
| 2018 | IoT-Based Remote Pain Monitoring System: From Device to Cloud PlatformabstractFacial expressions are among behavioral signs of pain that can be employed as an entry point to develop an automatic human pain assessment tool. Such a tool can be an alternative to the self-report method and particularly serve patients who are unable to self-report like patients in the intensive care unit and minors. In this paper, a wearable device with a biosensing facial mask is proposed to monitor pain intensity of a patient by utilizing facial surface electromyogram (sEMG). The wearable device works as a wireless sensor node and is integrated into an Internet of Things (IoT) system for remote pain monitoring. In the sensor node, up to eight channels of sEMG can be each sampled at 1000 Hz, to cover its full frequency range, and transmitted to the cloud server via the gateway in real time. In addition, both low energy consumption and wearing comfort are considered throughout the wearable device design for long-term monitoring. To remotely illustrate real-time pain data to caregivers, a mobile web application is developed for real-time streaming of high-volume sEMG data, digital signal processing, interpreting, and visualization. The cloud platform in the system acts as a bridge between the sensor node and web browser, managing wireless communication between the server and the web application. In summary, this study proposes a scalable IoT system for real-time biopotential monitoring and a wearable solution for automatic pain assessment via facial expressions. Geng Yang 0003, Mingzhe Jiang, Wei Ouyang 0001, Guangchao Ji, Haibo Xie, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
IEEE J. Biomed. Health Informatics | 6 |
| 2017 | Ultra-short-term analysis of heart rate variability for real-time acute pain monitoring with wearable electronicsabstractIn medical care, it is essential to assess and manage acute painful conditions adequately. Heart rate variability (HRV) analysis is based on the acquisition of electrocardiogram (ECG), which is available from both patient monitor and wearable device. As HRV analysis can reflect autonomic nervous system activity which is unconsciously regulated, HRV analysis in ultra-short-term is getting attention in indicating the reaction due to acute pain. Different HRV features in different window lengths are involved in pain monitoring studies as a signal index or part of a multi-parameter model. In this work, seven HRV features and median heart rate (HR) in ultra-short-term are evaluated for their competence in indicating experimental acute pain. Also, the choice of time window length in HRV analysis and its relation with pain detection are discussed. The results of the normalized HRV analysis from healthy volunteers show that the changes of lnRMSSD, pNN20 and median HR associated with the intensity of experimental electrical pain; and in the tests with experimental thermal pain, lnLF and ln(LF/HF) changed along with pain intensity. The fusion of the HRV features could tell pain from no pain. With either experimental pain stimulation, optimal time window length was observed around or larger than 40 seconds with better correlation analysis result and HRV feature fusion performance. Mingzhe Jiang, Riitta Mieronkoski, Amir-Mohammad Rahmani, Nora Hagelberg, Sanna Salanterä, Pasi Liljeberg |
BIBM | 3 |
| 2017 | Quality-configurable memory hierarchy through approximation: special sessionabstractThe memory subsystem is a major contributor to the overall performance and energy consumption of embedded computing platforms. The emergence of "killer" applications such as data-intensive recognition, mining, and synthesis (RMS) applications puts even more stress on the memory subsystem and exacerbates its energy consumption. Traditional mechanisms to ensure data integrity deploy overdesign (e.g., redundancy and error detection/correction) and/or guard-banding that consumes a significant part of the energy consumed in the memory subsystem. We explore opportunities for energy efficiency by exploiting the intrinsic tolerance of a vast class of approximate computing applications to some level of error in the on-chip memory hierarchy. We present two exemplars outlining the typical software and hardware mechanisms that are required for different components in the memory hierarchy, implemented in varying technologies such as SRAM and STT-MRAM. Majid Namaki-Shoushtari, Amir-Mohammad Rahmani, Nikil Dutt |
CASES | 2 |
| 2017 | From threads to events: Adapting a lightweight middleware for Contiki OSabstractInteroperability is one of the key requirements in the Internet of Things considering the diverse platforms, communication standards and specifications available today. Inherent resource constraints in the majority of IoT devices makes it very difficult to use existing solutions for interoperability, thus demanding new approaches. This paper presents the process of adapting a lightweight interoperability middleware for IoT, LISA, from RIOT to Contiki OS and evaluates memory and power overheads. The middleware follows a service oriented architecture and classifies devices according to available resources to assign different roles, such as Application, Service and Manager Nodes. These roles live in different tiers in a generic IoT architecture, where the Manager nodes are located in the intermediate Fog layer. To adapt to an event based kernel of Contiki, the middleware defines and handles a set of events that are used to communicate with the user application. A network of nodes is simulated to show the architecture promoted by the middleware and the results are presented. Uzair A. Noman, Behailu Negash, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
CCNC | 3 |
| 2017 | Smart energy efficient gateway for Internet of mobile thingsabstractInternet of Things (IoT) is a fast developing vision in which physical quantities are digitized, processed and analyzed. Internet of Mobile Things (IoMT) as one of new domains of IoT, due to mobility, requires a more demanding and rigorous solution in many aspects, especially in terms of energy efficiency. We propose a solution consisting of energy efficient and fast hardware platform for building IoMT Fog layer facilities. Experimental results are presented to prove superiority of the proposed hardware in several aspects to popular general purpose platforms. Igor Tcarenko, Yuxiang Huan, David Juhasz, Amir-Mohammad Rahmani, Zhuo Zou, Tomi Westerlund, Pasi Liljeberg, Lirong Zheng 0001, Hannu Tenhunen |
CCNC | 4 |
| 2017 | Self-awareness in remote health monitoring systems using wearable electronicsabstractIn healthcare, effective monitoring of patients plays a key role in detecting health deterioration early enough. Many signs of deterioration exist as early as 24 hours prior having a serious impact on the health of a person. As hospitalization times have to be minimized, in-home or remote early warning systems can fill the gap by allowing in-home care while having the potentially problematic conditions and their signs under surveillance and control. This work presents a remote monitoring and diagnostic system that provides a holistic perspective of patients and their health conditions. We discuss how the concept of self-awareness can be used in various parts of the system such as information collection through wearable sensors, confidence assessment of the sensory data, the knowledge base of the patient's health situation, and automation of reasoning about the health situation. Our approach to self-awareness provides (i) situation awareness to consider the impact of variations such as sleeping, walking, running, and resting, (ii) system personalization by reflecting parameters such as age, body mass index, and gender, and (iii) the attention property of self-awareness to improve the energy efficiency and dependability of the system via adjusting the priorities of the sensory data collection. We evaluate the proposed method using a full system demonstration. Arman Anzanpour, Iman Azimi, Maximilian Götzinger, Amir-Mohammad Rahmani, Nima Taherinejad, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DATE | 4 |
| 2017 | Autonomous Patient/Home Health Monitoring Powered by Energy HarvestingabstractThis paper presents the design of an autonomous smart patient/home health monitoring system. Both patient physiological parameters as well as room conditions are being monitored continuously to insure patient safety. The sensors are connected on an IoT regime, where the collected data is wirelessly transferred to a nearby gateway which performs preliminary data analysis, commonly referred to as fog computing, to make sure emergency personnel and healthcare providers are notified in case patient being monitored is at risk. To achieve power autonomy three energy harvesting sources are proposed, namely, solar, RF and thermal. The design of RF energy harvesting system is demonstrated, where novel multiband antenna is fabricated as well as an efficient RF- DC rectifier achieving maximum efficiency of 84%. Finally, the sensor node is tested with different type of sensors and settings while being solely powered by a Photovoltaic (PV) solar cell. Mai Ali, Tuan Nguyen Gia, Abd-Elhamid M. Taha, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen |
GLOBECOM | 4 |
| 2017 | QuARK: Quality-configurable approximate STT-MRAM cache by fine-grained tuning of reliability-energy knobsabstractEmerging STT-MRAM memories are promising alternatives for SRAM memories to tackle their low density and high static power consumption, but impose high energy consumption for reliable read/write operations. However, absolute data integrity is not required for many approximate computing applications, allowing energy savings with minimal quality loss. This paper proposes QuARK, a hardware/software approach for trading reliability of STT-MRAM caches for energy savings in the on-chip memory hierarchy of multi- and many-core systems running approximate applications. In contrast to SRAM-based cache-way-level actuators, QuARK utilizes fine-grained cache-line-level actuation knobs with different levels of reliability for individual read and write accesses which are unique to STT-MRAM and suitable for systems running multiple applications with mixed accuracy sensitivity, thus avoiding interapplication actuation interference. Our experimental results with a set of recognition, mining and synthesis (RMS) benchmarks demonstrate up to 40% energy savings over a fully-protected STT-MRAM cache, with negligible loss in the quality of the generated outputs. Amir Mahdi Hosseini Monazzah, Majid Namaki-Shoushtari, Seyed Ghassem Miremadi, Amir-Mohammad Rahmani, Nikil Dutt |
ISLPED | 4 |
| 2017 | Low-cost fog-assisted health-care IoT system with energy-efficient sensor nodesabstractA better lifestyle starts with a healthy heart. Unfortunately, millions of people around the world are either directly affected by heart diseases such as coronary artery disease and heart muscle disease (Cardiomyopathy), or are indirectly having heart-related problems like heart attack and/or heart rate irregularity. Monitoring and analyzing these heart conditions in some cases could save a life if proper actions are taken accordingly. A widely used method to monitor these heart conditions is to use ECG or electrocardiography. However, devices used for ECG are costly, energy inefficient, bulky, and mostly limited to the ambulatory environment. With the advancement and higher affordability of Internet of Things (IoT), it is possible to establish better health-care by providing real-time monitoring and analysis of ECG. In this paper, we present a low-cost health monitoring system that provides continuous remote monitoring of ECG together with automatic analysis and notification. The system consists of energy-efficient sensor nodes and a fog layer altogether taking advantage of IoT. The sensor nodes collect and wirelessly transmit ECG, respiration rate, and body temperature to a smart gateway which can be accessed by appropriate care-givers. In addition, the system can represent the collected data in useful ways, perform automatic decision making and provide many advanced services such as real-time notifications for immediate attention. Tuan Nguyen Gia, Mingzhe Jiang, Victor K. Sarker, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen |
IWCMC | 4 |
| 2017 | Special issue on energy efficient multi-core and many-core systems, Part II
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum |
J. Parallel Distributed Comput. | 1 |
| 2017 | Performance/Reliability-Aware Resource Management for Many-Cores in Dark Silicon EraabstractAggressive technology scaling has enabled the fabrication of many-core architectures while triggering challenges such as limited power budget and increased reliability issues, like aging phenomena. Dynamic power management and runtime mapping strategies can be utilized in such systems to achieve optimal performance while satisfying power constraints. However, lifetime reliability is generally neglected. We propose a novel lifetime reliability/performance-aware resource co-management approach for many-core architectures in the dark silicon era. The approach is based on a two-layered architecture, composed of a long-term runtime reliability controller and a short-term runtime mapping and resource management unit. The former evaluates the cores' aging status w.r.t. a target reference specified by the designer, and performs recovery actions on highly stressed cores by means of power capping. The aging status is utilized in runtime application mapping to maximize system performance while fulfilling reliability requirements and honoring the power budget. Experimental evaluation demonstrates the effectiveness of the proposed strategy, which outperforms most recent state-of-the-art contributions. M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
IEEE Trans. Computers | 3 |
| 2017 | HiCH: Hierarchical Fog-Assisted Computing Architecture for Healthcare IoTabstractThe Internet of Things (IoT) paradigm holds significant promises for remote health monitoring systems. Due to their life- or mission-critical nature, these systems need to provide a high level of availability and accuracy. On the one hand, centralized cloud-based IoT systems lack reliability, punctuality and availability (e.g., in case of slow or unreliable Internet connection), and on the other hand, fully outsourcing data analytics to the edge of the network can result in diminished level of accuracy and adaptability due to the limited computational capacity in edge nodes. In this paper, we tackle these issues by proposing a hierarchical computing architecture, HiCH, for IoT-based health monitoring systems. The core components of the proposed system are 1) a novel computing architecture suitable for hierarchical partitioning and execution of machine learning based data analytics, 2) a closed-loop management technique capable of autonomous system adjustments with respect to patient’s condition. HiCH benefits from the features offered by both fog and cloud computing and introduces a tailored management methodology for healthcare IoT systems. We demonstrate the efficacy of HiCH via a comprehensive performance assessment and evaluation on a continuous remote health monitoring case study focusing on arrhythmia detection for patients suffering from CardioVascular Diseases (CVDs). Iman Azimi, Arman Anzanpour, Amir-Mohammad Rahmani, Tapio Pahikkala, Marco Levorato, Pasi Liljeberg, Nikil Dutt |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient ApplicationsabstractPower capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Reliability-Aware Runtime Power Management for Many-Core Systems in the Dark Silicon EraabstractPower management of networked many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates considering network characteristics at runtime to achieve better performance while honoring the peak power upper bound. On the other hand, power management has a direct effect on chip temperature, which is the main driver of the aging effects. Therefore, alongside performance fulfillment, the controlling mechanism must also consider the current cores' reliability in its actuator manipulation to enhance the overall system lifetime in the long term. In this paper, we propose a multiobjective dynamic power management technique that uses current power consumption and other network characteristics including the reliability of the cores as the feedback while utilizing fine-grained voltage and frequency scaling and per-core power gating as the actuators. In addition, disturbance rejecter and reliability balancer are designed to help the controller to better smooth power consumption in the short term and reliability in the long term, respectively. Simulations of dynamic workloads and mixed criticality application profiles show that our method not only is effective in honoring the power budget while considerably boosting the system throughput, but also increases the overall system lifetime by minimizing aging effects by means of power consumption balancing. Amir-Mohammad Rahmani, M. H. Haghbayan, Antonio Miele, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | A lifetime-aware runtime mapping approach for many-core systems in the dark silicon era
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
DATE | 3 |
| 2016 | Approximation knob: power capping meets energy efficiencyabstractPower Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen |
ICCAD | 3 |
| 2016 | End-to-end security scheme for mobility enabled healthcare Internet of Things
Sanaz Rahimi Moosavi, Tuan Nguyen Gia, Ethiopia Nigussie, Amir-Mohammad Rahmani, Seppo Virtanen, Hannu Tenhunen, Jouni Isoaho |
Future Gener. Comput. Syst. | 4 |
| 2016 | Special issue on energy efficient multi-core and many-core systems, Part I
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum |
J. Parallel Distributed Comput. | 1 |
| 2016 | A Power-Aware Approach for Online Test Scheduling in Many-Core ArchitecturesabstractAggressive technology scaling triggers novel challenges to the design of multi-/many-core systems, such as limited power budget and increased reliability issues. Today's many-core systems employ dynamic power management and runtime mapping strategies trying to offer optimal performance while fulfilling power constraints. On the other hand, due to the reliability challenges, online testing techniques are becoming a necessity in current and near future technologies. However, state-of-the-art techniques are not aware of the other power/performance requirements. This paper proposes a power-aware non-intrusive online testing approach for many-core systems. The approach schedules software based self-test routines on the various cores during their idle periods, while honoring the power budget and limiting delays in the workload execution. A test criticality metric, based on a device aging model, is used to select cores to be tested at a time. Moreover, power and reliability issues related to the testing at different voltage and frequency levels are also handled. Extensive experimental results reveal that the proposed approach can i) efficiently test the cores within the available power budget causing a negligible performance penalty, ii) adapt the test frequency to the current cores' aging status, and iii) cover available voltage and frequency levels during the testing. M. H. Haghbayan, Amir-Mohammad Rahmani, Antonio Miele, Mohammad Fattah, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen |
IEEE Trans. Computers | 2 |
| 2015 | Smart e-Health Gateway: Bringing intelligence to Internet-of-Things based ubiquitous healthcare systemsabstractThere have been significant advances in the field of Internet of Things (IoT) recently. At the same time there exists an ever-growing demand for ubiquitous healthcare systems to improve human health and well-being. In most of IoT-based patient monitoring systems, especially at smart homes or hospitals, there exists a bridging point (i.e., gateway) between a sensor network and the Internet which often just performs basic functions such as translating between the protocols used in the Internet and sensor networks. These gateways have beneficial knowledge and constructive control over both the sensor network and the data to be transmitted through the Internet. In this paper, we exploit the strategic position of such gateways to offer several higher-level services such as local storage, real-time local data processing, embedded data mining, etc., proposing thus a Smart e-Health Gateway. By taking responsibility for handling some burdens of the sensor network and a remote healthcare center, a Smart e-Health Gateway can cope with many challenges in ubiquitous healthcare systems such as energy efficiency, scalability, and reliability issues. A successful implementation of Smart e-Health Gateways enables massive deployment of ubiquitous health monitoring systems especially in clinical environments. We also present a case study of a Smart e-Health Gateway called UTGATE where some of the discussed higher-level features have been implemented. Our proof-of-concept design demonstrates an IoT-based health monitoring system with enhanced overall system energy efficiency, performance, interoperability, security, and reliability. Amir-Mohammad Rahmani, Nanda Kumar Thanigaivelan, Tuan Nguyen Gia, Jose David Granados Vergara, Behailu Negash, Pasi Liljeberg, Hannu Tenhunen |
CCNC | 1 |
| 2015 | Power-aware online testing of manycore systems in the dark silicon era
M. H. Haghbayan, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Zainalabedin Navabi, Hannu Tenhunen |
DATE | 2 |
| 2015 | Dark silicon aware runtime mapping for many-core systems: A patterning approachabstractLimitation on power budget in many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. In such systems, an efficient run-time application mapping approach can considerably enhance resource utilization and mitigate the dark silicon phenomenon. In this paper, we propose a dark silicon aware runtime application mapping approach that patterns active cores alongside the inactive cores in order to evenly distribute power density across the chip. This approach leverages dark silicon to balance the temperature of active cores to provide higher power budget and better resource utilization, within a safe peak operating temperature. In contrast with exhaustive search based mapping approach, our agile heuristic approach has a negligible runtime overhead. Our patterning strategy yields a surplus power budget of up to 17% along with an improved throughput of up to 21% in comparison with other state-of-the-art run-time mapping strategies, while the surplus budget is as high as 40% compared to worst case scenarios. Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
ICCD | 3 |
| 2015 | Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approachabstractPower management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies. Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen |
ISLPED | 1 |
| 2015 | MapPro: Proactive Runtime Mapping for Dynamic Workloads by Quantifying Ripple Effect of Applications on Networks-on-ChipabstractIncreasing dynamic workloads running on NoC-based many-core systems necessitates efficient runtime mapping strategies. With an unpredictable nature of application profiles, selecting a rational region to map an incoming application is an NP-hard problem in view of minimizing congestion and maximizing performance. In this paper, we propose a proactive region selection strategy which prioritizes nodes that offer lower congestion and dispersion. Our proposed strategy, MapPro, quantitatively represents the propagated impact of spatial availability and dispersion on the network with every new mapped application. This allows us to identify a suitable region to accommodate an incoming application that results in minimal congestion and dispersion. We cluster the network into squares of different radii to suit applications of different sizes and proactively select a suitable square for a new application, eliminating the overhead caused with typical reactive mapping approaches. We evaluated our proposed strategy over different traffic patterns and observed gains of up to 41% in energy efficiency, 28% in congestion and 21% dispersion when compared to the state-of-the-art region selection methods. M. H. Haghbayan, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
NOCS | 3 |
| 2014 | Online testing of many-core systems in the Dark Silicon eraabstractAs the dark silicon era is about to embrace, it is not anymore possible to attain commensurate performance benefits by increasing the number of transistors due to thermal design power. Dark Silicon issue stresses that a fraction of silicon chip being able to switch in full frequency is dropping and designers will soon face the growing underutilization inherent in future technologies. On the other hand, by reducing the transistor size, susceptibility to internal defects drastically increases and large ranges of defects such as aging or transient faults will be shown up more frequently. In this paper, we propose an online test scheduling algorithm using software based self-test for dark silicon era to test dark cores while considering thermal design power of the system. As the dark area of the system is dynamic and reshapes at a runtime, the tested cores can be used by other applications in the near future. Empirical results show the effectiveness of the proposed algorithm in terms of applicability and fault coverage with a negligible negative impact on the system throughput. M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DDECS | 2 |
| 2014 | Dark silicon aware power management for manycore systems under dynamic workloadsabstractDark Silicon denotes the phenomenon that, due to thermal and power constraints, the fraction of transistors that can operate at full frequency is decreasing with each technology generation. We propose a PID (Proportional Integral Derivative) controller based dynamic power management method that considers an upper bound on power consumption (called the Thermal Design Power (TDP)). To avoid violation of the TDP constraint for manycore systems running highly dynamic workloads, it provides fine-grained DVFS (Dynamic Voltage and Frequency Scaling) including near-threshold operation. In addition, the method distinguishes applications with hard Real-Time, soft Real-Time and no Real-Time constraints and treats them with appropriate priorities. In simulations with dynamic workloads mixed-critical application profiles, we show that the method is effective in honoring the TDP bound and it can boost system throughput by over 43% compared to a naive TDP scheduling policy. M. H. Haghbayan, Amir-Mohammad Rahmani, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen |
ICCD | 2 |
| 2014 | Mixed-Criticality Run-Time Task Mapping for NoC-Based Many-Core SystemsabstractContiguous processor allocation improves both the network and the application performance, by decreasing the congestion probability among communication of different applications. Consequently, the average, standard deviation and worst-case latency of the network is decreased significantly. This makes the contiguous allocation a good solution for time-critical applications with bounded deadlines. On the other hand, non-contiguous allocation will increase the system throughput significantly. Isolated nodes are utilized and more applications can finish their job in a time unit. However, this will lead to poor network metrics, unsuitable for real-time applications. In this work, we combine these two approaches in order to manage workloads with mixed-critical characteristics. Real-time applications are mapped contiguously, while non-critical applications are allowed to get dispersed over the available system nodes. Results show over 50% improvement in worst-case latency and 100 times improvement in deadline misses. Mohammad Fattah, Amir-Mohammad Rahmani, Thomas Canhao Xu, Anil Kanduri, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 2 |
| 2014 | High-Performance and Fault-Tolerant 3D NoC-Bus Hybrid Architecture Using ARB-NET-Based Adaptive Monitoring PlatformabstractThe emerging three-dimensional integrated circuits (3D-ICs) achieve greater device integration and enhanced system performance at lower cost and reduced area footprint, thereby offering higher order of connectivity and greater design choices and possibilities. To exploit the intrinsic capability of reduced communication distances in 3D-ICs, three-dimensional NoC-bus hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored possibility to implement an efficient system-wide monitoring network. In this paper, an efficient three-dimensional NoC architecture is proposed which is optimized for system performance, power consumption, and reliability. The mechanism benefits from a congestion-aware and bus failure-tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we have integrated a low-cost monitoring platform on top of the three-dimensional NoC-Bus Hybrid mesh architecture that can be efficiently used for various system management purposes such as traffic monitoring, fault tolerance, and thermal management. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware and interlayer fault-tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of information generated within bus arbiters. Compared to recently proposed stacked mesh three-dimensional NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power, performance, and reliability improvements with a negligible hardware overhead. Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
IEEE Trans. Computers | 1 |
| 2014 | A task migration mechanism for distributed many-core operating systems
Simon Holmbacka, Mohammad Fattah, Wictor Lund, Amir-Mohammad Rahmani, Sébastien Lafond, Johan Lilius |
J. Supercomput. | 4 |
| 2014 | Special section on advances in methods for adaptive multicore systems
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Supercomput. | 1 |
| 2013 | Enhancing Performance of 3D Interconnection Networks using Efficient Multicast Communication ProtocolabstractThree-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. They also provide greater design flexibility by allowing heterogeneous integration. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement efficient multicast routings for 3D networks-on-chip. In this paper, we propose a novel multicast partitioning and routing strategy for the 3D NoC-Bus Hybrid mesh architectures to enhance the overall system performance and reduce the power consumption. The proposed architecture exploits the beneficial attribute of a single-hop (bus-based) interlayer communication of the 3D stacked mesh architecture to provide high-performance hardware multicast support. To this end, a customized partitioning method and an efficient routing algorithm are presented to reduce the average hop count and latency of the network. Compared to the recently proposed 3D NoC architectures being capable of supporting hardware multicasting, our extensive simulations with different traffic profiles reveal that our architecture using the proposed multicast routing strategy can help achieve significant performance improvements. Sanaz Rahimi Moosavi, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 2 |
| 2013 | Partial Virtual Channel Sharing: A Generic Methodology to Enhance Resource Management and Fault Tolerance in Networks-on-Chip
Khalid Latif 0002, Amir-Mohammad Rahmani, Ethiopia Nigussie, Tiberiu Seceleanu, Martin Radetzki, Hannu Tenhunen |
J. Electron. Test. | 2 |
| 2013 | Developing a power-efficient and low-cost 3D NoC using smart GALS-based vertical channels
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
J. Comput. Syst. Sci. | 1 |
| 2013 | Design space exploration of thermal-aware many-core systems
Kameswar Rao Vaddina, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila |
J. Syst. Archit. | 2 |
| 2012 | ARB-NET: A novel adaptive monitoring platform for stacked mesh 3D NoC architecturesabstractThe emerging three-dimensional integrated circuits (3D ICs) offer a promising solution to mitigate the barriers of interconnect scaling in modern systems. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is presented taking advantage from viable information generated within bus arbiters. Our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help achieving significant power and performance improvements compared to recently proposed stacked mesh 3D NoCs. Amir-Mohammad Rahmani, Khalid Latif 0002, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
ASP-DAC | 1 |
| 2012 | A Cluster-Based Core Protection Technique for Networks-on-ChipabstractPartial Virtual channel Sharing (PVS) architecture has been proposed to enhance the performance of Networks-on-Chip (NoC) based systems. In this paper, a cluster based processing core protection technique for NoC systems using PVS approach is presented. In case of network level faults, the processing core of faulty node can use any other router in the cluster for transmission or reception of data packets with proposed architecture. Simulation results show significant reduction in average packet latency at the expense of negligible area overhead. Khalid Latif 0002, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen, Tiberiu Seceleanu |
COMPSAC | 2 |
| 2012 | Designing a High Performance and Reliable Networks-on-Chip Using Network Interface Assisted Routing StrategyabstractPartial Virtual channel Sharing (PVS) architecture has been proposed to enhance the performance of Networks-on-Chip (NoC) based systems. In this paper, we present an efficient and reliable Network Interface (NI) assisted routing strategy for NoC using PVS architecture. For this purpose, NoC system is divided into clusters. Each cluster is a group of two nodes comprising Processing Elements (PE), switches, links, etc. Each PE in a cluster can inject data to the network through a router, which is closer to the destination. This helps to reduce the network load by reducing the average hop count of the network. The proposed architecture can recover the PE disconnected from the network due to network level faults by allowing the PE to transmit and receive the packets through the other router in the cluster. 5̅×6 crossbar is used for the proposed architecture which requires one more 5×1 multiplexer without increasing the critical path delay of the router as compared to the 5×5 crossbar. The proposed router has been simulated for uniform and negative exponential distribution (NED) traffic patterns. The simulation results show the significant reduction in average packet latency at the expense of negligible area overhead. Khalid Latif 0002, Amir-Mohammad Rahmani, Tiberiu Seceleanu, Hannu Tenhunen |
DSD | 2 |
| 2012 | Power and Thermal Analysis of Stacked Mesh 3D NoC Using AdaptiveXYZ Routing AlgorithmabstractThree-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement an integrated low-cost system-wide monitoring network. In this paper, a generic monitoring and management platform called ARB-NET is presented. Based on the ARB-NET monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is provided which takes advantage from viable information generated within the monitoring network. In addition, we address both the power and thermal issues of a stacked mesh 3D network on chips using AdaptiveXYZ routing. To this end, a thermal model of a 3D stacked NoC system in a modern flip-chip package is developed. Thermal and power analysis are performed in order to investigate the impact of the proposed adaptive routing from the power and thermal perspectives. Our experiments with a videoconference encoder as a real application show significant power, performance and peak temperature improvements compared to a typical stacked mesh 3D NoC. Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DSD | 1 |
| 2012 | Generic Monitoring and Management Infrastructure for 3D NoC-Bus Hybrid ArchitecturesabstractThree-dimensional integrated circuits (3D ICs) achieve enhanced system integration and improved performance at lower cost and reduced area footprint. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed which provides performance, power consumption, and area benefits. Besides its various advantages, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes such as traffic monitoring, thermal management and fault tolerance. The proposed generic monitoring and management infrastructure called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring and management platform, a fully congestion-aware and inter-layer fault tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of viable information generated using bus arbiter network. In addition, we propose a thermal monitoring and management strategy on top of our ARB-NET infrastructure. Compared to recently proposed stacked mesh 3D NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power and performance improvements while preserving the system reliability with negligible hardware overhead. Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
NOCS | 1 |
| 2012 | Design of a Reliable XOR-XNOR Circuit for Arithmetic Logic Units
Mouna Karmani, Chiraz Khedhiri, Belgacem Hamdi, Amir-Mohammad Rahmani, Ka Lok Man, Kaiyu Wan |
NPC | 4 |
| 2012 | An Efficient Hybridization Scheme for Stacked Mesh 3D NoC ArchitectureabstractThree-dimensional (3D) integration is a viable design paradigm to overcome the existing interconnect bottleneck in integrated systems and enhance system power/performance characteristics. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, stacked mesh 3D NoC architecture was proposed. However, this architecture suffers from naive and straightforward hybridization between NoC and bus media. In this paper, an efficient hybridization scheme is presented to enhance system performance, power consumption, and area of stacked mesh 3D NoC architectures. By utilizing a routing rule called LastZ the proposed hybridization scheme offers many advantages investigated in detail to emphasize the significant achievements. Our extensive simulations with synthetic and real benchmarks, including an integrated videoconference application show that compared to a typical 3D NoC-Bus Hybrid Mesh architecture, our hybridization scheme achieves significant power, performance, and area improvements. Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 1 |
| 2011 | Enhancing Performance of NoC-Based Architectures Using Heuristic Virtual-Channel Sharing ApproachabstractThis paper presents a novel virtual-channel (VC) sharing technique for NoC architecture. The proposed architecture improves the utilization of resources to enhance the performance with minimal overheads. A heuristic approach towards the proper VC sharing strategy is proposed, which is performed by an adaptive algorithm that configures the VC sharing based on link load parameters. Architectural design to realize the adaptive VC sharing in generic router is elaborated. The technique can be applied to any NoC architecture, including 3-D NoCs. Extensive quantitative experiments with synthetic and real benchmarks, including an integrated video conference application, demonstrate considerable improvement in area and power efficiency compared to existing VC-based 2D/3D NoC architectures. Khalid Latif 0002, Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Tiberiu Seceleanu, Pasi Liljeberg, Hannu Tenhunen |
COMPSAC | 2 |
| 2011 | Enhancing Performance Sustainability of Fault Tolerant Routing Algorithms in NoC-Based ArchitecturesabstractReliability of embedded systems and devices is becoming a challenge with technology scaling. To deal with the reliability issues, fault tolerant solutions are needed. The design paradigm for future System-on-Chip (SoC) implementation is Network-on-Chip (NoC). Fault tolerance in NoC can be achieved at many abstraction levels. Many fault tolerant architectures and routing algorithms have already been proposed for NoC but the utilization of resources, affected indirectly by faults is yet to be addressed. In this paper, we propose a NoC architecture, which sustains the overall system performance by utilizing resources, which cannot be used by other architectures under faults. An approach towards a proper virtual-channel (VC) sharing strategy is proposed, based on communication bandwidth requirements. The technique can be applied to any NoC architecture, including 3-D NoCs. Extensive quantitative experiments with synthetic benchmarks, including uniform, transpose and negative exponential distribution (NED), demonstrate considerable improvement in terms of performance sustainability under faulty conditions compared to existing VC-based NoC architectures. Khalid Latif 0002, Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Tiberiu Seceleanu, Pasi Liljeberg, Hannu Tenhunen |
DSD | 2 |
| 2011 | LastZ: An Ultra Optimized 3D Networks-on-Chip Architectureabstract3D IC technology enables NoC architectures to offer greater device integration and shorter interlayer interconnects. The primary 3D NoC architectures such as Symmetric 3D Mesh NoC could not exploit the beneficial feature of a negligible inter-layer distance in 3D chips. To cope with this, 3D NoC-Bus Hybrid architecture was proposed which is a hybrid between packet-switched network and a bus. This architecture is feasible providing both performance and area benefits, while still suffering from naive and straightforward hybridization between NoC and bus media. In this paper, an ultra optimized hybridization scheme is proposed to enhance system performance, power consumption, area and thermal issues of 3D NoC-Bus Hybrid Mesh. The scheme benefits from a rule called LastZ which enables ultra optimization of the inter-layer communication architecture. In addition, we present a wrapper to preserve the backward compatibility of the proposed architecture for connecting with the existing network interfaces. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. Our extensive simulations demonstrate significant area, power, and performance improvements compared to a typical 3D NoC-Bus Hybrid Mesh architecture. Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DSD | 1 |
| 2011 | Thermal Analysis of Job Allocation and Scheduling Schemes for 3D Stacked NoC'sabstractThree-dimensional technology offers greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In this work, a 3D thermal model of a stacked network-on-chip system is developed and thermal analysis is performed in order to analyze different job allocation and scheduling schemes using finite element simulations. The steady-state heat transfer analysis on the 3D stacked structure has been performed. We have analyzed the effect of variation of die power consumption, with and without hotspots, on peak temperatures in different layers of the stack. The optimal die placement solution is also provided based on the maximum temperature attained by the individual silicon dies. Kameswar Rao Vaddina, Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila |
DSD | 2 |
| 2011 | Congestion aware, fault tolerant, and thermally efficient inter-layer communication scheme for hybrid NoC-bus 3D architecturesabstractThree-dimensional IC technology offers greater device integration and shorter interlayer interconnects. In order to take advantage of these attributes, 3D stacked mesh architecture was proposed which is a hybrid between packet-switched network and a bus. Stacked mesh is a feasible architecture which provides both performance and area benefits, while suffering from inefficient intermediate buffers. In this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we hybridize the proposed adaptive routing with available algorithms to mitigate the thermal issues by herding most of the switching activities closer to the heat sink. Our extensive simulations with synthetic and real benchmarks, including the one with an integrated videoconference application, demonstrate significant power, performance, and peak temperature improvements compared to a typical stacked mesh 3D NoC. Amir-Mohammad Rahmani, Pasi Liljeberg, Khalid Latif 0002, Juha Plosila, Kameswar Rao Vaddina, Hannu Tenhunen |
NOCS | 1 |
| 2011 | PVS-NoC: Partial Virtual Channel Sharing NoC ArchitectureabstractA novel architecture aiming for ideal performance and overhead tradeoff, PVS-NoC (Partial VC Sharing NoC), is presented. Virtual channel (VC) is an efficient technique to improve network performance, while suffering from large silicon and power overhead. We propose sharing the VC buffers among dual inputs, which provides the performance advantage as conventional VC-based router with minimized overhead. We reason theoretically and demonstrate quantitatively the benefits of proposed architecture by comparing to state-of-the-art NoC routers, with various traffic patterns. Extensive experiments with synthetic and real benchmarks show significant area and power saving with similar performance compared to latest VC based NoC architectures. Khalid Latif 0002, Amir-Mohammad Rahmani, Liang Guang, Tiberiu Seceleanu, Hannu Tenhunen |
PDP | 2 |
| 2011 | A Stacked Mesh 3D NoC Architecture Enabling Congestion-Aware and Reliable Inter-layer CommunicationabstractIn this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. Stacked mesh is a feasible architecture which takes advantage of the short inter-layer wiring delays, while suffering from inefficient intermediate buffers. To cope with this, an inter-layer communication mechanism is developed to enhance the buffer utilization, load balancing, and system fault-tolerance. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm for vertical communication. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. In addition, a video conference encoder has been used as a real application for system analysis. Our extensive experiments show significant power and performance improvements compared to a typical stacked mesh 3D NoC. Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
PDP | 1 |
| 2010 | Developing reconfigurable FIFOs to optimize power/performance of Voltage/Frequency Island-based networks-on-chipabstractNetwork-on-chip architectures partitioned into several Voltage/Frequency Islands (VFIs) have been proposed to alleviate problems related to integration, excessive energy consumption and clock distribution. The architecture is composed of synchronous switches that communicate with each other using bi-synchronous FIFOs. However, these FIFOs are not needed if adjacent switches belong to the same clock domain. In this paper, a Reconfigurable Synchronous/Bi-Synchronous (RSBS) FIFO is proposed which can operate in either synchronous or bi-synchronous mode. The FIFO is scalable and synthesizable in synchronous standard cells and also a technique for mesochronous adaptation has been recommended. In addition, some techniques are suggested to show how the FIFO could be utilized in a VFI-based NoC. Our results reveal that compared to a non-reconfigurable system architecture, the RSBS FIFOs help to achieve up to 15% savings in average power consumption of NoC switches and 29% improvement in total average packet latency in the case of MPEG-4 encoder application. Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen |
DDECS | 1 |
| 2010 | Power-aware NoC router using central forecasting-based dynamic virtual channel allocationabstractIn this paper, we propose a high performance central dynamic virtual channel allocation mechanism for on-chip routers. This central management unit devotes each input port a number of virtual channels (VC) among a shared VC bank based on a traffic forecasting technique. The forecasting technique exploits the link and VC utilizations in predicting the traffic. Based on the predicted traffic, for each input port, the number of active virtual channels may be increased, decreased, or kept unchanged. The clock-gating power management technique is used to activate/deactivate the VCs. Simulation results using uniform and Negative Exponential Distribution (NED) traffic profiles show that a considerable power savings in the virtual channels and overall router power consumption may be achieved especially in low traffic loads. The area overhead of the technique is negligible. Amir-Mohammad Rahmani, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen |
ISCAS | 1 |
| 2010 | EDXY - A low cost congestion-aware routing algorithm for network-on-chips
Pejman Lotfi-Kamran, Amir-Mohammad Rahmani, Masoud Daneshtalab, Ali Afzali-Kusha, Zainalabedin Navabi |
J. Syst. Archit. | 2 |