Juha Plosila

dblp:37/5831 · DBLP profile ↗
← Back
120ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-4018-5495ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 76 · 5 since 2021Software engineering, systems software and programming languages · 13 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Theory of computation · 3Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Integrated Modeling And Simulation Of Mechanical, Computational, Communication, And Battery In Robots
abstract
In this paper, we propose an integrated simulation framework for a mobile robot to analyze the dependencies among system components, particularly electrical power and battery models, rather than treating them independently. The framework captures the joint effects of mechanical power, computational power, and computational load offloading on total energy consumption and long-term battery health. It models mechanical, computational, and communication energy while accounting for their dependencies. As a case study, we apply the holistic framework to an energy-aware mission planning problem in which the robot travels to a target location and returns to its starting point. Along the planned path, computational tasks may be offloaded to base stations (BSs) when communication coverage is available. Because the robot operates on battery power, energy efficiency directly influences routing decisions and battery aging. The spatial distribution of BSs affects offloading opportunities, motion planning, resource allocation, and overall energy consumption, ultimately impacting battery degradation. Results from the case study demonstrate that the proposed integration of computational, mechanical, and communication energy models into routing optimization reduces total energy consumption by approximately 13% and battery degradation by 24% compared to approaches that optimize only mechanical and computational energy without considering load offloading.
Hadis Mohammadi Kamizji, Eero Immonen, Juha Plosila, Hashem Haghbayan
ECMS3
2026 Adaptive Simulation Platform In NVIDIA Isaac Sim For Self-Aware Heterogeneous Robots
abstract
This paper presents an adaptive simulation framework, implemented in NVIDIA Isaac Sim, for self-aware heterogeneous robotic systems. The proposed platform enables each robot to maintain an internal predictive model of its own embodiment and its interaction with the surrounding environment, while explicitly accounting for both mechanical actuation and onboard computational processes. By treating mechanical and computational energy as internal state variables, the framework enables runtime estimation of the energetic consequences of robot actions. This capability is demonstrated across diverse robot morphologies, including a mobile humanoid, a fixed-base manipulator, and an aerial robot operating in a shared warehouse environment. Experimental results show that internal computational processes constitute a substantial portion of the overall energy budget and are tightly coupled with mechanical action, highlighting the importance of incorporating internal resource state into self-regulatory decision-making. The framework provides a reusable platform for studying resource-aware self-regulation and energy-informed decision-making, with extensions toward multi-robot and swarm systems.
Afrooz Naseri, Juha Plosila, M. H. Haghbayan
ECMS2
2025 Simulation Of Mechanical And Computational Power Consumption In Mobile Robots
abstract
This paper presents a simulation model to estimate the instantaneous power consumption of a mobile robot by taking into account both its mechanical and computational components. The simulation model is adaptable to be tuned based on the level of accuracy needed for estimating the power consumption for the robot and the simulation time penalty. This makes a multi-fidelity power estimation tool for the robot with the capability to run-time changing the fidelity according to environmental conditions and internal computational capabilities. Such multi-fidelity energy prediction is suitable for run-time predictive decision making in a wide range of usages such as training process in model-based Reinforcement Learning (RL) as well as decision making in Model Predictive Control (MPC). The experimental results show that the simulation accurately estimates energy consumption at different fidelity levels. Higher-fidelity models closely match real-world measurements, while lower-fidelity models trade some accuracy for faster predictions. Therefore, higher estimation precision comes at the cost of increased computation.
Afrooz Naseri, Sajad Shahsavari, Juha Plosila, M. H. Haghbayan
ECMS3
2025 Runtime Energy-Efficient Control Policy for Mobile Robots with Computing Workload and Battery Awareness
abstract
Energy efficiency is a fundamental goal in robotic control. Various components within a robot, such as mechanical systems, computational units, and sensors, consume energy, all powered by the battery unit. Each component features several actuators and individual controllers that optimize energy usage locally, often without regard to one another. In this paper, we highlight a significant phenomenon indicating a considerable dependency between the mechanical and computational parts of the robot as energy consumers and the battery state of charge (SOC) as the energy provider. We demonstrate that as the battery SOC fluctuates, the behavior of energy consumption also varies, necessitating a unified controller with awareness of this relationship. Motivated by this observation, we propose a battery-aware co-optimization strategy for the mechanical and computational units, leveraging configuration space exploration to optimize the motor speed and the CPU frequency under different environmental conditions and battery SOC levels. Experimental results demonstrate the effectiveness of our approach in extending the operational lifetime of a robot under varying battery SOC and workload conditions, enhancing the energy efficiency of a case study rover by up to 53.93% w.r.t. selected baselines and similar past approaches.
M. H. Haghbayan, Abdul Malik, Antonio Miele, Juha Plosila
IROS5
2025 A Coordinated Approach to Control Mechanical and Computing Resources in Mobile Robots
abstract
Energy management of mechanical and cyber parts in mobile robots consists of two processes operating concurrently at runtime. Both the two processes can significantly improve the robots' battery lifetime and further extend mission time. In each process, information on energy consumption of one of the two parts is captured and analyzed to manipulate various mechanical/computational actuators in a robot, such as motor speed and CPU voltage/frequency. In this article, we show that considering management of mechanical and computational segments separately does not necessarily result in an energy-optimal solution due to their co-dependence; as a consequence, a runtime co-management scheme is required. We propose a proactive energy optimization methodology in which dynamically trained internal models are utilized to predict the future energy consumption for the mechanical and computational parts of a mobile robot, and based on that, the optimal mechanical speed and CPU voltage/frequency are determined at runtime. The experimental results on a ground wheeled robot show up to 36.34% reduction in the overall energy consumption compared to the state-of-the-art methods.
Sajad Shahsavari, M. H. Haghbayan, Antonio Miele, Eero Immonen, Juha Plosila
IEEE Trans. Robotics5
2023 A Coupled Battery State-of-Charge and Voltage Model for Optimal Control Applications
abstract
Optimal control of electric vehicle (EV) batteries for maximal energy efficiency, safety and lifespan requires that the Battery Management System (BMS) has accurate real-time information on both the battery State-of-Charge (SoC) and its dynamics, i.e. long-term and short-term energy supply capacity, at all times. However, these quantities cannot be measured directly from the battery, and, in practice, only SoC estimation is typically carried out. In this article, we propose a novel parametric algebraic voltage model coupled to the well-known Manwell-McGowan dynamic Kinetic Battery Model (KiBaM), which is able to predict both battery SoC dynamics and its electrical response. Numerical simulations, based on laboratory measurements, are presented for prismatic Lithium-Titanate Oxide (LTO) battery cells. Such cells are prime candidates for modern heavy offroad EV applications.
Masoomeh Karami, Sajad Shahsavari, Eero Immonen, M. H. Haghbayan, Juha Plosila
DATE5
2023 A Light-Weight Model For Run-Time Battery SOC-SOH Estimation While Considering Aging
abstract
Batteries are becoming one important part to power varieties of devices including electro-mechanical robots and vehicles. Understanding the behaviour of the battery and its state of charge can help the control systems to significantly improve the decision-making and risk management at run-time, after the device starts its operation. Currently, there is an increased interest in tracking battery dynamics as a function of health in both academia and industry. In this paper, we propose a light-weight approach for modeling the state of charge of lithium-ion (Li-ion) batteries during the life-time of the system. We also consider the battery capacity of charge degradation over its usage. To do that, we use electrical equivalent circuit model (EECM) modeling as the basis for modeling the battery and add the aging model to it to consider the effect of battery usage in the long term. Experimental results show that our proposed technique successfully estimates the battery state of charge at different states of health for the National Aeronautics and Space Administration (NASA) randomized usage battery dataset in comparison with the state-of-the-art. The obtained estimation error in the worst case is 2.2%.
Mohsen Heydarzadeh, Eero Immonen, M. H. Haghbayan, Juha Plosila
ECMS4
2023 Run-Time Resource Management in CMPs Handling Multiple Aging Mechanisms
abstract
Run-time resource management is fundamental for efficient execution of workloads on Chip Multiprocessors. Application- and system-level requirements (e.g., on performance versus power versus lifetime reliability) are generally conflicting each other, and any decision on resource assignment, such as core allocation or frequency tuning, may positively affect some of them while penalizing some others. Resource assignment decisions can be perceived in few instants of time on performance and power consumption, but not on lifetime reliability. In fact, this latter changes very slowly based on the accumulation of effects of various decisions over a long time horizon. Moreover, aging mechanisms are various and have different causes; most of them, such as Electromigration (EM), are subject to temperature levels, while Thermal Cycling (TC) is caused mainly by temperature variations (both amplitude and frequency). Mitigating only EM may negatively affect TC and vice versa. We propose a resource orchestration strategy to balance the performance and power consumption constraints in the short-term and EM and TC aging in the long-term. Experimental results show that the proposed approach improves the average Mean Time To Failure at least by 17% and 20% w.r.t. EM and TC, respectively, while providing same performance level of the nominal counterpart and guaranteeing the power budget.
M. H. Haghbayan, Antonio Miele, Onur Mutlu, Juha Plosila
IEEE Trans. Computers4
2022 How To Run A World Record? A Reinforcement Learning Approach
abstract
Finding the optimal distribution of exerted effort by an athlete in competitive sports has been widely investigated in the fields of sport science, applied mathematics and optimal control. In this article, we propose a reinforcement learning-based solution to the optimal control problem in the running race application. Well-known mathematical model of Keller is used for numerically simulating the dynamics in runner's energy storage and motion. A feed-forward neural network is employed as the probabilistic controller model in continuous action space which transforms the current state (position, velocity and available energy) of the runner to the predicted optimal propulsive force that the runner should apply in the next time step. A logarithmic barrier reward function is designed to evaluate performance of simulated races as a continuous smooth function of runner's position and time. The neural network parameters, then, are identified by maximizing the expected reward using on-policy actor-critic policy-gradient RL algorithm. We trained the controller model for three race lengths: 400, 1500 and 10000 meters and found the force and velocity profiles that produce a near-optimal solution for the runner's problem. Results conform with Keller's theoretical findings with relative percent error of 0.59% and are comparable to real world records with relative percent error of 2.38%, while the same error for Keller's findings is 2.82%.
Sajad Shahsavari, Eero Immonen, Masoomeh Karami, M. H. Haghbayan, Juha Plosila
ECMS5
2021 MCX ? An Open-Source Framework For Digital Twins
abstract
This article describes ModelConductor-eXtended (MCX), which is an open-source software architecture for digital twins. The MCX framework facilitates co-execution of, and asynchronous data communication between, physical systems and their digital simulation models. MCX supports running FMUs (simulation models packaged according to the FMI specification) as well as machine learning models and customized models. We propose extensions to the previously published ModelConductor framework for higher performance and better scalability. The extensions include decoupling of the queue and the model computation module, utilization of a standard data transmission protocol and implementation of the facility to run time-consuming simulation models in a time synchronous manner. Additionally, three new validation case studies are presented. A performance evaluation shows that the extensions improve the average response time almost 4 times in three specific experiments.
Sajad Shahsavari, Eero Immonen, Mohammed Rabah, M. H. Haghbayan, Juha Plosila
ECMS5
2021 Hierarchical Fault Simulation of Deep Neural Networks on Multi-Core Systems
abstract
In this paper, a hierarchical fault simulation technique for neural networks is proposed, supporting both permanent and temporary faults. In the proposed technique, different levels of hierarchy are used, forming a mixed-level simulation environment. In such an environment, the pre-synthesis behavioral specification of the network and the post-synthesis gate-level model are co-simulated. To accelerate the fault simulation process, faults are injected in the gate-level specification of the selected neurons while the behavioral model in different levels of abstraction is used to simulate the remaining neurons. Further speedup is obtained through event-driven simulation and parallelization. Experimental results confirm the time efficiency of the proposed fault simulation technique.
Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Antonio Miele, Hannu Tenhunen, Juha Plosila
ETS6
2021 Energy-Efficient Mobile Robot Control via Run-time Monitoring of Environmental Complexity and Computing Workload
abstract
We propose an energy-efficient controller to minimize the energy consumption of a mobile robot by dynamically manipulating the mechanical and computational actuators of the robot. The mobile robot performs real-time vision-based applications based on an event-based camera. The actuators of the controller are CPU voltage/frequency for the computation part and motor voltage for the mechanical part. We show that independently considering speed control of the robot and voltage/frequency control of the CPU does not necessarily result in an energy-efficient solution. In fact, to obtain the highest efficiency, the computation and mechanical parts should be controlled together in synergy. We propose a fast hill-climbing optimization algorithm to allow the controller to find the best CPU/motor configuration at run-time and whenever the mobile robot is facing a new environment during its travel. Experimental results on a robot with Brushless DC Motors, Jetson TX2 board as the computing unit, and a DAVIS-346 event-based camera show that the proposed control algorithm can save battery energy by an average of 50.5%, 41%, and 30%, in low-complexity, medium-complexity, and high-complexity environments, over baselines.
Sherif Abdelmonem Sayed Mohamed, M. H. Haghbayan, Antonio Miele, Onur Mutlu, Juha Plosila
IROS5
2021 High-Performance Parallel Fault Simulation for Multi-Core Systems
abstract
Fault simulation is a time-consuming process that requires customized methods and techniques to accelerate it. Multi-threading and Multi-core approaches are two promising techniques that can be exploited to accelerate the fault simulation process by using different parts of the hardware at the same time. However, an efficient parallelization is obtained only by the refinement of software with respect to the hardware platform. In this paper, a parallel multi-thread fault simulation technique is proposed to accelerate the simulation process on multi-core platforms. In this approach, the gate input values are independently assigned to each thread. Each input value carries the information of several parallel simulation processes. This provides a multithread parallel fault simulation environment. The experimental results show that the proposed technique can efficiently use the hardware platform. In a single-core platform, the proposed technique can reduce the time by 25% while in a dual-core increasing the thread approximately halves the execution time.
Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Hamid Nejatollahi, Hannu Tenhunen, Juha Plosila
PDP6
2020 Thermal-Cycling-aware Dynamic Reliability Management in Many-Core System-on-Chip
abstract
Dynamic Reliability Management (DRM) is a common approach to mitigate aging and wear-out effects in multi- /many-core systems. State-of-the-art DRM approaches apply finegrained control on resource management to increase/balance the chip reliability while considering other system constraints, e.g., performance, and power budget. Such approaches, acting on various knobs such as workload mapping and scheduling, Dynamic Voltage/Frequency Scaling (DVFS) and Per-Core Power Gating (PCPG), demonstrated to work properly with the various aging mechanisms, such as electromigration, and Negative-Bias Temperature Instability (NBTI). However, we claim that they do not suffice for thermal cycling. Thus, we here propose a novel thermal-cycling-aware DRM approach for shared-memory many-core systems running multi-threaded applications. The approach applies a fine-grained control capable at reducing both temperature levels and variations. The experimental evaluations demonstrated that the proposed approach is able to achieve 39% longer lifetime than past approaches.
M. H. Haghbayan, Antonio Miele, Zhuo Zou, Hannu Tenhunen, Juha Plosila
DATE5
2020 Navigation System For Landing A Swarm Of Autonomous Drones On A Movable Surface
Anam Tahir, Jari M. Böling, M. H. Haghbayan, Juha Plosila
ECMS4
2020 Dynamic Resource-Aware Corner Detection for Bio-Inspired Vision Sensors
abstract
Event-based cameras are vision devices that transmit only brightness changes with low latency and ultra-low power consumption. Such characteristics make event-based cameras attractive in the field of localization and object tracking in resource-constrained systems. Since the number of generated events in such cameras is huge, the selection and filtering of the incoming events are beneficial from both increasing the accuracy of the features and reducing the computational load. In this paper, we present an algorithm to detect asynchronous corners form a stream of events in real-time on embedded systems. The algorithm is called the Three Layer Filtering-Harris or TLF-Harris algorithm. The algorithm is based on an events' filtering strategy whose purpose is 1) to increase the accuracy by deliberately eliminating some incoming events, i.e., noise and 2) to improve the real-time performance of the system, i.e., preserving a constant throughput in terms of input events per second, by discarding unnecessary events with a limited accuracy loss. An approximation of the Harris algorithm, in turn, is used to exploit its high-quality detection capability with a low-complexity implementation to enable seamless real-time performance on embedded computing platforms. The proposed algorithm is capable of selecting the best corner candidate among neighbors and achieves an average execution time savings of 59% compared with the conventional Harris score. Moreover, our approach outperforms the competing methods, such as eFAST, eHarris, and FA-Harris, in terms of real-time performance, and surpasses Arc* in terms of accuracy.
Sherif Abdelmonem Sayed Mohamed, Jawad Naveed Yasin, M. H. Haghbayan, Antonio Miele, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila
ICPR7
2019 Monocular visual odometry based on hybrid parameterization
abstract
Visual odometry (VO) is one of the most challenging techniques in computer vision for autonomous vehicle/vessels. In VO, the camera pose that also represents the robot pose in ego-motion is estimated analyzing the features and pixels extracted from the camera images. Different VO techniques mainly provide different trade-offs among the resources that are being considered for odometry, such as camera resolution, computation/communication capacity, power/energy consumption, and accuracy. In this paper, a hybrid technique is proposed for camera pose estimation by combining odometry based on triangulation using the long-term period of direct-based odometry and the short-term period of inverse depth mapping. Experimental results based on the EuRoC data set shows that the proposed technique significantly outperforms the traditional direct-based pose estimation method for Micro Aerial Vehicle (MAV), keeping its potential negative effect on performance negligible.
Sherif Abdelmonem Sayed Mohamed, M. H. Haghbayan, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila
ICMV5
2019 Energy-Aware VM Consolidation in Cloud Data Centers Using Utilization Prediction Model
abstract
Virtual Machine (VM) consolidation provides a promising approach to save energy and improve resource utilization in data centers. Many heuristic algorithms have been proposed to tackle the VM consolidation as a vector bin-packing problem. However, the existing algorithms have focused mostly on the number of active Physical Machines (PMs) minimization according to their current resource requirements and neglected the future resource demands. Therefore, they generate unnecessary VM migrations and increase the rate of Service Level Agreement (SLA) violations in data centers. To address this problem, we propose a VM consolidation approach that takes into account both the current and future utilization of resources. Our approach uses a regression-based model to approximate the future CPU and memory utilization of VMs and PMs. We investigate the effectiveness of virtual and physical resource utilization prediction in VM consolidation performance using Google cluster and PlanetLab real workload traces. The experimental results show, our approach provides substantial improvement over other heuristic and meta-heuristic algorithms in reducing the energy consumption, the number of VM migrations and the number of SLA violations.
Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Nguyen Trung Hieu, Hannu Tenhunen
IEEE Trans. Cloud Comput.4
2018 Parallel imperialist competitive algorithms
abstract
Summary The importance of optimization and NP‐problem solving cannot be overemphasized. The usefulness and popularity of evolutionary computing methods are also well established. There are various types of evolutionary methods; they are mostly sequential but some of them have parallel implementations as well. We propose a multi‐population method to parallelize the Imperialist Competitive Algorithm. The algorithm has been implemented with the Message Passing Interface on 2 computer platforms, and we have tested our method based on shared memory and message passing architectural models. An outstanding performance is obtained, demonstrating that the proposed method is very efficient concerning both speed and accuracy. In addition, compared with a set of existing well‐known parallel algorithms, our approach obtains more accurate results within a shorter time period.
Amin Majd, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila, Shahriar Lotfi, Hannu Tenhunen
Concurr. Comput. Pract. Exp.4
2017 A reliable weighted feature selection for auto medical diagnosis
abstract
Feature selection is a key step in data analysis. However, most of the existing feature selection techniques are serial and inefficient to be applied to massive data sets. We propose a feature selection method based on a multi-population weighted intelligent genetic algorithm to enhance the reliability of diagnoses in e-Health applications. The proposed approach, called PIGAS, utilizes a weighted intelligent genetic algorithm to select a proper subset of features that leads to a high classification accuracy. In addition, PIGAS takes advantage of multi-population implementation to further enhance accuracy. To evaluate the subsets of the selected features, the KNN classifier is utilized and assessed on UCI Arrhythmia dataset. To guarantee valid results, leave-one-out validation technique is employed. The experimental results show that the proposed approach outperforms other methods in terms of accuracy and efficiency. The results of the 16-class classification problem indicate an increase in the overall accuracy when using the optimal feature subset. Accuracy achieved being 99.70% indicating the potential of the algorithm to be utilized in a practical auto-diagnosis system. This accuracy was obtained using only half of features, as against an accuracy of66.76% using all the features.
Golnaz Sahebi, Amin Majd, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen
INDIN4
2017 Hierarchal Placement of Smart Mobile Access Points in Wireless Sensor Networks Using Fog Computing
abstract
Recent advances in computing and sensor technologies have facilitated the emergence of increasingly sophisticated and complex cyber-physical systems and wireless sensor networks. Moreover, integration of cyber-physical systems and wireless sensor networks with other contemporary technologies, such as unmanned aerial vehicles (i.e. drones) and fog computing, enables the creation of completely new smart solutions. By building upon the concept of a Smart Mobile Access Point (SMAP), which is a key element for a smart network, we propose a novel hierarchical placement strategy for SMAPs to improve scalability of SMAP based monitoring systems. SMAPs predict communication behavior based on information collected from the network, and select the best approach to support the network at any given time. In order to improve the network performance, they can autonomously change their positions. Therefore, placement of SMAPs has an important role in such systems. Initial placement of SMAPs is an NP problem. We solve it using a parallel implementation of the genetic algorithm with an efficient evaluation phase. The adopted hierarchical placement approach is scalable, it enables construction of arbitrarily large SMAP based systems.
Amin Majd, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen
PDP4
2016 An Approach for Smart Management of Big Data in the Fog Computing Context
abstract
In this paper, a new approach for tackling Big Data in Internet of Things (IoT) systems is presented. We approach the problem from the data perspective rather than only focusing on the computing platform. We design and develop a concept that we call "Smart Data". Taking advantage of a hierarchical fog computing system, we reshape the raw and passive form of the data generated by IoT sensors to intelligent and self-managed data cells that are able to evolve and become more meaningful information with reduced size. We believe that smart data will revolutionize the current perspective of data and will open many potential research directions to tackle emerging big data issues.
Farhoud Hosseinpour, Juha Plosila, Hannu Tenhunen
CloudCom2
2016 PICA: Multi-population Implementation of Parallel Imperialist Competitive Algorithms
abstract
The importance of optimization and NP-problems solving cannot be over emphasized. The usefulness and popularity of evolutionary computing methods are also well established. There are various types of evolutionary methods that are mostly sequential, and some others have parallel implementation. We propose a method to parallelize Imperialist Competitive Algorithm (Multi-Population). The algorithm has been implemented with MPI on two platforms and have tested our algorithms on a shared-memory and message passing architecture. An outstanding performance is obtained, which indicates that the method is efficient concern to speed and accuracy. In the second step, the proposed algorithm is compared with a set of existing well known parallel algorithms and is indicated that it obtains more accurate solutions in a lower time.
Amin Majd, Shahriar Lotfi, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila
PDP5
2016 A Power-Aware Approach for Online Test Scheduling in Many-Core Architectures
abstract
Aggressive technology scaling triggers novel challenges to the design of multi-/many-core systems, such as limited power budget and increased reliability issues. Today's many-core systems employ dynamic power management and runtime mapping strategies trying to offer optimal performance while fulfilling power constraints. On the other hand, due to the reliability challenges, online testing techniques are becoming a necessity in current and near future technologies. However, state-of-the-art techniques are not aware of the other power/performance requirements. This paper proposes a power-aware non-intrusive online testing approach for many-core systems. The approach schedules software based self-test routines on the various cores during their idle periods, while honoring the power budget and limiting delays in the workload execution. A test criticality metric, based on a device aging model, is used to select cores to be tested at a time. Moreover, power and reliability issues related to the testing at different voltage and frequency levels are also handled. Extensive experimental results reveal that the proposed approach can i) efficiently test the cores within the available power budget causing a negligible performance penalty, ii) adapt the test frequency to the current cores' aging status, and iii) cover available voltage and frequency levels during the testing.
M. H. Haghbayan, Amir-Mohammad Rahmani, Antonio Miele, Mohammad Fattah, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen
IEEE Trans. Computers5
2016 Polymorphic Configuration Architecture for CGRAs
abstract
In the era of platforms hosting multiple applications with arbitrary reconfiguration requirements, static configuration architectures are neither optimal nor desirable. The static reconfiguration architectures either incur excessive overheads or cannot support advanced features (like time-sharing and runtime parallelism). As a solution to this problem, we present a polymorphic configuration architecture (PCA) that provides each application with a configuration infrastructure tailored to its needs.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Ahmed Hemani, Kolin Paul, Juha Plosila, Peeter Ellervee, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Utilization Prediction Aware VM Consolidation Approach for Green Cloud Computing
abstract
Dynamic Virtual Machine (VM) consolidation is one of the most promising solutions to reduce energy consumption and improve resource utilization in data centers. Since VM consolidation problem is strictly NP-hard, many heuristic algorithms have been proposed to tackle the problem. However, most of the existing works deal only with minimizing the number of hosts based on their current resource utilization and these works do not explore the future resource requirements. Therefore, unnecessary VM migrations are generated and the rate of Service Level Agreement (SLA) violations are increased in data centers. To address this problem, our VM consolidation method which is formulated as a bin-packing problem considers both the current and future utilization of resources. The future utilization of resources is accurately predicted using a k-nearest neighbor regression based model. In this paper, we investigate the effectiveness of VM and host resource utilization predictions in the VM consolidation task using real workload traces. The experimental results show that our approach provides substantial improvement over other heuristic algorithms in reducing energy consumption, number of VM migrations and number of SLA violations.
Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
CLOUD4
2015 Power-aware online testing of manycore systems in the dark silicon era
M. H. Haghbayan, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Zainalabedin Navabi, Hannu Tenhunen
DATE5
2015 Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approach
abstract
Power management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies.
Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ISLPED6
2015 A Low-Overhead, Fully-Distributed, Guaranteed-Delivery Routing Algorithm for Faulty Network-on-Chips
abstract
This paper introduces a new, practical routing algorithm, Maze-routing, to tolerate faults in network-on-chips. The algorithm is the first to provide all of the following properties at the same time: 1) fully-distributed with no centralized component, 2) guaranteed delivery (it guarantees to deliver packets when a path exists between nodes, or otherwise indicate that destination is unreachable, while being deadlock and livelock free), 3) low area cost, 4) low reconfiguration overhead upon a fault. To achieve all these properties, we propose Maze-routing, a new variant of face routing in on-chip networks and make use of deflections in routing. Our evaluations show that Maze-routing has 16X less area overhead than other algorithms that provide guaranteed delivery. Our Maze-routing algorithm is also high performance: for example, when up to 5 links are broken, it provides 50% higher saturation throughput compared to the state-of-the-art.
Mohammad Fattah, Antti Airola, Rachata Ausavarungnirun, Nima Mirzaei, Pasi Liljeberg, Juha Plosila, Siamak Mohammadi, Tapio Pahikkala, Onur Mutlu, Hannu Tenhunen
NOCS6
2015 FIST: A Framework to Interleave Spiking Neural Networks on CGRAs
abstract
Coarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern embedded applications. In many application domains (e.g. robotics and cognitive embedded systems), the CGRAs are required to simultaneously host processing (e.g. Audio/video acquisition) and estimation (e.g. audio/video/image recognition) tasks. Recent works have revealed that the efficiency and scalability of the estimation algorithms can be significantly improved by using neural networks. However, existing CGRAs commonly employ homogeneous processing resources for both the tasks. To realize the best of both the worlds (conventional processing and neural networks), we present FIST. FIST allows the processing elements and the network to dynamically morph into either conventional CGRA or a neural network, depending on the hosted application. We have chosen the DRRA as a vehicle to study the feasibility and overheads of our approach. Synthesis results reveal that the proposed enhancements incur negligible overheads (4.4% area and 9.1% power) compared to the original DRRA cell.
Tuan Ngyen, Syed M. A. H. Jafri, Masoud Daneshtalab, Ahmed Hemani, Sergei Dytckov, Juha Plosila, Hannu Tenhunen
PDP6
2015 Parallel Implementation of Fuzzified Pattern Matching Algorithm on GPU
abstract
Approximate pattern discovery is one of the fundamental and challenging problems in computer science. Fast and high performance algorithms are highly demanded in many applications in bioinformatics and computational molecular biology, which are the domains that are mostly and directly benefit from any enhancement of pattern matching theoretical knowledge and solutions. This paper proposed an efficient GPU implementation of fuzzified Aho-Corasick algorithm using Levenshtein method and N-gram technique as a solution for approximate pattern matching problem.
Shima Soroushnia, Masoud Daneshtalab, Tapio Pahikkala, Juha Plosila
PDP4
2015 PDNOC: Partially diagonal network-on-chip for high efficiency multicore systems
abstract
Summary With the constantly increasing of number of cores in multicore processors, more emphasis should be paid to the on‐chip interconnect. Performance and power consumption of an on‐chip interconnect are directly affected by the network topology. Researchers have proposed various topologies to optimize these metrics. The efficiency can also be optimized by proper mapping of applications. Therefore in this paper, we propose a novel partially diagonal network‐on‐chip (PDNOC) design that takes advantage of both heterogeneous network topology and congestion‐aware application mapping. We analyse the partially diagonal network in terms of interconnect structure, area usage, power consumption, routing algorithm and implementation complexity. The key insight that enables the PDNOC is that most communication patterns in real‐world applications are hot‐spot and bursty. We implement a full system simulation environment using SPLASH‐2 benchmarks. Performance metrics of standard mesh, concentrated mesh, full diagonal mesh and four types of the proposed PDNOC are measured in terms of network latency, application execution time and energy delay product. Evaluation results show that on average, the proposed PDNOC designs provide up to 36% improvement in execution time over concentrated mesh, and 3.6× better energy delay product over fully connected diagonal network. PDNOC design with two adjacent PD networks is a better candidate for higher efficiency, while four PD networks provide better performance. Copyright © 2014 John Wiley & Sons, Ltd.
Thomas Canhao Xu, Ville Leppänen, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
Concurr. Comput. Pract. Exp.4
2015 Architecture and Implementation of Dynamic Parallelism, Voltage and Frequency Scaling (PVFS) on CGRAs
abstract
In the era of platforms hosting multiple applications with arbitrary performance requirements, providing a worst-case platform-wide voltage/frequency operating point is neither optimal nor desirable. As a solution to this problem, designs commonly employ dynamic voltage and frequency scaling (DVFS). DVFS promises significant energy and power reductions by providing each application with the operating point (and hence the performance) tailored to its needs. To further enhance the optimization potential, recent works interleave dynamic parallelism with conventional DVFS. The induced parallelism results in performance gains that allow an application to lower its operating point even further (thereby saving energy and power consumption). However, the existing works employ costly dedicated hardware (for synchronization) and rely solely on greedy algorithms to make parallelism decisions. To efficiently integrate parallelism with DVFS, compared to state-of-the-art, we exploit the reconfiguration (to reduce DVFS synchronization overheads) and enhance the intelligence of the greedy algorithm (to make optimal parallelism decisions). Specifically, our solution relies on dynamically reconfigurable isolation cells and an autonomous parallelism, voltage, and frequency selection algorithm. The dynamically reconfigurable isolation cells reduce the area overheads of DVFS circuitry by configuring the existing resources to provide synchronization. The autonomous parallelism, voltage, and frequency selection algorithm ensures high power efficiency by combining parallelism with DVFS. It selects that parallelism, voltage, and frequency trio which consumes minimum power to meet the deadlines on available resources. Synthesis and simulation results using various applications/algorithms (WLAN, MPEG4, FFT, FIR, matrix multiplication) show that our solution promises significant reduction in area and power consumption (23% and 51% ) compared to state-of-the-art.
Syed M. A. H. Jafri, Ozan Ozbag, Nasim Farahini, Kolin Paul, Ahmed Hemani, Juha Plosila, Hannu Tenhunen
ACM J. Emerg. Technol. Comput. Syst.6
2015 In-order delivery approach for 2D and 3D NoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Sergei Dytckov, Juha Plosila
J. Supercomput.4
2015 Using Ant Colony System to Consolidate VMs for Green Cloud Computing
abstract
High energy consumption of cloud data centers is a matter of great concern. Dynamic consolidation of Virtual Machines (VMs) presents a significant opportunity to save energy in data centers. A VM consolidation approach uses live migration of VMs so that some of the under-loaded Physical Machines (PMs) can be switched-off or put into a low-power mode. On the other hand, achieving the desired level of Quality of Service (QoS) between cloud providers and their users is critical. Therefore, the main challenge is to reduce energy consumption of data centers while satisfying QoS requirements. In this paper, we present a distributed system architecture to perform dynamic VM consolidation to reduce energy consumption of cloud data centers while maintaining the desired QoS. Since the VM consolidation problem is strictly NP-hard, we use an online optimization metaheuristic algorithm called Ant Colony System (ACS). The proposed ACS-based VM Consolidation (ACS-VMC) approach finds a near-optimal solution based on a specified objective function. Experimental results on real workload traces show that ACS-VMC reduces energy consumption while maintaining the required performance levels in a cloud data center. It outperforms existing VM consolidation approaches in terms of energy consumption, number of VM migrations, and QoS requirements concerning performance.
Fahimeh Farahnakian, Adnan Ashraf, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Ivan Porres, Hannu Tenhunen
IEEE Trans. Serv. Comput.5
2014 Energy-Aware Dynamic VM Consolidation in Cloud Data Centers Using Ant Colony System
abstract
As the scale of a cloud data center becomes larger and larger, the energy consumption of the data center also grows rapidly. Dynamic consolidation of Virtual Machines (VMs) presents a significant opportunity to save energy by turning off unused Physical Machines (PMs) in data centers. In this paper, we present a distributed controller to perform dynamic VM consolidation to improve the resource utilizations of PMs and to reduce their energy consumption. Moreover, we use the ant colony system to find a near-optimal VM placement solution based on the specified objective function. Experimental results on the real workload traces from more than a thousand PlanetLab VMs show that the proposed approach reduces energy consumption and maintains required performance levels in a large-scale data center.
Fahimeh Farahnakian, Adnan Ashraf, Pasi Liljeberg, Tapio Pahikkala, Juha Plosila, Ivan Porres, Hannu Tenhunen
IEEE CLOUD5
2014 Hierarchical Agent-Based Architecture for Resource Management in Cloud Data Centers
abstract
In order to resource management in a large-scale data center, we present a hierarchical agent-based architecture. In this architecture, multi agents cooperate together to minimize the number of active physical machines according to the current resource requirements. We proposed a local agent in each physical machine (PM) to determine the PM's status and a global agent to optimizes VM placement based on PM's status. Experimental results show the proposed architecture can minimize energy consumption while maintaining an acceptable QoS.
Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila
IEEE CLOUD4
2014 Adjustable contiguity of run-time task allocation in networked many-core systems
abstract
In this paper, we propose a run-time mapping algorithm, CASqA, for networked many-core systems. In this algorithm, the level of contiguousness of the allocated processors (α) can be adjusted in a fine-grained fashion. A strictly contiguous allocation (α = 0) decreases the latency and power dissipation of the network and improves the applications execution time. However, it limits the achievable throughput and increases the turnaround time of the applications. As a result, recent works consider non-contiguous allocation (α = 1) to improve the throughput traded off against applications execution time and network metrics. In contradiction, our experiments show that a higher throughput (by 3%) with improved network performance can be achieved when using intermediate α values. More precisely, up to 35% drop in the network costs can be gained by adjusting the level of contiguity compared to non-contiguous cases, while the achieved throughput is kept constant. Moreover, CASqA provides at least 32% energy saving in the network compared to other works.
Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
ASP-DAC3
2014 Hierarchical VM Management Architecture for Cloud Data Centers
abstract
Efficient energy use has become a critical issue for designing and managing of cloud data centers. Virtualization is a key technology for reducing energy cost and improving resource utilization in data centers. One of the challenges faced by virtualized data centers is to decide how to pack VMs on the least number of physical machines. This paper presents a VM management framework which is based on a multi-agent system to minimize energy consumption and Service Level Agreement (SLA) violations. The proposed agents are arranged in a three level hierarchical structure to perform VM assignment, VM placement and VM consolidation in a data center efficiently. Experimental results demonstrate that the framework achieves high quality solution in spite of its simplicity and scalability.
Fahimeh Farahnakian, Pasi Liljeberg, Tapio Pahikkala, Juha Plosila, Hannu Tenhunen
CloudCom4
2014 SHiFA: System-Level Hierarchy in Run-Time Fault-Aware Management of Many-Core Systems
abstract
A system-level approach to fault-aware resource management of many-core systems is proposed. The proposed approach, called SHiFA, is able to tolerate run-time faults at system level without any hardware overhead. In contrast to the existing system-level methods, network resources are also considered to be potentially faulty. Accordingly, applications are mapped onto healthy nodes of the system at run-time such that their interaction will not require the use of faulty elements. By utilizing the simple routing approach, results show 100% utilizability of PEs and 99.41% of successful mapping when up to 8 links are broken. SHiFA design is based on distributed operating systems, such that it is kept scalable for future many-core systems. A significant improvement in scalability properties is observed compared to the state-of-the-art distributed approaches.
Mohammad Fattah, Maurizio Palesi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DAC4
2014 Online testing of many-core systems in the Dark Silicon era
abstract
As the dark silicon era is about to embrace, it is not anymore possible to attain commensurate performance benefits by increasing the number of transistors due to thermal design power. Dark Silicon issue stresses that a fraction of silicon chip being able to switch in full frequency is dropping and designers will soon face the growing underutilization inherent in future technologies. On the other hand, by reducing the transistor size, susceptibility to internal defects drastically increases and large ranges of defects such as aging or transient faults will be shown up more frequently. In this paper, we propose an online test scheduling algorithm using software based self-test for dark silicon era to test dark cores while considering thermal design power of the system. As the dark area of the system is dynamic and reshapes at a runtime, the tested cores can be used by other applications in the near future. Empirical results show the effectiveness of the proposed algorithm in terms of applicability and fault coverage with a negligible negative impact on the system throughput.
M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DDECS4
2014 Parameterized AES-Based Crypto Processor for FPGAs
abstract
In this paper, we propose a parameterized crypto co-processor based on Advanced Encryption Standard (AES). This parameterized AES module is combined with a 32-bit general purpose 5-stage pipelined MIPS processor. The AES module used in this paper is fully pipelined. The processor fetches an instruction from the instruction memory and sends it to the decode stage. If the instruction is the crypto instruction it is pushed into the AES module during the decode stage. However if the instruction belongs to the MIPS processor, the remaining cycles will be completed on the MIPS processor. The parameterized AES module has different latencies on different rounds of AES according to the application requirements. The effects of different number of rounds on latency, memory, and area are studied and reported.
Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen, Sergei Dytckov, Giovanni Beltrame
DSD4
2014 Efficient STDP Micro-Architecture for Silicon Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are the closest approach to biological neurons in comparison with conventional artificial neural networks (ANN). SNNs are composed of neurons and synapses which are interconnected with a complex pattern. As communication in such massively parallel computational systems is getting critical, the network-on-chip (NoC) becomes a promising solution to provide a scalable and robust interconnection fabric. However, using NoC for large-scale SNNs arises a trade-off between scalability, throughput, neuron/router ratio (cluster size), and area overhead. In this paper, we tackle the trade-off using a clustering approach and try to optimize the synaptic resource utilization. An optimal cluster size can provide the lowest area overhead and power consumption. For the learning purposes, a phenomenon known as spike-timing-dependent plasticity (STDP) is utilized. The micro-architectures of the network, clusters, and the computational neurons are also described. The presented approach suggests a promising solution of integrating NoCs and STDP-based SNNs for the optimal performance based on the underlying application.
Sergei Dytckov, Masoud Daneshtalab, Masoumeh Ebrahimi, Hassan Anwar, Juha Plosila, Hannu Tenhunen
DSD5
2014 Morphable Compression Architecture for Efficient Configuration in CGRAs
abstract
Today, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads (up to 50% area of the overall platform). As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties (i.e. compression ratio and decoding time), and is therefore best suited for a particular class of applications (and situation). However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. The proposed architecture allows each application to enjoy a separate compression/decompression hierarchy (consisting of various types and implementations of hardware/software decoders) tailored to its needs. Thereby, our solution offers minimal memory while meeting the required configuration deadlines. Simulation results, using different applications (FFT, Matrix multiplication, and WLAN), reveal that the choice of compression hierarchy has a significant impact on compression ratio (from configware replication to 52%) and configuration cycles (from 33 nsec to 1.5 secs) for the tested applications. Synthesis results reveal that introducing adaptivity incurs negligible additional overheads (1%) compared to the overall platform area.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen
DSD7
2014 Customizable Compression Architecture for Efficient Configuration in CGRAs
abstract
Today, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads. As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties, and is therefore best suited for a particular class of applications. However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen
FCCM7
2014 TransPar: Transformation based dynamic Parallelism for low power CGRAs
abstract
Coarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern applications (e.g. 4G, CDMA, etc.). Recently proposed CGRAs offer runtime parallelism to reduce energy consumption (by lowering voltage/frequency). To implement the runtime parallelism, CGRAs commonly store multiple compile-time generated implementations of an application (with different degree of parallelism) and select the optimal version at runtime. However, the compile-time binding incurs excessive configuration memory overheads and/or is unable to parallelize an application even when sufficient resources are available. As a solution to this problem, we propose Transformation based dynamic Parallelism (TransPar). TransPar stores only a single implementation and applies a series for transformations to generate the bitstream for the parallel version. In addition, it also allows to displace and/or rotate an application to parallelize in resource constrained scenarios. By storing only a single implementation, TransPar offers significant reductions in configuration memory requirements (up to 73% for the tested applications), compared to state of the art compaction techniques. Simulation and synthesis results, using real applications, reveal that the additional flexibility allows up to 33% energy reduction compared to static memory based parallelism techniques. Gate level analysis reveals that TransPar incurs negligible silicon (0.2% of the platform) and timing (6 additional cycles per application) penalty.
Syed M. A. H. Jafri, Guilermo Serrano, Masoud Daneshtalab, Naeem Abbas, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen
FPL7
2014 Dark silicon aware power management for manycore systems under dynamic workloads
abstract
Dark Silicon denotes the phenomenon that, due to thermal and power constraints, the fraction of transistors that can operate at full frequency is decreasing with each technology generation. We propose a PID (Proportional Integral Derivative) controller based dynamic power management method that considers an upper bound on power consumption (called the Thermal Design Power (TDP)). To avoid violation of the TDP constraint for manycore systems running highly dynamic workloads, it provides fine-grained DVFS (Dynamic Voltage and Frequency Scaling) including near-threshold operation. In addition, the method distinguishes applications with hard Real-Time, soft Real-Time and no Real-Time constraints and treats them with appropriate priorities. In simulations with dynamic workloads mixed-critical application profiles, we show that the method is effective in honoring the TDP bound and it can boost system throughput by over 43% compared to a naive TDP scheduling policy.
M. H. Haghbayan, Amir-Mohammad Rahmani, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ICCD5
2014 Integration of AES on Heterogeneous Many-Core System
abstract
Increasing in the transistor density in a single chip makes it possible for many-core systems to utilize design space for implementing complex embedded systems. In this paper, we propose an architecture for heterogeneous many-core system to integrate block cipher algorithm which is based on Advanced Encryption Standard (AES). In order to implement AES as a crypto-core along with heterogeneous many-core system two different approaches are proposed. In the first approach, the platform is composed of an AES module, a crypto-core, a network interface, and an internal memory which are managed through a controller. In the second approach, the platform is composed of a Direct Memory Access (DMA), network interface, an internal memory, and a microprocessor in which the AES module is integrated as a crypto-core. Both approaches have been analyzed and compared in terms of area overhead and performance.
Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Marco Ramírez 0001, Juha Plosila, Hannu Tenhunen
PDP5
2014 Energy-Efficient Virtual Machines Consolidation in Cloud Data Centers Using Reinforcement Learning
abstract
Dynamic consolidation techniques optimize resource utilization and reduce energy consumption in Cloud data centers. They should consider the variability of the workload to decide when idle or underutilized hosts switch to sleep mode in order to minimize energy consumption. In this paper, we propose a Reinforcement Learning-based Dynamic Consolidation method (RL-DC) to minimize the number of active hosts according to the current resources requirement. The RL-DC utilizes an agent to learn the optimal policy for determining the host power mode by using a popular reinforcement learning method. The agent learns from past knowledge to decide when a host should be switched to the sleep or active mode and improves itself as the workload changes. Therefore, RL-DC does not require any prior information about workload and it dynamically adapts to the environment to achieve online energy and performance management. Experimental results on the real workload traces from more than a thousand PlanetLab virtual machines show that RL-DC minimizes energy consumption and maintains required performance levels.
Fahimeh Farahnakian, Pasi Liljeberg, Juha Plosila
PDP3
2014 Mixed-Criticality Run-Time Task Mapping for NoC-Based Many-Core Systems
abstract
Contiguous processor allocation improves both the network and the application performance, by decreasing the congestion probability among communication of different applications. Consequently, the average, standard deviation and worst-case latency of the network is decreased significantly. This makes the contiguous allocation a good solution for time-critical applications with bounded deadlines. On the other hand, non-contiguous allocation will increase the system throughput significantly. Isolated nodes are utilized and more applications can finish their job in a time unit. However, this will lead to poor network metrics, unsuitable for real-time applications. In this work, we combine these two approaches in order to manage workloads with mixed-critical characteristics. Real-time applications are mapped contiguously, while non-critical applications are allowed to get dispersed over the available system nodes. Results show over 50% improvement in worst-case latency and 100 times improvement in deadline misses.
Mohammad Fattah, Amir-Mohammad Rahmani, Thomas Canhao Xu, Anil Kanduri, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP6
2014 Multi Rectangle Modeling Approach for Application Mapping on a Many-Core System
abstract
The importance of first node selection in run-time resource management is shown in our previous work, SHiC. It is desired in SHiC to find the optimum node in an agile manner. Accordingly, the current mapping picture of the system is simplified to SHiC by modeling each application as a rectangle of occupied nodes. However, the algorithm performance can be influenced significantly with the rectangle model of each application. In this work, we introduce a precise description of our new accurate rectangle modeling algorithm. Moreover, we show that it is not sufficient to always model an application with only one rectangle, as dispersion of the allocated nodes is an irrepressible phenomenon. Accordingly, our algorithm enables to model a mapped application with several rectangles by tuning the model accuracy against the algorithm complexity. However, the algorithm is not in the critical path of the applications executions. Our results shows up to 5% reduction in power dissipation of the network.
Igor Tcarenko, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4
2014 Path-Based Partitioning Methods for 3D Networks-on-Chip with Minimal Adaptive Routing
abstract
Combining the benefits of 3D ICs and Networks-on-Chip (NoCs) schemes provides a significant performance gain in Chip Multiprocessors (CMPs) architectures. As multicast communication is commonly used in cache coherence protocols for CMPs and in various parallel applications, the performance of these systems can be significantly improved if multicast operations are supported at the hardware level. In this paper, we present several partitioning methods for the path-based multicast approach in 3D mesh-based NoCs, each with different levels of efficiency. In addition, we develop novel analytical models for unicast and multicast traffic to explore the efficiency of each approach. In order to distribute the unicast and multicast traffic more efficiently over the network, we propose the Minimal and Adaptive Routing (MAR) algorithm for the presented partitioning methods. The analytical and experimental results show that an advantageous method named Recursive Partitioning (RP) outperforms the other approaches. RP recursively partitions the network until all partitions contain a comparable number of switches and thus the multicast traffic is equally distributed among several subsets and the network latency is considerably decreased. The simulation results reveal that the RP method can achieve performance improvement across all workloads while performance can be further improved by utilizing the MAR algorithm. Nineteen percent average and 42 percent maximum latency reduction are obtained on SPLASH-2 and PARSEC benchmarks running on a 64-core CMP.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, José Flich, Hannu Tenhunen
IEEE Trans. Computers4
2014 High-Performance and Fault-Tolerant 3D NoC-Bus Hybrid Architecture Using ARB-NET-Based Adaptive Monitoring Platform
abstract
The emerging three-dimensional integrated circuits (3D-ICs) achieve greater device integration and enhanced system performance at lower cost and reduced area footprint, thereby offering higher order of connectivity and greater design choices and possibilities. To exploit the intrinsic capability of reduced communication distances in 3D-ICs, three-dimensional NoC-bus hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored possibility to implement an efficient system-wide monitoring network. In this paper, an efficient three-dimensional NoC architecture is proposed which is optimized for system performance, power consumption, and reliability. The mechanism benefits from a congestion-aware and bus failure-tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we have integrated a low-cost monitoring platform on top of the three-dimensional NoC-Bus Hybrid mesh architecture that can be efficiently used for various system management purposes such as traffic monitoring, fault tolerance, and thermal management. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware and interlayer fault-tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of information generated within bus arbiters. Compared to recently proposed stacked mesh three-dimensional NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power, performance, and reliability improvements with a negligible hardware overhead.
Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
IEEE Trans. Computers5
2014 Editorial: Special issue on design challenges for many-core processors
abstract
No abstract available.
Masoud Daneshtalab, Maurizio Palesi, Juha Plosila
ACM Trans. Embed. Comput. Syst.3
2014 Adaptive load balancing in learning-based approaches for many-core embedded systems
Fahimeh Farahnakian, Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila
J. Supercomput.5
2014 Special section on advances in methods for adaptive multicore systems
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Supercomput.3
2013 Private configuration environments (PCE) for efficient reconfiguration, in CGRAs
abstract
In this paper, we propose a polymorphic configuration architecture, that can be tailored to efficiently support reconfiguration needs of the applications at runtime. Today, CGRAs host multiple applications, running simultaneously on a single platform. Novel CGRAs allow each application to exploit late binding and time sharing for enhancing the power and area efficiency. These features require frequent reconfigurations, making reconfiguration time a bottleneck for time critical applications. Existing solutions to this problem either employ powerful configuration architectures or hide configuration latency (using configuration caching). However, both these methods incur significant costs when designed for worst-case reconfiguration needs. As an alternative to worst-case dedicated configuration mechanism, we exploit reconfiguration to provide each application its private configuration environment (PCE). PCE relies on a morphable configuration infrastructure, a distributed memory sub-system, and a set of PCE controllers. The PCE controllers customize the morphable configuration infrastructure and reserve portion of the a distributed memory sub-system, to act as a context memory for each application, separately. Thereby, each application enjoys its own configuration environment which is optimal in terms of configuration speed, memory requirements and energy. Simulation results using representative applications (WLAN and Matrix Multiplication) showed that PCE offers up to 58% reduction in memory requirements, compared to dedicated, worst case configuration architecture. Synthesis results show that the morphable reconfiguration architecture incurs negligible overheads ( 3% area and 4% power compared of a single processing element).
Muhammad Adeel Tajammul, Syed M. A. H. Jafri, Ahmed Hemani, Juha Plosila, Hannu Tenhunen
ASAP4
2013 MD: Minimal path-based fault-tolerant routing in on-Chip Networks
abstract
The communication requirements of many-core embedded systems are convened by the emerging Network-on-Chip (NoC) paradigm. As on-chip communication reliability is a crucial factor in many-core systems, the NoC paradigm should address the reliability issues. Using fault-tolerant routing algorithms to reroute packets around faulty regions will increase the packet latency and create congestion around the faulty region. On the other hand, the performance of NoC is highly affected by the network congestion. Congestion in the network can increase the delay of packets to route from a source to a destination, so it should be avoided. In this paper, a minimal and defect-resilient (MD) routing algorithm is proposed in order to route packets adaptively through the shortest paths in the presence of a faulty link, as long as a path exists. To avoid congestion, output channels can be adaptively chosen whenever the distance from the current to destination node is greater than one hop along both directions. In addition, an analytical model is presented to evaluate MD for two-faulty cases.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Farhad Mehdipour
ASP-DAC3
2013 Smart hill climbing for agile dynamic mapping in many-core systems
abstract
Stochastic hill climbing algorithm is adapted to rapidly find the appropriate start node in the application mapping of network-based many-core systems. Due to highly dynamic and unpredictable workload of such systems, an agile run-time task allocation scheme is required. The scheme is desired to map the tasks of an incoming application at run-time onto an optimum contiguous area of the available nodes. Contiguous and un-fragmented area mapping is to settle the communicating tasks in close proximity. Hence, the power dissipation, the congestion between different applications, and the latency of the system will be significantly reduced. To find an optimum region, we first propose an approximate model that quickly estimates the available area around a given node. Then the stochastic hill climbing algorithm is used as a search heuristic to find a node that has the required number of available nodes around it. Presented agile climber takes the steps using an adapted version of hill climbing algorithm named Smart Hill Climbing, SHiC, which takes the runtime status of the system into account. Finally, the application mapping is performed starting from the selected first node. Experiments show significant gain in the mapping contiguousness which results in better network latency and power dissipation, compared to state-of-the-art works.
Mohammad Fattah, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila
DAC4
2013 CARS: congestion-aware request scheduler for network interfaces in NoC-based manycore systems
abstract
Network congestion is a critical issue of memory parallelism in network-based manycore systems where multiple memories can be accessed simultaneously. Therefore, a congestion-aware method is necessitated to deal with the network congestion. In this paper, we present a streamlined method in order to reduce the network congestion. The idea is to use the global congestion information as a metric in network interfaces to reduce the congestion level of highly congested areas. Network interfaces connected to memory modules are equipped with an adaptive scheduler using the global congestion information to reduce additional traffic to congested areas. Experimental results with synthetic test cases demonstrate that the on-chip network utilizing the proposed adaptive scheduler presents up to 23% improvement in average latency.
Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen
DATE3
2013 Fault-tolerant routing algorithm for 3D NoC using Hamiltonian path strategy
abstract
While Networks-on-Chip (NoC) have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. In this paper, we address the problem of faulty elements by the means of routing algorithms. Commonly, fault-tolerant algorithms are complex due to supporting different fault models while preventing deadlock. When moving from 2D to 3D network, the complexity increases significantly due to the possibility of creating cycles within and between layers. In this paper, we take advantages of the Hamiltonian path to tolerate faults in the network. The presented approach is not only very simple but also able to support almost all one-faulty unidirectional links in 2D and 3D NoCs.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila
DATE3
2013 Enhanced fault-tolerant Network-on-Chip architecture using hierarchical agents
abstract
The reliability is a vital aspect in the design of Network-on-Chip (NoC) based systems because a fault in communication medium may cause an overall system failure. On the other hand, performance degradation is an inescapable consequence of fault-tolerant architectures. In this paper, we propose a fault-tolerant NoC architecture that attains higher performance by using low cost agents in a hierarchical manner. These agents which are distributed all over the network, collect, process, and distribute different fault information. Moreover, we propose an enhanced fault-tolerant and congestion-aware routing method that exploits the classified fault information related to the permanent faults that might occur inside the links, network interfaces and different parts of the routers. The experimental results reveal that the proposed NoC architecture imposes small area and power overheads.
Mojtaba Valinataj, Pasi Liljeberg, Juha Plosila
DDECS3
2013 Energy-Aware Fault-Tolerant CGRAs Addressing Application with Different Reliability Needs
abstract
In this paper, we propose a polymorphic fault tolerant architecture that can be tailored to efficiently support the reliability needs of multiple applications at run-time. Today, coarse-grained reconfigurable architectures (CGRAs) host multiple applications with potentially different reliability needs. Providing platform-wide worst-case (maximum) protection to all the applications is neither optimal nor desirable. To reduce the fault-tolerance overhead, adaptive fault-tolerance strategies have been proposed. The proposed techniques access the reliability requirements of each application and adjust the fault-tolerance intensity (and hence overhead), accordingly. However, existing flexible reliability schemes only allow to shift between different levels of modular redundancy (duplication, triplication, etc.) and deal with only a single class of faults (e.g. soft errors). To complement these strategies, we propose energy-aware fault-tolerance that, in addition to modular redundancy, can also provide low cost, sub-modular (e.g. residue mod 3) redundancy, to cater both permanent and temporary faults. Our solution relies on an agent based control layer and a configurable fault-tolerance data path. The control layer identifies the application class and configures the data path to provide the needed reliability. Simulation results using a few selected algorithms (FFT, matrix multiplication, and FIR filter) showed that the proposed method provides flexible protection with energy overhead ranging from 3.125% to 107% for different reliability levels. Synthesis results have confirmed that the proposed architecture significantly reduces the area overhead for self-checking (59.1%) and fault tolerant (7.1%) versions, compared to the state of the art adaptive reliability techniques.
Syed M. A. H. Jafri, Stanislaw J. Piestrak, Kolin Paul, Ahmed Hemani, Juha Plosila, Hannu Tenhunen
DSD5
2013 Generation of Structural VHDL Code with Library Components from Formal Event-B Models
abstract
We propose a design approach to integrating correct-by-construction formal modeling with hardware implementations in VHDL. Formal modeling is performed within the Event-B framework that supports the refinement approach, i.e., stepwise unfolding of system properties in a correct-by-construction manner. After an implement able model of a hardware system is derived, we apply an additional refinement step in order to introduce hardware library components in the form of functions. We show the mapping between these functions and corresponding library components such that structural, i.e., component-based, VHDL implementation is derived. The application of functions binds unrestricted data types and substitutes regular operations with function calls. The approach is presented through examples that illustrate the additional refinement step and the code generation. We show the advantages in terms of occupied area (2, 5% and 12, 5%) and performance (13, 7% and 15, 4%) of the descriptions that incorporate hardware library components.
Sergey Ostroumov, Leonidas Tsiopoulos, Kaisa Sere, Juha Plosila
DSD4
2013 Minimal-path fault-tolerant approach using connection-retaining structure in Networks-on-Chip
abstract
There are many fault-tolerant approaches presented both in off-chip and on-chip networks. Regardless of all varieties, there has always been a common assumption between them. Most of all known fault-tolerant methods are based on rerouting packets around faults. Rerouting might take place through nonminimal paths which affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. In this paper, we present a fault-tolerant approach based on using the shortest paths. This method maintains the performance of Networks-on-Chip in the presence of faults. To avoid using non-minimal paths, the router architecture is slightly modified. In the new form of architecture, there is an ability to connect the horizontal and orthogonal links of a faulty router such that healthy routers are kept connected to each other. Based on this architecture, a fault-tolerant routing algorithm is presented which is obviously much simpler than traditional fault-tolerant routing algorithms. According to this algorithm, only the shortest paths are used by packets in the presence of fault. This results retains the performance of NoCs in faulty situations. This algorithm is highly reliable, for an instance, the reliability is more than 99.5% when there are six faulty routers in an 8×8 mesh network.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen
NOCS3
2013 DyXYZ: Fully Adaptive Routing Algorithm for 3D NoCs
abstract
Traditional methods in 3D NoCs simply use a deterministic routing algorithm to deliver packets from a source to a destination node. However, deterministic methods are unable to distribute the traffic load over the network, which results in degrading the performance. In this paper, we present a fully adaptive routing algorithm for 3D NoCs, named DyXYZ. In DyXYZ, the congestion information at the input buffer of the neighboring routers is used as congestion metric to select among the output channels. This algorithm is proven to be deadlock free by using 4, 4, and 2 virtual channels along the X, Y, and Z dimensions, respectively.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen
PDP4
2013 High Performance Fault-Tolerant Routing Algorithm for NoC-Based Many-Core Systems
abstract
Networks-on-Chip (NoCs) has become a promising approach for the on-chip communication infrastructure of many-core Systems-on-Chip (SoCs). Faults may occur in the NoC both at the router and link level. There are many fault-tolerant approaches presented both in the off-chip and on-chip networks. Some approaches disable some healthy components in order to form a specific shape and others not. Regardless of all varieties, there has always been a common assumption among them. Most of all traditional fault-tolerant methods are based on rerouting packets around a faulty node or region. These approaches affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. The focus of this paper is to maintain the performance of NoC in the presence of faults. The presented method takes advantage of a fully adaptive routing algorithm using one and two virtual channels along the X and Y dimensions. This method is able to tolerate all cases of one-faulty node without losing the performance of NoC. According to the experimental results, this presented fault-tolerant routing algorithm is able to support up to six faulty nodes in the 8×8 mesh network by up to 98% reliability.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila
PDP3
2013 Enhancing Performance of 3D Interconnection Networks using Efficient Multicast Communication Protocol
abstract
Three-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. They also provide greater design flexibility by allowing heterogeneous integration. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement efficient multicast routings for 3D networks-on-chip. In this paper, we propose a novel multicast partitioning and routing strategy for the 3D NoC-Bus Hybrid mesh architectures to enhance the overall system performance and reduce the power consumption. The proposed architecture exploits the beneficial attribute of a single-hop (bus-based) interlayer communication of the 3D stacked mesh architecture to provide high-performance hardware multicast support. To this end, a customized partitioning method and an efficient routing algorithm are presented to reduce the average hop count and latency of the network. Compared to the recently proposed 3D NoC architectures being capable of supporting hardware multicasting, our extensive simulations with different traffic profiles reveal that our architecture using the proposed multicast routing strategy can help achieve significant performance improvements.
Sanaz Rahimi Moosavi, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4
2013 Cluster-based topologies for 3D Networks-on-Chip using advanced inter-layer bus architecture
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Comput. Syst. Sci.4
2013 Developing a power-efficient and low-cost 3D NoC using smart GALS-based vertical channels
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Comput. Syst. Sci.3
2013 A systematic reordering mechanism for on-chip networks using efficient congestion-aware method
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Syst. Archit.4
2013 Formal approach to agent-based dynamic reconfiguration in Networks-On-Chip
Sergey Ostroumov, Leonidas Tsiopoulos, Juha Plosila, Kaisa Sere
J. Syst. Archit.3
2013 Design space exploration of thermal-aware many-core systems
Kameswar Rao Vaddina, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila
J. Syst. Archit.5
2013 Optimal placement of vertical connections in 3D Network-on-Chip
Thomas Canhao Xu, Gert Schley, Pasi Liljeberg, Martin Radetzki, Juha Plosila, Hannu Tenhunen
J. Syst. Archit.5
2012 ARB-NET: A novel adaptive monitoring platform for stacked mesh 3D NoC architectures
abstract
The emerging three-dimensional integrated circuits (3D ICs) offer a promising solution to mitigate the barriers of interconnect scaling in modern systems. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is presented taking advantage from viable information generated within bus arbiters. Our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help achieving significant power and performance improvements compared to recently proposed stacked mesh 3D NoCs.
Amir-Mohammad Rahmani, Khalid Latif 0002, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
ASP-DAC5
2012 CATRA- congestion aware trapezoid-based routing algorithm for on-chip networks
abstract
Congestion occurs frequently in Networks-on-Chip when the packets demands exceed the capacity of network resources. Congestion-aware routing algorithms can greatly improve the network performance by balancing the traffic load in adaptive routing. Commonly, these algorithms either rely on purely local congestion information or take into account the congestion conditions of several nodes even though their statuses might be out-dated for the source node, because of dynamically changing congestion conditions. In this paper, we propose a method to utilize both local and non-local network information to determine the optimal path to forward a packet. The non-local information is gathered from the nodes that not only are more likely to be chosen as intermediate nodes in the routing path but also provide up-to-date information to a given node. Moreover, to collect and deliver the non-local information, a distributed propagation system is presented.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DATE4
2012 HLS-DoNoC: High-level simulator for dynamically organizational NoCs
abstract
A high-level simulator is presented for the design and analysis of dynamically organizational Networks-on-Chip (DoNoCs). The DoNoC is able to organize statically or dynamically different network nodes for run-time coarse and fine grained reconfiguration, in particular power management. As an important step in the design flow, a simulator for early-stage design exploration is the focus of the paper. Built upon classic wormhole-based NoC architecture, the simulator is capable of experimenting diverse run-time monitoring and reconfiguration methods. In particular, dynamic clusterization can be performed with inter-cluster interfaces properly configured at the run-time. The simulator is flit-level accurate, trace-driven, and easy-to-reconfigure. It supports both synchronous and ratiochronous timing, and can provide the communication performance and power/energy consumption. The paper demonstrates the usage of the simulator in the design of various cluster-based power management schemes.
Liang Guang, Ethiopia Nigussie, Juha Plosila, Jouni Isoaho, Hannu Tenhunen
DDECS3
2012 MAFA: Adaptive Fault-Tolerant Routing Algorithm for Networks-on-Chip
abstract
While Networks-on-Chip have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. This level of failure has architectural level ramifications, as it may cause an entire on-chip network to fail. Traditional fault-tolerant routing algorithms can overcome the faulty links or routers by rerouting packets around faulty regions. These approaches increase the packet latency and create congestion around the faulty region. In this paper, we present a novel fault-tolerant method that is able to route packets through shortest paths in the presence of faulty links, as long as a path exists. Although the same idea can be applied to a network with any number of virtual channels, we utilize two virtual channels to tolerate all one and two faulty links. Finally, the method is extended to support multiple faulty links by fully utilizing all allowable turns in the network.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen
DSD3
2012 Energy-Aware Fault-Tolerant Network-on-Chips for Addressing Multiple Traffic Classes
abstract
This paper presents an energy efficient architecture to provide on-demand fault tolerance to multiple traffic classes, running simultaneously on single network on chip (NoC) platform. Today, NoCs host multiple traffic classes with potentially different reliability needs. Providing platform-wide worst-case (maximum) protection to all the classes is neither optimal nor desirable. To reduce the overheads incurred by fault tolerance, various adaptive strategies have been proposed. The proposed techniques rely on individual packet fields and operating conditions to adjust the intensity and hence the overhead of fault tolerance. Presence of multiple traffic classes undermines the effectiveness of these methods. To complement the existing adaptive strategies, we propose on-demand fault tolerance, capable of providing required reliability, while significantly reducing the energy overhead. Our solution relies on a hierarchical agent based control layer and a reconfigurable fault tolerance data path. The control layer identifies the traffic class and directs the packet to the path providing the needed reliability. Simulation results using representative applications (matrix multiplication, FFT, wavefront, and HiperLAN) showed up to 95% decrease in energy consumption compared to traditional worst case methods. Synthesis results have confirmed a negligible additional overhead, for providing on-demand protection (up to 5.3% area), compared to the overall fault tolerance circuitry.
Syed M. A. H. Jafri, Liang Guang, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen
DSD5
2012 Power and Thermal Analysis of Stacked Mesh 3D NoC Using AdaptiveXYZ Routing Algorithm
abstract
Three-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement an integrated low-cost system-wide monitoring network. In this paper, a generic monitoring and management platform called ARB-NET is presented. Based on the ARB-NET monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is provided which takes advantage from viable information generated within the monitoring network. In addition, we address both the power and thermal issues of a stacked mesh 3D network on chips using AdaptiveXYZ routing. To this end, a thermal model of a 3D stacked NoC system in a modern flip-chip package is developed. Thermal and power analysis are performed in order to investigate the impact of the proposed adaptive routing from the power and thermal perspectives. Our experiments with a videoconference encoder as a real application show significant power, performance and peak temperature improvements compared to a typical stacked mesh 3D NoC.
Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DSD4
2012 NoC-AXI interface for FPGA-based MPSoC platforms
abstract
Streaming applications are a keystone in several emerging multimedia services like DVB-IPTV, VoD and on-line gaming. Due to the high computing requirements and real-time constraints inherent to this kind of applications multi-processor system-on-chip (MPSoCs) have been proposed as a solution. In addition, the FPGA technology has become popular among systems-on-chip (SoCs) designers due to its low development cost and short time to market. Here we present a FPGA-based MPSoC platform for streaming applications where the important component of this platform is the AXI interface.
Marco Ramírez 0001, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg
FPL3
2012 t(k)-SA: accelerated simulated annealing algorithm for application mapping on networks-on-chip
abstract
Simulated Annealing (SA) algorithm is a promising method for solving combinatorial optimization problems. The only limitation of applying the SA algorithm to application mapping problem on many-core networks-on-chip (NoCs) is its low speed. To alleviate this limitation, an accelerated SA algorithm called tk-SA algorithm is proposed in this work. The tk-SA algorithm starts the annealing process from a lower initial temperature tk with an optimized initial mapping solution. Based on the analysis of the typical behavior of the general SA algorithm, an efficient method is proposed for determining the temperature tk. Quantitative evaluations verify that the method is capable of obtaining an appropriate tk such that the tk-SA algorithm can reproduce the behavior of the full-range SA from temperature tk. Experimental results show that compared with a parameter-optimized SA algorithm, the proposed tk-SA algorithm achieves an average speedup of 1.55 without loss of solution quality.
Bo Yang 0009, Liang Guang, Tero Säntti, Juha Plosila
GECCO4
2012 Existing challenges and new opportunities in context-aware systems
abstract
Merging the features of Cloud computing, autonomic computing, pervasive computing, and mobile computing are now at its initial stage but the effort is visibly showing the benefits of these paradigms. A large number of applications can take advantage of this, including healthcare, traffic control, and social network applications. However, these applications are complex by nature and introduce several challenges of their own, for example, reliable sensing, accurate context recognition, scalability, security, and the challenge of dealing with previously unforeseen side-effects of adaptations. These challenges can be surmounted when researchers of diverse background come together and provide different views of the same problems and help each other understand the complex relationships between contending ideas. The Casemans 2012 workshop opens the necessary platform for researchers of ubiquitous computing, autonomic computing, and similar fields to address these issues.
Waltenegus Dargie, Juha Plosila, Vincenzo De Florio
UbiComp2
2012 Vertical and horizontal integration towards collective adaptive system: a visionary approach
abstract
Hybrid multi-domain computing systems are emerging. While the context-aware self-adaptive system models are under intensive research in individual computing domains, their integration into a collective adaptive system still remains a major challenge. This position paper visions a meet-in-the-middle approach, where horizontal integration is applied to sub-system models extracted from vertical integration. The integration relies on orthogonal behavior and execution models respectively capturing the functional and non-functional features of sub-systems. The construction towards guaranteed services can be achieved with composition of static (worst-case) execution models, while best-effort services can be constructed with statistical models. Given that each computing domain has, to some extent, formulated its own design flow of context-aware systems, the envisaged meet-in-the-middle integration approach maximizes the reuse of existing models and platforms, thus is promising for the highly-complex system design process.
Liang Guang, Ethiopia Nigussie, Juha Plosila, Hannu Tenhunen
UbiComp3
2012 CoNA: Dynamic application mapping for congestion reduction in many-core systems
abstract
Increasing the number of processors in a single chip toward network-based many-core systems requires a run-time task allocation algorithm. We propose an efficient mapping algorithm that assigns communicating tasks of incoming applications onto resources of a many-core system utilizing Network-on-Chip paradigm. In our contiguous neighborhood allocation (CoNA) algorithm, we target at the reduction of both internal and external congestion due to detrimental impact of congestion on the network performance. We approach the goal by keeping the mapped region contiguous and placing the communicating tasks in a close neighborhood. A completely synthesizable simulation environment where none of the system objects are assumed to be ideal is provided. Experiments show at least 40% gain in different mapping cost functions, as well as 16% reduction in average network latency compared to existing algorithms.
Mohammad Fattah, Marco Ramírez 0001, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila
ICCD5
2012 HARAQ: Congestion-Aware Learning Model for Highly Adaptive Routing Algorithm in On-Chip Networks
abstract
The occurrence of congestion in on-chip networks can severely degrade the performance due to increased message latency. In mesh topology, minimal methods can propagate messages over two directions at each switch. When shortest paths are congested, sending more messages through them can deteriorate the congestion condition considerably. In this paper, we present an adaptive routing algorithm for on-chip networks that provide a wide range of alternative paths between each pair of source and destination switches. Initially, the algorithm determines all permitted turns in the network including 180-degree turns on a single channel without creating cycles. The implementation of the algorithm provides the best usage of all allowable turns to route messages more adaptively in the network. On top of that, for selecting a less congested path, an optimized and scalable learning method is utilized. The learning method is based on local and global congestion information and can estimate the latency from each output channel to the destination region.
Masoumeh Ebrahimi, Masoud Daneshtalab, Fahimeh Farahnakian, Juha Plosila, Pasi Liljeberg, Maurizio Palesi, Hannu Tenhunen
NOCS4
2012 Generic Monitoring and Management Infrastructure for 3D NoC-Bus Hybrid Architectures
abstract
Three-dimensional integrated circuits (3D ICs) achieve enhanced system integration and improved performance at lower cost and reduced area footprint. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed which provides performance, power consumption, and area benefits. Besides its various advantages, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes such as traffic monitoring, thermal management and fault tolerance. The proposed generic monitoring and management infrastructure called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring and management platform, a fully congestion-aware and inter-layer fault tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of viable information generated using bus arbiter network. In addition, we propose a thermal monitoring and management strategy on top of our ARB-NET infrastructure. Compared to recently proposed stacked mesh 3D NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power and performance improvements while preserving the system reliability with negligible hardware overhead.
Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
NOCS5
2012 LEAR - A Low-Weight and Highly Adaptive Routing Method for Distributing Congestions in On-chip Networks
abstract
Congestion-aware routing algorithms can improve network throughput by avoiding packets to be routed through congested areas. In this paper, we propose a minimal/non-minimal routing algorithm to alleviate congestion in the network by making use of all available paths between sources and destinations. The simplicity of the proposed algorithm provides a cost and power efficient solution for Networks-on-Chip while the high degree of adaptive ness, achieved by using an additional virtual channel along the Y dimension, leads to an increased performance. In this method, different restrictions are imposed on the use of each virtual channel, so that the prohibited turns in one virtual channel are permitted in the other one. By fully exploiting of the eligible turns in the network, a large number of output channels can be provided by the proposed method. Based on this method, a packet is routed along the non-minimal path when the neighboring routers in the minimal path are congested.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4
2012 An Efficient Hybridization Scheme for Stacked Mesh 3D NoC Architecture
abstract
Three-dimensional (3D) integration is a viable design paradigm to overcome the existing interconnect bottleneck in integrated systems and enhance system power/performance characteristics. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, stacked mesh 3D NoC architecture was proposed. However, this architecture suffers from naive and straightforward hybridization between NoC and bus media. In this paper, an efficient hybridization scheme is presented to enhance system performance, power consumption, and area of stacked mesh 3D NoC architectures. By utilizing a routing rule called LastZ the proposed hybridization scheme offers many advantages investigated in detail to emphasize the significant achievements. Our extensive simulations with synthetic and real benchmarks, including an integrated videoconference application show that compared to a typical 3D NoC-Bus Hybrid Mesh architecture, our hybridization scheme achieves significant power, performance, and area improvements.
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP3
2012 Memory-Efficient On-Chip Network With Adaptive Interfaces
abstract
To achieve higher memory bandwidth in network-based multiprocessor architectures, multiple dynamic random access memories can be accessed simultaneously. In such architectures, not only resource utilization and latency are the critical issues but also a reordering mechanism is required to deliver the response transactions of concurrent memory accesses in-order. In this paper, we present a memory-efficient on-chip network architecture to cope with these issues efficiently. Each node of the network is equipped with a novel network interface (NI) to deal with out-of-order delivery, and a priority-based router to decrease the network latency. The proposed NI exploits a streamlined reordering mechanism to handle the in-order delivery and utilizes the advance extensible interface transaction-based protocol to maintain compatibility with existing intellectual property cores. To improve the memory utilization and reduce the memory latency, an optimized memory controller is integrated in the presented NI. Experimental results with synthetic test cases demonstrate that the proposed on-chip network architecture provides significant improvements in average network latency (16%), average memory access latency (19%), and average memory utilization (22%).
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2012 Semi-Serial On-Chip Link Implementation for Energy Efficiency and High Throughput
abstract
A high-throughput and low-energy semi-serial on-chip communication link based on novel design techniques and circuit solutions is presented. This self-timed link is designed using high-speed serialization/deserializtion and pulse dual-rail encoding techniques. The link also employs wave-pipelined differential pulse current-mode signaling to maintain the high speed data intake from the serializer. The energy efficiency of the proposed semi-serial link, which consists of bit-serial links in parallel, mainly comes from the sharing of the novel serializer's control circuit among the bit-serial links. In addition, the integration of pulse signaling with wave-pipelining, the use of a new low-complexity data validity detection technique, and the avoidance of data decoding logic also contribute to the power reduction. Furthermore, the formulated pulse dual-rail encoding provides an opportunity to implement pulse signaling at no cost. The ability to detect data validity at bit level allows acknowledgment per word without losing the delay-insensitivity of the transmission. The proposed semi-serial link is analyzed and compared with bit-serial and fully bit-parallel links for 64-bit data and communication distances of 1 to 8 mm. The semi-serial link which consists of eight bit-serial links provides 72.72 Gbps throughput with 286 fJ/bit energy dissipation for 8 mm transmission. It dissipates the lowest energy per bit compared to fully bit-parallel links while achieving the same throughput. The links are designed and simulated in Cadence Analog Spectre using 65-nm technology from STMicroelectronics.
Ethiopia Nigussie, Sampo Tuuna, Juha Plosila, Jouni Isoaho, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.3
2011 LastZ: An Ultra Optimized 3D Networks-on-Chip Architecture
abstract
3D IC technology enables NoC architectures to offer greater device integration and shorter interlayer interconnects. The primary 3D NoC architectures such as Symmetric 3D Mesh NoC could not exploit the beneficial feature of a negligible inter-layer distance in 3D chips. To cope with this, 3D NoC-Bus Hybrid architecture was proposed which is a hybrid between packet-switched network and a bus. This architecture is feasible providing both performance and area benefits, while still suffering from naive and straightforward hybridization between NoC and bus media. In this paper, an ultra optimized hybridization scheme is proposed to enhance system performance, power consumption, area and thermal issues of 3D NoC-Bus Hybrid Mesh. The scheme benefits from a rule called LastZ which enables ultra optimization of the inter-layer communication architecture. In addition, we present a wrapper to preserve the backward compatibility of the proposed architecture for connecting with the existing network interfaces. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. Our extensive simulations demonstrate significant area, power, and performance improvements compared to a typical 3D NoC-Bus Hybrid Mesh architecture.
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DSD3
2011 Thermal Analysis of Job Allocation and Scheduling Schemes for 3D Stacked NoC's
abstract
Three-dimensional technology offers greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In this work, a 3D thermal model of a stacked network-on-chip system is developed and thermal analysis is performed in order to analyze different job allocation and scheduling schemes using finite element simulations. The steady-state heat transfer analysis on the 3D stacked structure has been performed. We have analyzed the effect of variation of die power consumption, with and without hotspots, on peak temperatures in different layers of the stack. The optimal die placement solution is also provided based on the maximum temperature attained by the individual silicon dies.
Kameswar Rao Vaddina, Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila
DSD5
2011 Compact generic intermediate representation (CGIR) to enable late binding in coarse grained reconfigurable architectures
abstract
In the era of platforms hosting multiple applications, where inter-application communication and concurrency patterns are arbitrary, static compile time decision making is neither optimal nor desirable. As a part of solving this problem, we present a novel method for compactly representing multiple configuration bitstreams of a single application, with varying parallelisms, as a unique, compact, and customizable representation, called CGIR. The representation thus stored is unraveled at runtime to configure the device with optimal (e.g. in terms of energy) implementation. Our goal was to provide optimal decision making capability to the runtime resource manager (RTM) without compromising the runtime behavior or the memory requirements of the system. The presence of multiple binaries enhance optimality by providing the RTM with multiple implementations to choose from. The CGIR ensures minimal increase in memory requirements with the addition of each binary. The low-cost unraveling of CGIR guarantees the runtime behavior. We have chosen the dynamically reconfigurable resource array (DRRA) as a vehicle to study the feasibility of our approach. Simulation results using 16 point decimation in time fast Fourier transform (FFT) has showed massive (up to 18% for 2 versions, 33% for 3 versions) memory savings compared to state of the art. Formal evaluation shows that the savings increase with the increase in the number of implementations stored.
Syed M. A. H. Jafri, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen
FPT4
2011 Exploring partitioning methods for 3D Networks-on-Chip utilizing adaptive routing model
abstract
Three-Dimensional (3D) integration is a solution to the interconnect bottleneck in Two-Dimensional (2D) MultiProcessor System on Chip (MPSoC). 3D IC design improves performance and decreases power consumption by replacing long horizontal interconnects with shorter vertical ones. As the multicast communication is utilized commonly in various parallel applications, the performance can be significantly improved by supporting of multicast operations at the hardware level. In this paper, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method named Recursive Partitioning (RP) in which the network is recursively partitioned until all partitions contain comparable number of nodes. By this approach, the multicast traffic is distributed among several subsets and the network latency is considerably decreased. We also present Minimal Adaptive Routing (MAR) algorithm for the unicast and multicast traffic in 3D-mesh Networks-on-Chip (NoCs). The idea behind the MAR algorithm is utilizing the Hamiltonian path to provide a set of alternative paths.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
NOCS4
2011 Congestion aware, fault tolerant, and thermally efficient inter-layer communication scheme for hybrid NoC-bus 3D architectures
abstract
Three-dimensional IC technology offers greater device integration and shorter interlayer interconnects. In order to take advantage of these attributes, 3D stacked mesh architecture was proposed which is a hybrid between packet-switched network and a bus. Stacked mesh is a feasible architecture which provides both performance and area benefits, while suffering from inefficient intermediate buffers. In this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we hybridize the proposed adaptive routing with available algorithms to mitigate the thermal issues by herding most of the switching activities closer to the heat sink. Our extensive simulations with synthetic and real benchmarks, including the one with an integrated videoconference application, demonstrate significant power, performance, and peak temperature improvements compared to a typical stacked mesh 3D NoC.
Amir-Mohammad Rahmani, Pasi Liljeberg, Khalid Latif 0002, Juha Plosila, Kameswar Rao Vaddina, Hannu Tenhunen
NOCS4
2011 A Stacked Mesh 3D NoC Architecture Enabling Congestion-Aware and Reliable Inter-layer Communication
abstract
In this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. Stacked mesh is a feasible architecture which takes advantage of the short inter-layer wiring delays, while suffering from inefficient intermediate buffers. To cope with this, an inter-layer communication mechanism is developed to enhance the buffer utilization, load balancing, and system fault-tolerance. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm for vertical communication. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. In addition, a video conference encoder has been used as a real application for system analysis. Our extensive experiments show significant power and performance improvements compared to a typical stacked mesh 3D NoC.
Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4
2011 Agent-based on-chip network using efficient selection method
abstract
Congestion in on-chip networks may cause many drawbacks in multiprocessor systems including throughput reduction, increase in latency, and additional power consumption. Furthermore, conventional congestion control methods, employed for on-chip networks, cannot efficiently collect congestion information and distribute them over the on-chip network. In this paper, we present a novel structure for on-chip networks, named Agent-based Network-on-Chip (ANoC), to diagnose the congested areas. In addition to the presented structure, an efficient Congestion-Aware Selection (CAS) method is proposed to reduce overall network latency. CAS is capable of selecting an appropriate output channel to route packets along a less congested path. 29% average and 35% maximum latency reduction are achieved on SPLASH-2 and PARSEC benchmarks running on a 36-core Chip Multi-Processor.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
VLSI-SoC4
2010 Developing reconfigurable FIFOs to optimize power/performance of Voltage/Frequency Island-based networks-on-chip
abstract
Network-on-chip architectures partitioned into several Voltage/Frequency Islands (VFIs) have been proposed to alleviate problems related to integration, excessive energy consumption and clock distribution. The architecture is composed of synchronous switches that communicate with each other using bi-synchronous FIFOs. However, these FIFOs are not needed if adjacent switches belong to the same clock domain. In this paper, a Reconfigurable Synchronous/Bi-Synchronous (RSBS) FIFO is proposed which can operate in either synchronous or bi-synchronous mode. The FIFO is scalable and synthesizable in synchronous standard cells and also a technique for mesochronous adaptation has been recommended. In addition, some techniques are suggested to show how the FIFO could be utilized in a VFI-based NoC. Our results reveal that compared to a non-reconfigurable system architecture, the RSBS FIFOs help to achieve up to 15% savings in average power consumption of NoC switches and 29% improvement in total average packet latency in the case of MPEG-4 encoder application.
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DDECS3
2010 A fault-tolerant and congestion-aware routing algorithm for Networks-on-Chip
abstract
This paper presents a fault-tolerant routing algorithm for mesh-based Networks-on-Chip (NoC) with faulty links. It is a distributed, adaptive and congestion-aware routing algorithm where only two virtual channels are used for both adaptiveness and fault-tolerance. The proposed routing method has a multilevel fault-tolerance capability and therefore it is capable to tolerate more faulty links in more complicated faulty situations with additional hardware costs. The network performance, fault-tolerance capability and hardware overhead are evaluated through appropriate simulations. The experimental results show that the overall reliability of a Network-on-Chip is significantly enhanced against multiple link failures or partially faulty routers with only a small hardware overhead.
Mojtaba Valinataj, Siamak Mohammadi, Juha Plosila, Pasi Liljeberg
DDECS3
2010 Tree-model based mapping for energy-efficient and low-latency Network-on-Chip
abstract
With the NoC size growing constantly, efficient algorithms are needed to provide power/performance-aware task mapping on massively parallel systems. In this paper a novel tree-model based mapping algorithm is proposed, to achieve high energy efficiency and low latency on NoC platforms. A NoC is abstracted as a tree composed of a root node and median nodes at different levels. By mapping tasks starting from the root of the tree, our algorithm minimizes the communication cost and consequently reduces the energy consumption and network delay. Experimental results show that the run-time of our algorithm is decreased by 90% on average compared to the Greedy Incremental (GI) algorithm. Full system simulation also shows that for Radix traffic, compared to the original random mapping, the GI achieves 18.7% and 17.3% reduction in energy consumption and average network latency respectively, while our algorithm achieves 24.7% and 40.8% reduction respectively.
Bo Yang 0009, Thomas Canhao Xu, Tero Säntti, Juha Plosila
DDECS4
2010 Monitoring and reconfiguration techniques for power supply variation tolerant on-chip links
abstract
We present a runtime power supply and temperature variation tolerance technique for delay-insensitive on-chip communication links. Signal integrity of a current sensing receiver is affected by runtime supply voltage and temperature fluctuations. By monitoring receiver's input current and comparing it with receiver's reference current, the effect of variation on link reliability is detected. Receiver reconfiguration and request for retransmission is carried out when error is detected. This scheme makes the link adaptive to the effect of variations enabling continuous and reliable operation of the link. The power consumption due to monitoring is 489μW for a 2mm link, which is 1.63% of a 64-bit l-of-4 encoded link's power consumption. It requires only 5.83um additional active area. The circuits are designed and simulated in Cadence Analog Spectre using 65nm CMOS technology from STMicroelectronics.
Ethiopia Nigussie, Juha Plosila, Jouni Isoaho
ISCAS2
2010 A Low-Latency and Memory-Efficient On-chip Network
abstract
Using multiple SDRAMs in MPSoCs and NoCs to increase memory parallelism is very common nowadays. In-order delivery, resource utilization, and latency are the most critical issues in such architectures. In this paper, we present a novel network interface architecture to cope with these issues efficiently. The proposed network interface exploits a resourceful reordering mechanism to handle the in-order delivery and to increase the resource utilization. A brilliant memory controller is efficiently integrated into this network interface to improve the memory utilization and reduce both memory and network latencies. In addition, to bring compatibility with existing IP cores the proposed network interface utilizes AXI transaction based protocol. Experimental results with synthetic test cases demonstrate that the proposed architecture gives significant improvements in average network latency (12%), average memory access latency (19%), and average memory utilization (22%).
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
NOCS4
2010 A High-Performance Network Interface Architecture for NoCs Using Reorder Buffer Sharing
abstract
Increasing memory parallelism in MPSoCs to provide higher memory bandwidth is achieved by accessing multiple memories simultaneously. Inasmuch as the response transactions of concurrent memory accesses must be in-order, a reordering mechanism is required. To our knowledge the resource utilization of conventional reordering mechanisms is low. In this paper, we present a novel network interface architecture for on-chip networks to increase the resource utilization and to improve overall performance. Also, based on the proposed architecture, a hybrid network interface is presented to integrate both memory and processor in a tile. The proposed architecture exploits AXI transaction based protocol to be compatible with existing IP cores. Experimental results with synthetic test cases demonstrate that the proposed architecture outperforms the conventional architecture in terms of latency. Also, the cost of the presented architecture is evaluated with UMC 0.09 ¿ m technology.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4
2010 Self-Adaptive System for Addressing Permanent Errors in On-Chip Interconnects
abstract
We present a self-contained adaptive system for detecting and bypassing permanent errors in on-chip interconnects. The proposed system reroutes data on erroneous links to a set of spare wires without interrupting the data flow. To detect permanent errors at runtime, a novel in-line test (ILT) method using spare wires and a test pattern generator is proposed. In addition, an improved syndrome storing-based detection (SSD) method is presented and compared to the ILT method. Each detection method (ILT and SSD) is integrated individually into the noninterrupting adaptive system, and a case study is performed to compare them with Hamming and Bose-Chaudhuri-Hocquenghem (BCH) code implementations. In the presence of permanent errors, the probability of correct transmission in the proposed systems is improved by up to 140% over the standalone Hamming code. Furthermore, our methods achieve up to 38% area, 64% energy, and 61% latency improvements over the BCH implementation at comparable error performance.
Teijo Lehtonen, David Wolpert 0001, Pasi Liljeberg, Juha Plosila, Paul Ampadu
IEEE Trans. Very Large Scale Integr. Syst.4
2009 An efficent dynamic multicast routing protocol for distributing traffic in NOCs
abstract
Nowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms.
Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen
DATE6
2009 Self-timed thermal sensing and monitoring of multicore systems
abstract
As the number of cores increases thermal challenges increase, thereby degrading the performance and reliability of the system. We approach this challenge with a self-timed thermal monitoring method which is based on the use of thermal sensors. Since leakage currents are sensitive to temperature and increase with scaling, we propose the use of a leakage current based thermal sensing for monitoring purposes. In this work we have implemented a novel thermal sensing circuit in 65 nm CMOS technology, which converts analog temperature information into digital form. We have also proposed a novel thermal sensing and monitoring interconnection network structure based on self-timed signaling, comprising of an encoder/transmitter and decoder/ receiver. We have performed power supply noise, additive noise on sensor input signal and dynamic power supply voltage variation analysis on the thermal sensing circuit and show that it is robust enough under different operating temperatures.
Kameswar Rao Vaddina, Ethiopia Nigussie, Pasi Liljeberg, Juha Plosila
DDECS4
2008 A novel hardware acceleration scheme for java method calls
abstract
This paper presents a novel strategy for accelerating the method calls in the REALJava co-processor. The hardware assisted virtual machine architecture is described shortly to provide context for the method call acceleration. The strategy is realized as an FPGA prototype. It allows measurements of real life performance increase, and validates the concept. The system is intended to be used in embedded environments, limiting the CPU performance and memory available to the virtual machine. The co-processor is designed in a highly modular fashion, especially separating the communication from the actual core. This modularity of the design makes the co-processor more reusable and allows system level scalability. This work is a part of a project focusing on design of an advanced Java co-processor for Java intensive SoC applications.
Tero Säntti, Joonas Tyystjärvi, Juha Plosila
ISCAS3
2007 Fault Tolerance Analysis of NoC Architectures
abstract
The paper presents an approach for analyzing and improving fault tolerance aspects in NoC architectures. This is a necessary step to be taken in order to implement reliable systems in future nanoscale technologies. Several NoC architectures and the router structures as well as the network interface needed for them are presented and compared for their fault tolerance, area and performance. The results indicate that a network structure built from simple 3-port routers provides better fault tolerance than a structure based on more complex multiport routers, and that the area overhead can be kept moderate
Teijo Lehtonen, Pasi Liljeberg, Juha Plosila
ISCAS3
2007 Current Mode On-Chip Interconnect using Level-Encoded Two-Phase Dual-Rail Encoding
abstract
We present delay variations insensitive long on-chip interconnect implementation based on current mode signaling and level-encoded two-phase dual-rail (LEDR) encoding. LEDR encoding is chosen over the normal two-phase dual-rail encoding because its completion detection and decoding circuitry is faster and much simpler since detection is level based rather than transition. The communication latency of this interconnect at global lengths of the wires reduces by half compared to conventional voltage mode LEDR interconnect. This is due to current mode signaling, making it possible to achieve high speed without pipelining and/or using repeaters. Performance simulation shows that at 5 mm wire length the throughput of this interconnect is 1 Gbps per one dual-rail wire pair. The effect of crosstalk on signal propagation delay is analyzed using four-bit parallel data transfer with the worst-case switching pattern and transmission line model which have both capacitive and inductive coupling. The interconnect circuitry is designed and simulated using Cadence Analog Spectre and Hspice with 130 nm CMOS technology.
Ethiopia Nigussie, Juha Plosila, Jouni Isoaho
ISCAS2
2006 Time Aware Modelling and Analysis of Multiclocked VLSI Systems
Tomi Westerlund, Juha Plosila
ICFEM2
2006 An approach for analysing and improving fault tolerance in radio architectures
abstract
We present an approach for analysing and improving fault-tolerance aspects in radio architectures. This is a necessary step to be taken in order to implement reliable radio systems in future nanoscale technologies. We present problem formulation, optimisation approach and implementation methodology. We are adding fault tolerance at architecture level by taking advantage of existing parallel structures and using a spare module approach in order to minimise hardware overhead needed. These issues have been analysed and demonstrated using two radio case studies: a UMTS MIMO and a GSM diversity receivers
Teijo Lehtonen, Pekka Rantala, P. Isomaki, Juha Plosila, Jouni Isoaho
ISCAS4
2006 Full-duplex link implementation using dual-rail encoding and multiple-valued current-mode logic
abstract
In this paper we present the circuit implementation of a new asynchronous on-chip link structure, where two modules placed on the opposite sides of the link can exchange data simultaneously. The link uses a special communication protocol called 2-color 1-phase in which the number of communication actions per transfer is only one, making it potentially faster than the conventional handshake-based protocols. The transceiver circuits are designed using multiple-valued current-mode logic, linear summation is implemented by wiring without active devices simplifying the resulting circuitry. By using 90mV voltage swing the power consumption of the link is 16mW for 178ps propagation delay and 2mm interconnect length. The circuit is designed and simulated using Cadence Analog Spectre with a 0.13mum CMOS technology
Ethiopia Nigussie, Juha Plosila, Jouni Isoaho
ISCAS2
2006 Implementing a Self-Timed Low-Power Java Accelerator for Network-on-Chip Applications
abstract
This paper presents an advanced self-timed Java accelerator core which has extremely low power consumption while providing sufficient performance for even the most demanding real-time telecommunication and multimedia applications. The goal is that the accelerator can be directly attached to any general-purpose processor core running some Java-intensive application software. Asynchronous self-timed circuit technology, where timing is based on local handshakes between circuit blocks instead of a global clock signal, provides a promising platform for obtaining a highly modular low-power Java accelerator implementation
Juha Plosila, Lu Yan, Kaisa Sere
PDCAT2
2005 Modelling and Refinement of an On-Chip Communication Architecture
Juha Plosila, Pasi Liljeberg, Jouni Isoaho
ICFEM1
2005 On-chip Debug for an Asynchronous Java Accelerator
abstract
The solution to debug a problem in a deeply embedded system is to integrate the debug and communication module inside the chip. In this paper, we propose an on-chip in-circuit emulation (ICE) architecture for debugging an asynchronous Java accelerator core which can be integrated with any existing processor and operating system. The operation of this ICE module and the debug strategy of the Java accelerator are specifically designed for asynchronous implementation. They not only facilitate the system development but also provide a manufacture test method for asynchronous chips.
Juha Plosila, Lu Yan, Kaisa Sere
PDCAT2
2005 Asynchronous system synthesis
Juha Plosila, Kaisa Sere, Marina Waldén
Sci. Comput. Program.1
2004 Constituent Elements of a Correctness-Preserving UML Design Approach
Tiberiu Seceleanu, Juha Plosila
IFM2
2004 Self-timed communication platform for implementing high-performance systems-on-chip
Pasi Liljeberg, Juha Plosila, Jouni Isoaho
Integr.2
2002 Specification of an Asynchronous On-chip Bus
Juha Plosila, Tiberiu Seceleanu
ICFEM1