Hannu Tenhunen

dblp:68/5555 · DBLP profile ↗
← Back
149ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0003-1959-6513ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 107 · 4 since 2021Software engineering, systems software and programming languages · 13Applied, interdisciplinary, general and emerging computing · 10Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2Theory of computation · 2Artificial intelligence and machine learning · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 DAI-NET: Toward communication-aware collaborative training for the industrial edge
abstract
The industrial edge generates an abundance of spatially distributed and dynamic data that needs to remain on-site for privacy and security reasons. Collaborative training at the edge can leverage this data to refine pre-trained models locally for specific industrial tasks and environments and have them adapt to local changes for enhanced performance, agility, and resilience. However, communication between the devices during training is a key bottleneck and is not modelled by existing frameworks such as MxNet, PyTorch and TensorFlow. This paper introduces DAI-NET, a co-simulation framework for examining communication and its associated costs, and provides results from an implementation using Python, OMNET++ and INET. To validate it and showcase its utility, the developed platform is applied in the analysis of (i) the performance and cost of collaboratively training a Multilayer Perceptron model, and (ii) the influence of computational heterogeneity. Communication costs generated during the training are captured at the device and system levels. In computationally heterogeneous clusters , the root cause of stragglers is exposed. In addition, the key performance contributors are identified to be a cluster’s computation capability and the variation in the relative computation capabilities of its devices. This study is particularly useful for Artificial Intelligence of Things (AIoT) systems, whose bandwidth and energy resources are limited. It lends the way for more practical research on communication-efficient algorithms, network protocols and architectures for the AIoT edge.
Christine Mwase, Yi Jin 0007, Tomi Westerlund, Hannu Tenhunen, Zhuo Zou
Future Gener. Comput. Syst.4
2022 Communication-efficient distributed AI strategies for the IoT edge
Christine Mwase, Yi Jin 0007, Tomi Westerlund, Hannu Tenhunen, Zhuo Zou
Future Gener. Comput. Syst.4
2021 Hierarchical Fault Simulation of Deep Neural Networks on Multi-Core Systems
abstract
In this paper, a hierarchical fault simulation technique for neural networks is proposed, supporting both permanent and temporary faults. In the proposed technique, different levels of hierarchy are used, forming a mixed-level simulation environment. In such an environment, the pre-synthesis behavioral specification of the network and the post-synthesis gate-level model are co-simulated. To accelerate the fault simulation process, faults are injected in the gate-level specification of the selected neurons while the behavioral model in different levels of abstraction is used to simulate the remaining neurons. Further speedup is obtained through event-driven simulation and parallelization. Experimental results confirm the time efficiency of the proposed fault simulation technique.
Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Antonio Miele, Hannu Tenhunen, Juha Plosila
ETS5
2021 High-Performance Parallel Fault Simulation for Multi-Core Systems
abstract
Fault simulation is a time-consuming process that requires customized methods and techniques to accelerate it. Multi-threading and Multi-core approaches are two promising techniques that can be exploited to accelerate the fault simulation process by using different parts of the hardware at the same time. However, an efficient parallelization is obtained only by the refinement of software with respect to the hardware platform. In this paper, a parallel multi-thread fault simulation technique is proposed to accelerate the simulation process on multi-core platforms. In this approach, the gate input values are independently assigned to each thread. Each input value carries the information of several parallel simulation processes. This provides a multithread parallel fault simulation environment. The experimental results show that the proposed technique can efficiently use the hardware platform. In a single-core platform, the proposed technique can reduce the time by 25% while in a dual-core increasing the thread approximately halves the execution time.
Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Hamid Nejatollahi, Hannu Tenhunen, Juha Plosila
PDP5
2021 Real time performance analysis of secure IoT protocols for microgrid communication
Aron Kondoro, Imed Ben Dhaou, Hannu Tenhunen, Nerey H. Mvungi
Future Gener. Comput. Syst.3
2020 Thermal-Cycling-aware Dynamic Reliability Management in Many-Core System-on-Chip
abstract
Dynamic Reliability Management (DRM) is a common approach to mitigate aging and wear-out effects in multi- /many-core systems. State-of-the-art DRM approaches apply finegrained control on resource management to increase/balance the chip reliability while considering other system constraints, e.g., performance, and power budget. Such approaches, acting on various knobs such as workload mapping and scheduling, Dynamic Voltage/Frequency Scaling (DVFS) and Per-Core Power Gating (PCPG), demonstrated to work properly with the various aging mechanisms, such as electromigration, and Negative-Bias Temperature Instability (NBTI). However, we claim that they do not suffice for thermal cycling. Thus, we here propose a novel thermal-cycling-aware DRM approach for shared-memory many-core systems running multi-threaded applications. The approach applies a fine-grained control capable at reducing both temperature levels and variations. The experimental evaluations demonstrated that the proposed approach is able to achieve 39% longer lifetime than past approaches.
M. H. Haghbayan, Antonio Miele, Zhuo Zou, Hannu Tenhunen, Juha Plosila
DATE4
2020 Lightweight Security Algorithms for Resource-constrained IoT-based Sensor Nodes
abstract
With the constant improvement of electronics and development by research community, professionals and enthusiasts around the world, Internet of Things (IoT) based devices have seen a massive increase. These devices are now connected to our daily life in multiple ways and facilitate smooth operation of large, autonomous and semi-autonomous systems in different sectors. The communication among these systems needs to be done in a secure manner. However, as most of the IoT devices have very limited processing capability and energy source, all cryptography algorithms are not able to run on all devices. In addition, depending on the required data performance, it can be desirable to use one specific type of algorithm over others. In this paper, we analyze popularly used lightweight algorithms in terms of operational latency by running them on multiple widely used embedded modules. In addition, we measure power consumption while running an algorithm to realize its impact on battery life as an example. Finally, we discuss design-time considerations to help designers to select an appropriate cryptography algorithm for different applications.
Victor K. Sarker, Tuan Nguyen Gia, Hannu Tenhunen, Tomi Westerlund
ICC3
2020 Dynamic Resource-Aware Corner Detection for Bio-Inspired Vision Sensors
abstract
Event-based cameras are vision devices that transmit only brightness changes with low latency and ultra-low power consumption. Such characteristics make event-based cameras attractive in the field of localization and object tracking in resource-constrained systems. Since the number of generated events in such cameras is huge, the selection and filtering of the incoming events are beneficial from both increasing the accuracy of the features and reducing the computational load. In this paper, we present an algorithm to detect asynchronous corners form a stream of events in real-time on embedded systems. The algorithm is called the Three Layer Filtering-Harris or TLF-Harris algorithm. The algorithm is based on an events' filtering strategy whose purpose is 1) to increase the accuracy by deliberately eliminating some incoming events, i.e., noise and 2) to improve the real-time performance of the system, i.e., preserving a constant throughput in terms of input events per second, by discarding unnecessary events with a limited accuracy loss. An approximation of the Harris algorithm, in turn, is used to exploit its high-quality detection capability with a low-complexity implementation to enable seamless real-time performance on embedded computing platforms. The proposed algorithm is capable of selecting the best corner candidate among neighbors and achieves an average execution time savings of 59% compared with the conventional Harris score. Moreover, our approach outperforms the competing methods, such as eFAST, eHarris, and FA-Harris, in terms of real-time performance, and surpasses Arc* in terms of accuracy.
Sherif Abdelmonem Sayed Mohamed, Jawad Naveed Yasin, M. H. Haghbayan, Antonio Miele, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila
ICPR6
2019 Monocular visual odometry based on hybrid parameterization
abstract
Visual odometry (VO) is one of the most challenging techniques in computer vision for autonomous vehicle/vessels. In VO, the camera pose that also represents the robot pose in ego-motion is estimated analyzing the features and pixels extracted from the camera images. Different VO techniques mainly provide different trade-offs among the resources that are being considered for odometry, such as camera resolution, computation/communication capacity, power/energy consumption, and accuracy. In this paper, a hybrid technique is proposed for camera pose estimation by combining odometry based on triangulation using the long-term period of direct-based odometry and the short-term period of inverse depth mapping. Experimental results based on the EuRoC data set shows that the proposed technique significantly outperforms the traditional direct-based pose estimation method for Micro Aerial Vehicle (MAV), keeping its potential negative effect on performance negligible.
Sherif Abdelmonem Sayed Mohamed, M. H. Haghbayan, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila
ICMV4
2019 Implementation of a Fuel Estimation Algorithm on SoC FPGA
abstract
This paper proposes hardware architecture for a fuel estimation algorithm suitable for SoC FPGA. The architecture utilizes 32-point single precision floating point representation. A resource constraint scheduling algorithm is elaborated to synthesize an area efficient architecture. The floating point arithmetic is implemented using commercial IP. Synthesis results for the Virtex-6 FPGA family reveal that the proposed architecture consumes 1655 slices, has a latency of 71 ns and consumes 0.15 µJ. Additionally, the present work describes a software implementation of the fuel estimation algorithm using Zynq-7000 SoC.
Imed Ben Dhaou, Faisal Mahroogi, Hannu Tenhunen
ISCAS3
2019 Energy efficient fog-assisted IoT system for monitoring diabetic patients with cardiovascular disease
Tuan Nguyen Gia, Imed Ben Dhaou, Mai Ali, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen
Future Gener. Comput. Syst.7
2019 Towards an interoperable Internet of Things through a web of virtual things at the Fog layer
Behailu Negash, Tomi Westerlund, Hannu Tenhunen
Future Gener. Comput. Syst.3
2019 Energy-Aware VM Consolidation in Cloud Data Centers Using Utilization Prediction Model
abstract
Virtual Machine (VM) consolidation provides a promising approach to save energy and improve resource utilization in data centers. Many heuristic algorithms have been proposed to tackle the VM consolidation as a vector bin-packing problem. However, the existing algorithms have focused mostly on the number of active Physical Machines (PMs) minimization according to their current resource requirements and neglected the future resource demands. Therefore, they generate unnecessary VM migrations and increase the rate of Service Level Agreement (SLA) violations in data centers. To address this problem, we propose a VM consolidation approach that takes into account both the current and future utilization of resources. Our approach uses a regression-based model to approximate the future CPU and memory utilization of VMs and PMs. We investigate the effectiveness of virtual and physical resource utilization prediction in VM consolidation performance using Google cluster and PlanetLab real workload traces. The experimental results show, our approach provides substantial improvement over other heuristic and meta-heuristic algorithms in reducing the energy consumption, the number of VM migrations and the number of SLA violations.
Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Nguyen Trung Hieu, Hannu Tenhunen
IEEE Trans. Cloud Comput.6
2018 A 3D Tiled Low Power Accelerator for Convolutional Neural Network
abstract
It remains a challenge to run Deep Learning in devices with stringent power budget in the Internet-of-Things. This paper presents a low-power accelerator for processing Convolutional Neural Networks on the embedded devices. The power reduction is realized by exploring data reuse in three different aspects, with regards to convolution, filter and input features. A systolic-like data flow is proposed and applied to rows of Processing Elements (PEs), which facilitate reusing the data during convolution. Reuse of input features and filters is achieved by arranging the PE array in a 3D tiled architecture, whose dimension is 3 × 14 × 4. Local storage within PEs is therefore reduced and only cost 17.75 kB, which is 20% of the state-of-the-art. With dedicated delay chains in each PE, this accelerator is reconfigurable to suit various parameter settings of convolutional layers. Evaluated in UMC 65 nm low leakage process, the accelerator can reach a peak performance of 84 GOPS and consume only 136 mW at 250 Mhz.
Yuxiang Huan, Jiawei Xu 0002, Lirong Zheng 0001, Hannu Tenhunen, Zhuo Zou
ISCAS4
2018 Parallel imperialist competitive algorithms
abstract
Summary The importance of optimization and NP‐problem solving cannot be overemphasized. The usefulness and popularity of evolutionary computing methods are also well established. There are various types of evolutionary methods; they are mostly sequential but some of them have parallel implementations as well. We propose a multi‐population method to parallelize the Imperialist Competitive Algorithm. The algorithm has been implemented with the Message Passing Interface on 2 computer platforms, and we have tested our method based on shared memory and message passing architectural models. An outstanding performance is obtained, demonstrating that the proposed method is very efficient concerning both speed and accuracy. In addition, compared with a set of existing well‐known parallel algorithms, our approach obtains more accurate results within a shorter time period.
Amin Majd, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila, Shahriar Lotfi, Hannu Tenhunen
Concurr. Comput. Pract. Exp.6
2018 IoT-Based Remote Pain Monitoring System: From Device to Cloud Platform
abstract
Facial expressions are among behavioral signs of pain that can be employed as an entry point to develop an automatic human pain assessment tool. Such a tool can be an alternative to the self-report method and particularly serve patients who are unable to self-report like patients in the intensive care unit and minors. In this paper, a wearable device with a biosensing facial mask is proposed to monitor pain intensity of a patient by utilizing facial surface electromyogram (sEMG). The wearable device works as a wireless sensor node and is integrated into an Internet of Things (IoT) system for remote pain monitoring. In the sensor node, up to eight channels of sEMG can be each sampled at 1000 Hz, to cover its full frequency range, and transmitted to the cloud server via the gateway in real time. In addition, both low energy consumption and wearing comfort are considered throughout the wearable device design for long-term monitoring. To remotely illustrate real-time pain data to caregivers, a mobile web application is developed for real-time streaming of high-volume sEMG data, digital signal processing, interpreting, and visualization. The cloud platform in the system acts as a bridge between the sensor node and web browser, managing wireless communication between the server and the web application. In summary, this study proposes a scalable IoT system for real-time biopotential monitoring and a wearable solution for automatic pain assessment via facial expressions.
Geng Yang 0003, Mingzhe Jiang, Wei Ouyang 0001, Guangchao Ji, Haibo Xie, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen
IEEE J. Biomed. Health Informatics8
2017 From threads to events: Adapting a lightweight middleware for Contiki OS
abstract
Interoperability is one of the key requirements in the Internet of Things considering the diverse platforms, communication standards and specifications available today. Inherent resource constraints in the majority of IoT devices makes it very difficult to use existing solutions for interoperability, thus demanding new approaches. This paper presents the process of adapting a lightweight interoperability middleware for IoT, LISA, from RIOT to Contiki OS and evaluates memory and power overheads. The middleware follows a service oriented architecture and classifies devices according to available resources to assign different roles, such as Application, Service and Manager Nodes. These roles live in different tiers in a generic IoT architecture, where the Manager nodes are located in the intermediate Fog layer. To adapt to an event based kernel of Contiki, the middleware defines and handles a set of events that are used to communicate with the user application. A network of nodes is simulated to show the architecture promoted by the middleware and the results are presented.
Uzair A. Noman, Behailu Negash, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen
CCNC5
2017 Smart energy efficient gateway for Internet of mobile things
abstract
Internet of Things (IoT) is a fast developing vision in which physical quantities are digitized, processed and analyzed. Internet of Mobile Things (IoMT) as one of new domains of IoT, due to mobility, requires a more demanding and rigorous solution in many aspects, especially in terms of energy efficiency. We propose a solution consisting of energy efficient and fast hardware platform for building IoMT Fog layer facilities. Experimental results are presented to prove superiority of the proposed hardware in several aspects to popular general purpose platforms.
Igor Tcarenko, Yuxiang Huan, David Juhasz, Amir-Mohammad Rahmani, Zhuo Zou, Tomi Westerlund, Pasi Liljeberg, Lirong Zheng 0001, Hannu Tenhunen
CCNC9
2017 Autonomous Patient/Home Health Monitoring Powered by Energy Harvesting
abstract
This paper presents the design of an autonomous smart patient/home health monitoring system. Both patient physiological parameters as well as room conditions are being monitored continuously to insure patient safety. The sensors are connected on an IoT regime, where the collected data is wirelessly transferred to a nearby gateway which performs preliminary data analysis, commonly referred to as fog computing, to make sure emergency personnel and healthcare providers are notified in case patient being monitored is at risk. To achieve power autonomy three energy harvesting sources are proposed, namely, solar, RF and thermal. The design of RF energy harvesting system is demonstrated, where novel multiband antenna is fabricated as well as an efficient RF- DC rectifier achieving maximum efficiency of 84%. Finally, the sensor node is tested with different type of sensors and settings while being solely powered by a Photovoltaic (PV) solar cell.
Mai Ali, Tuan Nguyen Gia, Abd-Elhamid M. Taha, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen
GLOBECOM7
2017 A reliable weighted feature selection for auto medical diagnosis
abstract
Feature selection is a key step in data analysis. However, most of the existing feature selection techniques are serial and inefficient to be applied to massive data sets. We propose a feature selection method based on a multi-population weighted intelligent genetic algorithm to enhance the reliability of diagnoses in e-Health applications. The proposed approach, called PIGAS, utilizes a weighted intelligent genetic algorithm to select a proper subset of features that leads to a high classification accuracy. In addition, PIGAS takes advantage of multi-population implementation to further enhance accuracy. To evaluate the subsets of the selected features, the KNN classifier is utilized and assessed on UCI Arrhythmia dataset. To guarantee valid results, leave-one-out validation technique is employed. The experimental results show that the proposed approach outperforms other methods in terms of accuracy and efficiency. The results of the 16-class classification problem indicate an increase in the overall accuracy when using the optimal feature subset. Accuracy achieved being 99.70% indicating the potential of the algorithm to be utilized in a practical auto-diagnosis system. This accuracy was obtained using only half of features, as against an accuracy of66.76% using all the features.
Golnaz Sahebi, Amin Majd, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen
INDIN5
2017 Low-latency hardware architecture for cipher-based message authentication code
abstract
Cipher-based message authentication code, CMAC, is a NIST approved standard for checking message integrity and authentication. This work presents a low-latency AES architecture for CMAC. The architecture uses intensive parallel processing per round and takes advantage of the BRAM present in modern FPGA. Experimental results show that for typical IoT application, the proposed architecture has a latency of 10 clock cycles, consumes 1355 slices, 2 BRAMs and achieves a throughput of 3.8Gbps.
Imed Ben Dhaou, Tuan Nguyen Gia, Pasi Liljeberg, Hannu Tenhunen
ISCAS4
2017 Low-cost fog-assisted health-care IoT system with energy-efficient sensor nodes
abstract
A better lifestyle starts with a healthy heart. Unfortunately, millions of people around the world are either directly affected by heart diseases such as coronary artery disease and heart muscle disease (Cardiomyopathy), or are indirectly having heart-related problems like heart attack and/or heart rate irregularity. Monitoring and analyzing these heart conditions in some cases could save a life if proper actions are taken accordingly. A widely used method to monitor these heart conditions is to use ECG or electrocardiography. However, devices used for ECG are costly, energy inefficient, bulky, and mostly limited to the ambulatory environment. With the advancement and higher affordability of Internet of Things (IoT), it is possible to establish better health-care by providing real-time monitoring and analysis of ECG. In this paper, we present a low-cost health monitoring system that provides continuous remote monitoring of ECG together with automatic analysis and notification. The system consists of energy-efficient sensor nodes and a fog layer altogether taking advantage of IoT. The sensor nodes collect and wirelessly transmit ECG, respiration rate, and body temperature to a smart gateway which can be accessed by appropriate care-givers. In addition, the system can represent the collected data in useful ways, perform automatic decision making and provide many advanced services such as real-time notifications for immediate attention.
Tuan Nguyen Gia, Mingzhe Jiang, Victor K. Sarker, Amir-Mohammad Rahmani, Tomi Westerlund, Pasi Liljeberg, Hannu Tenhunen
IWCMC7
2017 Hierarchal Placement of Smart Mobile Access Points in Wireless Sensor Networks Using Fog Computing
abstract
Recent advances in computing and sensor technologies have facilitated the emergence of increasingly sophisticated and complex cyber-physical systems and wireless sensor networks. Moreover, integration of cyber-physical systems and wireless sensor networks with other contemporary technologies, such as unmanned aerial vehicles (i.e. drones) and fog computing, enables the creation of completely new smart solutions. By building upon the concept of a Smart Mobile Access Point (SMAP), which is a key element for a smart network, we propose a novel hierarchical placement strategy for SMAPs to improve scalability of SMAP based monitoring systems. SMAPs predict communication behavior based on information collected from the network, and select the best approach to support the network at any given time. In order to improve the network performance, they can autonomously change their positions. Therefore, placement of SMAPs has an important role in such systems. Initial placement of SMAPs is an NP problem. We solve it using a parallel implementation of the genetic algorithm with an efficient evaluation phase. The adopted hierarchical placement approach is scalable, it enables construction of arbitrarily large SMAP based systems.
Amin Majd, Golnaz Sahebi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen
PDP5
2017 Special issue on energy efficient multi-core and many-core systems, Part II
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum
J. Parallel Distributed Comput.4
2017 Performance/Reliability-Aware Resource Management for Many-Cores in Dark Silicon Era
abstract
Aggressive technology scaling has enabled the fabrication of many-core architectures while triggering challenges such as limited power budget and increased reliability issues, like aging phenomena. Dynamic power management and runtime mapping strategies can be utilized in such systems to achieve optimal performance while satisfying power constraints. However, lifetime reliability is generally neglected. We propose a novel lifetime reliability/performance-aware resource co-management approach for many-core architectures in the dark silicon era. The approach is based on a two-layered architecture, composed of a long-term runtime reliability controller and a short-term runtime mapping and resource management unit. The former evaluates the cores' aging status w.r.t. a target reference specified by the designer, and performs recovery actions on highly stressed cores by means of power capping. The aging status is utilized in runtime application mapping to maximize system performance while fulfilling reliability requirements and honoring the power budget. Experimental evaluation demonstrates the effectiveness of the proposed strategy, which outperforms most recent state-of-the-art contributions.
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen
IEEE Trans. Computers5
2017 Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient Applications
abstract
Power capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.6
2017 Reliability-Aware Runtime Power Management for Many-Core Systems in the Dark Silicon Era
abstract
Power management of networked many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates considering network characteristics at runtime to achieve better performance while honoring the peak power upper bound. On the other hand, power management has a direct effect on chip temperature, which is the main driver of the aging effects. Therefore, alongside performance fulfillment, the controlling mechanism must also consider the current cores' reliability in its actuator manipulation to enhance the overall system lifetime in the long term. In this paper, we propose a multiobjective dynamic power management technique that uses current power consumption and other network characteristics including the reliability of the cores as the feedback while utilizing fine-grained voltage and frequency scaling and per-core power gating as the actuators. In addition, disturbance rejecter and reliability balancer are designed to help the controller to better smooth power consumption in the short term and reliability in the long term, respectively. Simulations of dynamic workloads and mixed criticality application profiles show that our method not only is effective in honoring the power budget while considerably boosting the system throughput, but also increases the overall system lifetime by minimizing aging effects by means of power consumption balancing.
Amir-Mohammad Rahmani, M. H. Haghbayan, Antonio Miele, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.6
2016 An Approach for Smart Management of Big Data in the Fog Computing Context
abstract
In this paper, a new approach for tackling Big Data in Internet of Things (IoT) systems is presented. We approach the problem from the data perspective rather than only focusing on the computing platform. We design and develop a concept that we call "Smart Data". Taking advantage of a hierarchical fog computing system, we reshape the raw and passive form of the data generated by IoT sensors to intelligent and self-managed data cells that are able to evolve and become more meaningful information with reduced size. We believe that smart data will revolutionize the current perspective of data and will open many potential research directions to tackle emerging big data issues.
Farhoud Hosseinpour, Juha Plosila, Hannu Tenhunen
CloudCom3
2016 A lifetime-aware runtime mapping approach for many-core systems in the dark silicon era
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen
DATE5
2016 Approximation knob: power capping meets energy efficiency
abstract
Power Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen
ICCAD7
2016 End-to-end security scheme for mobility enabled healthcare Internet of Things
Sanaz Rahimi Moosavi, Tuan Nguyen Gia, Ethiopia Nigussie, Amir-Mohammad Rahmani, Seppo Virtanen, Hannu Tenhunen, Jouni Isoaho
Future Gener. Comput. Syst.6
2016 Special issue on energy efficient multi-core and many-core systems, Part I
Amir-Mohammad Rahmani, Pasi Liljeberg, José Luis Ayala, Hannu Tenhunen, Alexander V. Veidenbaum
J. Parallel Distributed Comput.4
2016 A Power-Aware Approach for Online Test Scheduling in Many-Core Architectures
abstract
Aggressive technology scaling triggers novel challenges to the design of multi-/many-core systems, such as limited power budget and increased reliability issues. Today's many-core systems employ dynamic power management and runtime mapping strategies trying to offer optimal performance while fulfilling power constraints. On the other hand, due to the reliability challenges, online testing techniques are becoming a necessity in current and near future technologies. However, state-of-the-art techniques are not aware of the other power/performance requirements. This paper proposes a power-aware non-intrusive online testing approach for many-core systems. The approach schedules software based self-test routines on the various cores during their idle periods, while honoring the power budget and limiting delays in the workload execution. A test criticality metric, based on a device aging model, is used to select cores to be tested at a time. Moreover, power and reliability issues related to the testing at different voltage and frequency levels are also handled. Extensive experimental results reveal that the proposed approach can i) efficiently test the cores within the available power budget causing a negligible performance penalty, ii) adapt the test frequency to the current cores' aging status, and iii) cover available voltage and frequency levels during the testing.
M. H. Haghbayan, Amir-Mohammad Rahmani, Antonio Miele, Mohammad Fattah, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen
IEEE Trans. Computers7
2016 Polymorphic Configuration Architecture for CGRAs
abstract
In the era of platforms hosting multiple applications with arbitrary reconfiguration requirements, static configuration architectures are neither optimal nor desirable. The static reconfiguration architectures either incur excessive overheads or cannot support advanced features (like time-sharing and runtime parallelism). As a solution to this problem, we present a polymorphic configuration architecture (PCA) that provides each application with a configuration infrastructure tailored to its needs.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Ahmed Hemani, Kolin Paul, Juha Plosila, Peeter Ellervee, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.7
2015 Utilization Prediction Aware VM Consolidation Approach for Green Cloud Computing
abstract
Dynamic Virtual Machine (VM) consolidation is one of the most promising solutions to reduce energy consumption and improve resource utilization in data centers. Since VM consolidation problem is strictly NP-hard, many heuristic algorithms have been proposed to tackle the problem. However, most of the existing works deal only with minimizing the number of hosts based on their current resource utilization and these works do not explore the future resource requirements. Therefore, unnecessary VM migrations are generated and the rate of Service Level Agreement (SLA) violations are increased in data centers. To address this problem, our VM consolidation method which is formulated as a bin-packing problem considers both the current and future utilization of resources. The future utilization of resources is accurately predicted using a k-nearest neighbor regression based model. In this paper, we investigate the effectiveness of VM and host resource utilization predictions in the VM consolidation task using real workload traces. The experimental results show that our approach provides substantial improvement over other heuristic algorithms in reducing energy consumption, number of VM migrations and number of SLA violations.
Fahimeh Farahnakian, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
CLOUD5
2015 Smart e-Health Gateway: Bringing intelligence to Internet-of-Things based ubiquitous healthcare systems
abstract
There have been significant advances in the field of Internet of Things (IoT) recently. At the same time there exists an ever-growing demand for ubiquitous healthcare systems to improve human health and well-being. In most of IoT-based patient monitoring systems, especially at smart homes or hospitals, there exists a bridging point (i.e., gateway) between a sensor network and the Internet which often just performs basic functions such as translating between the protocols used in the Internet and sensor networks. These gateways have beneficial knowledge and constructive control over both the sensor network and the data to be transmitted through the Internet. In this paper, we exploit the strategic position of such gateways to offer several higher-level services such as local storage, real-time local data processing, embedded data mining, etc., proposing thus a Smart e-Health Gateway. By taking responsibility for handling some burdens of the sensor network and a remote healthcare center, a Smart e-Health Gateway can cope with many challenges in ubiquitous healthcare systems such as energy efficiency, scalability, and reliability issues. A successful implementation of Smart e-Health Gateways enables massive deployment of ubiquitous health monitoring systems especially in clinical environments. We also present a case study of a Smart e-Health Gateway called UTGATE where some of the discussed higher-level features have been implemented. Our proof-of-concept design demonstrates an IoT-based health monitoring system with enhanced overall system energy efficiency, performance, interoperability, security, and reliability.
Amir-Mohammad Rahmani, Nanda Kumar Thanigaivelan, Tuan Nguyen Gia, Jose David Granados Vergara, Behailu Negash, Pasi Liljeberg, Hannu Tenhunen
CCNC7
2015 Power-aware online testing of manycore systems in the dark silicon era
M. H. Haghbayan, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Zainalabedin Navabi, Hannu Tenhunen
DATE7
2015 Dark silicon aware runtime mapping for many-core systems: A patterning approach
abstract
Limitation on power budget in many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. In such systems, an efficient run-time application mapping approach can considerably enhance resource utilization and mitigate the dark silicon phenomenon. In this paper, we propose a dark silicon aware runtime application mapping approach that patterns active cores alongside the inactive cores in order to evenly distribute power density across the chip. This approach leverages dark silicon to balance the temperature of active cores to provide higher power budget and better resource utilization, within a safe peak operating temperature. In contrast with exhaustive search based mapping approach, our agile heuristic approach has a negligible runtime overhead. Our patterning strategy yields a surplus power budget of up to 17% along with an improved throughput of up to 21% in comparison with other state-of-the-art run-time mapping strategies, while the surplus budget is as high as 40% compared to worst case scenarios.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
ICCD6
2015 Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approach
abstract
Power management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies.
Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ISLPED8
2015 A Low-Overhead, Fully-Distributed, Guaranteed-Delivery Routing Algorithm for Faulty Network-on-Chips
abstract
This paper introduces a new, practical routing algorithm, Maze-routing, to tolerate faults in network-on-chips. The algorithm is the first to provide all of the following properties at the same time: 1) fully-distributed with no centralized component, 2) guaranteed delivery (it guarantees to deliver packets when a path exists between nodes, or otherwise indicate that destination is unreachable, while being deadlock and livelock free), 3) low area cost, 4) low reconfiguration overhead upon a fault. To achieve all these properties, we propose Maze-routing, a new variant of face routing in on-chip networks and make use of deflections in routing. Our evaluations show that Maze-routing has 16X less area overhead than other algorithms that provide guaranteed delivery. Our Maze-routing algorithm is also high performance: for example, when up to 5 links are broken, it provides 50% higher saturation throughput compared to the state-of-the-art.
Mohammad Fattah, Antti Airola, Rachata Ausavarungnirun, Nima Mirzaei, Pasi Liljeberg, Juha Plosila, Siamak Mohammadi, Tapio Pahikkala, Onur Mutlu, Hannu Tenhunen
NOCS10
2015 MapPro: Proactive Runtime Mapping for Dynamic Workloads by Quantifying Ripple Effect of Applications on Networks-on-Chip
abstract
Increasing dynamic workloads running on NoC-based many-core systems necessitates efficient runtime mapping strategies. With an unpredictable nature of application profiles, selecting a rational region to map an incoming application is an NP-hard problem in view of minimizing congestion and maximizing performance. In this paper, we propose a proactive region selection strategy which prioritizes nodes that offer lower congestion and dispersion. Our proposed strategy, MapPro, quantitatively represents the propagated impact of spatial availability and dispersion on the network with every new mapped application. This allows us to identify a suitable region to accommodate an incoming application that results in minimal congestion and dispersion. We cluster the network into squares of different radii to suit applications of different sizes and proactively select a suitable square for a new application, eliminating the overhead caused with typical reactive mapping approaches. We evaluated our proposed strategy over different traffic patterns and observed gains of up to 41% in energy efficiency, 28% in congestion and 21% dispersion when compared to the state-of-the-art region selection methods.
M. H. Haghbayan, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
NOCS6
2015 FIST: A Framework to Interleave Spiking Neural Networks on CGRAs
abstract
Coarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern embedded applications. In many application domains (e.g. robotics and cognitive embedded systems), the CGRAs are required to simultaneously host processing (e.g. Audio/video acquisition) and estimation (e.g. audio/video/image recognition) tasks. Recent works have revealed that the efficiency and scalability of the estimation algorithms can be significantly improved by using neural networks. However, existing CGRAs commonly employ homogeneous processing resources for both the tasks. To realize the best of both the worlds (conventional processing and neural networks), we present FIST. FIST allows the processing elements and the network to dynamically morph into either conventional CGRA or a neural network, depending on the hosted application. We have chosen the DRRA as a vehicle to study the feasibility and overheads of our approach. Synthesis results reveal that the proposed enhancements incur negligible overheads (4.4% area and 9.1% power) compared to the original DRRA cell.
Tuan Ngyen, Syed M. A. H. Jafri, Masoud Daneshtalab, Ahmed Hemani, Sergei Dytckov, Juha Plosila, Hannu Tenhunen
PDP7
2015 PDNOC: Partially diagonal network-on-chip for high efficiency multicore systems
abstract
Summary With the constantly increasing of number of cores in multicore processors, more emphasis should be paid to the on‐chip interconnect. Performance and power consumption of an on‐chip interconnect are directly affected by the network topology. Researchers have proposed various topologies to optimize these metrics. The efficiency can also be optimized by proper mapping of applications. Therefore in this paper, we propose a novel partially diagonal network‐on‐chip (PDNOC) design that takes advantage of both heterogeneous network topology and congestion‐aware application mapping. We analyse the partially diagonal network in terms of interconnect structure, area usage, power consumption, routing algorithm and implementation complexity. The key insight that enables the PDNOC is that most communication patterns in real‐world applications are hot‐spot and bursty. We implement a full system simulation environment using SPLASH‐2 benchmarks. Performance metrics of standard mesh, concentrated mesh, full diagonal mesh and four types of the proposed PDNOC are measured in terms of network latency, application execution time and energy delay product. Evaluation results show that on average, the proposed PDNOC designs provide up to 36% improvement in execution time over concentrated mesh, and 3.6× better energy delay product over fully connected diagonal network. PDNOC design with two adjacent PD networks is a better candidate for higher efficiency, while four PD networks provide better performance. Copyright © 2014 John Wiley & Sons, Ltd.
Thomas Canhao Xu, Ville Leppänen, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
Concurr. Comput. Pract. Exp.5
2015 Special Issue on Emerging Many-Core Systems for Exascale Computing
abstract
No abstract available.
Masoud Daneshtalab, Farhad Mehdipour, Zhiyi Yu, Hannu Tenhunen
ACM J. Emerg. Technol. Comput. Syst.4
2015 Architecture and Implementation of Dynamic Parallelism, Voltage and Frequency Scaling (PVFS) on CGRAs
abstract
In the era of platforms hosting multiple applications with arbitrary performance requirements, providing a worst-case platform-wide voltage/frequency operating point is neither optimal nor desirable. As a solution to this problem, designs commonly employ dynamic voltage and frequency scaling (DVFS). DVFS promises significant energy and power reductions by providing each application with the operating point (and hence the performance) tailored to its needs. To further enhance the optimization potential, recent works interleave dynamic parallelism with conventional DVFS. The induced parallelism results in performance gains that allow an application to lower its operating point even further (thereby saving energy and power consumption). However, the existing works employ costly dedicated hardware (for synchronization) and rely solely on greedy algorithms to make parallelism decisions. To efficiently integrate parallelism with DVFS, compared to state-of-the-art, we exploit the reconfiguration (to reduce DVFS synchronization overheads) and enhance the intelligence of the greedy algorithm (to make optimal parallelism decisions). Specifically, our solution relies on dynamically reconfigurable isolation cells and an autonomous parallelism, voltage, and frequency selection algorithm. The dynamically reconfigurable isolation cells reduce the area overheads of DVFS circuitry by configuring the existing resources to provide synchronization. The autonomous parallelism, voltage, and frequency selection algorithm ensures high power efficiency by combining parallelism with DVFS. It selects that parallelism, voltage, and frequency trio which consumes minimum power to meet the deadlines on available resources. Synthesis and simulation results using various applications/algorithms (WLAN, MPEG4, FFT, FIR, matrix multiplication) show that our solution promises significant reduction in area and power consumption (23% and 51% ) compared to state-of-the-art.
Syed M. A. H. Jafri, Ozan Ozbag, Nasim Farahini, Kolin Paul, Ahmed Hemani, Juha Plosila, Hannu Tenhunen
ACM J. Emerg. Technol. Comput. Syst.7
2015 Using Ant Colony System to Consolidate VMs for Green Cloud Computing
abstract
High energy consumption of cloud data centers is a matter of great concern. Dynamic consolidation of Virtual Machines (VMs) presents a significant opportunity to save energy in data centers. A VM consolidation approach uses live migration of VMs so that some of the under-loaded Physical Machines (PMs) can be switched-off or put into a low-power mode. On the other hand, achieving the desired level of Quality of Service (QoS) between cloud providers and their users is critical. Therefore, the main challenge is to reduce energy consumption of data centers while satisfying QoS requirements. In this paper, we present a distributed system architecture to perform dynamic VM consolidation to reduce energy consumption of cloud data centers while maintaining the desired QoS. Since the VM consolidation problem is strictly NP-hard, we use an online optimization metaheuristic algorithm called Ant Colony System (ACS). The proposed ACS-based VM Consolidation (ACS-VMC) approach finds a near-optimal solution based on a specified objective function. Experimental results on real workload traces show that ACS-VMC reduces energy consumption while maintaining the required performance levels in a cloud data center. It outperforms existing VM consolidation approaches in terms of energy consumption, number of VM migrations, and QoS requirements concerning performance.
Fahimeh Farahnakian, Adnan Ashraf, Tapio Pahikkala, Pasi Liljeberg, Juha Plosila, Ivan Porres, Hannu Tenhunen
IEEE Trans. Serv. Comput.7
2014 Energy-Aware Dynamic VM Consolidation in Cloud Data Centers Using Ant Colony System
abstract
As the scale of a cloud data center becomes larger and larger, the energy consumption of the data center also grows rapidly. Dynamic consolidation of Virtual Machines (VMs) presents a significant opportunity to save energy by turning off unused Physical Machines (PMs) in data centers. In this paper, we present a distributed controller to perform dynamic VM consolidation to improve the resource utilizations of PMs and to reduce their energy consumption. Moreover, we use the ant colony system to find a near-optimal VM placement solution based on the specified objective function. Experimental results on the real workload traces from more than a thousand PlanetLab VMs show that the proposed approach reduces energy consumption and maintains required performance levels in a large-scale data center.
Fahimeh Farahnakian, Adnan Ashraf, Pasi Liljeberg, Tapio Pahikkala, Juha Plosila, Ivan Porres, Hannu Tenhunen
IEEE CLOUD7
2014 Adjustable contiguity of run-time task allocation in networked many-core systems
abstract
In this paper, we propose a run-time mapping algorithm, CASqA, for networked many-core systems. In this algorithm, the level of contiguousness of the allocated processors (α) can be adjusted in a fine-grained fashion. A strictly contiguous allocation (α = 0) decreases the latency and power dissipation of the network and improves the applications execution time. However, it limits the achievable throughput and increases the turnaround time of the applications. As a result, recent works consider non-contiguous allocation (α = 1) to improve the throughput traded off against applications execution time and network metrics. In contradiction, our experiments show that a higher throughput (by 3%) with improved network performance can be achieved when using intermediate α values. More precisely, up to 35% drop in the network costs can be gained by adjusting the level of contiguity compared to non-contiguous cases, while the achieved throughput is kept constant. Moreover, CASqA provides at least 32% energy saving in the network compared to other works.
Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
ASP-DAC4
2014 Hierarchical VM Management Architecture for Cloud Data Centers
abstract
Efficient energy use has become a critical issue for designing and managing of cloud data centers. Virtualization is a key technology for reducing energy cost and improving resource utilization in data centers. One of the challenges faced by virtualized data centers is to decide how to pack VMs on the least number of physical machines. This paper presents a VM management framework which is based on a multi-agent system to minimize energy consumption and Service Level Agreement (SLA) violations. The proposed agents are arranged in a three level hierarchical structure to perform VM assignment, VM placement and VM consolidation in a data center efficiently. Experimental results demonstrate that the framework achieves high quality solution in spite of its simplicity and scalability.
Fahimeh Farahnakian, Pasi Liljeberg, Tapio Pahikkala, Juha Plosila, Hannu Tenhunen
CloudCom5
2014 SHiFA: System-Level Hierarchy in Run-Time Fault-Aware Management of Many-Core Systems
abstract
A system-level approach to fault-aware resource management of many-core systems is proposed. The proposed approach, called SHiFA, is able to tolerate run-time faults at system level without any hardware overhead. In contrast to the existing system-level methods, network resources are also considered to be potentially faulty. Accordingly, applications are mapped onto healthy nodes of the system at run-time such that their interaction will not require the use of faulty elements. By utilizing the simple routing approach, results show 100% utilizability of PEs and 99.41% of successful mapping when up to 8 links are broken. SHiFA design is based on distributed operating systems, such that it is kept scalable for future many-core systems. A significant improvement in scalability properties is observed compared to the state-of-the-art distributed approaches.
Mohammad Fattah, Maurizio Palesi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DAC5
2014 Online testing of many-core systems in the Dark Silicon era
abstract
As the dark silicon era is about to embrace, it is not anymore possible to attain commensurate performance benefits by increasing the number of transistors due to thermal design power. Dark Silicon issue stresses that a fraction of silicon chip being able to switch in full frequency is dropping and designers will soon face the growing underutilization inherent in future technologies. On the other hand, by reducing the transistor size, susceptibility to internal defects drastically increases and large ranges of defects such as aging or transient faults will be shown up more frequently. In this paper, we propose an online test scheduling algorithm using software based self-test for dark silicon era to test dark cores while considering thermal design power of the system. As the dark area of the system is dynamic and reshapes at a runtime, the tested cores can be used by other applications in the near future. Empirical results show the effectiveness of the proposed algorithm in terms of applicability and fault coverage with a negligible negative impact on the system throughput.
M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DDECS5
2014 Parameterized AES-Based Crypto Processor for FPGAs
abstract
In this paper, we propose a parameterized crypto co-processor based on Advanced Encryption Standard (AES). This parameterized AES module is combined with a 32-bit general purpose 5-stage pipelined MIPS processor. The AES module used in this paper is fully pipelined. The processor fetches an instruction from the instruction memory and sends it to the decode stage. If the instruction is the crypto instruction it is pushed into the AES module during the decode stage. However if the instruction belongs to the MIPS processor, the remaining cycles will be completed on the MIPS processor. The parameterized AES module has different latencies on different rounds of AES according to the application requirements. The effects of different number of rounds on latency, memory, and area are studied and reported.
Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen, Sergei Dytckov, Giovanni Beltrame
DSD5
2014 Efficient STDP Micro-Architecture for Silicon Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are the closest approach to biological neurons in comparison with conventional artificial neural networks (ANN). SNNs are composed of neurons and synapses which are interconnected with a complex pattern. As communication in such massively parallel computational systems is getting critical, the network-on-chip (NoC) becomes a promising solution to provide a scalable and robust interconnection fabric. However, using NoC for large-scale SNNs arises a trade-off between scalability, throughput, neuron/router ratio (cluster size), and area overhead. In this paper, we tackle the trade-off using a clustering approach and try to optimize the synaptic resource utilization. An optimal cluster size can provide the lowest area overhead and power consumption. For the learning purposes, a phenomenon known as spike-timing-dependent plasticity (STDP) is utilized. The micro-architectures of the network, clusters, and the computational neurons are also described. The presented approach suggests a promising solution of integrating NoCs and STDP-based SNNs for the optimal performance based on the underlying application.
Sergei Dytckov, Masoud Daneshtalab, Masoumeh Ebrahimi, Hassan Anwar, Juha Plosila, Hannu Tenhunen
DSD6
2014 Morphable Compression Architecture for Efficient Configuration in CGRAs
abstract
Today, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads (up to 50% area of the overall platform). As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties (i.e. compression ratio and decoding time), and is therefore best suited for a particular class of applications (and situation). However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. The proposed architecture allows each application to enjoy a separate compression/decompression hierarchy (consisting of various types and implementations of hardware/software decoders) tailored to its needs. Thereby, our solution offers minimal memory while meeting the required configuration deadlines. Simulation results, using different applications (FFT, Matrix multiplication, and WLAN), reveal that the choice of compression hierarchy has a significant impact on compression ratio (from configware replication to 52%) and configuration cycles (from 33 nsec to 1.5 secs) for the tested applications. Synthesis results reveal that introducing adaptivity incurs negligible additional overheads (1%) compared to the overall platform area.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen
DSD8
2014 Customizable Compression Architecture for Efficient Configuration in CGRAs
abstract
Today, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads. As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties, and is therefore best suited for a particular class of applications. However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen
FCCM8
2014 TransPar: Transformation based dynamic Parallelism for low power CGRAs
abstract
Coarse Grained Reconfigurable Architectures (CGRAs) are emerging as enabling platforms to meet the high performance demanded by modern applications (e.g. 4G, CDMA, etc.). Recently proposed CGRAs offer runtime parallelism to reduce energy consumption (by lowering voltage/frequency). To implement the runtime parallelism, CGRAs commonly store multiple compile-time generated implementations of an application (with different degree of parallelism) and select the optimal version at runtime. However, the compile-time binding incurs excessive configuration memory overheads and/or is unable to parallelize an application even when sufficient resources are available. As a solution to this problem, we propose Transformation based dynamic Parallelism (TransPar). TransPar stores only a single implementation and applies a series for transformations to generate the bitstream for the parallel version. In addition, it also allows to displace and/or rotate an application to parallelize in resource constrained scenarios. By storing only a single implementation, TransPar offers significant reductions in configuration memory requirements (up to 73% for the tested applications), compared to state of the art compaction techniques. Simulation and synthesis results, using real applications, reveal that the additional flexibility allows up to 33% energy reduction compared to static memory based parallelism techniques. Gate level analysis reveals that TransPar incurs negligible silicon (0.2% of the platform) and timing (6 additional cycles per application) penalty.
Syed M. A. H. Jafri, Guilermo Serrano, Masoud Daneshtalab, Naeem Abbas, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen
FPL8
2014 Dark silicon aware power management for manycore systems under dynamic workloads
abstract
Dark Silicon denotes the phenomenon that, due to thermal and power constraints, the fraction of transistors that can operate at full frequency is decreasing with each technology generation. We propose a PID (Proportional Integral Derivative) controller based dynamic power management method that considers an upper bound on power consumption (called the Thermal Design Power (TDP)). To avoid violation of the TDP constraint for manycore systems running highly dynamic workloads, it provides fine-grained DVFS (Dynamic Voltage and Frequency Scaling) including near-threshold operation. In addition, the method distinguishes applications with hard Real-Time, soft Real-Time and no Real-Time constraints and treats them with appropriate priorities. In simulations with dynamic workloads mixed-critical application profiles, we show that the method is effective in honoring the TDP bound and it can boost system throughput by over 43% compared to a naive TDP scheduling policy.
M. H. Haghbayan, Amir-Mohammad Rahmani, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ICCD7
2014 Integration of AES on Heterogeneous Many-Core System
abstract
Increasing in the transistor density in a single chip makes it possible for many-core systems to utilize design space for implementing complex embedded systems. In this paper, we propose an architecture for heterogeneous many-core system to integrate block cipher algorithm which is based on Advanced Encryption Standard (AES). In order to implement AES as a crypto-core along with heterogeneous many-core system two different approaches are proposed. In the first approach, the platform is composed of an AES module, a crypto-core, a network interface, and an internal memory which are managed through a controller. In the second approach, the platform is composed of a Direct Memory Access (DMA), network interface, an internal memory, and a microprocessor in which the AES module is integrated as a crypto-core. Both approaches have been analyzed and compared in terms of area overhead and performance.
Hassan Anwar, Masoud Daneshtalab, Masoumeh Ebrahimi, Marco Ramírez 0001, Juha Plosila, Hannu Tenhunen
PDP6
2014 Mixed-Criticality Run-Time Task Mapping for NoC-Based Many-Core Systems
abstract
Contiguous processor allocation improves both the network and the application performance, by decreasing the congestion probability among communication of different applications. Consequently, the average, standard deviation and worst-case latency of the network is decreased significantly. This makes the contiguous allocation a good solution for time-critical applications with bounded deadlines. On the other hand, non-contiguous allocation will increase the system throughput significantly. Isolated nodes are utilized and more applications can finish their job in a time unit. However, this will lead to poor network metrics, unsuitable for real-time applications. In this work, we combine these two approaches in order to manage workloads with mixed-critical characteristics. Real-time applications are mapped contiguously, while non-critical applications are allowed to get dispersed over the available system nodes. Results show over 50% improvement in worst-case latency and 100 times improvement in deadline misses.
Mohammad Fattah, Amir-Mohammad Rahmani, Thomas Canhao Xu, Anil Kanduri, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP7
2014 Multi Rectangle Modeling Approach for Application Mapping on a Many-Core System
abstract
The importance of first node selection in run-time resource management is shown in our previous work, SHiC. It is desired in SHiC to find the optimum node in an agile manner. Accordingly, the current mapping picture of the system is simplified to SHiC by modeling each application as a rectangle of occupied nodes. However, the algorithm performance can be influenced significantly with the rectangle model of each application. In this work, we introduce a precise description of our new accurate rectangle modeling algorithm. Moreover, we show that it is not sufficient to always model an application with only one rectangle, as dispersion of the allocated nodes is an irrepressible phenomenon. Accordingly, our algorithm enables to model a mapped application with several rectangles by tuning the model accuracy against the algorithm complexity. However, the algorithm is not in the critical path of the applications executions. Our results shows up to 5% reduction in power dissipation of the network.
Igor Tcarenko, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP5
2014 Path-Based Partitioning Methods for 3D Networks-on-Chip with Minimal Adaptive Routing
abstract
Combining the benefits of 3D ICs and Networks-on-Chip (NoCs) schemes provides a significant performance gain in Chip Multiprocessors (CMPs) architectures. As multicast communication is commonly used in cache coherence protocols for CMPs and in various parallel applications, the performance of these systems can be significantly improved if multicast operations are supported at the hardware level. In this paper, we present several partitioning methods for the path-based multicast approach in 3D mesh-based NoCs, each with different levels of efficiency. In addition, we develop novel analytical models for unicast and multicast traffic to explore the efficiency of each approach. In order to distribute the unicast and multicast traffic more efficiently over the network, we propose the Minimal and Adaptive Routing (MAR) algorithm for the presented partitioning methods. The analytical and experimental results show that an advantageous method named Recursive Partitioning (RP) outperforms the other approaches. RP recursively partitions the network until all partitions contain a comparable number of switches and thus the multicast traffic is equally distributed among several subsets and the network latency is considerably decreased. The simulation results reveal that the RP method can achieve performance improvement across all workloads while performance can be further improved by utilizing the MAR algorithm. Nineteen percent average and 42 percent maximum latency reduction are obtained on SPLASH-2 and PARSEC benchmarks running on a 64-core CMP.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, José Flich, Hannu Tenhunen
IEEE Trans. Computers6
2014 High-Performance and Fault-Tolerant 3D NoC-Bus Hybrid Architecture Using ARB-NET-Based Adaptive Monitoring Platform
abstract
The emerging three-dimensional integrated circuits (3D-ICs) achieve greater device integration and enhanced system performance at lower cost and reduced area footprint, thereby offering higher order of connectivity and greater design choices and possibilities. To exploit the intrinsic capability of reduced communication distances in 3D-ICs, three-dimensional NoC-bus hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored possibility to implement an efficient system-wide monitoring network. In this paper, an efficient three-dimensional NoC architecture is proposed which is optimized for system performance, power consumption, and reliability. The mechanism benefits from a congestion-aware and bus failure-tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we have integrated a low-cost monitoring platform on top of the three-dimensional NoC-Bus Hybrid mesh architecture that can be efficiently used for various system management purposes such as traffic monitoring, fault tolerance, and thermal management. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware and interlayer fault-tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of information generated within bus arbiters. Compared to recently proposed stacked mesh three-dimensional NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power, performance, and reliability improvements with a negligible hardware overhead.
Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
IEEE Trans. Computers6
2014 Special section on advances in methods for adaptive multicore systems
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Supercomput.4
2013 Private configuration environments (PCE) for efficient reconfiguration, in CGRAs
abstract
In this paper, we propose a polymorphic configuration architecture, that can be tailored to efficiently support reconfiguration needs of the applications at runtime. Today, CGRAs host multiple applications, running simultaneously on a single platform. Novel CGRAs allow each application to exploit late binding and time sharing for enhancing the power and area efficiency. These features require frequent reconfigurations, making reconfiguration time a bottleneck for time critical applications. Existing solutions to this problem either employ powerful configuration architectures or hide configuration latency (using configuration caching). However, both these methods incur significant costs when designed for worst-case reconfiguration needs. As an alternative to worst-case dedicated configuration mechanism, we exploit reconfiguration to provide each application its private configuration environment (PCE). PCE relies on a morphable configuration infrastructure, a distributed memory sub-system, and a set of PCE controllers. The PCE controllers customize the morphable configuration infrastructure and reserve portion of the a distributed memory sub-system, to act as a context memory for each application, separately. Thereby, each application enjoys its own configuration environment which is optimal in terms of configuration speed, memory requirements and energy. Simulation results using representative applications (WLAN and Matrix Multiplication) showed that PCE offers up to 58% reduction in memory requirements, compared to dedicated, worst case configuration architecture. Synthesis results show that the morphable reconfiguration architecture incurs negligible overheads ( 3% area and 4% power compared of a single processing element).
Muhammad Adeel Tajammul, Syed M. A. H. Jafri, Ahmed Hemani, Juha Plosila, Hannu Tenhunen
ASAP5
2013 CARS: congestion-aware request scheduler for network interfaces in NoC-based manycore systems
abstract
Network congestion is a critical issue of memory parallelism in network-based manycore systems where multiple memories can be accessed simultaneously. Therefore, a congestion-aware method is necessitated to deal with the network congestion. In this paper, we present a streamlined method in order to reduce the network congestion. The idea is to use the global congestion information as a metric in network interfaces to reduce the congestion level of highly congested areas. Network interfaces connected to memory modules are equipped with an adaptive scheduler using the global congestion information to reduce additional traffic to congested areas. Experimental results with synthetic test cases demonstrate that the on-chip network utilizing the proposed adaptive scheduler presents up to 23% improvement in average latency.
Masoud Daneshtalab, Masoumeh Ebrahimi, Juha Plosila, Hannu Tenhunen
DATE4
2013 Energy-Aware Fault-Tolerant CGRAs Addressing Application with Different Reliability Needs
abstract
In this paper, we propose a polymorphic fault tolerant architecture that can be tailored to efficiently support the reliability needs of multiple applications at run-time. Today, coarse-grained reconfigurable architectures (CGRAs) host multiple applications with potentially different reliability needs. Providing platform-wide worst-case (maximum) protection to all the applications is neither optimal nor desirable. To reduce the fault-tolerance overhead, adaptive fault-tolerance strategies have been proposed. The proposed techniques access the reliability requirements of each application and adjust the fault-tolerance intensity (and hence overhead), accordingly. However, existing flexible reliability schemes only allow to shift between different levels of modular redundancy (duplication, triplication, etc.) and deal with only a single class of faults (e.g. soft errors). To complement these strategies, we propose energy-aware fault-tolerance that, in addition to modular redundancy, can also provide low cost, sub-modular (e.g. residue mod 3) redundancy, to cater both permanent and temporary faults. Our solution relies on an agent based control layer and a configurable fault-tolerance data path. The control layer identifies the application class and configures the data path to provide the needed reliability. Simulation results using a few selected algorithms (FFT, matrix multiplication, and FIR filter) showed that the proposed method provides flexible protection with energy overhead ranging from 3.125% to 107% for different reliability levels. Synthesis results have confirmed that the proposed architecture significantly reduces the area overhead for self-checking (59.1%) and fault tolerant (7.1%) versions, compared to the state of the art adaptive reliability techniques.
Syed M. A. H. Jafri, Stanislaw J. Piestrak, Kolin Paul, Ahmed Hemani, Juha Plosila, Hannu Tenhunen
DSD6
2013 Minimal-path fault-tolerant approach using connection-retaining structure in Networks-on-Chip
abstract
There are many fault-tolerant approaches presented both in off-chip and on-chip networks. Regardless of all varieties, there has always been a common assumption between them. Most of all known fault-tolerant methods are based on rerouting packets around faults. Rerouting might take place through nonminimal paths which affect the performance significantly not only by taking longer paths but also by creating hotspot around a fault. In this paper, we present a fault-tolerant approach based on using the shortest paths. This method maintains the performance of Networks-on-Chip in the presence of faults. To avoid using non-minimal paths, the router architecture is slightly modified. In the new form of architecture, there is an ability to connect the horizontal and orthogonal links of a faulty router such that healthy routers are kept connected to each other. Based on this architecture, a fault-tolerant routing algorithm is presented which is obviously much simpler than traditional fault-tolerant routing algorithms. According to this algorithm, only the shortest paths are used by packets in the presence of fault. This results retains the performance of NoCs in faulty situations. This algorithm is highly reliable, for an instance, the reliability is more than 99.5% when there are six faulty routers in an 8×8 mesh network.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen
NOCS4
2013 DyXYZ: Fully Adaptive Routing Algorithm for 3D NoCs
abstract
Traditional methods in 3D NoCs simply use a deterministic routing algorithm to deliver packets from a source to a destination node. However, deterministic methods are unable to distribute the traffic load over the network, which results in degrading the performance. In this paper, we present a fully adaptive routing algorithm for 3D NoCs, named DyXYZ. In DyXYZ, the congestion information at the input buffer of the neighboring routers is used as congestion metric to select among the output channels. This algorithm is proven to be deadlock free by using 4, 4, and 2 virtual channels along the X, Y, and Z dimensions, respectively.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen
PDP6
2013 Enhancing Performance of 3D Interconnection Networks using Efficient Multicast Communication Protocol
abstract
Three-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. They also provide greater design flexibility by allowing heterogeneous integration. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement efficient multicast routings for 3D networks-on-chip. In this paper, we propose a novel multicast partitioning and routing strategy for the 3D NoC-Bus Hybrid mesh architectures to enhance the overall system performance and reduce the power consumption. The proposed architecture exploits the beneficial attribute of a single-hop (bus-based) interlayer communication of the 3D stacked mesh architecture to provide high-performance hardware multicast support. To this end, a customized partitioning method and an efficient routing algorithm are presented to reduce the average hop count and latency of the network. Compared to the recently proposed 3D NoC architectures being capable of supporting hardware multicasting, our extensive simulations with different traffic profiles reveal that our architecture using the proposed multicast routing strategy can help achieve significant performance improvements.
Sanaz Rahimi Moosavi, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP5
2013 Partial Virtual Channel Sharing: A Generic Methodology to Enhance Resource Management and Fault Tolerance in Networks-on-Chip
Khalid Latif 0002, Amir-Mohammad Rahmani, Ethiopia Nigussie, Tiberiu Seceleanu, Martin Radetzki, Hannu Tenhunen
J. Electron. Test.6
2013 Cluster-based topologies for 3D Networks-on-Chip using advanced inter-layer bus architecture
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Comput. Syst. Sci.5
2013 Developing a power-efficient and low-cost 3D NoC using smart GALS-based vertical channels
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Comput. Syst. Sci.4
2013 A systematic reordering mechanism for on-chip networks using efficient congestion-aware method
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
J. Syst. Archit.5
2013 Fuzzy-based Adaptive Routing Algorithm for Networks-on-Chip
Masoumeh Ebrahimi, Hannu Tenhunen, Masoud Dehyadegari
J. Syst. Archit.2
2013 A development and verification framework for the SegBus platform
Moazzam Fareed Niazi, Tiberiu Seceleanu, Hannu Tenhunen
J. Syst. Archit.3
2013 Optimal placement of vertical connections in 3D Network-on-Chip
Thomas Canhao Xu, Gert Schley, Pasi Liljeberg, Martin Radetzki, Juha Plosila, Hannu Tenhunen
J. Syst. Archit.6
2013 A Hybrid Low Power Biopatch for Body Surface Potential Measurement
abstract
This paper presents a wearable biopatch prototype for body surface potential measurement. It combines three key technologies, including mixed-signal system on chip (SoC) technology, inkjet printing technology, and anisotropic conductive adhesive (ACA) bonding technology. An integral part of the biopatch is a low-power low-noise SoC. The SoC contains a tunable analog front end, a successive approximation register analog-to-digital converter, and a reconfigurable digital controller. The electrodes, interconnections, and interposer are implemented by inkjet-printing the silver ink precisely on a flexible substrate. The reliability of printed traces is evaluated by static bending tests. ACA is used to attach the SoC to the printed structures and form the flexible hybrid system. The biopatch prototype is light and thin with a physical size of 16 cm × 16 cm. Measurement results show that low-noise concurrent electrocardiogram signals from eight chest points have been successfully recorded using the implemented biopatch.
Geng Yang 0003, Jian Chen 0001, Jia Mao, Hannu Tenhunen, Lirong Zheng 0001
IEEE J. Biomed. Health Informatics5
2012 ARB-NET: A novel adaptive monitoring platform for stacked mesh 3D NoC architectures
abstract
The emerging three-dimensional integrated circuits (3D ICs) offer a promising solution to mitigate the barriers of interconnect scaling in modern systems. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. Besides its various advantages in terms of area, power consumption, and performance, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes. The proposed generic monitoring platform called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is presented taking advantage from viable information generated within bus arbiters. Our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help achieving significant power and performance improvements compared to recently proposed stacked mesh 3D NoCs.
Amir-Mohammad Rahmani, Khalid Latif 0002, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
ASP-DAC6
2012 A Cluster-Based Core Protection Technique for Networks-on-Chip
abstract
Partial Virtual channel Sharing (PVS) architecture has been proposed to enhance the performance of Networks-on-Chip (NoC) based systems. In this paper, a cluster based processing core protection technique for NoC systems using PVS approach is presented. In case of network level faults, the processing core of faulty node can use any other router in the cluster for transmission or reception of data packets with proposed architecture. Simulation results show significant reduction in average packet latency at the expense of negligible area overhead.
Khalid Latif 0002, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen, Tiberiu Seceleanu
COMPSAC4
2012 CATRA- congestion aware trapezoid-based routing algorithm for on-chip networks
abstract
Congestion occurs frequently in Networks-on-Chip when the packets demands exceed the capacity of network resources. Congestion-aware routing algorithms can greatly improve the network performance by balancing the traffic load in adaptive routing. Commonly, these algorithms either rely on purely local congestion information or take into account the congestion conditions of several nodes even though their statuses might be out-dated for the source node, because of dynamically changing congestion conditions. In this paper, we propose a method to utilize both local and non-local network information to determine the optimal path to forward a packet. The non-local information is gathered from the nodes that not only are more likely to be chosen as intermediate nodes in the routing path but also provide up-to-date information to a given node. Moreover, to collect and deliver the non-local information, a distributed propagation system is presented.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DATE5
2012 A multi-parameter bio-electric ASIC sensor with integrated 2-wire data transmission protocol for wearable healthcare system
abstract
This paper presents a fully integrated application specific integrated circuit (ASIC) sensor for the recording of multiple bio-electric signals. It consists of an analog front-end circuit with tunable bandwidth and programmable gain, a 6-input 8-bit successive approximation register analog to digital converter (SAR ADC), and a reconfigurable digital core. The ASIC is fabricated in a 0.18-µm 1P6M CMOS technology, occupies an area of 1.5 × 3.0 mm2, and totally consumes a current of 16.7 µA from a 1.2 V supply. Incorporated with the ASIC, an Intelligent Electrode can be dynamically configured for on-site measurement of different bio-signals. A 2-wire data transmission protocol is also integrated on chip. It enables the serial connection over a group of Intelligent Electrodes, thus minimizes the number of connecting cables. A wearable healthcare system is built upon a printed Active Cable and a scalable number of Intelligent Electrodes. The system allows synchronous processing of maximum 14-channel bio-signals. The ASIC performance has been successfully verified in in-vivo bio-electric recording experiments.
Geng Yang 0003, Jian Chen 0001, Fredrik Jonsson, Hannu Tenhunen, Lirong Zheng 0001
DATE4
2012 HLS-DoNoC: High-level simulator for dynamically organizational NoCs
abstract
A high-level simulator is presented for the design and analysis of dynamically organizational Networks-on-Chip (DoNoCs). The DoNoC is able to organize statically or dynamically different network nodes for run-time coarse and fine grained reconfiguration, in particular power management. As an important step in the design flow, a simulator for early-stage design exploration is the focus of the paper. Built upon classic wormhole-based NoC architecture, the simulator is capable of experimenting diverse run-time monitoring and reconfiguration methods. In particular, dynamic clusterization can be performed with inter-cluster interfaces properly configured at the run-time. The simulator is flit-level accurate, trace-driven, and easy-to-reconfigure. It supports both synchronous and ratiochronous timing, and can provide the communication performance and power/energy consumption. The paper demonstrates the usage of the simulator in the design of various cluster-based power management schemes.
Liang Guang, Ethiopia Nigussie, Juha Plosila, Jouni Isoaho, Hannu Tenhunen
DDECS5
2012 Designing a High Performance and Reliable Networks-on-Chip Using Network Interface Assisted Routing Strategy
abstract
Partial Virtual channel Sharing (PVS) architecture has been proposed to enhance the performance of Networks-on-Chip (NoC) based systems. In this paper, we present an efficient and reliable Network Interface (NI) assisted routing strategy for NoC using PVS architecture. For this purpose, NoC system is divided into clusters. Each cluster is a group of two nodes comprising Processing Elements (PE), switches, links, etc. Each PE in a cluster can inject data to the network through a router, which is closer to the destination. This helps to reduce the network load by reducing the average hop count of the network. The proposed architecture can recover the PE disconnected from the network due to network level faults by allowing the PE to transmit and receive the packets through the other router in the cluster. 5̅×6 crossbar is used for the proposed architecture which requires one more 5×1 multiplexer without increasing the critical path delay of the router as compared to the 5×5 crossbar. The proposed router has been simulated for uniform and negative exponential distribution (NED) traffic patterns. The simulation results show the significant reduction in average packet latency at the expense of negligible area overhead.
Khalid Latif 0002, Amir-Mohammad Rahmani, Tiberiu Seceleanu, Hannu Tenhunen
DSD4
2012 MAFA: Adaptive Fault-Tolerant Routing Algorithm for Networks-on-Chip
abstract
While Networks-on-Chip have been increasing in popularity with industry and academia, it is threatened by the decreasing reliability of aggressively scaled transistors. This level of failure has architectural level ramifications, as it may cause an entire on-chip network to fail. Traditional fault-tolerant routing algorithms can overcome the faulty links or routers by rerouting packets around faulty regions. These approaches increase the packet latency and create congestion around the faulty region. In this paper, we present a novel fault-tolerant method that is able to route packets through shortest paths in the presence of faulty links, as long as a path exists. Although the same idea can be applied to a network with any number of virtual channels, we utilize two virtual channels to tolerate all one and two faulty links. Finally, the method is extended to support multiple faulty links by fully utilizing all allowable turns in the network.
Masoumeh Ebrahimi, Masoud Daneshtalab, Juha Plosila, Hannu Tenhunen
DSD4
2012 Energy-Aware Fault-Tolerant Network-on-Chips for Addressing Multiple Traffic Classes
abstract
This paper presents an energy efficient architecture to provide on-demand fault tolerance to multiple traffic classes, running simultaneously on single network on chip (NoC) platform. Today, NoCs host multiple traffic classes with potentially different reliability needs. Providing platform-wide worst-case (maximum) protection to all the classes is neither optimal nor desirable. To reduce the overheads incurred by fault tolerance, various adaptive strategies have been proposed. The proposed techniques rely on individual packet fields and operating conditions to adjust the intensity and hence the overhead of fault tolerance. Presence of multiple traffic classes undermines the effectiveness of these methods. To complement the existing adaptive strategies, we propose on-demand fault tolerance, capable of providing required reliability, while significantly reducing the energy overhead. Our solution relies on a hierarchical agent based control layer and a reconfigurable fault tolerance data path. The control layer identifies the traffic class and directs the packet to the path providing the needed reliability. Simulation results using representative applications (matrix multiplication, FFT, wavefront, and HiperLAN) showed up to 95% decrease in energy consumption compared to traditional worst case methods. Synthesis results have confirmed a negligible additional overhead, for providing on-demand protection (up to 5.3% area), compared to the overall fault tolerance circuitry.
Syed M. A. H. Jafri, Liang Guang, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen
DSD6
2012 Power and Thermal Analysis of Stacked Mesh 3D NoC Using AdaptiveXYZ Routing Algorithm
abstract
Three-dimensional integrated circuits (3D ICs) offer greater device integration, reduced signal delay and reduced interconnect power. It also provides greater design flexibility by allowing heterogeneous integration. However, 3D technology exacerbates the on-chip thermal issues and increases packaging and cooling costs. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed. This architecture provides a seemingly significant platform to implement an integrated low-cost system-wide monitoring network. In this paper, a generic monitoring and management platform called ARB-NET is presented. Based on the ARB-NET monitoring platform, a fully congestion-aware adaptive routing algorithm named AdaptiveXYZ is provided which takes advantage from viable information generated within the monitoring network. In addition, we address both the power and thermal issues of a stacked mesh 3D network on chips using AdaptiveXYZ routing. To this end, a thermal model of a 3D stacked NoC system in a modern flip-chip package is developed. Thermal and power analysis are performed in order to investigate the impact of the proposed adaptive routing from the power and thermal perspectives. Our experiments with a videoconference encoder as a real application show significant power, performance and peak temperature improvements compared to a typical stacked mesh 3D NoC.
Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DSD5
2012 Vertical and horizontal integration towards collective adaptive system: a visionary approach
abstract
Hybrid multi-domain computing systems are emerging. While the context-aware self-adaptive system models are under intensive research in individual computing domains, their integration into a collective adaptive system still remains a major challenge. This position paper visions a meet-in-the-middle approach, where horizontal integration is applied to sub-system models extracted from vertical integration. The integration relies on orthogonal behavior and execution models respectively capturing the functional and non-functional features of sub-systems. The construction towards guaranteed services can be achieved with composition of static (worst-case) execution models, while best-effort services can be constructed with statistical models. Given that each computing domain has, to some extent, formulated its own design flow of context-aware systems, the envisaged meet-in-the-middle integration approach maximizes the reuse of existing models and platforms, thus is promising for the highly-complex system design process.
Liang Guang, Ethiopia Nigussie, Juha Plosila, Hannu Tenhunen
UbiComp4
2012 HARAQ: Congestion-Aware Learning Model for Highly Adaptive Routing Algorithm in On-Chip Networks
abstract
The occurrence of congestion in on-chip networks can severely degrade the performance due to increased message latency. In mesh topology, minimal methods can propagate messages over two directions at each switch. When shortest paths are congested, sending more messages through them can deteriorate the congestion condition considerably. In this paper, we present an adaptive routing algorithm for on-chip networks that provide a wide range of alternative paths between each pair of source and destination switches. Initially, the algorithm determines all permitted turns in the network including 180-degree turns on a single channel without creating cycles. The implementation of the algorithm provides the best usage of all allowable turns to route messages more adaptively in the network. On top of that, for selecting a less congested path, an optimized and scalable learning method is utilized. The learning method is based on local and global congestion information and can estimate the latency from each output channel to the destination region.
Masoumeh Ebrahimi, Masoud Daneshtalab, Fahimeh Farahnakian, Juha Plosila, Pasi Liljeberg, Maurizio Palesi, Hannu Tenhunen
NOCS7
2012 Generic Monitoring and Management Infrastructure for 3D NoC-Bus Hybrid Architectures
abstract
Three-dimensional integrated circuits (3D ICs) achieve enhanced system integration and improved performance at lower cost and reduced area footprint. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, 3D NoC-Bus Hybrid mesh architecture was proposed which provides performance, power consumption, and area benefits. Besides its various advantages, this architecture has a unique and hitherto previously unexplored way to implement an efficient system-wide monitoring network. In this paper, an integrated low-cost monitoring platform for 3D stacked mesh architectures is proposed which can be efficiently used for various system management purposes such as traffic monitoring, thermal management and fault tolerance. The proposed generic monitoring and management infrastructure called ARB-NET utilizes bus arbiters to exchange the monitoring information directly with each other without using the data network. As a test case, based on the proposed monitoring and management platform, a fully congestion-aware and inter-layer fault tolerant routing algorithm named AdaptiveXYZ is presented taking advantage of viable information generated using bus arbiter network. In addition, we propose a thermal monitoring and management strategy on top of our ARB-NET infrastructure. Compared to recently proposed stacked mesh 3D NoCs, our extensive simulations with synthetic and real benchmarks reveal that our architecture using the AdaptiveXYZ routing can help in achieving significant power and performance improvements while preserving the system reliability with negligible hardware overhead.
Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
NOCS6
2012 LEAR - A Low-Weight and Highly Adaptive Routing Method for Distributing Congestions in On-chip Networks
abstract
Congestion-aware routing algorithms can improve network throughput by avoiding packets to be routed through congested areas. In this paper, we propose a minimal/non-minimal routing algorithm to alleviate congestion in the network by making use of all available paths between sources and destinations. The simplicity of the proposed algorithm provides a cost and power efficient solution for Networks-on-Chip while the high degree of adaptive ness, achieved by using an additional virtual channel along the Y dimension, leads to an increased performance. In this method, different restrictions are imposed on the use of each virtual channel, so that the prohibited turns in one virtual channel are permitted in the other one. By fully exploiting of the eligible turns in the network, a large number of output channels can be provided by the proposed method. Based on this method, a packet is routed along the non-minimal path when the neighboring routers in the minimal path are congested.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP5
2012 An Efficient Hybridization Scheme for Stacked Mesh 3D NoC Architecture
abstract
Three-dimensional (3D) integration is a viable design paradigm to overcome the existing interconnect bottleneck in integrated systems and enhance system power/performance characteristics. In order to exploit the intrinsic capability of reducing the wire length in 3D ICs, stacked mesh 3D NoC architecture was proposed. However, this architecture suffers from naive and straightforward hybridization between NoC and bus media. In this paper, an efficient hybridization scheme is presented to enhance system performance, power consumption, and area of stacked mesh 3D NoC architectures. By utilizing a routing rule called LastZ the proposed hybridization scheme offers many advantages investigated in detail to emphasize the significant achievements. Our extensive simulations with synthetic and real benchmarks, including an integrated videoconference application show that compared to a typical 3D NoC-Bus Hybrid Mesh architecture, our hybridization scheme achieves significant power, performance, and area improvements.
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4
2012 Memory-Efficient On-Chip Network With Adaptive Interfaces
abstract
To achieve higher memory bandwidth in network-based multiprocessor architectures, multiple dynamic random access memories can be accessed simultaneously. In such architectures, not only resource utilization and latency are the critical issues but also a reordering mechanism is required to deliver the response transactions of concurrent memory accesses in-order. In this paper, we present a memory-efficient on-chip network architecture to cope with these issues efficiently. Each node of the network is equipped with a novel network interface (NI) to deal with out-of-order delivery, and a priority-based router to decrease the network latency. The proposed NI exploits a streamlined reordering mechanism to handle the in-order delivery and utilizes the advance extensible interface transaction-based protocol to maintain compatibility with existing intellectual property cores. To improve the memory utilization and reduce the memory latency, an optimized memory controller is integrated in the presented NI. Experimental results with synthetic test cases demonstrate that the proposed on-chip network architecture provides significant improvements in average network latency (16%), average memory access latency (19%), and average memory utilization (22%).
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2012 Bio-Patch Design and Implementation Based on a Low-Power System-on-Chip and Paper-Based Inkjet Printing Technology
abstract
This paper presents the prototype implementation of a Bio-Patch using fully integrated low-power System-on-Chip (SoC) sensor and paper-based inkjet printing technology. The SoC sensor is featured with programmable gain and bandwidth to accommodate a variety of bio-signals. It is fabricated in a 0.18-ìm standard CMOS technology, with a total power consumption of 20 ìW from a 1.2 V supply. Both the electrodes and interconnections are implemented by printing conductive nano-particle inks on a flexible photo paper substrate using inkjet printing technology. A Bio-Patch prototype is developed by integrating the SoC sensor, a soft battery, printed electrodes and interconnections on a photo paper substrate. The Bio-Patch can work alone or operate along with other patches to establish a wired network for synchronous multiple-channel bio-signals recording. The measurement results show that electrocardiogram and electromyogram are successfully measured in in-vivo tests using the implemented Bio-Patch prototype.
Geng Yang 0003, Matti Mäntysalo, Jian Chen 0001, Hannu Tenhunen, Lirong Zheng 0001
IEEE Trans. Inf. Technol. Biomed.5
2012 Semi-Serial On-Chip Link Implementation for Energy Efficiency and High Throughput
abstract
A high-throughput and low-energy semi-serial on-chip communication link based on novel design techniques and circuit solutions is presented. This self-timed link is designed using high-speed serialization/deserializtion and pulse dual-rail encoding techniques. The link also employs wave-pipelined differential pulse current-mode signaling to maintain the high speed data intake from the serializer. The energy efficiency of the proposed semi-serial link, which consists of bit-serial links in parallel, mainly comes from the sharing of the novel serializer's control circuit among the bit-serial links. In addition, the integration of pulse signaling with wave-pipelining, the use of a new low-complexity data validity detection technique, and the avoidance of data decoding logic also contribute to the power reduction. Furthermore, the formulated pulse dual-rail encoding provides an opportunity to implement pulse signaling at no cost. The ability to detect data validity at bit level allows acknowledgment per word without losing the delay-insensitivity of the transmission. The proposed semi-serial link is analyzed and compared with bit-serial and fully bit-parallel links for 64-bit data and communication distances of 1 to 8 mm. The semi-serial link which consists of eight bit-serial links provides 72.72 Gbps throughput with 286 fJ/bit energy dissipation for 8 mm transmission. It dissipates the lowest energy per bit compared to fully bit-parallel links while achieving the same throughput. The links are designed and simulated in Cadence Analog Spectre using 65-nm technology from STMicroelectronics.
Ethiopia Nigussie, Sampo Tuuna, Juha Plosila, Jouni Isoaho, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.5
2012 Modeling of Energy Dissipation in RLC Current-Mode Signaling
abstract
In this paper, energy dissipation in resistance-inductance-capacitance (RLC) current-mode signaling is modeled. The energy dissipation is derived separately for driver, wire, and receiver termination. The effects of rise time and clock cycle are included. A realizable Π-model for the driving-point impedance of an RLC current-mode transmission line is derived. The output current of an RLC current-mode transmission line is also derived. The model is extended to multiple parallel coupled interconnects with inductive and capacitive coupling between them. The model is verified by comparing it to HSPICE in 65-nm technology and applied to differential current-mode signaling.
Sampo Tuuna, Ethiopia Nigussie, Jouni Isoaho, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Enhancing Performance of NoC-Based Architectures Using Heuristic Virtual-Channel Sharing Approach
abstract
This paper presents a novel virtual-channel (VC) sharing technique for NoC architecture. The proposed architecture improves the utilization of resources to enhance the performance with minimal overheads. A heuristic approach towards the proper VC sharing strategy is proposed, which is performed by an adaptive algorithm that configures the VC sharing based on link load parameters. Architectural design to realize the adaptive VC sharing in generic router is elaborated. The technique can be applied to any NoC architecture, including 3-D NoCs. Extensive quantitative experiments with synthetic and real benchmarks, including an integrated video conference application, demonstrate considerable improvement in area and power efficiency compared to existing VC-based 2D/3D NoC architectures.
Khalid Latif 0002, Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Tiberiu Seceleanu, Pasi Liljeberg, Hannu Tenhunen
COMPSAC6
2011 Evaluating Sustainability, Environmental Assessment and Toxic Emissions during Manufacturing Process of RFID Based Systems
abstract
The present state of the art research in the direction of embedded systems demonstrate that analysis of life-cycle, sustainability and environmental assessment have not been a core focus for researchers. To maximize a researcher's contribution in formulating environmentally friendly products, devising green manufacturing processes and services, there is a strong need to enhance life-cycle awareness and sustainability understandings among embedded systems researchers, so that the next generation of engineers will be able to realize the goal of a sustainable life-cycle. In this work an attempt has been made to investigate and evaluate the life-cycle management and environmental assessment in fabricating processes of the RFID based systems. We have chosen a general life cycle assessment approach which involves the collection and evaluation of quantitative data on the inputs and outputs of materials and energy associated with the RFID based systems. Based on the developed generic models, we have obtained the results in terms of environmental emissions for a production of paper substrate printed RFID antennas. We also make an attempt to raise several sustainability issues and quantify the toxic emissions during the manufacturing process.
Rajeev Kumar Kanth, Pasi Liljeberg, Hannu Tenhunen, Qiansu Wan, Yasar Amin, Botao Shao, Qiang Chen 0014, Lirong Zheng 0001
DASC3
2011 Optimal number and placement of Through Silicon Vias in 3D Network-on-Chip
abstract
In this paper, we analyze the performance impact of different number of Through Silicon Vias (TSVs) in 3D Network-on-Chip (NoC). The adoption of a 3D NoC design depends on the performance and manufacturing cost of the chip. Therefore, a study of the placement of the TSV, that connects different layers of a 3D chip, is crucial. A 64-core 3D NoC is modeled based on state-of-the-art 2D chips. We discuss the number of TSVs required for a 3D NoC. Different placements of layer-layer connections are explored. We present benchmark results using a cycle accurate full system simulator based on realistic workloads. Experiments show that under different workloads, the average network latencies in two configurations (full and quarter connection) are reduced by 14.78% and 7.38% respectively, compared with the one-eighth connection design. The improvement of performance is a trade-off of manufacturing cost. Our analysis and experiment results provide a guideline for selecting optimal number of TSVs in 3D NoCs.
Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen
DDECS3
2011 Enhancing Performance Sustainability of Fault Tolerant Routing Algorithms in NoC-Based Architectures
abstract
Reliability of embedded systems and devices is becoming a challenge with technology scaling. To deal with the reliability issues, fault tolerant solutions are needed. The design paradigm for future System-on-Chip (SoC) implementation is Network-on-Chip (NoC). Fault tolerance in NoC can be achieved at many abstraction levels. Many fault tolerant architectures and routing algorithms have already been proposed for NoC but the utilization of resources, affected indirectly by faults is yet to be addressed. In this paper, we propose a NoC architecture, which sustains the overall system performance by utilizing resources, which cannot be used by other architectures under faults. An approach towards a proper virtual-channel (VC) sharing strategy is proposed, based on communication bandwidth requirements. The technique can be applied to any NoC architecture, including 3-D NoCs. Extensive quantitative experiments with synthetic benchmarks, including uniform, transpose and negative exponential distribution (NED), demonstrate considerable improvement in terms of performance sustainability under faulty conditions compared to existing VC-based NoC architectures.
Khalid Latif 0002, Amir-Mohammad Rahmani, Kameswar Rao Vaddina, Tiberiu Seceleanu, Pasi Liljeberg, Hannu Tenhunen
DSD6
2011 LastZ: An Ultra Optimized 3D Networks-on-Chip Architecture
abstract
3D IC technology enables NoC architectures to offer greater device integration and shorter interlayer interconnects. The primary 3D NoC architectures such as Symmetric 3D Mesh NoC could not exploit the beneficial feature of a negligible inter-layer distance in 3D chips. To cope with this, 3D NoC-Bus Hybrid architecture was proposed which is a hybrid between packet-switched network and a bus. This architecture is feasible providing both performance and area benefits, while still suffering from naive and straightforward hybridization between NoC and bus media. In this paper, an ultra optimized hybridization scheme is proposed to enhance system performance, power consumption, area and thermal issues of 3D NoC-Bus Hybrid Mesh. The scheme benefits from a rule called LastZ which enables ultra optimization of the inter-layer communication architecture. In addition, we present a wrapper to preserve the backward compatibility of the proposed architecture for connecting with the existing network interfaces. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. Our extensive simulations demonstrate significant area, power, and performance improvements compared to a typical 3D NoC-Bus Hybrid Mesh architecture.
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DSD4
2011 Compact generic intermediate representation (CGIR) to enable late binding in coarse grained reconfigurable architectures
abstract
In the era of platforms hosting multiple applications, where inter-application communication and concurrency patterns are arbitrary, static compile time decision making is neither optimal nor desirable. As a part of solving this problem, we present a novel method for compactly representing multiple configuration bitstreams of a single application, with varying parallelisms, as a unique, compact, and customizable representation, called CGIR. The representation thus stored is unraveled at runtime to configure the device with optimal (e.g. in terms of energy) implementation. Our goal was to provide optimal decision making capability to the runtime resource manager (RTM) without compromising the runtime behavior or the memory requirements of the system. The presence of multiple binaries enhance optimality by providing the RTM with multiple implementations to choose from. The CGIR ensures minimal increase in memory requirements with the addition of each binary. The low-cost unraveling of CGIR guarantees the runtime behavior. We have chosen the dynamically reconfigurable resource array (DRRA) as a vehicle to study the feasibility of our approach. Simulation results using 16 point decimation in time fast Fourier transform (FFT) has showed massive (up to 18% for 2 versions, 33% for 3 versions) memory savings compared to state of the art. Formal evaluation shows that the savings increase with the increase in the number of implementations stored.
Syed M. A. H. Jafri, Ahmed Hemani, Kolin Paul, Juha Plosila, Hannu Tenhunen
FPT5
2011 A Minimal Average Accessing Time Scheduler for Multicore Processors
Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen
ICA3PP (2)3
2011 Exploring partitioning methods for 3D Networks-on-Chip utilizing adaptive routing model
abstract
Three-Dimensional (3D) integration is a solution to the interconnect bottleneck in Two-Dimensional (2D) MultiProcessor System on Chip (MPSoC). 3D IC design improves performance and decreases power consumption by replacing long horizontal interconnects with shorter vertical ones. As the multicast communication is utilized commonly in various parallel applications, the performance can be significantly improved by supporting of multicast operations at the hardware level. In this paper, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method named Recursive Partitioning (RP) in which the network is recursively partitioned until all partitions contain comparable number of nodes. By this approach, the multicast traffic is distributed among several subsets and the network latency is considerably decreased. We also present Minimal Adaptive Routing (MAR) algorithm for the unicast and multicast traffic in 3D-mesh Networks-on-Chip (NoCs). The idea behind the MAR algorithm is utilizing the Hamiltonian path to provide a set of alternative paths.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
NOCS5
2011 Congestion aware, fault tolerant, and thermally efficient inter-layer communication scheme for hybrid NoC-bus 3D architectures
abstract
Three-dimensional IC technology offers greater device integration and shorter interlayer interconnects. In order to take advantage of these attributes, 3D stacked mesh architecture was proposed which is a hybrid between packet-switched network and a bus. Stacked mesh is a feasible architecture which provides both performance and area benefits, while suffering from inefficient intermediate buffers. In this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm called AdaptiveZ for vertical communication. In addition, we hybridize the proposed adaptive routing with available algorithms to mitigate the thermal issues by herding most of the switching activities closer to the heat sink. Our extensive simulations with synthetic and real benchmarks, including the one with an integrated videoconference application, demonstrate significant power, performance, and peak temperature improvements compared to a typical stacked mesh 3D NoC.
Amir-Mohammad Rahmani, Pasi Liljeberg, Khalid Latif 0002, Juha Plosila, Kameswar Rao Vaddina, Hannu Tenhunen
NOCS6
2011 PVS-NoC: Partial Virtual Channel Sharing NoC Architecture
abstract
A novel architecture aiming for ideal performance and overhead tradeoff, PVS-NoC (Partial VC Sharing NoC), is presented. Virtual channel (VC) is an efficient technique to improve network performance, while suffering from large silicon and power overhead. We propose sharing the VC buffers among dual inputs, which provides the performance advantage as conventional VC-based router with minimized overhead. We reason theoretically and demonstrate quantitatively the benefits of proposed architecture by comparing to state-of-the-art NoC routers, with various traffic patterns. Extensive experiments with synthetic and real benchmarks show significant area and power saving with similar performance compared to latest VC based NoC architectures.
Khalid Latif 0002, Amir-Mohammad Rahmani, Liang Guang, Tiberiu Seceleanu, Hannu Tenhunen
PDP5
2011 A Stacked Mesh 3D NoC Architecture Enabling Congestion-Aware and Reliable Inter-layer Communication
abstract
In this paper, an efficient architecture to optimize system performance, power consumption, and reliability of stacked mesh 3D NoC is proposed. Stacked mesh is a feasible architecture which takes advantage of the short inter-layer wiring delays, while suffering from inefficient intermediate buffers. To cope with this, an inter-layer communication mechanism is developed to enhance the buffer utilization, load balancing, and system fault-tolerance. The mechanism benefits from a congestion-aware and bus failure tolerant routing algorithm for vertical communication. To estimate the efficiency of the proposed architecture, the system has been simulated using uniform, hotspot 10%, and Negative Exponential Distribution (NED) traffic patterns. In addition, a video conference encoder has been used as a real application for system analysis. Our extensive experiments show significant power and performance improvements compared to a typical stacked mesh 3D NoC.
Amir-Mohammad Rahmani, Khalid Latif 0002, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP5
2011 Agent-based on-chip network using efficient selection method
abstract
Congestion in on-chip networks may cause many drawbacks in multiprocessor systems including throughput reduction, increase in latency, and additional power consumption. Furthermore, conventional congestion control methods, employed for on-chip networks, cannot efficiently collect congestion information and distribute them over the on-chip network. In this paper, we present a novel structure for on-chip networks, named Agent-based Network-on-Chip (ANoC), to diagnose the congested areas. In addition to the presented structure, an efficient Congestion-Aware Selection (CAS) method is proposed to reduce overall network latency. CAS is capable of selecting an appropriate output channel to route packets along a less congested path. 29% average and 35% maximum latency reduction are achieved on SPLASH-2 and PARSEC benchmarks running on a 36-core Chip Multi-Processor.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
VLSI-SoC5
2011 A generic adaptive path-based routing method for MPSoCs
Masoud Daneshtalab, Masoumeh Ebrahimi, Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen
J. Syst. Archit.5
2010 On signalling over Through-Silicon Via (TSV) interconnects in 3-D Integrated Circuits
abstract
This paper discusses signal integrity (SI) issues and signalling techniques for Through Silicon Via (TSV) interconnects in 3-D Integrated Circuits (ICs). Field-solver extracted parasitics of TSVs have been employed in Spice simulations to investigate the effect of each parasitic component on performance metrics such as delay and crosstalk and identify a reduced-order electrical model that captures all relevant effects. We show that in dense TSV structures voltage-mode (VM) signalling does not lend itself to achieving high data-rates, and that current-mode (CM) signalling is more effective for high throughput signalling as well as jitter reduction. Data rates, energy consumption and coupled noise for the different signalling modes are extracted.
Roshan Weerasekera, Matt Grange, Dinesh Pamunuwa, Hannu Tenhunen
DATE4
2010 Partitioning methods for unicast/multicast traffic in 3D NoC architecture
abstract
As the scale of integration grows, the interconnection problem becomes one of the major design considerations of Multi Processor System on Chip (MPSoC). In recent years, many researchers have conducted studies on 3D IC designs stacking multiple layers on top of each other. In order to decrease the transmission delay of unicast/multicast messages in a network based multicore system, the network is divided into several partitions. In this paper, we first introduce a novel idea of balanced partitioning that allows the network to be partitioned effectively. Then, we propose a set of partitioning approaches each with a different level of efficiency. In addition, we present an advantageous method based on the idea of balanced partitioning to provide a high degree of parallelism with a considerable reduction of packet delay in unicast/multicast traffic. Simulations are provided to evaluate and compare the performance of proposed methods.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen
DDECS4
2010 Developing reconfigurable FIFOs to optimize power/performance of Voltage/Frequency Island-based networks-on-chip
abstract
Network-on-chip architectures partitioned into several Voltage/Frequency Islands (VFIs) have been proposed to alleviate problems related to integration, excessive energy consumption and clock distribution. The architecture is composed of synchronous switches that communicate with each other using bi-synchronous FIFOs. However, these FIFOs are not needed if adjacent switches belong to the same clock domain. In this paper, a Reconfigurable Synchronous/Bi-Synchronous (RSBS) FIFO is proposed which can operate in either synchronous or bi-synchronous mode. The FIFO is scalable and synthesizable in synchronous standard cells and also a technique for mesochronous adaptation has been recommended. In addition, some techniques are suggested to show how the FIFO could be utilized in a VFI-based NoC. Our results reveal that compared to a non-reconfigurable system architecture, the RSBS FIFOs help to achieve up to 15% savings in average power consumption of NoC switches and 29% improvement in total average packet latency in the case of MPEG-4 encoder application.
Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DDECS4
2010 Power-aware NoC router using central forecasting-based dynamic virtual channel allocation
abstract
In this paper, we propose a high performance central dynamic virtual channel allocation mechanism for on-chip routers. This central management unit devotes each input port a number of virtual channels (VC) among a shared VC bank based on a traffic forecasting technique. The forecasting technique exploits the link and VC utilizations in predicting the traffic. Based on the predicted traffic, for each input port, the number of active virtual channels may be increased, decreased, or kept unchanged. The clock-gating power management technique is used to activate/deactivate the VCs. Simulation results using uniform and Negative Exponential Distribution (NED) traffic profiles show that a considerable power savings in the virtual channels and overall router power consumption may be achieved especially in low traffic loads. The area overhead of the technique is negligible.
Amir-Mohammad Rahmani, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen
ISCAS4
2010 A Low-Latency and Memory-Efficient On-chip Network
abstract
Using multiple SDRAMs in MPSoCs and NoCs to increase memory parallelism is very common nowadays. In-order delivery, resource utilization, and latency are the most critical issues in such architectures. In this paper, we present a novel network interface architecture to cope with these issues efficiently. The proposed network interface exploits a resourceful reordering mechanism to handle the in-order delivery and to increase the resource utilization. A brilliant memory controller is efficiently integrated into this network interface to improve the memory utilization and reduce both memory and network latencies. In addition, to bring compatibility with existing IP cores the proposed network interface utilizes AXI transaction based protocol. Experimental results with synthetic test cases demonstrate that the proposed architecture gives significant improvements in average network latency (12%), average memory access latency (19%), and average memory utilization (22%).
Masoud Daneshtalab, Masoumeh Ebrahimi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
NOCS5
2010 A High-Performance Network Interface Architecture for NoCs Using Reorder Buffer Sharing
abstract
Increasing memory parallelism in MPSoCs to provide higher memory bandwidth is achieved by accessing multiple memories simultaneously. Inasmuch as the response transactions of concurrent memory accesses must be in-order, a reordering mechanism is required. To our knowledge the resource utilization of conventional reordering mechanisms is low. In this paper, we present a novel network interface architecture for on-chip networks to increase the resource utilization and to improve overall performance. Also, based on the proposed architecture, a hybrid network interface is presented to integrate both memory and processor in a tile. The proposed architecture exploits AXI transaction based protocol to be compatible with existing IP cores. Experimental results with synthetic test cases demonstrate that the proposed architecture outperforms the conventional architecture in terms of latency. Also, the cost of the presented architecture is evaluated with UMC 0.09 ¿ m technology.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP5
2010 HAMUM - A Novel Routing Protocol for Unicast and Multicast Traffic in MPSoCs
abstract
Many parallel applications in MPSoCs take advantage of multicast communication. Several multicast schemes such as path-based, tree-based, and unicast-based have been proposed in interconnection networks. Path-based multicast scheme has been proven to be more efficient than the other schemes in on-chip interconnection network. A new adaptive routing model based on Hamiltonian path for both the multicast and unicast traffics, called Hamiltonian Adaptive Multicast and Unicast Model (HAMUM), is presented. Results obtained in both multicast and mixed traffic models show that the proposed adaptive algorithm for multicast aspect has lower latency and power dissipation compared to previously proposed path-based multicasting algorithms with less than 0.5% hardware overhead. Additionally, for the unicast aspect the proposed adaptive model outperforms the other unicast turn models.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen
PDP4
2010 Hierarchical agent monitoring design approach towards self-aware parallel systems-on-chip
abstract
Hierarchical agent framework is proposed to construct a monitoring layer towards self-aware parallel systems-on-chip (SoCs). With monitoring services as a new design dimension, systems are capable of observing and reconfiguring themselves dynamically at all levels of granularity, based on application requirements and platform conditions. Agents with hierarchical priorities work adaptively and cooperatively to maintain and improve system performance in the presence of variations and faults. Function partitioning of agents and hierarchical monitoring operations on parallel SoCs are analyzed. Applying the design approach on the Network-on-Chip (NoC) platform demonstrates the design process and benefits using the novel approach.
Liang Guang, Ethiopia Nigussie, Pekka Rantala, Jouni Isoaho, Hannu Tenhunen
ACM Trans. Embed. Comput. Syst.5
2009 An efficent dynamic multicast routing protocol for distributing traffic in NOCs
abstract
Nowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms.
Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen
DATE7
2009 An Adaptive Unicast/Multicast Routing Algorithm for MPSoCs
abstract
Several parallel applications in MPSoCs take advantage of multicast communication. Path-based multicast scheme has been proven to be more efficient than the others multicast schemes in on-chip interconnection network. We present a new adaptive path based model for both the multicast and unicast wormhole routing protocols. The proposed model under mixed traffic models has lower latency than the previous path-based methods with negligible hardware overhead.
Masoumeh Ebrahimi, Masoud Daneshtalab, Pasi Liljeberg, Hannu Tenhunen
DSD4
2009 Architectural Exploration of Per-Core DVFS for Energy-Constrained On-Chip Networks
abstract
A feasible and scalable per-core DVFS architecture for on-chip network is presented. The supplies are dynamically adjusted at a very fine granularity based on the local traffic status. The adoption of multiple voltage supply networks and power selecting transistors provides the architecture with scalability and feasibility superior to existing similar techniques. With high-level simulation using 65 nm power model obtained from widely-acknowledged tools, the effectiveness of the technique is demonstrated with quantitative analysis of energy overhead and latency penalty. Under various traffic patterns, the average flit energy is reduced considerably, ranging from 45% to 60%, with moderately increased but stable transmission latency.
Alexander Wei Yin, Liang Guang, Ethiopia Nigussie, Pasi Liljeberg, Jouni Isoaho, Hannu Tenhunen
DSD6
2009 Scalability of network-on-chip communication architecture for 3-D meshes
abstract
Design constraints imposed by global interconnect delays as well as limitations in integration of disparate technologies make 3D chip stacks an enticing technology solution for massively integrated electronic systems. The scarcity of vertical interconnects however imposes special constraints on the design of the communication architecture. This article examines the performance and scalability of different communication topologies for 3D network-on-chips (NoC) using through-silicon-vias (TSV) for inter-die connectivity. Cycle accurate RTL-level simulations are conducted for two communication schemes based on a 7-port switch and a centrally arbitrated vertical bus using different traffic patterns. The scalability of the 3D NoC is examined under both communication architectures and compared to 2D NoC structures in terms of throughput and latency in order to quantify the variation of network performance with the number of nodes and derive key design guidelines.
Awet Yemane Weldezion, Matt Grange, Dinesh Pamunuwa, Zhonghai Lu, Axel Jantsch, Roshan Weerasekera, Hannu Tenhunen
NOCS7
2009 Explorations of Honeycomb Topologies for Network-on-Chip
abstract
Rectangular mesh and torus are the mostly used topologies in network-on-chip (NoC) based systems. In this paper, we quantitatively illustrate that the honeycomb topology is an advantageous design alternative in terms of network cost which is one of the most important parameters that reflects both network performance and implementation cost. Comparing with the rectangular mesh and torus, honeycomb mesh and torus topologies lead to 40% decrease of the network cost. Then we explore the NoC related topological properties of both honeycomb mesh and torus topologies. By transforming the honeycomb topologies into rectangular brick shapes, we demonstrate that the honeycomb topologies are feasible to be implemented with rectangular devices. We also propose a 3D honeycomb topology since 3D IC has become an emerging and promising technique. Another contribution of this paper is the proposal of deadlock free routing algorithms. Based on either the concept of turn model or the logical network, deadlock free routing for all the discussed honeycomb topologies can be achieved.
Alexander Wei Yin, Thomas Canhao Xu, Pasi Liljeberg, Hannu Tenhunen
NPC4
2009 Two-Dimensional and Three-Dimensional Integration of Heterogeneous Electronic Systems Under Cost, Performance, and Technological Constraints
abstract
Present day market demand for high-performance high-density portable hand-held applications has shifted the focus from 2-D planar system-on-a-chip-type single-chip solutions to alternatives such as tiled silicon and single-level embedded modules as well as 3-D die stacks. Among the various choices, finding an optimal solution for system implementation deals usually with cost, performance, power, thermal, and technological tradeoff analyses at the system conceptual level. It has been estimated that decisions made in the first 20% of the design cycle influence up to 80% of the final product cost. In this paper, we discuss realistic metrics appropriate for performance and cost tradeoff analyses both at the system conceptual level in the early stages of the design cycle and in the implementation phase, for verification. In order to validate the proposed metrics and methodology, two ubiquitous electronic systems are analyzed under various implementation schemes and the performance tradeoffs discussed. This case study is used to highlight the importance of a cost and performance tradeoff analysis early in the design flow.
Roshan Weerasekera, Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2008 On Analysis and Synthesis of (n, k)-Non-Linear Feedback Shift Registers
abstract
Non-linear feedback shift registers (NLFSRs) have been proposed as an alternative to Linear Feedback Shift Registers (LFSRs) for generating pseudo-random sequences for stream ciphers. In this paper, we introduce (n,k)-NLFSRs which can be considered a generalization of the Galois type of LFSR. In an (n,fc)-NLFSR, the feedback can be taken from any of the n bits, and the next state functions can be any Boolean function of up to k variables. Our motivation for considering this type NLFSRs is that their Galois configuration makes it possible to compute each next state function in parallel, thus increasing the speed of output sequence generation. Thus, for stream cipher application where the encryption speed is important, (n,k)-NLFSRs may be a better alternative than the traditional Fibonacci ones. We derive a number of properties of (n,k)- NLFSRs. First, we demonstrate that they are capable of generating output sequences with good statistical properties which cannot be generated by the Fibonacci type of NLFSRs. Second, we show that the period of the output sequence of an (n,k)-NLFSR is not necessarily equal to the length of the largest cycle of its states. Third, we compute the period of an (n,k)-NLFSR constructed from several parallel NLFSRs whose outputs are XOR-ed and show how to maximize this period. We also present an algorithm for estimating the length of cycles of states of (n,k)-NLFSRs which uses binary decision diagrams for representing the set of states and the transition relation on this set.
Elena Dubrova, Maxim Teslenko, Hannu Tenhunen
DATE3
2008 Modeling of On-Chip Bus Switching Current and Its Impact on Noise in Power Supply Grid
abstract
In this paper, an analytical model for the current draw of an on-chip bus is presented. The model is combined with an on-chip power supply grid model in order to analyze noise caused by switching buses in a power supply grid. The bus is modeled as distributed resistance-inductance-capacitance (RLC) lines that are capacitively and inductively coupled to each other. Different switching patterns and driver skewing times are also included in the model. The power supply grid is modeled as a network ofRLCsegments. The model is verified by comparing it to HSPICE. The error was below 8%. The model is applied to determine the influence of driver skewing times on maximum power supply noise.
Sampo Tuuna, Lirong Zheng 0001, Jouni Isoaho, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Minimal-Power, Delay-Balanced Smart Repeaters for Global Interconnects in the Nanometer Regime
abstract
A smart repeater is proposed for driving capacitively-coupled, global-length on-chip interconnects that alters its drive strength dynamically to match the relative bit pattern on the wires and thus the effective capacitive load. This is achieved by partitioning the driver into main and assistant drivers; for a higher effective load capacitance both drivers switch, while for a lower effective capacitance the assistant driver is quiet. In a UMC 0.18-mum technology the potential energy saving is around 10% and the reduction in jitter 20%, in comparison to a traditional repeater for typical global wire lengths. It is also shown that the average energy saving for nanometer technologies is in the range of 20% to 25%. The driver architecture exploits the fact that as feature sizes decrease, the capacitive load per transistor shrinks, whereas global wire loads remain relatively unchanged. Hence, the smaller the technology, the greater the potential saving.
Roshan Weerasekera, Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.4
2007 Towards a Design Methodology for Multiprocessor Platforms
abstract
We discuss a design methodology for SegBus, a multicore segmented bus platform. The methodology supports the modeling of the platform at several abstraction levels, enabling the designer to focus only on the relevant aspects of the architecture at a given development stage. We employ the unified modeling language (UML) as a specification language for theSegBusplatform and we customize its elements to serve our specific purposes via the profiling mechanism. The approach enables us to take advantage of graphical models of the platform and of automated refinements of these models towards implementation.
Dragos Truscan, Tiberiu Seceleanu, Hannu Tenhunen, Johan Lilius
COMPSAC (1)3
2007 Novel Agent-Based Management for Fault-Tolerance in Network-on-Chip
abstract
We introduce a novel agent-based reconfiguring concept for futures network-on-chip (NoC) systems. The necessary properties to increase architecture level fault tolerance are introduced. The system control is modeled as multi-level agent hierarchy that is able to increase application fault-tolerance and performance with autonomous reactions of agents. The agent technology adds a system level intelligence level to the traditional NoC system design. The architecture and functions of this system are described on conceptual level. Communication and reconfiguring data flows are presented as study cases. Principles of reconfiguration of a NoC on faulty environment are demonstrated and simulated. Probability of reconfiguration success is measured with different latency requirements and amount of redundancy by Monte Carlo simulations. The effect of network topology in reconfiguration of a faulty mesh was also under research in the simulations.
Pekka Rantala, Jouni Isoaho, Hannu Tenhunen
DSD3
2007 Extending systems-on-chip to the third dimension: performance, cost and technological tradeoffs
abstract
Because of the today’s market demand for high- performance, high-density portable hand-held applications, elec- tronic system design technology has shifted the focus from 2-D planar SoC single-chip solutions to different alternative options as tiled silicon and single-level embedded modules as well as 3- D integration. Among the various choices, finding an optimal solution for system implementation dealt usually with cost, performance and other technological trade-off analysis at the system conceptual level. It has been identified that the decisions made within the first 20% of the total design cycle time will ultimately result upto 80% of the final product cost. In this paper, we discuss appropriate and realistic metric for performance and cost trade-off analysis both at system conceptual level (up-front in the design phase) and at implementation phase for verification in the three-dimensional integration. In order to validate the methodology, two ubiquitous electronic systems are analyzed under various implementation schemes and discuss the pros and cons of each of them.
Roshan Weerasekera, Lirong Zheng 0001, Dinesh Pamunuwa, Hannu Tenhunen
ICCAD4
2007 A Novel Passive Tag with Asymmetric Wireless Link for RFID and WSN Applications
abstract
In this paper, we present a radio-powered module with asymmetric wireless link utilizing ultra wideband radio system for RFID and wireless sensor applications. Our contribution includes using two different standards in uplink and downlink. Such as conventional RFIDs, incoming RF signal transmitted by reader is used to power the internal circuitry and receive the data. However, in upstream link, an IR-UWB transmitter is utilized. Unlike traditional RFID systems, due to great advantages of UWB communication, this tag is very robust to multi-path fading and collision problem and it is more secure against eavesdropping or jamming. The module consists of a power scavenging unit, a RF receiver, an IR-UWB transmitter, digital baseband controller, and an embedded UWB antenna are designed for integration on liquid-crystal polymer (LCP) substrate, using 0.18mum CMOS process technology.
Majid Baghaei Nejad, Zhuo Zou, Hannu Tenhunen, Lirong Zheng 0001
ISCAS3
2006 Analytical model for crosstalk and intersymbol interference in point-to-point buses
abstract
In this paper, an analytical model to estimate crosstalk noise and intersymbol interference on capacitively and inductively coupled point-to-point on-chip buses is derived. The derived closed-form equation for output voltage enables the usage of the model in computer-aided design (CAD) tools for complex systems where high simulation speed is essential. The model also combines together properties such as inductive coupling, initial conditions, signal rise time, input phases, and bit sequences that have not been included in a single closed-form model before. The model is compared to HSPICE and previous models. The model and HSPICE are in good agreement with each other
Sampo Tuuna, Jouni Isoaho, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 Computing attractors in dynamic networks
Elena Dubrova, Maxim Teslenko, Hannu Tenhunen
IADIS AC3
2005 Modeling delay and noise in arbitrarily coupled RC trees
abstract
Closed-form equations for second-order transfer functions of general arbitrarily coupled resistance-capacitance (RC) trees with multiple drivers are reported. The models allow precise delay and noise calculations for systems of coupled interconnects with guaranteed stability and represent the minimum complexity associated with this class of circuits. Their accuracy is extensively compared against other relevant models and is found to be better or comparable to more expensive models. All results are derived from a theoretical approach, and their physical basis is examined. The simplicity, accuracy, and generality of the models make them suitable for use in early signal integrity analyses of complex systems and incremental physical optimization.
Dinesh Pamunuwa, Shauki Elassaad, Hannu Tenhunen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 A study on the implementation of 2-D mesh-based networks-on-chip in the nanometre regime
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch, Hannu Tenhunen
Integr.6
2004 Special issue on networks on chip
Axel Jantsch, Johnny Öberg, Hannu Tenhunen
J. Syst. Archit.3
2004 Interconnect intellectual property for Network-on-Chip (NoC)
Lirong Zheng 0001, Hannu Tenhunen
J. Syst. Archit.3
2004 Efficient library characterization for high-level power estimation
abstract
This paper describes LP-DSM, which is an algorithm used for efficient library characterization in high-level power estimation. LP-DSM characterizes the power consumption of building blocks using the entropy of primary inputs and primary outputs. The experimental results showed that over a wide range of benchmark circuits implemented using full custom design in 0.35-/spl mu/m 3.3 V CMOS process the statistical performance (mean and maximum error) of LP-DSM is comparable or sometimes better than most of the published algorithms. Moreover, it was found that LP-DSM has the lowest prediction sum of squares, which makes it an efficient tool for power prediction. Furthermore, the complexity of the LP-DSM is linear in relation to the number of primary inputs (O(NI)), whereas state of the art published library characterization algorithms have a complexity of O(NI/sup 2/).
Imed Ben Dhaou, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Analytic Modeling of Interconnects for Deep Sub-Micron Circuits
Dinesh Pamunuwa, Shauki Elassaad, Hannu Tenhunen
ICCAD3
2003 Maximizing throughput over parallel wire structures in the deep submicrometer regime
abstract
In a parallel multiwire structure, the exact spacing and size of the wires determine both the resistance and the distribution of the capacitance between the ground plane and the adjacent signal carrying conductors, and have a direct effect on the delay. Using closed-form equations that map the geometry to the wire parasitics and empirical switch factor based delay models that show how repeaters can be optimized to compensate for dynamic effects, we devise a method of analysis for optimizing throughput over a given metal area. This analysis is used to show that there is a clear optimum configuration for the wires which maximizes the total bandwidth. Additionally, closed form equations are derived, the roots of which give close to optimal solutions. It is shown that for wide buses, the optimal wire width and spacing are independent of the total width of the bus, allowing easy optimization of on-chip buses. Our analysis and results are valid for lossy interconnects as are typical of wires in submicron technologies.
Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.3
2002 The Case for Fine-Grained Re-configurable Architectures: An Analysis of Conceived Performance
Tuomas Valtonen, Jouni Isoaho, Hannu Tenhunen
FPL3
2000 A fifth-order comb decimation filter for multi-standard transceiver applications
abstract
In multi-standard transceivers programmable decimation filters are required to perform channel select filtering at baseband since the channel bandwidths, sampling rates, and CNR requirements are different. This paper presents a low power fifth-order comb decimation filter with programmable decimation ratios (16 and 8) and sampling rates (12.8 MHz and 44.8 MHz) for GSM and DECT applications. The non-recursive architecture is employed for the comb filter and low power VLSI implementation techniques are developed.
Yonghong Gao, Lihong Jia, Hannu Tenhunen
ISCAS3
2000 New metrics for architectural level power performance evaluation
abstract
In this paper, we present new metrics to evaluate the performance of VLSI circuits with regards to the power consumption in the architectural level from the system point of view. One advantage of the new metrics is that the metrics not only calculate the power consumption itself but also evaluate the "low power possibility" of the architectures from the system point of view; another advantage is that the metrics consider the effects of the process technology and it can be used to evaluate the power consumption performance in the future advanced VLSI technology. Specially, we point out that reducing power supply voltage to reduce the power consumption will become less and less efficient in the deep submicron regime and thus the architecture-driven voltage scaling low power design approach will not work as well as in the previous technology.
Lihong Jia, Yonghong Gao, Hannu Tenhunen
ISCAS3
2000 Combating digital noise in high speed ULSI circuits using binary BCH encoding
abstract
Increased integration in deep submicron (DSM) technologies has caused very high increases in the RLC parasitics which affect the coupling of noise to signals propagating over interconnect. Error free transmission on-chip will no longer be guaranteed, This paper examines the issue of high speed signaling in DSM and proposes the use of particular BCH codes to improve the bit error rate in the face of noise. We conclude from our results that it is possible to achieve a considerable coding gain by choosing the code properly.
Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen
ISCAS3
2000 Efficient and accurate modeling of power supply noise on distributed on-chip power networks
abstract
In this paper, we propose an efficient and accurate modeling technique for power supply noise estimation over on-chip power lines which are modeled as distributed RCL networks. With this model, peak noise on the power lines that includes on-chip resistive and inductive voltage drops, switching noise on packages, and on-chip decoupling effects, can be computed very efficiently and accurately. The model is verified by SPICE simulations.
Lirong Zheng 0001, Hannu Tenhunen
ISCAS3
1999 Phase noise in sampling and its importance to wideband multicarrier base station receivers
abstract
Future base stations for the narrowband cellular standards, e.g. GSM and D-AMPS, will deploy wideband multicarrier receivers. In these receivers the phase noise of the sampling stage is crucial in order to meet the blocking performance specified in the standards. An expression relating the single sideband phase noise power density to carrier ratio of a sampled signal to that of the sampling clock is derived. The implications of the theory for the clock local oscillator (LO) and clock drive amplifier for a GSM-1900 receiver are shown.
Patrik Eriksson, Hannu Tenhunen
ICASSP2
1998 Implementation aspects for noncoherent tracking based on a time-discrete delay-locked loop
abstract
A direct sequence spread spectrum receiver with a square root raised cosine chip-matched filter (CMF) and a noncoherent early-late (NCEL) time-discrete delay-locked loop (DLL) is analyzed. The receiver is studied in regard of the chip over-sampling factor /spl lambda/ and of parameters relating the CMF. The NCEL discriminator's S-curve is derived and used in a linearized loop model from which the DLL's tracking performance is analyzed. The bit error rate simulations in which /spl lambda/ and CMF parameters are varied, can then be done using DLLs which theoretically all have the same tracking capabilities. The results show performance degradations between 0.1 and 0.95 dB for different /spl lambda/ and CMFs when compared to ideal detection.
Henrik Olson, Hannu Tenhunen
PIMRC2
1995 Noise Suppression System Integration Using an Analog Allpass Filter Bank
abstract
A new method to enhance the quality of a telephone connection by suppressing noise and disturbances in speech signals is presented. As the noise is evenly spread over the entire signal band, a nonlinear filtering algorithm is required to improve the signal-to-noise-ratio. The system is being realized as an analog full custom ASIC using switched capacitor (SC)-substructures and CMOS-technology. The advantages of the analog implementation include low power consumption, small integration area and inexpensiveness of the chip. In this paper, the noise suppression algorithm is described and the SC-circuit implementation is presented.
M. Rinne, Tiina Jarske, Hannu Tenhunen, Olli Vainio, Yrjö Neuvo
ISCAS3
1995 VLSI implementation of DS-CDMA receiver using asynchronous design techniques
abstract
This paper describes the implementation of an asynchronous self-timed direct sequence spread spectrum radio receiver. The designed receiver is planned to be used in MINT [l] (Mobile InterNet Router) radio interface for broadband wireless datacommunication. The receiver has been implemented in a 0.8 pm CMOS technology using a standard cell library. The final design contains approximately 100 OOO transistors.
Bengt Oelmann, Henk Martijn, Hannu Tenhunen
PIMRC3
1994 The Walkstation transceiver design
abstract
The design of flexible and efficient future mobile communication systems is a major challenge. The Walkstation Project involves researchers from different areas in order to find a solution via a global system approach. An important task is the investigation of new digital, highly integrated radio interfaces with low cost, small size and low power consumption based on direct sequence CDMA. The simplicity of the design of both the analog and digital parts will allow low power operation and small area in a eventual BiCMOS implementation.>
Daniel Kerek, Hannu Tenhunen, Gerald Q. Maguire Jr., Frank Reichert
VTC2
1991 Methods and Algorithms for Converting IC Designs Between Incompatible Design Systems
abstract
Methods and algorithms are described for converting designs from a geometrical database design system to a design system which is based on basic electrical objects. The feasibility of these algorithms is demonstrated using GDS II and L language as examples. The performance of the system is such that it can be used in design verification by converting mask layouts to a format suitable for simulation.>
Eero Pajarre, Tapani Ritoniemi, Hannu Tenhunen
ICCD3