Axel Jantsch

dblp:38/5744 · DBLP profile ↗
← Back
149ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0003-2251-0004ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 128 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 33 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Computer networks · 2 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 A Conformal Prediction-Based Framework for CPU Load Forecasting: A Black-Box Approach
abstract
To address safety concerns in industrial systems, we propose a framework for forecasting CPU load with respect to a predetermined threshold, allowing customers to add tasks from a predefined library. Existing tools, akin to Windows Task Manager, provide limited insights due to their aggregate nature and high computational overhead. Our approach uses conformal prediction for rapid uncertainty-aware forecasts and Shapley value analysis to quantify individual task contributions to the CPU load. This proof-of-concept framework improves system safety assessment by addressing key research questions in load prediction and validation, paving the way for refined measurement methodologies in industrial applications.
Edin Jelacic, Cristina Cerschi Seceleanu, Peter Backeman, Ning Xiong 0001, Tiberiu Seceleanu, Axel Jantsch
COMPSAC6
2025 Efficient Edge Inference via Entropy and Magnitude-Aware Feature Map Pruning in Partitioned CNNs
abstract
Edge–to–cloud inference for vision based models often stalls on the activation data that must cross the link between an IoT device and its server. This paper integrates Time-dependent Clustering Loss (TCL) with a lightweight, layer-specific novel hybrid L2 + Entropy channel-pruning rule applied exactly at the partition boundary. TCL aligns feature maps to discrete levels, enabling aggressive post-training quantization, while the hybrid pruning criterion removes up to 90 % of channels without noticeable accuracy loss.Experiments on ResNet50, EfficientNetV2-S, and YOLOv10n running on a Jetson Orin Nano node demonstrates clear gains across three realistic uplink capacities. At 100 Mbit s−1, partitioned inference completes in 1.2 ms (ResNet50), 1.1 ms (EfficientNetV2-S) and 2.8 ms (YOLOv10n), yielding approximately 6×, 7× and 18× while speed-ups over an all-server baseline still surpassing full on-device execution. With uplinks of 300 Mbit s−1and 500 Mbit s−1, the communication penalty shrinks, yet the proposed pipeline retains more than 2× acceleration over all-edge processing and limits accuracy degradation to a 1–2 % point drop.Because pruning is confined to a single layer and relies only on post-training activation statistics, integration is simple and incurs negligible overhead. The combined TCL quantization and L2 + Entropy pruning strategy therefore offers a practical, bandwidth-aware solution for deploying modern CNNs in IoT-edge scenarios.
Eiraj Saqib, Oscar Artur Bernd Berg, Isaac Sánchez Leal, Irida Shallari, Silvia Krug, Axel Jantsch, Mattias O'Nils
ICMLA6
2025 Efficient and interpretable raw audio classification with diagonal state space models
abstract
Abstract State Space Models have achieved good performance on long sequence modeling tasks such as raw audio classification. Their definition in continuous time allows for discretization and operation of the network at different sampling rates. However, this property has not yet been utilized to decrease the computational demand on a per-layer basis. We propose a family of hardware-friendly S-Edge models with a layer-wise downsampling approach to adjust the temporal resolution between individual layers. Applying existing methods from linear control theory allows us to analyze state/memory dynamics and provides an understanding of how and where to downsample. Evaluated on the Google Speech Command dataset, our autoregressive/causal S-Edge models range from 8–141k parameters at 90–95% test accuracy in comparison to a causal S5 model with 208k parameters at 95.8% test accuracy. Using our C++17 header-only implementation on an ARM Cortex-M4F the largest model requires 103 sec. inference time with 95.19% test accuracy, and the smallest model with 88.01% test accuracy, requires 0.29 sec. Our solutions cover a design space that spans 17x in model size, 358x in inference latency, and 7.18 percentage points in accuracy.
Matthias Bittner, Daniel Schnöll, Matthias Wess, Axel Jantsch
Mach. Learn.4
2025 Evaluation of Drift Detection Algorithms in the Condition Monitoring Domain
abstract
In condition monitoring, early detection of process signal drifts indicating, e.g., equipment degradation is crucial. exponentially weighted moving average (EWMA), cumulative sum (CUSUM), and discrete average block (DAB)-based drift detectors are statistical and commonly used methods. Each has benefits and limitations, suited to different data types. However, EWMA and CUSUM are fixed mean drift detectors, limiting their applicability and adaptability. This article explores adding dynamic behavior to drift detection methods. We use a wide range of synthetic data based on a real-world manufacturing process. The investigated parameter space includes standard deviation, drift rates, and outliers. Besides, each algorithm has some tuning parameters that define its behavior. Two metrics validate experiments against labeled data. Based on our observations, EWMA performs better for drift detection on average, but CUSUM is superior in detecting very small drifts. Furthermore, we derive guidelines for the choice and application of drift detection in practice.
Alireza Estaji, Maximilian Götzinger, Benedikt Tutzer, Stefan Kollmann, Thilo Sauter, Axel Jantsch
IEEE Trans. Ind. Informatics6
2024 Special Session: Estimation and Optimization of DNNs for Embedded Platforms
abstract
Several state of the art estimation and optimization techniques for CNNs and LLMs on embedded devices are summarized. For LLMs an Activation-aware Weight Quantization and on-the-fly dequantization techniques is presented. For CNNs various pruning algorithms and an integrated optimization and implementation flow is discussed. To estimate inference latency of CNNs on specific hardware platforms, three different techniques are reviewed: A mixed analytic-stochastic model, an analytic model based on step-wise linear functions, and a method that uses a detailed architecture description of the hardware.
Axel Jantsch, Song Han 0003, Lin Meng 0001, Oliver Bringmann 0001, Haotian Tang, Shang Yang, Matthias Wess, Martin Lechner
CODES+ISSS1
2024 Resource Management of Automotive Engine Control Units
abstract
Embedded systems play a crucial role in the contemporary automotive industry, where advanced vehicles integrate numerous electronic control units responsible for diverse functionalities. This paper specifically focuses on the engine control unit within the power-train. With increasing ambitions on reducing emission values and a growing functional environment, Robert Bosch AG anticipates that the resources of the current generation of Engine Control Units (ECUs) will reach their limits within the next 3–5 years. The current development workflow relies on resource optimizations with an iterative expert-based approach, success is achieved through significant manual effort. In this work, we propose a methodology that leverages an existing measure based on task splitting from the early stages of ECU development. Our goal is to streamline the design process for engineers and accelerate ECUs development through a cutting-edge software architecture workflow, aiming to minimize manual effort and expedite project timelines. The proposed approach was integrated and validated on a real-world control unit of the power-train, demonstrating highly significant results, with savings of around${50}\% \pm 10\%$. This considerable reduction in time can be attributed to the increased investment required during the architectural design stage, which subsequently reduces the labor-intensive manual integration effort. Additionally, there has been an achieved reduction of approximately 10% in core load.
Istvan Andras Gergely, Sebastian Rausch, Nahla A. El-Araby, Axel Jantsch
VLSI-SoC4
2023 Self-awareness in Cyber-Physical Systems: Recent Developments and Open Challenges
abstract
Self-aware computing systems enable computing systems to reflect on their actions and behavior. This becomes even more relevant in Cyber-Physical Systems where computing systems have to control and interact with elements in the real world. This paper reports on recent advances made in computational self-awareness for cyber-physical systems.
Lukas Esterle, Nikil Dutt, Christian Gruhl, Peter R. Lewis 0001, Lucio Marcenaro, Carlo S. Regazzoni, Axel Jantsch
DATE7
2023 Multispectral Feature Fusion for Deep Object Detection on Embedded NVIDIA Platforms
abstract
Multispectral images can improve object detection systems' performance due to their complementary information, especially in adverse environmental conditions. To use multispec-tral image data in deep-learning-based object detectors, a fusion of the information from the individual spectra, e.g., inside the neural network, is necessary. This paper compares the impact of general fusion schemes in the backbone of the YOLOv4 object detector. We focus on optimizing these fusion approaches for an NVIDIA Jetson AGX Xavier and elaborating on their impact on the device in physical metrics. We optimize six different fusion architectures in the network's backbone for the TensorRT framework and compare their inference time, power consumption, and object detection performance. Our results show that multispectral fusion approaches with little design effort can benefit resource usage and object detection metrics compared to individual networks.
Thomas Kotrba, Martin Lechner, Omair Sarwar, Axel Jantsch
DATE4
2023 VADAR: A Vision-based Anomaly Detection Algorithm for Railroads
abstract
Detecting damages and anomalies on railroads is a tedious and expensive task. This paper proposes the Vision-based Anomaly Detection Algorithm for Railroads (VADAR), which can find rail damages and foreign objects on the trackbed in monochrome images captured by a train-mounted camera system. VADAR analyzes the input image with three Autoencoders (AEs), a segmentation network, and a one-class classifier. The detection of unknown anomalies justifies our architecture's advantage, i.e., no anomalies are necessary for training VADAR. In experiments with a dataset of over 218,000 images, VADAR achieves a detection accuracy of 95% and a recall rate of 70% for smaller and up to 100% for bigger instances of several anomaly classes. Compared with a state-of-the-art approach which is based on more expensive equipment, VADAR achieves accuracy and recall rates (for anomalies of particular interest) of about 22pps and up to 45pps higher, respectively. With a setting that achieves 83.5% accuracy, VADAR's recall rate outperforms the state-of-the-art approach for every anomaly class and object size.
David Breuss, Maximilian Götzinger, Jenny Vuong, Clemens Reisner, Axel Jantsch
DSD5
2023 Fast, Quantization Aware DNN Training for Efficient HW Implementation
abstract
Quantization of Deep Neural Networks is a central technique to reduce the computation load in embedded devices. Even in quantized Deep Neural Networks (DNNs), the scaler/rescaler following a convolution or dense layer often requires a high bit width multiplication and a shift. Previous work has proposed to remove the multiplier by restricting the quantization method. We propose a Quantisation Aware Training (QAT) approach, which explicitly models the rescaler during training, eliminating the limitations of quantization functions and achieving a 30–35% improvement in training time and a significant reduction in memory requirements compared to the state-of-the-art. GitHub: https://github.com/embedded-machine-learning/FastQATforPOTRescaler
Daniel Schnöll, Matthias Wess, Matthias Bittner, Maximilian Götzinger, Axel Jantsch
DSD5
2023 Energy Profiling of DNN Accelerators
abstract
This paper introduces a novel methodology for assessing the energy efficiency of neural network accelerators at both layer and network granularity. The approach involves extracting per-layer timing reports from recorded power profiles. The power and energy consumption of three prominent neural network accelerators, namely the Intel Neural Compute Stick 2, the Coral Edge TPU, and the NXP i.MX8M Plus is evaluated for three different Deep Neural Networks (DNNs) using this method. The study investigates the relationship between decreasing sampling frequencies and the average error, as well as the detailed energy consumption of individual DNN layers and layer types. The findings reveal that latency outperforms the number of operations per layer as a predictor for both overall and dynamic energy, with errors of 10 % and 100 % respectively. The main conclusions are: a sampling frequency of 200 kHz is necessary to achieve an average error of 5 %; the number of operations is an inadequate predictor of energy consumption; and specific hardware settings significantly influence power and energy consumption, emphasizing the need for their consideration in estimation.
Matthias Wess, Dominik Dallinger, Daniel Schnöll, Matthias Bittner, Maximilian Götzinger, Axel Jantsch
DSD6
2023 Forecasting Critical Overloads based on Heterogeneous Smart Grid Simulation
abstract
Climate change mitigation poses a great challenge for our society. The need to reduce greenhouse gas emissions facilitates the expansion of renewable energy sources and electromobility. This transition is an already ongoing process, and with the worldwide increasing energy consumption, we face the need for automatic control and monitoring of the future electrical grid. To ensure a calculable and stable Low Voltage grid we need reliable load forecasting in order to avoid critical overloads and potential financial losses. This paper presents a novel concept for forecasting critical overloads based on an LSTM recurrent neural network. Our algorithm was tested using a one-year simulation of a rural Low Voltage grid section containing a grid-friendly energy community. Our results show the successful detection of 29 overloads within 12 simulated weeks. We reach a recall of 100% and a precision of 85%. Furthermore, we proved the ability of our LSTM to forecast two weeks with an MAE of 12.41 kW for the month of July. When optimizing the weather forecast data, we can lower this to 6.89 kW.
Matthias Bittner, Daniel Hauer, Christian Stippel, Katharina Scheucher, Robin Sudhoff, Axel Jantsch
ICMLA6
2022 Run Time Power and Accuracy Management with Approximate Circuits
abstract
The ever-expanding need for low-power devices can be approached by implementing approximate computing methods. A restrictive energy budget is met by dropping the concept of fully exact or entirely deterministic computations. We propose a methodology to trade off accuracy with run-time power consumption through Dynamic Partial Reconfiguration (DPR) of Field Programmable Gate Arrays (FPGAs). Optimization is done by switching between predefined design configurations and combining exact and approximate versions of the most power-consuming circuit blocks. We designed a dynamic reconfiguration manager to select and configure the FPGA with the appropriate partial bitstream. The reconfiguration is executed automatically at run time according to the system power state and the accuracy requirement of the running application. The experimental results show that the proposed mechanism can achieve between 10% and 58% power reduction with a maximum error of 0.35 and an average error range of 0.1 beside negligible reconfiguration energy cost.
Nahla A. El-Araby, David Frismuth, Nilson Neves Filho, Axel Jantsch
VLSI-SoC4
2022 Confidence-Enhanced Early Warning Score Based on Fuzzy Logic
abstract
Abstract Cardiovascular diseases are one of the world’s major causes of loss of life. The vital signs of a patient can indicate this up to 24 hours before such an incident happens. Healthcare professionals use Early Warning Score (EWS) as a common tool in healthcare facilities to indicate the health status of a patient. However, the chance of survival of an outpatient could be increased if a mobile EWS system would monitor them during their daily activities to be able to alert in case of danger. Because of limited healthcare professional supervision of this health condition assessment, a mobile EWS system needs to have an acceptable level of reliability - even if errors occur in the monitoring setup such as noisy signals and detached sensors. In earlier works, a data reliability validation technique has been presented that gives information about the trustfulness of the calculated EWS. In this paper, we propose an EWS system enhanced with the self-aware property confidence, which is based on fuzzy logic. In our experiments, we demonstrate that - under adverse monitoring circumstances (such as noisy signals, detached sensors, and non-nominal monitoring conditions) - our proposed Self-Aware Early Warning Score (SA-EWS) system provides a more reliable EWS than an EWS system without self-aware properties.
Maximilian Götzinger, Arman Anzanpour, Iman Azimi, Nima Taherinejad, Axel Jantsch, Amir-Mohammad Rahmani, Pasi Liljeberg
Mob. Networks Appl.5
2021 MLComp: A Methodology for Machine Learning-based Performance Estimation and Adaptive Selection of Pareto-Optimal Compiler Optimization Sequences
abstract
Embedded systems have proliferated in various consumer and industrial applications with the evolution of Cyber-Physical Systems and the Internet of Things. These systems are subjected to stringent constraints so that embedded software must be optimized for multiple objectives simultaneously, namely reduced energy consumption, execution time, and code size. Compilers offer optimization phases to improve these metrics. However, proper selection and ordering of them depends on multiple factors and typically requires expert knowledge. State-of-the-art optimizers facilitate different platforms and applications case by case, and they are limited by optimizing one metric at a time, as well as requiring a time-consuming adaptation for different targets through dynamic profiling. To address these problems, we propose the novel MLComp methodology, in which optimization phases are sequenced by a Reinforcement Learning-based policy. Training of the policy is supported by Machine Learning-based analytical models for quick performance estimation, thereby drastically reducing the time spent for dynamic profiling. In our framework, different Machine Learning models are automatically tested to choose the best-fitting one. The trained Performance Estimator model is leveraged to efficiently devise Reinforcement Learning-based multi-objective policies for creating quasi-optimal phase sequences. Compared to state-of-the-art estimation models, our Performance Estimator model achieves lower relative error (<2%) with up to 50x faster training time over multiple platforms and application domains. Our Phase Selection Policy improves execution time and energy consumption of a given code by up to 12% and 6%, respectively. The Performance Estimator and the Phase Selection Policy can be trained efficiently for any target platform and application domain.
Alessio Colucci, David Juhasz, Martin Mosbeck, Alberto Marchisio, Semeen Rehman, Manfred Kreutzer, Günther Nadbath, Axel Jantsch, Muhammad Shafique 0001
DATE8
2021 MELODI: An Online Platform for Mass Education of Digital Design - HDL to Remote FPGA
abstract
Learning and teaching digital hardware design involves significant efforts on both sides, teachers and students. Hardware Description Languages (HDLs) simplify the design process, where the created designs can be tested either in simulators or using real hardware. For the latter, Field Programmable Gate Arrays (FPGAs) play a crucial role in facilitating and speeding up the prototyping process. Mass E-Learning of design, test, and prototyping Digital hardware (MELODI) is developed for scalable teaching of HDL to a large number of students. The primary goals are to a) minimize the requirements for students and b) reduce the resources required at the university. MELODI provides a complete HDL workflow, including real remote hardware prototyping on FPGAs without the need for any tool at the students’ side.
Friedrich Bauer, Felix Braun, Daniel Hauer, Axel Jantsch, Markus D. Kobelrausch, Martin Mosbeck, Nima Taherinejad, Philipp-Sebastian Vogt
FPL4
2020 Evaluation of Reinforcement Learning Methods for a Self-learning System
David Bechtold, Alexander Wendt, Axel Jantsch
ICAART (2)3
2020 Cognitive Architectures for Process Monitoring - an Analysis
abstract
In smart manufacturing, the demand increases to be able to monitor and adjust process execution during production. It has led to a shift towards distributed, modular automation. A solution seems to be to use a cognitive architecture. It turns out that their generality often makes them unsuitable for specific industrial problems. In this paper, we propose a cognition-inspired architecture design for health monitoring tasks. The problem class is represented by a conveyor belt use case. We then discuss to what extent this architecture matches common implementations of cognitive theories by following a generalized cognitive process.
Alexander Wendt, Stefan Kollmann, Aleksey Bratukhin, Alireza Estaji, Thilo Sauter, Axel Jantsch
INDIN6
2020 Embodied Self-Aware Computing Systems
abstract
Embodied self-aware computing systems are embedded in a physical environment with a rich set of sensors and actuators to interact both with their environment and with their own embodiment. Through this interaction, they learn about their situation, their own state, and their performance. Although they are application specific like traditional embedded systems (ESs), they are significantly more flexible, robust, and autonomous; they can adapt to a wide range of environmental variation and can cope with deterioration and shortcomings of their own performance. As such, embodied self-aware computing systems are an evolution of traditional embedded and cyber-physical systems into the direction of more autonomy, robustness, and flexibility. When traditional ESs operate in a changing world by demanding unchanging and fully characterized computing resources, embodied self-aware computing systems adapt to a changing world and changing computing resources. This article surveys the methods and methodologies used for embodied self-aware computing systems structured along with the faculties of: 1) sensory observation and abstraction; 2) self-aware assessment; and 3) hierarchical goals and control. The discussion is exemplified by application cases in the areas of systems-on-chip, control systems, health monitoring, and condition monitoring in industrial production systems.
Henry Hoffmann, Axel Jantsch, Nikil Dutt
Proc. IEEE2
2020 Machine Learning for Power, Energy, and Thermal Management on Multicore Processors: A Survey
abstract
Due to the high integration density and roadblock of voltage scaling, modern multicore processors experience higher power densities than previous technology scaling nodes. When unattended, this issue might lead to temperature hot spots, that in turn may cause nonuniform aging, accelerate chip failure, impair reliability, and reduce the performance of the system. This paper presents an overview of several research efforts that propose to use machine learning (ML) techniques for power and thermal management on single-core and multicore processors. Traditional power and thermal management techniques rely on a certain a-priori knowledge of the chip's thermal model, as well as information of the workloads/applications to be executed (e.g., transient and average power consumption). Nevertheless, these a-priori information is not always available, and even if it is, it cannot reflect the spatial and temporal uncertainties and variations that come from the environment, the hardware, or from the workloads/applications. Contrarily, techniques based on ML can potentially adapt to varying system conditions and workloads, learning from past events in order to improve themselves as the environment changes, resulting in improved management decisions.
Santiago Pagani, Sai Manoj Pudukotai Dinakarrao, Axel Jantsch, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Self-aware Cyber-Physical Systems
abstract
In this article, we make the case for the new class of Self-aware Cyber-physical Systems. By bringing together the two established fields of cyber-physical systems and self-aware computing, we aim at creating systems with strongly increased yet managed autonomy, which is a main requirement for many emerging and future applications and technologies. Self-aware cyber-physical systems are situated in a physical environment and constrained in their resources, and they understand their own state and environment and, based on that understanding, are able to make decisions autonomously at runtime in a self-explanatory way. In an attempt to lay out a research agenda, we bring up and elaborate on five key challenges for future self-aware cyber-physical systems: (i) How can we build resource-sensitive yet self-aware systems? (ii) How to acknowledge situatedness and subjectivity? (iii) What are effective infrastructures for implementing self-awareness processes? (iv) How can we verify self-aware cyber-physical systems and, in particular, which guarantees can we give? (v) What novel development processes will be required to engineer self-aware cyber-physical systems? We review each of these challenges in some detail and emphasize that addressing all of them requires the system to make a comprehensive assessment of the situation and a continual introspection of its own state to sensibly balance diverse requirements, constraints, short-term and long-term objectives. Throughout, we draw on three examples of cyber-physical systems that may benefit from self-awareness: a multi-processor system-on-chip, a Mars rover, and an implanted insulin pump. These three very different systems nevertheless have similar characteristics: limited resources, complex unforeseeable environmental dynamics, high expectations on their reliability, and substantial levels of risk associated with malfunctioning. Using these examples, we discuss the potential role of self-awareness in both highly complex and rather more simple systems, and as a main conclusion we highlight the need for research on above listed topics.
Kirstie L. Bellman, Christopher Landauer, Nikil Dutt, Lukas Esterle, Andreas Herkersdorf, Axel Jantsch, Nima Taherinejad, Peter R. Lewis 0001, Marco Platzner, Kalle Tammemäe
ACM Trans. Cyber Phys. Syst.6
2020 Introduction to the Special Issue on Self-Aware Cyber-physical Systems
abstract
No abstract available.
Axel Jantsch, Peter R. Lewis 0001, Nikil Dutt
ACM Trans. Cyber Phys. Syst.1
2020 Modular and Distributed Management of Many-Core SoCs
abstract
Many-Core Systems-on-Chip increasingly require Dynamic Multi-objective Management (DMOM) of resources. DMOM uses different management components for objectives and resources to implement comprehensive and self-adaptive system resource management. DMOMs are challenging because they require a scalable and well-organized framework to make each component modular, allowing it to be instantiated or redesigned with a limited impact on other components. This work evaluates two state-of-the-art distributed management paradigms and, motivated by their drawbacks, proposes a new one called Management Application (MA) , along with a DMOM framework based on MA. MA is a distributed application, specific for management, where each task implements a management role. This paradigm favors scalability and modularity because the management design assumes different and parallel modules, decoupled from the OS. An experiment with a task mapping case study shows that MA reduces the overhead of management resources (-61.5%), latency (-66%), and communication volume (-96%) compared to state-of-the-art per-application management. Compared to cluster-based management (CBM) implemented directly as part of the OS, MA is similar in resources and communication volume, increasing only the mapping latency (+16%). Results targeting a complete DMOM control loop addressing up to three different objectives show the scalability regarding system size and adaptation frequency compared to CBM, presenting an overall management latency reduction of 17.2% and an overall monitoring messages’ latency reduction of 90.2%.
Marcelo Ruaro, Anderson C. Sant'Ana, Axel Jantsch, Fernando Gehm Moraes
ACM Trans. Comput. Syst.3
2019 TrojanZero: Switching Activity-Aware Design of Undetectable Hardware Trojans with Zero Power and Area Footprint
abstract
Conventional Hardware Trojan (HT) detection techniques are based on the validation of integrated circuits to determine changes in their functionality, and on non-invasive side-channel analysis to identify the variations in their physical parameters. In particular, almost all the proposed side-channel power-based detection techniques presume that HTs are detectable because they only add gates to the original circuit with a noticeable increase in power consumption. This paper demonstrates how undetectable HTs can be realized with zero impact on the power and area footprint of the original circuit. Towards this, we propose a novel concept of TrojanZero and a systematic methodology for designing undetectable HTs in the circuits, which conceals their existence by gate-level modifications. The crux is to salvage the cost of the HT from the original circuit without being detected using standard testing techniques. Our methodology leverages the knowledge of transition probabilities of the circuit nodes to identify and safely remove expendable gates, and embeds malicious circuitry at the appropriate locations with zero power and area overheads when compared to the original circuit. We synthesize these designs and then embed in multiple ISCAS85 benchmarks using a 65nm technology library, and perform a comprehensive power and area characterization. Our experimental results demonstrate that the proposed TrojanZero designs are undetectable by the state-of-the-art power-based detection methods.
Imran Hafeez Abbassi, Faiq Khalid, Semeen Rehman, Awais M. Kamboh, Axel Jantsch, Siddharth Garg, Muhammad Shafique 0001
DATE5
2019 Goal-Driven Autonomy for Efficient On-chip Resource Management: Transforming Objectives to Goals
abstract
Run-time resource allocation of heterogeneous multi-core systems is challenging with varying workloads and limited power and energy budgets. User interaction within these systems changes the performance requirements, often conflicting with concurrent applications' objective and system constraints. Current resource allocation approaches focus on optimizing fixed objective, ignoring the variation in system and applications' objective at run-time. For an efficient resource allocation, the system has to operate autonomously by formulating a hierarchy of goals. We present goal-driven autonomy (GDA) for on-chip resource allocation decisions, which allows systems to generate and prioritize goals in response to the workload and system dynamic variation. We implemented a proof-of-concept resource management framework that integrates the proposed goal management control to meet power, performance and user requirements simultaneously. Experimental results on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) show the effectiveness of GDA in efficient resource allocation in comparison with existing fixed objective policies.
Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DATE5
2019 Dynamic Computation Migration at the Edge: Is There an Optimal Choice?
abstract
In the era of Fog computing where one can decide to compute certain time-critical tasks at the edge of the network, designers often encounter a question whether the sensor layer provides the optimal response time for a service, or the Fog layer, or their combination. In this context, minimizing the total response time using computation migration is a communication-computation co-optimization problem as the response time does not depend only on the computational capacity of each side. In this paper, we aim at investigating this question and addressing it in certain situations. We formulate this question as a static or dynamic computation migration problem depending on whether certain communication and computation characteristics of the underlying system is known at design-time or not. We first propose a static approach to find the optimal computation migration strategy using models known at design-time. We then make a more realistic assumption that several sources of variation can affect the system's response latency (e.g., the change in computation time, bandwidth, transmission channel reliability, etc.), and propose a dynamic computation migration approach which can adaptively identify the latency optimal computation layer at runtime. We evaluate our solution using a case-study of artificial neural network based arrhythmia classification using a simulation environment as well as a real test-bed.
Sina Shahhosseini, Iman Azimi, Arman Anzanpour, Axel Jantsch, Pasi Liljeberg, Nikil Dutt, Amir-Mohammad Rahmani
ACM Great Lakes Symposium on VLSI4
2019 MemGANs: Memory Management for Energy-Efficient Acceleration of Complex Computations in Hardware Architectures for Generative Adversarial Networks
abstract
Generative Adversarial Networks (GANs) have gained importance because of their tremendous unsupervised learning capability and enormous applications in data generation, for example, text to image synthesis, synthetic medical data generation, video generation, and artwork generation. Hardware acceleration for GANs become challenging due to the intrinsic complex computational phases, which require efficient data management during the training and inference. In this work, we propose a distributed on-chip memory architecture, which aims at efficiently handling the data for complex computations involved in GANs, such as strided convolution or transposed convolution. We also propose a controller that improves the computational efficiency by pre-arranging the data from either the off-chip memory or the computational units before storing it in the on-chip memory. Our architectural enhancement supports to achieve 3.65x performance improvement in state-of-the-art, and reduces the number of read accesses and write accesses by 85% and 75%, respectively.
Muhammad Abdullah Hanif, Muhammad Zuhaib Akbar, Semeen Rehman, Axel Jantsch, Muhammad Shafique 0001
ISLPED5
2019 Distributed SDN architecture for NoC-based many-core SoCs
abstract
In the Software-Defined Networking (SDN) paradigm, routers are generic and programmable forwarding units that transmit packets according to a given policy defined by a software controller. Recent research has shown the potential of such a communication concept for NoC management, resulting in hardware complexity reduction, management flexibility, real-time guarantees, and self-adaptation. However, a centralized SDN controller is a bottleneck for large-scale systems.
Marcelo Ruaro, Nedison Velloso, Axel Jantsch, Fernando Gehm Moraes
NOCS3
2019 Efficient Design-for-Test Approach for Networks-on-Chip
abstract
To achieve high reliability in on-chip networks, it is necessary to test the network continuously with Built-in Self-Tests (BIST) so that the faults can be detected quickly and the number of affected packets can be minimized. However, BIST causes significant performance loss due to data dependencies. We propose EsyTest, a comprehensive test strategy with minimized influence on system performance. EsyTest tests the data path and the control path separately. The data path test starts periodically, but the actual test performs in the free time slots to avoid deactivating the router for testing. A reconfigurable router architecture and an adaptive fault-tolerant routing algorithm are proposed to guarantee the access to the processing core when the associated router is under test. During the whole test procedure of the network, all processing cores are accessible, and thus the system performance is maintained during the test. At the same time, EsyTest provides a full test coverage for the NoC and a better hardware compatibility comparing with the existing test strategies. Under the PARSEC benchmark and different test frequencies, the execution time increases less than 5 percent at the cost of 9.9 percent more area and 4.6 percent more power in comparison with the execution where no test procedure is applied.
Junshi Wang, Masoumeh Ebrahimi, Letian Huang, Qiang Li 0021, Guangjun Li, Axel Jantsch
IEEE Trans. Computers7
2019 Self-Adaptive QoS Management of Computation and Communication Resources in Many-Core SoCs
abstract
Providing quality of service (QoS) for many-core systems with dynamic application admission is challenging due to the high amount of resources to manage and the unpredictability of computation and communication events. Related works propose a self-adaptive QoS mechanism concerned either in communication or computation resources, lacking, however, a comprehensive QoS management of both. Assuming a many-core system with QoS monitoring, runtime circuit-switching establishment, task migration, and a soft real-time task scheduler, this work fills this gap by proposing a novel self-adaptive QoS management. The contribution of this proposal comes with the following features in the QoS management: ( i ) comprehensiveness, by covering communication and computation resources; ( ii ) online, adopting the ODA (Observe, Decide, Act) runtime closed-loop adaptation; and ( iii ) reactive and proactive decisions, by using a dynamic application profile extraction technique, which enables the QoS management to be aware of the profile of running applications, allowing it to take proactive decisions based on a prediction analysis. The proposed QoS management adopts a decentralized organization by partitioning the system in clusters, each one managed by a dedicated processor, making the proposal scalable. Results show that the proactive feature accurately extracts the applications’ profile, and can prevent future QoS violations. The synergy of reactive and proactive decisions was able to sustain QoS, reducing the deadline miss rate by 99.5% with a severe disturbance in communication and computation levels, and avoiding deadline misses up to 70% of system utilization.
Marcelo Ruaro, Axel Jantsch, Fernando Gehm Moraes
ACM Trans. Embed. Comput. Syst.2
2018 SPECTR: Formal Supervisory Control and Coordination for Many-core Systems Resource Management
abstract
Resource management strategies for many-core systems need to enable sharing of resources such as power, processing cores, and memory bandwidth while coordinating the priority and significance of system- and application-level objectives at runtime in a scalable and robust manner. State-of-the-art approaches use heuristics or machine learning for resource management, but unfortunately lack formalism in providing robustness against unexpected corner cases. While recent efforts deploy classical control-theoretic approaches with some guarantees and formalism, they lack scalability and autonomy to meet changing runtime goals. We present SPECTR, a new resource management approach for many-core systems that leverages formal supervisory control theory (SCT) to combine the strengths of classical control theory with state-of-the-art heuristic approaches to efficiently meet changing runtime goals. SPECTR is a scalable and robust control architecture and a systematic design flow for hierarchical control of many-core systems. SPECTR leverages SCT techniques such as gain scheduling to allow autonomy for individual controllers. It facilitates automatic synthesis of the high-level supervisory controller and its property verification. We implement SPECTR on an Exynos platform containing ARM»s big.LITTLE-based heterogeneous multi-processor (HMP) and demonstrate that SPECTR»s use of SCT is key to managing multiple interacting resources (e.g., chip power and processing cores) in the presence of competing objectives (e.g., satisfying QoS vs. power capping). The principles of SPECTR are easily applicable to any resource type and objective as long as the management problem can be modeled using dynamical systems theory (e.g., difference equations), discrete-event dynamic systems, or fuzzy dynamics.
Amir-Mohammad Rahmani, Bryan Donyanavard, Tiago Rogério Mück, Kasra Moazzemi, Axel Jantsch, Onur Mutlu, Nikil Dutt
ASPLOS5
2018 Resource Management for Mixed-Criticality Systems on Multi-core Platforms with Focus on Communication
abstract
The System-on-Chip revolution has not impacted the field of safety-critical systems as widely as the rest of the electronics market. However, it is nowadays acknowledged that sharing resources between applications of different criticality is a key leverage to reduce costs and improve performance. Certification of systems hosting such mixed-criticality application sets requires sufficient isolation between criticality levels, i.e. a low-critical task must not cause a fault in a high-critical task, especially a temporal fault, which is likely when several co-running tasks compete for a shared resource. This digest describes implementation schemes which can be included, at both hardware and software levels, in mixed-critical systems to enforce such isolation.
Robin Arbaud, David Juhasz, Axel Jantsch
DSD3
2018 Trends in On-chip Dynamic Resource Management
abstract
The Complexity of emerging multi/many-core architectures and diversity of modern workloads demands coordinated dynamic resource management methods. We introduce a classification for these methods capturing the utilized resources and metrics. In this work, we use this classification to survey the key efforts in dynamic resource management. We first cover heuristic and optimization methods used to manage resources such as power, energy, temperature, Quality-of-Service (QoS) and reliability of the system. We then identify some of the machine learning based methods used in tuning architectural parameters in computer systems. In many cases, resource managers need to enforce design constraints during runtime with a certain level of guarantee. Hence, we also study the trend in deploying formal control theoretic approaches in order to achieve efficient and robust dynamic resource management.
Kasra Moazzemi, Anil Kanduri, David Juhasz, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DSD7
2018 ADDHard: Arrhythmia Detection with Digital Hardware by Learning ECG Signal
abstract
Anomaly detection in Electrocardiogram (ECG) signals facilitates the diagnosis of cardiovascular diseases i.e., arrhythmias. Existing methods, although fairly accurate, demand a large number of computational resources. Based on the pre-processing of ECG signal, we present a low-complex digital hardware implementation (ADDHard) for arrhythmia detection. ADDHard has the advantages of low-power consumption and a small foot print. ADDHard is suitable especially for resource constrained systems such as body wearable devices. Its implementation was tested with the MIT-BIH arrhythmia database and achieved an accuracy of 97.28% with a specificity of 98.25% on average.
Sai Manoj Pudukotai Dinakarrao, Axel Jantsch
ACM Great Lakes Symposium on VLSI2
2018 Applicability of Context-Aware Health Monitoring to Hydraulic Circuits
abstract
Monitoring is an important aspect of operation and maintenance in virtually every industrial system. However, the extent and methods of monitoring vastly vary in different systems, from fully automated to fully manual. One of the challenges of automated monitoring is the tediousness of, and the extent of engineering time and effort required to develop necessary models or machine learning algorithms for the units to be monitored. Model-free monitoring, on the other hand, can save resources and efforts substantially. However, more often than not they have a very limited scope and application. Such a system is needed, for example, to monitor entire Heating, Ventilation and Air Conditioning (HVAC) systems, consisting of different types of sensors such as temperature, pressure, humidity or flow sensors. Recently, we proposed the Context-Aware Health Monitoring (CAH) system for model-free monitoring of any injective-function black-box, and it was tested successfully on an AC motor. In this paper, we evaluate the CAH system for an entirely different industrial use-case, that is, a hydraulic circuit. The results show the potential for considerable benefits in monitoring HVAC systems. Moreover, in the light of applying CAH to different use-cases which may potentially need a different setup of parameters, we performed a sensitivity analysis on the values of different parameters in the system. The results show the robustness of CAH with regard to the values of these parameters.
Maximilian Götzinger, Edwin Willegger, Nima Taherinejad, Axel Jantsch, Thilo Sauter, Thomas Glatzl, P. Lilieberg
IECON4
2018 Enhancement of Classification of Small Data Sets Using Self-awareness - An Iris Flower Case-Study
abstract
In big-data (Deep) Neural Network (NN) algorithm is often used for classification. However, such a massive mine of data is not always available and a shortage of training data can significantly deteriorate the performance of NNs and other classifiers. Therefore, we propose a self-aware multiple classifier system suitable for “Small-Data” cases. This algorithm uses self-awareness to switch between classifiers to improve its performance. We tested the algorithm for the classification of iris flower species using the Iris standard database. Compared to NN, our algorithm showed up to 17% classification success rate improvement with up to 10 times smaller standard deviation.
Hedyeh A. Kholerdi, Nima Taherinejad, Axel Jantsch
ISCAS3
2018 adBoost: Thermal Aware Performance Boosting Through Dark Silicon Patterning
abstract
Increasing power densities of many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. Efficient management of critical interlinked parameters - power, performance and temperature can improve resource utilization and mitigate dark silicon. In this paper, we present a run-time resource management system for thermal aware performance boosting using a dark silicon aware run-time application mapping strategy. The mapping policy patterns inactive cores among active cores for relatively lower and even distribution of operating temperatures. This provides enough thermal headroom for boosting the frequency of active cores upon performance surges and allows sustained boosting periods, improving the performance further. We design a controller for thermal aware performance boosting that decides on efficient allocation utilization of power budget and thermal headroom obtained from patterning. Our strategy yields up to 37 percent better throughput, 29 percent lower waiting time and up to 2 x longer boosting periods, in comparison with other state-of-the-art run-time mapping policies.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Muhammad Shafique 0001, Axel Jantsch, Pasi Liljeberg
IEEE Trans. Computers5
2018 Weighted Quantization-Regularization in DNNs for Weight Memory Minimization Toward HW Implementation
abstract
Deployment of deep neural networks on hardware platforms is often constrained by limited on-chip memory and computational power. The proposed weight quantization offers the possibility of optimizing weight memory alongside transforming the weights to hardware friendly data types. We apply dynamic fixed point (DFP) and power-of-two (Po2) quantization in conjunction with layer-wise precision scaling to minimize the weight memory. To alleviate accuracy degradation due to precision scaling, we employ quantization-aware fine-tuning. For fine-tuning, quantization-regularization (QR) and weighted QR are introduced to force the trained quantization by adding the distance of the weights to the desired quantization levels as a regularization term to the loss-function. While DFP quantization performs better when allowing different bit-widths for each layer, Po2 quantization in combination with retraining allows higher compression rates for equal bit-width quantization. The techniques are verified on an all-convolutional network. With accuracy degradation of 0.10% points, for DFP with layer-wise precision scaling we achieve compression ratios of 7.34 for CIFAR-10, 4.7 for CIFAR-100, and 9.33 for SVHN dataset.
Matthias Wess, Sai Manoj Pudukotai Dinakarrao, Axel Jantsch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Toggle MUX: How X-Optimism Can Lead to Malicious Hardware
abstract
To highlight a potential threat to hardware security, we propose a methodology to derive a trigger signal from the behavior of Verilog simulation models of field-programmable gate array (FPGA) primitives that behave X-optimistic. We demonstrate our methodology with an example trigger that is implemented using Xilinx 7 Series FPGAs. Experimental results show that it is easily possible to create a trigger signal that is '0' in simulation (pre- and post-synthesis), and '1' in hardware. We show that this kind of trigger is neither detectable by formal equivalence checks, nor by recent Trojan detection techniques. As a countermeasure, we propose to carefully reconsider the utilization of X-optimism in FPGA simulation models.
Christian Krieg, Clifford Wolf, Axel Jantsch, Tanja Zseby
DAC3
2017 Self-awareness in remote health monitoring systems using wearable electronics
abstract
In healthcare, effective monitoring of patients plays a key role in detecting health deterioration early enough. Many signs of deterioration exist as early as 24 hours prior having a serious impact on the health of a person. As hospitalization times have to be minimized, in-home or remote early warning systems can fill the gap by allowing in-home care while having the potentially problematic conditions and their signs under surveillance and control. This work presents a remote monitoring and diagnostic system that provides a holistic perspective of patients and their health conditions. We discuss how the concept of self-awareness can be used in various parts of the system such as information collection through wearable sensors, confidence assessment of the sensory data, the knowledge base of the patient's health situation, and automation of reasoning about the health situation. Our approach to self-awareness provides (i) situation awareness to consider the impact of variations such as sleeping, walking, running, and resting, (ii) system personalization by reflecting parameters such as age, body mass index, and gender, and (iii) the attention property of self-awareness to improve the energy efficiency and dependability of the system via adjusting the priorities of the sensory data collection. We evaluate the proposed method using a full system demonstration.
Arman Anzanpour, Iman Azimi, Maximilian Götzinger, Amir-Mohammad Rahmani, Nima Taherinejad, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DATE7
2017 SAMBA: A self-aware health monitoring architecture for distributed industrial systems
abstract
In the context of Industry 4.0, constantly evolving shop floors generate the need for a highly adaptive and autonomous automation system with lean maintenance, minimum downtime, maximum reliability, and resilience. Future Manufacturing Execution Systems (MESs) will be more complex and dynamic as well as distributed physically and logically. This makes it very difficult, if not impossible, for the conventional centralized architectures to effectively control these vibrant Cyber-Physical Production Systems (CPPSs). To address these issues, we propose Self-Aware health Monitoring and Bio-inspired coordination for distributed Automation systems (SAMBA), an architecture which tackles these challenges. SAMBA increases the ability of the system to intelligently adapt to rapidly changing environment and conditions of future CPPSs.
Lydia Siafara, Hedyeh A. Kholerdi, Aleksey Bratukhin, Nima Taherinejad, Alexander Wendt, Axel Jantsch, Albert Treytl, Thilo Sauter
IECON6
2017 Non-blocking BIST for continuous reliability monitoring of Networks-on-Chip
abstract
To achieve high reliability in on-chip networks, frequent runs of Built-in Self-Test allow the detection of and recovery from faults before they affect packets and the system functionality. However, to test routers, wrappers isolate cores from the network which leads to execution blocking and performance loss. In this paper, we propose a design-for-test reconfigurable router with two alternative bypassing channels. The router architecture allows maintaining the connection between cores and the network during the testing procedure by utilizing the bypassing channels. With the help of an adaptive routing algorithm and a testing strategy, networks can be fully tested at a high testing frequency with <;15% increase of execution time.
Junshi Wang, Letian Huang, Masoumeh Ebrahimi, Qiang Li 0021, Guangjun Li, Axel Jantsch
ISCAS6
2017 Neural network based ECG anomaly detection on FPGA and trade-off analysis
abstract
This paper presents FPGA-based ECG arrhythmia detection using an Artificial Neural Network (ANN). The objective is to implement a neural network based machine learning algorithm on FPGA to detect anomalies in ECG signals, with a better performance and accuracy, compared to statistical methods. An implementation with Principal Component Analysis (PCA) for feature reduction and a multi-layer perceptron (MLP) for classification, proved superior to other algorithms. For implementation on FPGA, the effects of several parameters and simplification on performance, accuracy and power consumption were studied. Piecewise linear approximation for activation functions and fixed point implementation were effective methods to reduce the amount of needed resources. The resulting neural network with twelve inputs and six neurons in the hidden layer, achieved, in spite of the simplifications, the same overall accuracy as simulations with floating point number representation. An accuracy of 99.82% was achieved on average for the MIT-BIH database.
Matthias Wess, Sai Manoj Pudukotai Dinakarrao, Axel Jantsch
ISCAS3
2017 Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient Applications
abstract
Power capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Reliability-Aware Runtime Power Management for Many-Core Systems in the Dark Silicon Era
abstract
Power management of networked many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates considering network characteristics at runtime to achieve better performance while honoring the peak power upper bound. On the other hand, power management has a direct effect on chip temperature, which is the main driver of the aging effects. Therefore, alongside performance fulfillment, the controlling mechanism must also consider the current cores' reliability in its actuator manipulation to enhance the overall system lifetime in the long term. In this paper, we propose a multiobjective dynamic power management technique that uses current power consumption and other network characteristics including the reliability of the cores as the feedback while utilizing fine-grained voltage and frequency scaling and per-core power gating as the actuators. In addition, disturbance rejecter and reliability balancer are designed to help the controller to better smooth power consumption in the short term and reliability in the long term, respectively. Simulations of dynamic workloads and mixed criticality application profiles show that our method not only is effective in honoring the power budget while considerably boosting the system throughput, but also increases the overall system lifetime by minimizing aging effects by means of power consumption balancing.
Amir-Mohammad Rahmani, M. H. Haghbayan, Antonio Miele, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Approximation knob: power capping meets energy efficiency
abstract
Power Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen
ICCAD5
2016 Malicious LUT: a stealthy FPGA trojan injected and triggered by the design flow
abstract
We present a novel type of Trojan trigger targeted at the field-programmable gate array (FPGA) design flow. Traditional triggers base on rare events, such as rare values or sequences. While in most cases these trigger circuits are able to hide a Trojan attack, exhaustive functional simulation and testing will reveal the Trojan due to violation of the specification. Our trigger behaves functionally and formally equivalent to the hardware description language (HDL) specification throughout the entire FPGA design flow, until the design is written by the place-and-route tool as bitstream configuration file. From then, Trojan payload is always on. We implement the trigger signal using a 4-input lookup table (LUT), each of the inputs connecting to the same signal. This lets us directly address the least significant bit (LSB) and most significant bit (MSB) of the LUT. With the remaining 14 bits, we realize a “magic” unary operation. This way, we are able to implement 16 different Triggers. We demonstrate the attack with a simple example and discuss the effectiveness of the recent detection techniques unused circuit identification (UCI), functional analysis for nearly-unused circuit identification (FANCI) and VeriTrust in order to reveal our trigger.
Christian Krieg, Clifford Wolf, Axel Jantsch
ICCAD3
2016 Non-Blocking Testing for Network-on-Chip
abstract
To achieve high reliability in on-chip networks, it is necessary to test the network as frequently as possible to detect physical failures before they lead to system-level failures. A main obstacle is that the circuit under test has to be isolated, resulting in network cuts and packet blockage which limit the testing frequency. To address this issue, we propose a comprehensive network-level approach which could test multiple routers simultaneously at high speed without blocking or dropping packets. We first introduce a reconfigurable router architecture allowing the cores to keep their connections with the network while the routers are under test. A deadlock-free and highly adaptive routing algorithm is proposed to support reconfigurations for testing. In addition, a testing sequence is defined to allow testing multiple routers to avoid dropping of packets. A procedure is proposed to control the behavior of the affected packets during the transition of a router from the normal to the testing mode and vice versa. This approach neither interrupts the execution of applications nor has a significant impact on the execution time. Experiments with the PARSEC benchmarks on an 8x8 NoC-based chip multiprocessors show only 3 percent execution time increase with four routers simultaneously under test.
Letian Huang, Junshi Wang, Masoumeh Ebrahimi, Masoud Daneshtalab, Xiaofan Zhang 0004, Guangjun Li, Axel Jantsch
IEEE Trans. Computers7
2016 Toward Smart Embedded Systems: A Self-aware System-on-Chip (SoC) Perspective
abstract
Embedded systems must address a multitude of potentially conflicting design constraints such as resiliency, energy, heat, cost, performance, security, etc., all in the face of highly dynamic operational behaviors and environmental conditions. By incorporating elements of intelligence, the hope is that the resulting “smart” embedded systems will function correctly and within desired constraints in spite of highly dynamic changes in the applications and the environment, as well as in the underlying software/hardware platforms. Since terms related to “smartness” (e.g., self-awareness, self-adaptivity, and autonomy) have been used loosely in many software and hardware computing contexts, we first present a taxonomy of “self-x” terms and use this taxonomy to relate major “smart” software and hardware computing efforts. A major attribute for smart embedded systems is the notion of self-awareness that enables an embedded system to monitor its own state and behavior, as well as the external environment, so as to adapt intelligently. Toward this end, we use a System-on-Chip perspective to show how the CyberPhysical System-on-Chip (CPSoC) exemplar platform achieves self-awareness through a combination of cross-layer sensing, actuation, self-aware adaptations, and online learning. We conclude with some thoughts on open challenges and research directions.
Nikil Dutt, Axel Jantsch, Santanu Sarma
ACM Trans. Embed. Comput. Syst.2
2016 Weighted Round Robin Configuration for Worst-Case Delay Optimization in Network-on-Chip
abstract
We propose an approach for computing the end-to-end delay bound of individual variable bit-rate flows in an First Input First Output multiplexer with aggregate scheduling under weighted round robin (WRR) policy. To this end, we use a network calculus to derive per-flow end-to-end equivalent service curves employed for computing least upper delay bounds (LUDBs) of the individual flows. Since the real-time applications are going to meet guaranteed services with lower delay bounds, we optimize the weights in WRR policy to minimize the LUDBs while satisfying the performance constraints. We formulate two constrained delay optimization problems, namely, minimize-delay and multiobjective optimization. Multiobjective optimization has both the total delay bounds and their variance as the minimization objectives. The proposed optimizations are solved using a genetic algorithm. A video object plane decoder case study exhibits a 15.4% reduction of the total worst case delays and a 40.3% reduction on the variance of delays when compared with round robin policy. The optimization algorithm has low run-time complexity, enabling quick exploration of the large design spaces. We conclude that an appropriate weight allocation can be a valuable instrument for the delay optimization in on-chip network designs.
Fahimeh Jafari, Axel Jantsch, Zhonghai Lu
IEEE Trans. Very Large Scale Integr. Syst.2
2015 A packet-switched interconnect for many-core systems with BE and RT service
Runan Ma, Zhida Hui, Axel Jantsch
DATE3
2015 Self-Aware Cyber-Physical Systems-on-Chip
abstract
Self-awareness has a long history in biology, psychology, medicine, and more recently in engineering and computing, where self-aware features are used to enable adaptivity to improve a system's functional value, performance and robustness. With complex many-core Systems-on-Chip (SoCs) facing the conflicting requirements of performance, resiliency, energy, heat, cost, security, etc. - in the face of highly dynamic operational behaviors coupled with process, environment, and workload variabilities - there is an emerging need for self-awareness in these complex SoCs. Unlike traditional MultiProcessor Systems-on-Chip (MPSoCs), self-aware SoCs must deploy an intelligent co-design of the control, communication, and computing infrastructure that interacts with the physical environment in real-time in order to modify the system's behavior so as to adaptively achieve desired objectives and Quality-of-Service (QoS). Self-aware SoCs require a combination of ubiquitous sensing and actuation, health-monitoring, and statistical model-building to enable the SoC's adaptation over time and space. After defining the notion of self-awareness in computing, this paper presents the Cyber-Physical System-on-Chip (CPSoC) concept as an exemplar of a self-aware SoC that intrinsically couples on-chip and cross-layer sensing and actuation using a sensor-actuator rich fabric to enable self-awareness.
Nikil Dutt, Axel Jantsch, Santanu Sarma
ICCAD2
2015 Dark silicon aware runtime mapping for many-core systems: A patterning approach
abstract
Limitation on power budget in many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. In such systems, an efficient run-time application mapping approach can considerably enhance resource utilization and mitigate the dark silicon phenomenon. In this paper, we propose a dark silicon aware runtime application mapping approach that patterns active cores alongside the inactive cores in order to evenly distribute power density across the chip. This approach leverages dark silicon to balance the temperature of active cores to provide higher power budget and better resource utilization, within a safe peak operating temperature. In contrast with exhaustive search based mapping approach, our agile heuristic approach has a negligible runtime overhead. Our patterning strategy yields a surplus power budget of up to 17% along with an improved throughput of up to 21% in comparison with other state-of-the-art run-time mapping strategies, while the surplus budget is as high as 40% compared to worst case scenarios.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
ICCD5
2015 Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approach
abstract
Power management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies.
Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ISLPED7
2015 MapPro: Proactive Runtime Mapping for Dynamic Workloads by Quantifying Ripple Effect of Applications on Networks-on-Chip
abstract
Increasing dynamic workloads running on NoC-based many-core systems necessitates efficient runtime mapping strategies. With an unpredictable nature of application profiles, selecting a rational region to map an incoming application is an NP-hard problem in view of minimizing congestion and maximizing performance. In this paper, we propose a proactive region selection strategy which prioritizes nodes that offer lower congestion and dispersion. Our proposed strategy, MapPro, quantitatively represents the propagated impact of spatial availability and dispersion on the network with every new mapped application. This allows us to identify a suitable region to accommodate an incoming application that results in minimal congestion and dispersion. We cluster the network into squares of different radii to suit applications of different sizes and proactively select a suitable square for a new application, eliminating the overhead caused with typical reactive mapping approaches. We evaluated our proposed strategy over different traffic patterns and observed gains of up to 41% in energy efficiency, 28% in congestion and 21% dispersion when compared to the state-of-the-art region selection methods.
M. H. Haghbayan, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
NOCS5
2015 Highway in TDM NoCs
abstract
TDM (Time Division Multiplexing) is a well-known technique to provide QoS guarantees in NoCs. However, unused time slots commonly exist in TDM NoCs. In the paper, we propose a TDM highway technique which can enhance the slot utilization of TDM NoCs. A TDM highway is an express TDM connection composed of special buffer queues, called highway channels (HWCs). It can enhance the throughput and reduce data transfer delay of the connection, while keeping the quality of service (QoS) guarantee on minimum bandwidth and in-order packet delivery. We have developed a dynamic and repetitive highway setup policy which has no dependency on particular TDM NoC techniques and no overhead on traffic flows. As a result, highways can be efficiently established and utilized in various TDM NoCs.
Shaoteng Liu, Zhonghai Lu, Axel Jantsch
NOCS3
2015 A Routing-Level Solution for Fault Detection, Masking, and Tolerance in NoCs
abstract
Faults may occur in numerous locations of a router in a NoC platform. Compared with the faults in the data path, faults in the control path may cause more severe effects which may result in crashing the entire system. Most of the current efforts in literature focus on disabling a router when a fault is detected. Considering this level of coarse-granularity, the functioning parts of a router have to be unnecessarily disabled which may severely affect the performance or functionality of the on-chip network. To cope with this problem, in this paper we propose a mechanism to tolerate faults in the control path which largely avoid disabling a router as long as the fault is not severe. This mechanism is called DMT, standing for three distinguishing characteristics of the proposed method as fault Detection, fault Masking and fault Tolerance. The proposed mechanism can efficiently detect the faults expressed as illegal turns while it has the capability to tolerate faults without a prior knowledge on where and why a fault has happened.
Xiaofan Zhang 0004, Masoumeh Ebrahimi, Letian Huang, Guangjun Li, Axel Jantsch
PDP5
2015 MultiCS: Circuit switched NoC with multiple sub-networks and sub-channels
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
J. Syst. Archit.2
2015 Least Upper Delay Bound for VBR Flows in Networks-on-Chip with Virtual Channels
abstract
Real-time applications such as multimedia and gaming require stringent performance guarantees, usually enforced by a tight upper bound on the maximum end-to-end delay. For FIFO multiplexed on-chip packet switched networks we consider worst-case delay bounds for Variable Bit-Rate (VBR) flows with aggregate scheduling, which schedules multiple flows as an aggregate flow. VBR Flows are characterized by a maximum transfer size ( L ), peak rate ( p ), burstiness (σ), and average sustainable rate (ρ). Based on network calculus, we present and prove theorems to derive per-flow end-to-end Equivalent Service Curves (ESC), which are in turn used for computing Least Upper Delay Bounds (LUDBs) of individual flows. In a realistic case study we find that the end-to-end delay bound is up to 46.9% more accurate than the case without considering the traffic peak behavior. Likewise, results also show similar improvements for synthetic traffic patterns. The proposed methodology is implemented in C++ and has low run-time complexity, enabling quick evaluation for large and complex SoCs.
Fahimeh Jafari, Zhonghai Lu, Axel Jantsch
ACM Trans. Design Autom. Electr. Syst.3
2014 Parallel probe based dynamic connection setup in TDM NoCs
abstract
We propose a Time-Division Multiplexing (TDM) based connection oriented NoC with a novel double time-wheel router architecture combined with a run-time parallel probing setup method. In comparison with traditional TDM connection setup methods, our design has the following advantages: (1) it allocates paths and time slots at run-time; (2) it is fast with predictable and bounded setup latency; (3) it avoids additional resources (no auxiliary network or central processor to find and manage connections); (4) it is fully distributed and therefore it scales nicely with network size. Compared to a packet based setup method, our probe based design can reduce path setup delay by 81% and increase network load by 110% in an 8×8 mesh, while avoiding the auxiliary network. Compared to a centralized method, our solution can double the success rate, while eliminating the central resource for path setup and reducing the wire overhead. Synthesis results suggest that our design is faster and smaller than all comparable solutions.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
DATE2
2014 Dark silicon aware power management for manycore systems under dynamic workloads
abstract
Dark Silicon denotes the phenomenon that, due to thermal and power constraints, the fraction of transistors that can operate at full frequency is decreasing with each technology generation. We propose a PID (Proportional Integral Derivative) controller based dynamic power management method that considers an upper bound on power consumption (called the Thermal Design Power (TDP)). To avoid violation of the TDP constraint for manycore systems running highly dynamic workloads, it provides fine-grained DVFS (Dynamic Voltage and Frequency Scaling) including near-threshold operation. In addition, the method distinguishes applications with hard Real-Time, soft Real-Time and no Real-Time constraints and treats them with appropriate priorities. In simulations with dynamic workloads mixed-critical application profiles, we show that the method is effective in honoring the TDP bound and it can boost system throughput by over 43% compared to a naive TDP scheduling policy.
M. H. Haghbayan, Amir-Mohammad Rahmani, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ICCD6
2014 Performance and network power evaluation of tightly mixed SRAM NUCA for 3D Multi-core Network on Chips
abstract
Last level cache (LLC) is crucial for the performance of chip multiprocessors (CMPs), while power is a significant design concern for 3D CMPs. In this paper, we focus on the SRAM-based Non-Uniform Cache Architecture (NUCA) for 3D Multi-core Network-on-Chip (McNoC) systems. A tightly mixed SRAM NUCA for 3D mesh NoC is presented and analyzed. We evaluate the performance and network power with benchmarks based on a full system simulation framework. Experiment results on 16-core 3D NoC systems show that the tightly mixed NUCA could provide up to 31.71% and on average 5.95% performance improvement compared to a base 3D NUCA scheme. The tightly mixed 3D NUCA NoC can reduce network power consumption in 1.07%-15.74% and 9.64% on average compared to a baseline 3D NoCs. Our analysis and experimental results provide a guideline to design efficient 3D NoCs with stacking NUCA.
Li Li 0003, Zhonghai Lu, Axel Jantsch, Minglun Gao
ISCAS4
2014 A Fair and Maximal Allocator for Single-Cycle On-Chip Homogeneous Resource Allocation
abstract
Traditional allocators for network-on-chip (NoC) routers suffer from either poor-matching quality or limited fairness. We propose a waterfall (WTF) allocator targeting homogeneous resource allocation, which provides single-cycle maximal matching while guaranteeing strong fairness based on the round-robin principle. It can be implemented with a loop-free structure. In 90 nm technology, the allocator operates at about 1 GHz clock frequency. We compare WTF with wave-front, separable-input-first, and separable-output-first allocators and find that it is at least 10% smaller, has 50% less delay under high load, and uses 3% less power than any of these alternatives. Also, WTF is at least as fair or clearly fairer. We also find that in a 4×4 circuit switched NoC the use of WTF gives up to 20% higher network performance.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Analysis and Evaluation of Circuit Switched NoC and Packet Switched NoC
abstract
Circuit switched NoC has, compared to packet switching, a longer setup time, guaranteed throughput and latency, higher clock frequency, lower HW complexity, and higher energy efficiency. Depending on packet size and throughput requirements they exhibit better or worse performance. In this paper we designed a circuit switched NoC and compared that with packet switched NoC. By speculation and analysis, we propose that, as packet size increases, performance decreases for packet switched NoC, while it increases for circuit switched NoC. By close examination on the router architecture, we suggest that circuit switched NoC can operate at a higher clock frequency than packet switched NoC, and thus at zero load above a certain packet size circuit switched NoC could be better than packet switched NoC in packet delay. Experiment results support our intuitions and analysis. We find the cross-over point, above which circuit switching has lower latency, is around 30 flits/packet under low load and 60-70 flits/packet under high network load.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
DSD2
2013 Scalability Analysis of Memory Consistency Models in NoC-Based Distributed Shared Memory SoCs
abstract
We analyze the scalability of six memory consistency models in network-on-chip (NoC)-based distributed shared memory multicore systems: 1) protected release consistency (PRC); 2) release consistency (RC); 3) weak consistency (WC); 4) partial store ordering (PSO); 5) total store ordering (TSO); and 6) sequential consistency (SC). Their realizations are based on a transaction counter and an address-stack-based approach. The scalability analysis is based on different workloads mapped on various sizes of networks using different problem sizes. For the experiments, we use Nostrum NoC-based configurable multicore platform with a 2-D mesh topology and a deflection routing algorithm. Under the synthetic workloads, the average execution time for the PRC, RC, WC, PSO, and TSO models in the 8 × 8 network (64-cores) is reduced by 32.3%, 28.3%, 20.1%, 13.8%, and 9.9% over the SC model, respectively. For the application workloads, as the network size grows, the average execution time under these relaxed memory models decreases with respect to the SC model depending on the application and its match to the architecture. The performance improvement of the PRC and RC models over the SC model tends to be higher than 50% as observed in the experiments, when the system is further scaled up. The area cost in the network interface for the relaxed memory models is increased by less than 4% over the SC model.
Abdul Naeem, Axel Jantsch, Zhonghai Lu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 Addressing Transient and Permanent Faults in NoC With Efficient Fault-Tolerant Deflection Router
abstract
Continuing decrease in the feature size of integrated circuits leads to increases in susceptibility to transient and permanent faults. This paper proposes a fault-tolerant solution for a bufferless network-on-chip, including an on-line fault-diagnosis mechanism to detect both transient and permanent faults, a hybrid automatic repeat request, and forward error correction link-level error control scheme to handle transient faults and a reinforcement-learning-based fault-tolerant deflection routing (FTDR) algorithm to tolerate permanent faults without deadlock and livelock. A hierarchical-routing-table-based algorithm (FTDR-H) is also presented to reduce the area overhead of the FTDR router. Synthesized results show that, compared with the FTDR router, the FTDR-H router can reduce the area by 27% in an 88 network. Simulation results demonstrate that under synthetic workloads, in the presence of permanent link faults, the throughput of an 8 8 network with FTDR and FTDR-H algorithms are 14% and 23% higher on average than that with the fault-on-neighbor (FoN) aware deflection routing algorithm and the cost-based deflection routing algorithm, respectively. Under real application workloads, the FTDR-H algorithm achieves 20% less hop counts on average than that of the FoN algorithm. For transient faults, the performance of the FTDR router can achieve graceful degradation even at a high fault rate. We also implement the fault-tolerant deflection router which can achieve 400 MHz in TSMC 65-nm technology.
Chaochao Feng, Zhonghai Lu, Axel Jantsch, Minxuan Zhang, Zuocheng Xing
IEEE Trans. Very Large Scale Integr. Syst.3
2013 An Analytical Latency Model for Networks-on-Chip
abstract
We propose an analytical model based on queueing theory for delay analysis in a wormhole-switched network-on-chip (NoC). The proposed model takes as input an application communication graph, a topology graph, a mapping vector, and a routing matrix, and estimates average packet latency and router blocking time. It works for arbitrary network topology with deterministic routing under arbitrary traffic patterns. This model can estimate per-flow average latency accurately and quickly, thus enabling fast design space exploration of various design parameters in NoC designs. Experimental results show that the proposed analytical model can predict the average packet latency more than four orders of magnitude faster than an accurate simulation, while the computation error is less than 10% in non-saturated networks for different system-on-chip platforms.
Abbas Eslami Kiasari, Zhonghai Lu, Axel Jantsch
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Analytical approaches for performance evaluation of networks-on-chip
abstract
This tutorial reviews four popular mathematical formalisms -- dataflow analysis, schedulability analysis, network calculus, and queueing theory -- and how they have been applied to the analysis of Network-on-Chip (NoC) performance. We review the basic concepts and results of each formalism and provide examples of how they have been used in on-chip communication performance analysis. The tutorial also discusses the respective strengths and weaknesses of each formalism, their suitability for a specific purpose, and the attempts that have been made to bridge these analytical approaches. Finally, we conclude the tutorial by discussing open research issues.
Abbas Eslami Kiasari, Axel Jantsch, Marco Bekooij, Alan Burns 0001, Zhonghai Lu
CASES2
2012 Worst-case delay analysis of Variable Bit-Rate flows in network-on-chip with aggregate scheduling
abstract
Aggregate scheduling in routers merges several flows into one aggregate flow. We propose an approach for computing the end-to-end delay bound of individual flows in a FIFO multiplexer under aggregate scheduling. A synthetic case study exhibits that the end-to-end delay bound is up to 33.6% tighter than the case without considering the traffic peak behavior.
Fahimeh Jafari, Axel Jantsch, Zhonghai Lu
DATE2
2012 Parallel probing: Dynamic and constant time setup procedure in circuit switching NoC
abstract
We propose a circuit switching Network-on-chip with a parallel probe searching setup method, which can search the entire network in constant time, only dependent on the network size but independent of the network load. Under a specific search policy, the setup procedure is guaranteed to terminate in time 3D+6 cycles, where D is the geometric distance between source and destination. If a path can be found, the method succeeds in 3D+6 cycles; if a path cannot be found, it fails in maximum 3D+6 cycles. Compared to previous work, our method can reduce the setup time and enhance the success rate of setups. Our experiments show that compared with a sequential probe searching method, this method can reduce the search time by up to 20%. Compared with a centralized channel allocator method, this method can enhance the success rate by up to 20%.
Shaoteng Liu, Axel Jantsch, Zhonghai Lu
DATE2
2012 Architecture Support and Comparison of Three Memory Consistency Models in NoC Based Systems
abstract
We propose a novel hardware support for three relaxed memory models, Release Consistency (RC), Partial Store Ordering (PSO) and Total Store Ordering (TSO) in Network-on-Chip (NoC) based distributed shared memory multicore systems. The RC model is realized by using a Transaction Counter and an Address Stack based approach to enforce the required global orders on the shared memory operations. The PSO and TSO models are realized by using a Write Transaction Counter and a Write Address Stack based approach to enforce the required global orders on the shared memory operations. In the experiments, we use a configurable platform based on a 2D mesh NoC using deflection routing policy. The results show that under synthetic workloads, the average execution time for the RC, PSO and TSO models in 8×8 network (64 cores) is reduced by 35.8%, 22.7% and 16.5% over the sequential consistency (SC) model, respectively. The average speedup for the RC, PSO and TSO models in 8×8 network under different application workloads is increased by 34.3%, 10.6% and 8.9% over the SC model, respectively. The area cost for the TSO, PSO and RC models is increased by less than 2% over the SC model at the interface to the processor.
Abdul Naeem, Axel Jantsch, Zhonghai Lu
DSD2
2012 System-level evaluation of sensor networks deployment strategies: Coverage, lifetime and cost
abstract
In wireless sensor networks, sensor nodes can be organized either randomly or deterministically according to regular deployment patterns. Due to the trade-offs between performance and cost, evaluating the advantages and disadvantages of node deployment strategies are fundamental issues to be solved. In this paper, we present a framework for analyzing node deployment schemes in terms of three performance metrics: coverage, lifetime, and cost. Based on the proposed coverage analysis model, energy model and cost model, we compare the performance of two node deployment schemes: rectangle mesh and uniformly random. The results show that the rectangle mesh scheme is generally better than the uniformly random scheme in terms of coverage and network lifetime. Our method can be used to evaluate the benefits of different deployment schemes and thus provide guidelines for network designers.
Huimin She, Zhonghai Lu, Axel Jantsch
IWCMC3
2012 Performance Analysis of Reconfigurations in Adaptive Real-Time Streaming Applications
abstract
We propose a performance analysis framework for adaptive real-time synchronous data flow streaming applications on runtime reconfigurable FPGAs. As the main contribution, we present a constraint based approach to capture both streaming application execution semantics and the varying design concerns during reconfigurations. With our event models constructed as cumulative functions on data streams, we exploit a novel compile-time analysis framework based on iterative timing phases. Finally, we implement our framework on a public domain constraint solver, and illustrate its capabilities in the analysis of design trade-offs due to reconfigurations with experiments.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
ACM Trans. Embed. Comput. Syst.3
2011 Power-efficient tree-based multicast support for Networks-on-Chip
abstract
In this paper, a novel hardware support for multicast on mesh Networks-on-Chip (NoC) is proposed. It supports multicast routing on any shape of tree-based paths. Two power-efficient tree-based multicast routing algorithms, Optimized tree (OPT) and Left-XY-Right-Optimized tree (LXYROPT) are also proposed. XY tree-based (XYT) algorithm and multiple unicast copies (MUC) are also implemented on the router as baselines. Along with the increase of the destination size, compared with MUC, OPT and LXYROPT achieve a remarkable improvement in both latency and throughput while the average power consumption is reduced by 50% and 45%, respectively. Compared with XYT, OPT is 10% higher in latency but gains 17% saving in power consumption. LXYROPT is 3% lower in latency and 8% lower in power consumption. In some cases, OPT and LXYROPT give power saving up to 70% less than the XYT.
Wenmin Hu, Zhonghai Lu, Axel Jantsch, Hengzhu Liu
ASP-DAC3
2011 Realization and performance comparison of sequential and weak memory consistency models in network-on-chip based multi-core systems
abstract
This paper studies realization and performance comparison of the sequential and weak consistency models in the network-on-chip (NoC) based distributed shared memory (DSM) multi-core systems. Memory consistency constrains the order of shared memory operations for the expected behavior of the multi-core systems. Both the consistency models are realized in the NoC based multi-core systems. The performance of the two consistency models are compared for various sizes of networks using regular mesh topologies and deflection routing algorithm. The results show that the weak consistency improves the performance by 46.17% and 33.76% on average in the code and consistency latencies over the sequential consistency model, due to relaxation in the program order, as the system grows from single core to 64 cores.
Abdul Naeem, Zhonghai Lu, Axel Jantsch
ASP-DAC4
2011 Realization and Scalability of Release and Protected Release Consistency Models in NoC Based Systems
abstract
This paper studies the realization and scalability of release and protected release consistency models in Network-on-Chip (NoC) based Distributed Shared Memory (DSM) multi-core systems. The protected release consistency (PRC) model is proposed as an extension of the release consistency (RC) model and provides further relaxation in the shared memory operations. The realization schemes of RC and PRC models use a transaction counter in each node of the NoC based multi-core (McNoC) systems. Further, we study the scalability of these RC and PRC models and evaluate their performance in the McNoC platform. A configurable NoC based platform with 2D mesh topology and deflection routing algorithm is used in the tests. We experiment both with synthetic and application workloads. The performance of the RC and PRC models are compared using sequential consistency (SC) as the baseline. The experiments show that the average code execution time for the PRC model in 8x8 network (64 cores) is reduced by 30.5% over SC, and by 6.5% over RC model. Average data execution time in the 8x8 network for the PRC model is reduced by almost 37% over SC and by 8.8% over RC. The increase in area for the PRC of RC is about 880 gates in the network interface ( 1.7% ).
Abdul Naeem, Axel Jantsch, Zhonghai Lu
DSD2
2011 Modeling the computational efficiency of 2-D and 3-D silicon processors for early-chip planning
abstract
Hierarchical models from physical to system-level are proposed for architectural exploration of high-performance silicon systems to quantify the performance and cost trade offs for 2-D and 3-D IC implementations. We show that 3-D systems can reduce interconnect delay and energy by up to an order of magnitude over 2-D, with an increase of 20-30% in performance-per-watt for every doubling of stack height. Contrary to previous analysis, the improved energy efficiency is achievable at a favorable cost. The models are packaged as a standalone tool and can provide fast estimation of coarse-grain performance and cost limitations for a variety of processing systems to be used at the early chip-planning phase of the design cycle.
Matt Grange, Axel Jantsch, Roshan Weerasekera, Dinesh Pamunuwa
ICCAD2
2011 Output process of variable bit-rate flows in on-chip networks based on aggregate scheduling
abstract
In NoCs often several flows are merged into one aggregate flow due to heavy resource sharing. For strengthening formal performance analysis, we propose an improved model for an output flow of a FIFO multiplexer under aggregate scheduling. The model of the aggregate flow is formally proven and can serve as the basis for a stringent worst case delay and buffer analysis.
Fahimeh Jafari, Axel Jantsch, Zhonghai Lu
ICCD2
2011 Optimal network architectures for minimizing average distance in k-ary n-dimensional mesh networks
abstract
A general expression for the average distance for meshes of any dimension and radix, including unequal radices in different dimensions, valid for any traffic pattern under zero-load condition is formulated rigorously to allow its calculation without network-level simulations. The average distance expression is solved analytically for uniform random traffic and for a set of local random traffic patterns. Hot spot traffic patterns are also considered and the formula is empirically validated by cycle true simulations for uniform random, local, and hot spot traffic. Moreover, a methodology to attain closed-form solutions for other traffic patterns is detailed. Furthermore, the model is applied to guide design decisions. Specifically, we show that the model can predict the optimal 3-D topology for uniform and local traffic patterns. It can also predict the optimal placement of hot spots in the network. The fidelity of the approach in suggesting the correct design choices even for loaded and congested networks is surprising. For those cases we studied empirically it is 100%.
Matt Grange, Roshan Weerasekera, Dinesh Pamunuwa, Axel Jantsch, Awet Yemane Weldezion
NOCS4
2011 Stochastic coverage in event-driven sensor networks
abstract
One of the primary tasks of sensor networks is to detect events in a field of interest (FoI). To quantify how well events are detected in such networks, coverage of events is a fundamental problem to be studied. However, traditional studies mostly focus on analyzing the coverage of the FoI, which is usually called area coverage. In this paper, we propose an analytic method to evaluate the performance of event coverage in sensor networks with randomly deployed sensor nodes and stochastic event occurrences. We provide formulas to calculate the probabilities of event coverage and event missing. The numerical results show how these two probabilities change with the sensor and event densities. Moreover, simulations are conducted to validate the analytic method. This method can provide guidelines for determining the amount of sensor nodes to achieve a certain level of coverage in event-driven sensor networks.
Huimin She, Zhonghai Lu, Axel Jantsch, Dian Zhou, Lirong Zheng 0001
PIMRC3
2011 Network-on-Chip multicasting with low latency path setup
abstract
A low-latency path setup approach with multiple setup packets for parallel set is presented. It reduces the header overhead compared to multiaddress encoding. Further, we propose four variants of deadlock-free multicast routing algorithms using different subpath generation methods, different destination partitioning, and channel sharing strategies. Experimental results show that the quatuor partitions path-like tree outperforms other algorithms.
Wenmin Hu, Zhonghai Lu, Axel Jantsch, Hengzhu Liu, Botao Zhang 0002, Dongpei Liu
VLSI-SoC3
2011 3-D integration and the limits of silicon computation
abstract
The intrinsic computational efficiency (ICE) of silicon defines the upper limit of the amount of computation within a given technology and power envelope. The effective computational efficiency (ECE) and the effective computational density (ECD) of silicon, by taking computation, memory and communication into account, offer a more realistic upper bound for computation of a given technology. Among other factors, they consider how distributed the memory is, how much area is occupied by computation, memory and interconnect, and the geometric properties of 3-D stacked technology with through silicon vias (TSV) as vertical links. We use the ECE and ECD to study the limits of performance under different memory distribution, power, thermal and cost constraints for various 2-D and 3-D topologies, in current and future technology nodes.
Dinesh Pamunuwa, Matt Grange, Roshan Weerasekera, Axel Jantsch
VLSI-SoC4
2011 Modeling and analysis of Rayleigh fading channels using stochastic network calculus
abstract
Deterministic network calculus (DNC) is not suitable for deriving performance guarantees for wireless networks due to their inherently random behaviors. In this paper, we develop a method for Quality of Service (QoS) analysis of wireless channels subject to Rayleigh fading based on stochastic network calculus. We provide closed-form stochastic service curve for the Rayleigh fading channel. With this service curve, we derive stochastic delay and backlog bounds. Simulation results verify that the bounds are reasonably tight. Moreover, through numerical experiments, we show the method is not only capable of deriving stochastic performance bounds, but also can provide guidelines for designing transmission strategies in wireless networks.
Huimin She, Zhonghai Lu, Axel Jantsch, Dian Zhou, Lirong Zheng 0001
WCNC3
2010 Constrained global scheduling of streaming applications on MPSoCs
abstract
We present a global scheduling framework for synchronous data flow (SDF) streaming applications on MPSoCs, based on optimized computation and contention-free routing. The global scheduling of processors computing and communication transactions are formulated as constraint based problem, to avoid the scheduling overhead in TDMA-like heuristic schemes. A public domain constraint solver is exploited to solve the NP-complete scheduling efficiently, together with problem specific constraint modeling techniques. Experimental results show that the proposed framework can achieve a high predictable application throughput with minimized buffer cost. For instance, for applications in communication domain, higher throughput (up to 87%) has been observed with less buffer cost, compared to scenarios considering the heuristic scheduling overhead.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
ASP-DAC3
2010 Supporting Distributed Shared Memory on multi-core Network-on-Chips using a dual microcoded controller
abstract
Supporting Distributed Shared Memory (DSM) is essential for multi-core Network-on-Chips for the sake of reusing huge amount of legacy code and easy programmability. We propose a microcoded controller as a hardware module in each node to connect the core, the local memory and the network. The controller is programmable where the DSM functions such as virtual-to-physical address translation, memory access and synchronization etc. are realized using microcode. To enable concurrent processing of memory requests from the local and remote cores, our controller features two mini-processors, one dealing with requests from the local core and the other from remote cores. Synthesis results suggest that the controller consumes 51k gates for the logic and can run up to 455 MHz in 130 nm technology. To evaluate its performance, we use synthetic and application workloads. Results show that, when the system size is scaled up, the delay overhead incurred by the controller may become less significant when compared with the network delay. In this way, the delay efficiency of our DSM solution is close to hardware solutions on average but still have all the flexibility of software solutions.
Zhonghai Lu, Axel Jantsch, Shuming Chen
DATE3
2010 Optimal regulation of traffic flows in networks-on-chip
abstract
We have proposed (¿, ¿)-based flow regulation to reduce delay and backlog bounds in SoC architectures, where ¿ bounds the traffic burstiness and ¿ the traffic rate. The regulation is conducted per-flow for its peak rate and traffic burstiness. In this paper, we optimize these regulation parameters in networks on chips where many flows may have conflicting regulation requirements. We formulate an optimization problem for minimizing total buffers under performance constraints. We solve the problem with the interior point method. Our case study results exhibit 48% reduction of total buffers and 16% reduction of total latency for the proposed problem. The optimization solution has low run-time complexity, enabling quick exploration of large design space.
Fahimeh Jafari, Zhonghai Lu, Axel Jantsch, Mohammad Hossein Yaghmaee Moghaddam
DATE3
2010 FPGA-based adaptive computing for correlated multi-stream processing
abstract
In conventional static implementations for correlated streaming applications, computing resources may be in-efficiently utilized since multiple stream processors may supply their sub-results at asynchronous rates for result correlation or synchronization. To enhance the resource utilization efficiency, we analyze multi-streaming models and implement an adaptive architecture based on FPGA Partial Reconfiguration (PR) technology. The adaptive system can intelligently schedule and manage various processing modules during run-time. Experimental results demonstrate up to 78.2% improvement in throughput-per-unit-area on unbalanced processing of correlated streams, as well as only 0.3% context switching overhead in the overall processing time in the worst-case.
Ming Liu 0011, Zhonghai Lu, Wolfgang Kuehn, Axel Jantsch
DATE4
2010 Pareto efficient design for reconfigurable streaming applications on CPU/FPGAs
abstract
We present a Pareto efficient design method for multi-dimensional optimization of run-time reconfigurable streaming applications on CPU/FPGA platforms, which automatically allocates applications with optimized buffer requirement and software/hardware implementation cost. At the same time, application performance is guaranteed with sustainable throughput during run-time reconfigurations. As the main contribution, we formulate the constraint based application allocation, scheduling, and reconfiguration analysis, and propose a design Pareto-point calculation flow. A public domain solver - Gecode is used in solutions finding. The capability of our method has been exemplified by two cases studies on applications from media and communication domains.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
DATE3
2010 Theorem proving techniques for the formal verification of NoC communications with non-minimal adaptive routing
abstract
This paper focuses on the formal verification of communications in Networks on Chip. We describe how an enhanced version of the GeNoC proof methodology has been applied to the Nostrum NoC which encompasses various non-trivial features such as a deflective non-minimal routing algorithm. We demonstrate how the features of the Nostrum protocol layers can be captured by the current version of GeNoC that enables a step-by-step formalization of communication operations while taking various protocol details into consideration. We prove that packets arrive properly and that packets are never lost. Also, we prove the soundness of the Nostrum data link, network and transport layers.
Amr Helmy, Laurence Pierre, Axel Jantsch
DDECS3
2010 HetMoC: Heterogeneous Modelling in SystemC
Jun Zhu 0011, Ingo Sander, Axel Jantsch
FDL3
2010 Scalability of weak consistency in NoC based multicore architectures
abstract
In Multicore Network-on-Chip, it is preferable to realize distributed but shared memory (DSM) in order to reuse the huge amount of legacy code and easy programming. Within DSM systems, memory consistency is a critical issue since it affects not only performance but also the correctness of programs. In this paper, we investigate the scalability of the weak consistency model, which may be implemented using a transaction counter. The experimental results compare synchronization latencies for various network sizes, topologies and lock positions in the network. Average synchronization latency rises exponentially for mesh and torus topologies as the network size grows. However, torus improves the synchronization latency in comparison to mesh. For mesh topology network average synchronization latency is also slightly affected by the lock position with respect to the network center.
Abdul Naeem, Zhonghai Lu, Axel Jantsch
ISCAS4
2010 A Worst Case Performance Model for TDM Virtual Circuit in NoCs
Axel Jantsch
NPC2
2010 Buffer Optimization in Network-on-Chip Through Flow Regulation
abstract
For network-on-chip (NoC) designs, optimizing buffers is an essential task since buffers are a major source of cost and power consumption. This paper proposes flow regulation and has defined a regulation spectrum as a means for system-on-chip architects to control delay and backlog bounds. The regulation is performed per flow for its peak rate and burstiness. However, many flows may have conflicting regulation requirements due to interferences with each other. Based on the regulation spectrum, this paper optimizes the regulation parameters aiming for buffer optimization. Three timing-constrained buffer optimization problems are formulated, namely, buffer size minimization, buffer variance minimization, and multiobjective optimization, which has both buffer size and variance as minimization objectives. Minimizing buffer variance is also important because it affects the modularity of routers and network interfaces. A realistic case study exhibits 62.8% reduction of total buffers, 84.3% reduction of total latency, and 94.4% reduction on the sum of variances of buffers. Likewise, the experimental results demonstrate similar improvements in the case of synthetic traffic patterns. The optimization algorithm has low run-time complexity, enabling quick exploration of large design spaces. This paper concludes that optimal flow regulation can be a highly valuable instrument for buffer optimization in NoC designs.
Fahimeh Jafari, Zhonghai Lu, Axel Jantsch, Mohammad Hossein Yaghmaee Moghaddam
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 Guest Editorial: Special Section on the ACM/IEEE Symposium on Networks-on-Chip 2009
abstract
The four papers in this special section are extended versions of papers presented at the 3rd ACM/IEEE Symposium on Networks-on-Chip (NOCS) in San Diego, CA, in 2009.
Radu Marculescu, Axel Jantsch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 Flow regulation for on-chip communication
abstract
We propose (σ, ρ)-based flow regulation as a design instrument for System-on-Chip (SoC) architects to control quality-of-service and achieve cost-effective communication, where σ bounds the traffic burstiness and ρ the traffic rate. This regulation changes the burstiness and timing of traffic flows, and can be used to decrease delay and reduce buffer requirements in the SoC infrastructure. In this paper, we define and analyze the regulation spectrum, which bounds the upper and lower limits of regulation. Experiments on a Network-on-Chip (NoC) with guaranteed service demonstrate the benefits of regulation. We conclude that flow regulation may exert significant positive impact on communication performance and buffer requirements.
Zhonghai Lu, Mikael Millberg, Axel Jantsch, Alistair C. Bruce, Pieter van der Wolf, Tomas Henriksson
DATE3
2009 Priority based forced requeue to reduce worst-case latencies for bursty traffic
abstract
In this paper we introduce priority based forced requeue to decrease worst-case latencies in NoCs offering best effort services. Forced Requeue is to prematurely lift out low priority packets from the network and requeue them outside using priority queues. The first benefit of this approach, applicable to any NoC offering best effort services, is that packets that have not yet entered the network now compete with packets inside the network and hence tighter bounds on admission times can be given. The second benefit - which is more specific to deflective routing as in the Nostrum NoC - is that packet ldquoreshufflingrdquo dramatically reduces the latency inside the network for bursty traffic due to a lowered risk of collisions at the exit of the network. This paper studies the Forced Requeuing on a mesh with varying burst sizes and traffic scenarios. The experimental results show a 50% reduction in worst-case latency from a system perspective thanks to a reshaped latency distribution whilst keeping the average latency the same.
Mikael Millberg, Axel Jantsch
DATE2
2009 Buffer minimization of real-time streaming applications scheduling on hybrid CPU/FPGA architectures
abstract
We address the problem of real-time streaming applications scheduling on hybrid CPU/FPGA architectures. The main contribution is a two-step approach to minimize the buffer requirement for streaming applications with throughput guarantees. A novel declarative way of constraint based scheduling for real-time hybrid SW/HW systems is proposed, while the application throughput is guaranteed by periodic phases in execution. We use a voice-band modem application to exemplify the scheduling capabilities of our method. The experimental results show the advantages of our techniques in both less buffer requirement and higher throughput guarantees compared to the traditional PAPS method.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
DATE3
2009 Run-time Partial Reconfiguration speed investigation and architectural design space exploration
abstract
Run-time partial reconfiguration (PR) speed is significant in applications especially when fast IP core switching is required. In this paper, we propose to use direct memory access (DMA), master (MST) burst, and a dedicated block RAM (BRAM) cache respectively to reduce the reconfiguration time. Based on the Xilinx PR technology and the Internal Configuration Access Port (ICAP) primitive in the FPGA fabric, we discuss multiple design architectures and thoroughly investigate their performance with measurements for different partial bitstream sizes. Compared to the reference OPB HWICAP and XPS HWICAP designs, experimental results show that DMA HWICAP and MST HWICAP reduce the reconfiguration time by one order of magnitude, with little resource consumption overhead. The BRAM HWICAP design can even approach the reconfiguration speed limit of the ICAP primitive at the cost of large block RAM utilization.
Ming Liu 0011, Wolfgang Kuehn, Zhonghai Lu, Axel Jantsch
FPL4
2009 High-level estimation and trade-off analysis for adaptive real-time systems
abstract
We propose a novel design estimation method for adaptive streaming applications to be implemented on a partially reconfigurable FPGA. Based on experimental results we enable accurate design cost estimates at an early design stage. Given the size and computation time of a set of configurations, which can be derived through logic synthesis, our method gives estimates for configuration parameters, such as bitstream sizes, computation and reconfiguration times. To fulfil the system's throughput requirements, the required FIFO buffer sizes are then calculated using a hybrid analysis approach based on integer linear programming and simulation. Finally, we are able to calculate the total design cost as the sum of the costs for the FPGA area, the required configuration memory and the FIFO buffers. We demonstrate our method by analysing non-obvious trade-offs for a static and dynamic implementation of adaptivity.
Ingo Sander, Jun Zhu 0011, Axel Jantsch, Andreas Herrholz, Philipp A. Hartmann, Wolfgang Nebel
IPDPS3
2009 Scalability of network-on-chip communication architecture for 3-D meshes
abstract
Design constraints imposed by global interconnect delays as well as limitations in integration of disparate technologies make 3D chip stacks an enticing technology solution for massively integrated electronic systems. The scarcity of vertical interconnects however imposes special constraints on the design of the communication architecture. This article examines the performance and scalability of different communication topologies for 3D network-on-chips (NoC) using through-silicon-vias (TSV) for inter-die connectivity. Cycle accurate RTL-level simulations are conducted for two communication schemes based on a 7-port switch and a centrally arbitrated vertical bus using different traffic patterns. The scalability of the 3D NoC is examined under both communication architectures and compared to 2D NoC structures in terms of throughput and latency in order to quantify the variation of network performance with the number of nodes and derive key design guidelines.
Awet Yemane Weldezion, Matt Grange, Dinesh Pamunuwa, Zhonghai Lu, Axel Jantsch, Roshan Weerasekera, Hannu Tenhunen
NOCS5
2009 Analytical Evaluation of Retransmission Schemes in Wireless Sensor Networks
abstract
Retransmission has been adopted as one of the most popular schemes for improving transmission reliability in wireless sensor networks. Many previous works have been done on reliable transmission issues in experimental ways, however, there still lack of analytical techniques to evaluate these solutions. Based on the traffic model, service model and energy model, we propose an analytical method to analyze the delay and energy metrics of two categories of retransmission schemes: hop-by- hop retransmission (HBH) and end-to-end retransmission (ETE). With the experiment results, the maximum packet transfer delay and energy efficiency of these two scheme are compared in several scenarios. Moreover, the analytical results of transfer delay are validated through simulations. Our experiments demonstrate that HBH has less energy consumption at the cost of larger transfer delay compared with ETE. With the same target success probability, ETE is superior on the delay metric for low bit-error- rate (BER) cases, while HBH is superior for high BER cases.
Huimin She, Zhonghai Lu, Axel Jantsch, Dian Zhou, Lirong Zheng 0001
VTC Spring3
2008 Heterogeneous System-level Specification Using SystemC
abstract
Provides an abstract of the tutorial presentation and a brief professional biography of the presenter. The complete presentation was not made available for publication as part of the conference proceedings.
Eugenio Villar, Axel Jantsch, Christoph Grimm 0001, Tim Kogel
DATE2
2008 System-on-an-FPGA Design for Real-time Particle Track Recognition and Reconstruction in Physics Experiments
abstract
In particle physics experiments, the momenta of charged particles are studied by observing their deflection in a magnetic field. Dedicated detectors measure the particle tracks and complex algorithms are required for track recognition and reconstruction. This CPU-intensive task is usually implemented as off-line software running on PC clusters. In this paper, we present a system-on-chip design for the track recognition and reconstruction based on modern FPGA technologies. The basic principle of the algorithm is ported from software into the FPGA fabric. The fundamental architecture of the tracking processor is described in detail. Working as processing engines in compute nodes, the tracking processor contributes to recognize potential track candidates in real-time and promotes the selection efficiency of the data acquisition and trigger system. Our design study shows that the tracking module can be integrated in a single Xilinx Virtex-4 FX60 FPGA. The processing capability of the design is about 16.7K sub-events per second per module with our experimental setup, which achieves 20 times speedup compared to the software implementation.
Ming Liu 0011, Wolfgang Kuehn, Zhonghai Lu, Axel Jantsch
DSD4
2008 Energy efficient streaming applications with guaranteed throughput on MPSoCs
abstract
In this paper we present a design space exploration flow to achieve energy efficiency for streaming applications on MPSoCs while meeting the specified throughput constraints. The public domain simulators Sim-Panalyzer and Cacti are used to estimate the energy dissipations of the parameterized architectural components. As the main contributions, we schedule the streaming applications on a multi-clock synchronous modeling framework, guarantee the application timing properties by throughput analysis, and customize both processor voltage-frequency levels and memory sizes in the design space to optimize the application pipeline parallelism for energy efficiency. Two widely used heuristic algorithms (i.e., greedy and taboo search) are used during the design optimization process. Our experiments show an energy reduction of 21% without any loss in application throughput compared with an ad-hoc approach.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
EMSOFT3
2008 ATCA-based computation platform for data acquisition and triggering in particle physics experiments
abstract
An ATCA-based computation platform for data acquisition and trigger applications in nuclear and particle physics experiments has been developed. Each compute node (CN) which appears as a field replaceable unit (FRU) in an ATCA shelf, features 5 Xilinx Virtex-4 FX60 FPGAs and up to 10 GBytes DDR2 memory. Connectivity is provided with 8 optical links and 5 Gigabit Ethernet ports, which are mounted on each board to receive data from detectors and forward results to outer shelves or PC farms with attached mass storage. Fast point-to-point on-board interconnections between FPGAs as well as the full-mesh shelf backplane provide flexibility and high bandwidth to partition algorithms and correlate results among them. The system represents a highly reconfigurable and scalable solution for multiple applications.
Ming Liu 0011, Johannes Lang, Tiago Perez, Wolfgang Kuehn, Dapeng Jin, Zhen'An Liu, Zhonghai Lu, Axel Jantsch
FPL12
2008 Modeling Communication with Synchronized Environments
Tiberiu Seceleanu, Axel Jantsch
Fundam. Informaticae2
2008 Application and Verification of Local Nonsemantic-Preserving Transformations in System Design
abstract
Due to the increasing abstraction gap between the initial system model and a final implementation, the verification of the respective models against each other is a formidable task. This paper addresses the verification problem by proposing a stepwise application of combined refinement and verification activities in the context of synchronous model of computation. An implementation model is developed from the system model by applying predefined design transformations which are as follows: 1) semantic preserving or 2) nonsemantic preserving. Nonsemantic-preserving transformations introduce lower level implementation details, which are necessary to yield an efficient implementation. Our approach divides the verification tasks into two activities: 1) the local correctness of a refined block is checked by using formal verification tools and predefined properties, which are developed for each nonsemantic-preserving transformation, and 2) the global influence of the refinement to the entire system is studied through static analysis. We illustrate the design refinement and verification approach with three transformations: 1) a communication refinement mapping a synchronous channel to an asynchronous one including a handshake mechanism; 2) a computation refinement, which introduces resource sharing in a combinational computation block; and 3) a synchronization demanding refinement, where an algorithm analyzes the influence of a local refinement to the temporal properties of the entire system and restores the system's correct temporal behavior if necessary.
Tarvo Raudvere, Ingo Sander, Axel Jantsch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2008 A multiprocessor system-on-chip for real-time biomedical monitoring and analysis: ECG prototype architectural design space exploration
abstract
In this article we focus on multiprocessor system-on-chip (MPSoC) architectures for human heart electrocardiogram (ECG) real time analysis as a hardware/software (HW/SW) platform offering an advance relative to state-of-the-art solutions. This is a relevant biomedical application with good potential market, since heart diseases are responsible for the largest number of yearly deaths. Hence, it is a good target for an application-specific system-on-chip (SoC) and HW/SW codesign. We investigate a symmetric multiprocessor architecture based on STMicroelectronics VLIW DSPs that process in real time 12-lead ECG signals. This architecture improves upon state-of-the-art SoC designs for ECG analysis in its ability to analyze the full 12 leads in real time, even with high sampling frequencies, and its ability to detect heart malfunction for the whole ECG signal interval. We explore the design space by considering a number of hardware and software architectural options. Comparing our design with present-day solutions from an SoC and application point-of-view shows that our platform can be used in real time and without failures.
Iyad Al Khatib, Francesco Poletti, Davide Bertozzi, Luca Benini, Mohamed Bechara, Hasan Khalifeh, Axel Jantsch, Rustam Nabiev
ACM Trans. Design Autom. Electr. Syst.7
2008 TDM Virtual-Circuit Configuration for Network-on-Chip
abstract
In network-on-chip (NoC), time-division-multiplexing (TDM) virtual circuits (VCs) have been proposed to satisfy the quality-of-service requirements of applications. TDM VC is a connection-oriented communication service by which two or more connections take turns to share buffers and link bandwidth using dedicated time slots. In the paper, we first give a formulation of the multinode VC configuration problem for arbitrary NoC topologies. A multinode VC allows multiple source and destination nodes on it. Then we address the two problems of path selection and slot allocation for TDM VC configuration. For the path selection, we use a backtracking algorithm to explore the path diversity, constructively searching the solution space. In the slot allocation phase, overlapped VCs must be configured such that no conflict occurs and their bandwidth requirements are satisfied. We define the concept of a logical network (LN) as an infinite set of associated (time slot, buffer) pairs with respect to a buffer on a given VC. Based on this concept, we develop and prove theorems that constitute sufficient and necessary conditions to establish conflict-free VCs. They are applicable for networks where all nodes operate with the same clock frequency but allowing different phases. Using these theorems, slot allocation for VCs is a procedure of assigning VCs to different LNs. TDM VC configuration can thus be predictable and correct-by-construction. Our experiments on synthetic and real applications validate the effectiveness and efficiency of our approach.
Zhonghai Lu, Axel Jantsch
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Layered Switching for Networks on Chip
abstract
We present and evaluate a novel switching mechanism called layered switching. Conceptually, the layered switching implements wormhole on top of virtual cut-through switching. To show the feasibility of layered switching, as well as to confirm its advantages, we conducted an RTL implementation study based on a canonical wormhole architecture. Synthesis results show that our strategy suggests negligible degradation in hardware speed (1%) and area overhead (7%). Simulation results demonstrate that it achieves higher throughput than wormhole alone while significantly reducing the buffer space required at network nodes when compared with virtual cut-through.
Zhonghai Lu, Ming Liu 0011, Axel Jantsch
DAC3
2007 Increasing NoC Performance and Utilisation using a Dual Packet Exit Strategy
abstract
When designing a network the use of buffers is inevitable. Buffers are used at the entry point, inside and at the exits of the network. The usage of these buffers significantly changes the performance of the system as a whole. In order to enhance the buffer utilisation the concept of letting more than one packet exit the network at every switch each clock cycle is introduced -DualPacketExit(DPE). The approach is tried on a 4 x 4 and a 6 x 6 mesh. We demonstrate the buffers used in combination with different routing strategies for best effort performance. The result we present shows a 50 % reduction in terms of worst case latency and a 30 % reduction in terms of average latency as well as an increased throughput both from a system and network perspective. We define the termOperationalEfficiencyas a measure of the network efficiency and show that it increases by roughly 20 % with the DPE technique.
Mikael Millberg, Axel Jantsch
DSD2
2007 The ANDRES Project: Analysis and Design of Run-Time Reconfigurable, Heterogeneous Systems
abstract
Today's heterogeneous embedded systems combine components from different domains, such as software, analogue hardware and digital hardware. The design and implementation of these systems is still a complex and error-prone task due to the different Models of Computations (MoCs), design languages and tools associated with each of the domains. Though making such systems adaptive is technologically feasible, most of the current design methodologies do not explicitely support adaptive architectures. This paper present the ANDRES project. The main objective of ANDRES is the development of a seamless design flow for adaptive heterogeneous embedded systems (AHES) based on the modelling language SystemC. Using domain-specific modelling extensions and libraries, ANDRES will provide means to efficiently use and exploit adaptivity in embedded system design. The design flow is completed by a methodology tools for automatic hardware and software synthesis for adaptive architectures.
Andreas Herrholz, Frank Oppenheimer, Philipp A. Hartmann, Andreas Schallenberg, Wolfgang Nebel, Christoph Grimm 0001, Markus Damm, Jan Haase 0001, Florian Brame, Fernando Herrera, Eugenio Villar, Ingo Sander, Axel Jantsch, Anne-Marie Fouilliart, Marcos Martínez
FPL13
2007 Hardware/Software Co-design of a General-Purpose Computation Platform in Particle Physics
abstract
In this paper we present a hardware/software co-design based computation platform for online data processing in particle physics experiments. Our goal is to ease and accelerate the development and make it universal and scalable for multiple applications, on the premise of guaranteeing high communicating and processing capabilities. The entire computation network consists of quite a few interconnected compute nodes, each of which has multiple FPGAs to implement specific algorithms for data processing. High-speed communication features including RocketIO multi-gigabit transceiver and Gigabit Ethernet are supported by FPGAs to construct internal and external connections. An embedded Linux operating system is fitted on the PowerPC CPU core inside the Xilinx Virtex-4 FX FPGA. Thus programmers can access hardware resources via device drivers and write application programs to manage the system from the high level. Furthermore measurements have been executed using the development board to investigate both communicating and processing performances of the system. Results show that the computation platform is able to communicate at a UDP/IP data rate of around 400 Mbps per Ethernet link, and the event selection engine could process an event stream of 148.1 MBytes/s at an interesting event rate of 25%.
Ming Liu 0011, Wolfgang Kuehn, Zhonghai Lu, Axel Jantsch, Tiago Perez, Zhen'An Liu
FPT4
2007 A synchronization algorithm for local temporal refinements in perfectly synchronous models with nested feedback loops
abstract
Due to the abstract and simple computation and communication mechanism in the synchronous computational model it is easy to simulate synchronous systems and to apply formal verification methods. In synchronous models, a local temporal refinement that increases the delay in a single computation block may affect the functionality of the entire model. To preserve the system's functionality after temporal refinements we provide a synchronization algorithm that applies also to models with nested feedback loops. The algorithm adds pure delay elements to the model in order to balance the delay caused by refinement and to assure concurrent data arrival at computation blocks. It is done so that the refined model stays latency equivalent to the original model. The advantages of our approach are that (a) we remain fully within the synchronous model of computation, (b) we preserve the functionality of the existing computation blocks, and (c) we do not require additional computation resources, wrapper circuits or schedulers.
Tarvo Raudvere, Ingo Sander, Axel Jantsch
ACM Great Lakes Symposium on VLSI3
2007 Slot allocation using logical networks for TDM virtual-circuit configuration for network-on-chip
abstract
Configuring Time-Division-Multiplexing (TDM) Virtual Circuits (VCs) for network-on-chip must guarantee conflict freedom for overlapping VCs besides allocating sufficient time slots to them. These requirements are fulfilled in the slot allocation phase. In the paper, we define the concept of a logical network (LN). Based on this concept, we develop and prove theorems that constitute sufficient and necessary conditions to establish conflict-free VCs. Using these theorems, slot allocation for VCs becomes a procedure of computing LNs and then assigning VCs to different LNs. TDM VC configuration can thus be predictable and correct-by-construction. We have integrated this slot allocation method into our multi-node VC configuration program and applied the program to an industrial application.
Zhonghai Lu, Axel Jantsch
ICCAD2
2007 An Analytical Approach for Dimensioning Mixed Traffic Networks
abstract
We present an analytical method for analyzing and dimensioning a network based communication architecture. The method is based on the classic (sigma, p) network calculus. We use a TDMA approach for creating logically separated networks which makes statistical methods possible for calculations on best effort traffic, and supports implementation of guaranteed bandwidth services by using virtual circuits with looped containers
Per Badlund, Axel Jantsch
NOCS2
2007 Towards Open Network-on-Chip Benchmarks
abstract
Measuring and comparing performance, cost, and other features of advanced communication architectures for complex multi core/multiprocessor systems on chip is a significant challenge which has hardly been addressed so far. This document outlines the top-level view on a system of benchmarks for networks on chip (NoC), which intends to cover a wide spectrum of NoC design aspects, from application modeling to performance evaluation and post-manufacturing test and reliability. For performance benchmarking, requirements and features are described for application programs, synthetic micro-benchmarks, and abstract benchmark applications. Then, it proposes ways to measure and benchmark reliability, fault tolerance and testability of the on-chip communication fabric. This paper introduces the main concepts and ideas for benchmarking NoCs in a systematic and comparable way. It will be followed up by a report that will define a benchmark framework and the syntax of interfaces for benchmark programs that will allow the community to build-up a benchmark suite
Cristian Grecu, André Ivanov, Partha Pratim Pande, Axel Jantsch, Erno Salminen, Ümit Y. Ogras, Radu Marculescu
NOCS4
2007 A Study of NoC Exit Strategies
abstract
The throughput of a network is limited due to several interacting components. Analysing simulation results made it clear that the component that was worth attacking was the exit bandwidth between the network and the connected resources. The obvious approach is to increase this bandwidth; the benefit is a higher throughput of the network and a significant lowering of the buffer requirements at the entry points of the network; this because worst case scenarios now happens at a higher injection rate. The result we present shows significant differences in throughput as well as in average and worst case latency
Mikael Millberg, Axel Jantsch
NOCS2
2007 EWD: A metamodeling driven customizable multi-MoC system modeling framework
abstract
We present the EWD design environment and methodology, a modeling and simulation framework suited for complex and heterogeneous embedded systems with varying degrees of expressibility and modeling fidelity. This environment promotes the use of multiple models of computation (MoCs) to support heterogeneity and metamodeling for conformance tests of syntactic and static semantics during the process of modeling. Therefore, EWD is a multiple MoC modeling and simulation framework that ensures conformance of the MoC formalisms during model construction using a metamodeling approach. In addition, EWD provides a suite of translation tools that generate executable models for two simulation frameworks to demonstrate its language-independent modeling framework. The EWD methodology uses the Generic Modeling Environment for customization of the MoC-specific modeling syntax into a visual representation. To embed the execution semantics of the MoCs into the models, we have built parsing and translation tools that leverage an XML-based interoperability language. This interoperability language is then translated into executable Standard ML or Haskell models that can also be analyzed by existing simulation frameworks such as SML-Sys or ForSyDe. In summary, EWD is a metamodeling driven multitarget design environment with multi-MoC modeling capability.
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla, Axel Jantsch
ACM Trans. Design Autom. Electr. Syst.4
2006 A multiprocessor system-on-chip for real-time biomedical monitoring and analysis: architectural design space exploration
abstract
In this paper we focus on MPSoC architectures for human heart ECG real-time monitoring and analysis. This is a very relevant bio-medical application, with a huge potential market, hence it is an ideal target for an application-specific SoC implementation. We investigate a symmetric multi-processor architecture based on STMicroelectronics VLIW DSPs that process in real-time 12-lead ECG signals. This architecture improves upon state-of-the-art SoC designs for ECG analysis in its ability to analyze the full 12 leads in real-time, even with high sampling frequencies, and ability to detect heart malfunction. We explore the design space by considering a number of hardware and software architectural options.
Iyad Al Khatib, Francesco Poletti, Davide Bertozzi, Luca Benini, Mohamed Bechara, Hasan Khalifeh, Axel Jantsch, Rustam Nabiev
DAC7
2006 Adaptive Power Management for the On-Chip Communication Network
abstract
An on-chip communication network is most power efficient when it operates just below the saturation point. For any given traffic load the network can be operated in this region by adjusting frequency and voltage. For a deflective routing network we propose the design of a central controller for dynamic frequency and voltage scaling. Given history information including the load and frequency in the network, the controller adjusts the frequency and voltage such that the network operates just below the saturation point. We provide control mechanisms for continuous and discrete frequency ranges. With a discrete frequency range and taking into account voltage switching delays, we evaluate the control mechanism under stochastic, smoothly varying and very bursty traffic. Experiments demonstrate that adaptive control is very effective in minimizing power consumption at reasonable performance. Compared with a fixed high frequency network, the adaptively controlled network is significantly more power efficient. We compare it to fixed frequency networks, which are either too slow exhibiting unbounded delays, or are dimensioned for the worst case with very high frequency and are very power hungry.
Guang Liang, Axel Jantsch
DSD2
2006 Towards Performance-Oriented Pattern-Based Refinement of Synchronous Models onto NoC Communication
abstract
We present a performance-oriented refinement approach that refines a perfectly synchronous communicationmodel onto Network-on-Chip (NoC) communication. We first identify four basic forms of NoC process interaction patterns at the process level, namely, producer-consumer, peers, client-server, and multicast. We propose a threestep top-down refinement method: channel refinement, protocol refinement and channel mapping. For the producer-consumer pattern, we describe it in detail. In channel refinement, we deal with interfacing multiple clock domains and use a stochastic process to model channel delay and jitter. In protocol refinement, we show how to refine communication towards application requirements such as reliability and throughput. In channel mapping, we discuss channel convergence and channel merge arising from channel overlapping. All the refinements have been conducted and validated as an integral design phase towards implementation in ForSyDe, a formal system-level design methodology based on a synchronous model of computation.
Zhonghai Lu, Ingo Sander, Axel Jantsch
DSD3
2006 A High Level Power Model for the Nostrum NoC
abstract
We propose a power model for the Nostrum NoC. For this purpose an empirical power model of links and switches has been formulated and validated with the synopsys power compiler. The model, which from now on will be called Nos-HPM (Nostrum high-level power model) allows a fast power analysis and is accurate within 5%. System simulations with Nos-HPM run up to 500 times faster than with power compiler for a 4 times 4 network. We find a maximum power consumption of 0.7 W for a 4 times 4 mesh and 3.5 W for an 8 times 8 mesh, both implemented in 0.18mum UPC CMOS technology. In the worst case the average energy per cycle for a 128-bit packet is 508 pJ, while it is 20 pJ for a payload byte. The power consumption of all the links is equivalent or slightly higher than the power consumption of all the switches. A comparison between our results and some related work is also presented
Sandro Penolazzi, Axel Jantsch
DSD2
2006 Flexible Bus and NoC Performance Analysis with Configurable Synthetic Workloads
abstract
We present a flexible method for bus and network on chip performance analysis, which is based on the adaptation of workload models to resemble various applications. Our analysis method assists in the selection of a communication infrastructure early in the design process. The method uses (1) synthetic workload models which are similar to timed Petri nets and (2) the b-model for self-similar workloads. This allows the exploration of larger portions of the design space than possible with traditional stochastic models. The method is illustrated with tutorial examples where both a NoC and a bus based platform are analyzed
Rikard Thid, Ingo Sander, Axel Jantsch
DSD3
2006 Evaluation of on-chip networks using deflection routing
abstract
Deflection routing is being proposed for networks on chips since it is simple and adaptive. A deflection switch can be much smaller and faster than a wormhole or virtual cut-through switch. A deflection-routed network has three orthogonal characteristics: topology, routing algorithm and deflection policy. In this paper we evaluate deflection networks with different topologies such as mesh, torus and Manhattan Street Network, different routing algorithms such as random, dimension XY, delta XY and minimum deflection, as well as different deflection policies such as non-priority, weighted priority and straight-through policies. Our results suggest that the performance of a deflection network is more sensitive to its topology than the other two parameters. It is less sensitive to its routing algorithm, but a routing algorithm should be minimal. A priority-based deflection policy that uses global and history-related criterion can achieve both better average-case and worst-case performance than a non-priority or priority policy that uses local and stateless criterion. These findings are important since they can guide designers to make right decisions on the deflection network architecture, for instance, selecting a routing algorithm or deflection policy which has potentially low cost and high speed for hardware implementation.
Zhonghai Lu, Mingchen Zhong, Axel Jantsch
ACM Great Lakes Symposium on VLSI3
2006 An algorithm for electing cluster heads based on maximum residual energy
abstract
One of the main challenges in wireless sensor networks is to obtain long system lifetime. We propose an algorithm for electing the cluster head node based on the maximum residual energy for the purpose of even distribution of energy consumption in the overall network and obtaining the longest network lifetime. To maintain the original performance of the network, the lifetime is suggested to be expressed as to both the maximum last node dying time and the minimum time difference between the last node dying and the first node dying. The key parameter — the electing coefficient (θ) was obtained and evaluated. The optimal θ value is related to number of nodes, energy consumption of cluster members (ECCM), and energy consumption of the cluster head (ECCH). θ descends when number of nodes and ECCM decrease, and when ECCH increases. However, when energy consumptions of the cluster head and cluster members change proportionally, θ seems to be affected slightly. Results show that network lifetime can be prolonged when cluster heads are elected with the optimal θ value.
Axel Jantsch
IWCMC2
2005 Feasibility analysis of messages for on-chip networks using wormhole routing
abstract
The feasibility of a message in a network concerns if its timing property can be satisfied without jeopardizing any messages already in the network to meet their timing properties. We present a novel feasibility analysis for real-time (RT) and nonreal-time (NT) messages in wormhole-routed networks on chip. For RT messages, we formulate a contention tree that captures contentions in the network. For coexisting RT and NT messages, we propose a simple bandwidth partitioning method that allows us to analyze their feasibility independently.
Zhonghai Lu, Axel Jantsch, Ingo Sander
ASP-DAC2
2005 Refinement of Perfectly Synchronous Communication Model
Zhonghai Lu, Ingo Sander, Axel Jantsch
FDL3
2005 Modelling Environment for Heterogeneous Systems based on MoCs
Deepak Mathaikutty, Hiren D. Patel, Sandeep K. Shukla, Axel Jantsch
FDL4
2005 System level verification of digital signal processing applications based on the polynomial abstraction technique
abstract
Polynomial abstraction has been developed for data abstraction of sequential circuits, where the functionality can be expressed as polynomials. The method, based on the fundamental theorem of algebra, abstracts a possibly infinite domain of input values, into a much smaller and finite one, whose size is calculated according to the degree of the respective polynomial. The abstract model preserves the system's control and data properties, which can be verified by model checking. Experiments show that our approach does not only allow an automatic verification, but also gives considerably better results than existing methods.
Tarvo Raudvere, Ashish Kumar Singh, Ingo Sander, Axel Jantsch
ICCAD4
2004 System design for DSP applications in transaction level modeling paradigm
abstract
In this paper, we systematically define three transaction level models (TLMs), which reside at different levels of abstraction between the functional and the implementation model of a DSP system. We also show a unique language support to build the TLMs. Our results show that the abstract TLMs can be built and simulated much faster than the implementation model at the expense of a reasonable amount of simulation accuracy.
Abhijit K. Deb, Axel Jantsch, Johnny Öberg
DAC2
2004 System Design for DSP Applications Using the MASIC Methodology
abstract
Expensive top-down iterations are often required in the design cycle of complex DSP systems. In this paper, we introduce two levels of abstraction in the design flow by systematically categorizing the architectural decisions. As a result, the top-down iteration loop is broken. We also present a technique to capture and inject the architectural decisions such that the system models can be created and simulated efficiently. The concepts are illustrated by a realistic speech processing example, which is implemented using the AMBA on-chip architecture. Our methodology offers a smooth path from the functional modeling phase to the implementation level, facilitates the reuse of HW and SW components, and enjoys existing tool support at the backend.
Abhijit K. Deb, Axel Jantsch, Johnny Öberg
DATE2
2004 Guaranteed Bandwidth Using Looped Containers in Temporally Disjoint Networks within the Nostrum Network on Chip
abstract
In today's emerging network-on-chips, there is a need for different traffic classes with different quality-of-service guarantees. Within our NoC architecture nostrum, we have implemented a service of guaranteed bandwidth (GB), and latency, in addition to the already existing service of best-effort (BE) packet delivery. The guaranteed bandwidth is accessed via virtual circuits (VC). The VCs are implemented using a combination of two concepts that we call 'Looped Containers' and 'Temporally Disjoint Networks'. The looped containers are used to guarantee access to the network-independently of the current network load without dropping packets; and the TDNs are used in order to achieve several VCs, plus ordinary BE traffic, in the network. The TDNs are a consequence of the deflective routing policy used, and gives rise to an explicit time-division-multiplexing within the network. To prove our concept an HDL implementation has been synthesised and simulated. The cost in terms of additional hardware needed, as well as additional bandwidth is very low-less than 2 percent in both cases! Simulations showed that ordinary BE traffic is practically unaffected by the VCs.
Mikael Millberg, Erland Nilsson, Rikard Thid, Axel Jantsch
DATE4
2004 Polynomial Abstraction for Verification of Sequentially Implemented Combinational Circuits
abstract
Today's integrated circuits with increasing complexity cause the well known state space explosion problem in verification tools. In order to handle this problem a much simpler abstract model of the design has to be created for verification. We introduce the polynomial abstraction technique, which efficiently simplifies the verification task of sequential design blocks whose functionality can be expressed as a polynomial. Through our technique, the domains of possible values of data input signals can be reduced. This is done in such a way that the abstract model is still valid for model checking of the design functionality in terms of the system's control and data properties. We incorporate polynomial abstraction into the ForSyDe methodology, for the verification of clock domain design refinements.
Tarvo Raudvere, Ashish Kumar Singh, Ingo Sander, Axel Jantsch
DATE4
2004 A study on the implementation of 2-D mesh-based networks-on-chip in the nanometre regime
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch, Hannu Tenhunen
Integr.5
2004 Special issue on networks on chip
Axel Jantsch, Johnny Öberg, Hannu Tenhunen
J. Syst. Archit.1
2004 System modeling and transformational design refinement in ForSyDe [formal system design]
abstract
The scope of the formal system design (ForSyDe) methodology is high-level modeling and refinement of systems-on-a-chip and embedded systems. Starting with a formal specification model, that captures the functionality of the system at a high abstraction level, it provides formal design-transformation methods for a transparent refinement process of the system model into an implementation model that is optimized for synthesis. The main contribution of this paper is the ForSyDe modeling technique and the formal treatment of transformational design refinement. We introduce process constructors, that cleanly separate the computation part of a process from the synchronization and communication part. We develop the characteristic function for each process type and use it to define semantic preserving and design decision transformations. In a study of a digital equalizer example, we illustrate the modeling and refinement process and focus in particular on refinement of the clock domain, communication refinement, and resource sharing.
Ingo Sander, Axel Jantsch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Simulation and Analysis of Embedded DSP Systems Using MASIC Methodology
Abhijit K. Deb, Johnny Öberg, Axel Jantsch
DATE3
2003 Load Distribution with the Proximity Congestion Awareness in a Network on Chip
Erland Nilsson, Mikael Millberg, Johnny Öberg, Axel Jantsch
DATE4
2003 Development and Application of Design Transformations in ForSyDe
Ingo Sander, Axel Jantsch, Zhonghai Lu
DATE2
2003 NoCs: A new Contract between Hardware and Software
abstract
Future single chip systems will resemble more traditional computer networks than traditional central processors. The main reasons for this trend are the infeasibility of global synchrony on a single chip, the necessity of reuse of existing hardware and software components as much as possible, and the heterogeneity and irregularity of system functions and features. The consequences of this trend are far reaching and imply the shift in concern from computation and sequential algorithms to concurrency, communication and interaction in every aspect of design and development of hardware and software. Based on an analysis of current trends we suggest that there is an opportunity for defining an interface between applications and Network-on-Chip (NoC) platform implementations with significant benefits for both worlds. We analyze the desirable properties of such an interface by means of studying a particular NoC platform, the Nostrum. We draw the general conclusion that such an interface, which we also call contract, has to include (1) description of functionality, (2) description of communication semantics and performance, and (3) mapping of task to resources.
Axel Jantsch
DSD1
2003 Layout, Performance and Power Trade-Offs in Mesh-Based Network-on-Chip Architectures
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch
VLSI-SOC5
2002 Transformation based communication and clock domain refinement for system design
abstract
The ForSyDe methodology has been developed for system level design. In this paper we present formal transformation methods for the refinement of an abstract and formal system model into an implementation model. The methodology defines two classes of design transformations: (1) semantic-preserving transformations and (2) design decisions. In particular we present and illustrate communication and clock domain refinement by way of a digital equalizer system.
Ingo Sander, Axel Jantsch
DAC2
2001 Grammar-based design of embedded systems
Johnny Öberg, Mattias O'Nils, Axel Jantsch, Adam Postula, Ahmed Hemani
J. Syst. Archit.3
2001 Modeling of mixed control and dataflow systems in MASCOT
abstract
The Matlab and SDL Codesign Technique (MASCOT) method integrates modeling of dataflow and control dominated parts at the system level. Based on the established languages specification and description language (SDL) and Matlab, MASCOT provides a modeling and simulation technique which realizes the communication and synchronization between the two domains. Moreover, it offers modeling guidelines for a disciplined and efficient way of using the technique. Most of the tedious details of modeling synchronization and communication is handled automatically and is transparent to the user. Consequently, the user can focus on the application and on the important tradeoffs to be made at the system level.
Per Bjuréus, Axel Jantsch
IEEE Trans. Very Large Scale Integr. Syst.2
2000 MASCOT: A Specification and Cosimulation Method Integrating Data and Control Flow
abstract
We integrate data and control flow at the system specification level, using the two specialized and well established languages Matlab and SDL. For this we provide a modeling technique, which integrates the timing concepts and allows synchronization of vector-based computation with event based state transition. The technique is supported by a library of wrappers and communication functions, which has been implemented to make cosimulation easy to use and almost transparent to the user. A methodology formulates the rules to use the modeling technique, to partition the system, and to select communication modes. A complex industrial example illustrates the modeling technique and the methodology, and shows the efficiency of the Matlab-SDL cosimulation.
Per Bjuréus, Axel Jantsch
DATE2
2000 Composite Signal Flow: A Computational Model Combining Events, Sampled Streams, and Vectors
abstract
The composite signal flow model of computation targets systems with significant control and data processing parts. It builds on the data flow and synchronous data flow models and extends them to include three signal types: nonperiodic signals, sampled signals, and vectorized sampled signals. Vectorized sampled signals are used to represent vectors and computations on vectors. Several conversion processes are introduced to facilitate synchronization and communication with these signals. We discuss the severe implications, that these processes have on the causal behaviour of the system. We illustrate the model and its usefulness with three applications. A co-modelling and co-simulation environment combining Matlab and SDL; a high level timing analysis as a consequence of the operations on vectors; conditions for a parallel, distributed simulation.
Axel Jantsch, Per Bjuréus
DATE1
1999 The Rugby Model: A Conceptual Frame for the Study of Modelling, Analysis and Synthesis Concepts of Electronic Systems
abstract
We propose a conceptual framework, called the Rugby Model, in which designs, design processes and design tools can be studied. It is an extension of the Y chart and adds two dimensions for design representation, namely Data and Tune. The behavioural domain of Y chart is replaced by a more restricted domain called Computation. The structural and physical domains of Y chart are merged into a more general domain called Communication. A fifth dimension deals with design manipulations and transformations at three abstraction levels. The model shall establish a common understanding of modelling and design process concepts for communication and education in the community. In a case study we illustrate how a design can be characterized with the concepts the Rugby model.
Axel Jantsch, Shashi Kumar, Ahmed Hemani
DATE1
1999 Operating System Sensitive Device Driver Synthesis from Implementation Independent Protocol Specification
abstract
We present a method for generation of the software part of a HW/SW interface (i.e. the device drivers), which separates the behaviour of the interface from the architecture dependent parts. We do this by modelling the behaviour in ProGram (a grammar based protocol specification language) and capture the processor and OS kernel parts in separate libraries. By separating the behaviour from the architectural specific parts, compared to other approaches up to 50% development time can be saved the first time the component is used, and up to 98% for each time the interfaced component is reused.
Mattias O'Nils, Axel Jantsch
DATE2