Maurizio Palesi

dblp:45/1025 · DBLP profile ↗
← Back
101ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0003-3129-0664ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 81 · 10 first-author · 18 since 2021Software engineering, systems software and programming languages · 11 · 5 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 POSTER: Reinforcement Learning-based QoS-aware Online Scheduling for Multi-Tenant DNN Inference on Heterogeneous Accelerators
abstract
Deep Neural Networks (DNNs) are increasingly deployed through cloud services such as Inference-as-a-Service, where multiple tenants submit deadline-constrained inference requests to shared hardware infrastructures.
Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti
CF3
2026 Topology and Reliability Aware Qubit Mapping for Quantum Core Systems
abstract
Scaling quantum processors motivates multicore (modular) quantum architectures, where inter-core teleportation and state transfer are substantially slower and less reliable than intra-core operations. A key compilation/runtime challenge in such systems is inter-core qubit mapping: assigning logical qubits to cores across circuit timeslices to minimize non-local transfers under core-capacity and network constraints. Existing approaches such as FGP-rOEE, QUBO-based mapping, and Hungarian Qubit Assignment (HQA) reduce inter-core interactions, but are often evaluated under simplified interconnect assumptions and do not explicitly optimize for topology and reliability-dependent transfer costs. In this work, we present TR-HQA, a topology and reliability aware extension of HQA that minimizes expected inter-core transfer cost over a weighted interconnect graph, where edge weights capture hop distance and link success probability. We also propose GTR-QA, a greedy approximation that reduces placement time while maintaining competitive transfer cost. Using a multicore mapping simulator built on a pytket front-end, we evaluate TR-HQA and GTR-QA across circuit families (structured benchmarks and random circuits) and multiple interconnect topologies. Our results show that TR-HQA consistently reduces expected inter-core communication cost compared to topology-agnostic baselines, while GTR-QA provides a favorable quality–runtime trade-off. Across standard quantum benchmarks (QFT, Cuccaro Adder, and Draper Adder), TR-HQA reduces inter-core communication cost by 2.15 × on average (1.70 × –2.67 ×) and improves runtime by 11.8 × on average (1.02 × –30 ×).
Rajeswari Suance P. S, Akansh Khandelwal, Surankan De, Satyajit Das, John Jose, Maurizio Palesi
CF6
2026 Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI Workloads
abstract
Energy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts.
Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci
DATE10
2026 Timechain-level modeling and analysis of the bitcoin lightning network
Davide Patti, Salvatore Monteleone, Enrico Russo 0002, Maurizio Palesi
Comput. Networks4
2026 Assessing the Role of Communication in Modular Multi-Core Quantum Systems
abstract
The scalability of quantum computing is constrained by the physical and architectural limitations of monolithic quantum processors. Modular multi-core quantum architectures, which interconnect multiple quantum cores (QCs) via classical and quantum-coherent links, offer a promising alternative to address these challenges. However, transitioning to a modular architecture introduces communication overhead, where classical communication plays a crucial role in executing quantum algorithms by transmitting measurement outcomes and synchronizing operations across QCs. Understanding the impact of classical communication on execution time is therefore essential for optimizing system performance. In this work, we introduce qcomm , an open-source simulator designed to evaluate the role of classical communication in modular quantum computing architectures. qcomm provides a high-level execution and timing model that captures the interplay between quantum gate execution, entanglement distribution, teleportation protocols, and classical communication latency. We conduct an extensive experimental analysis to quantify the impact of classical communication bandwidth, interconnect types, and quantum circuit mapping strategies on overall execution time. Furthermore, we assess classical communication overhead when executing real quantum benchmarks mapped onto a cryogenically-controlled multi-core quantum system. Our results show that, while classical communication is generally not the dominant contributor to execution time, its impact becomes increasingly relevant in optimized scenarios—such as improved quantum technology, large-scale interconnects, or communication-aware circuit mappings. These findings provide useful insights for the design of scalable modular quantum architectures and highlight the importance of evaluating classical communication as a performance-limiting factor in future systems.
Maurizio Palesi, Enrico Russo 0002, Giuseppe Ascia, Hamaad Rafique, Davide Patti, Vincenzo Catania, Sergi Abadal, Abhijit Das 0002, Pau Escofet, Eduard Alarcón, Carmen G. Almudéver
ACM Trans. Design Autom. Electr. Syst.1
2025 A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
abstract
Graph Neural Networks (GNNs) have shown significant promise in various domains, such as recommendation systems, bioinformatics, and network analysis. However, the irregularity of graph data poses unique challenges for efficient computation, leading to the development of specialized GNN accelerator architectures that surpass traditional CPU and GPU performance. Despite this, the structural diversity of input graphs results in varying performance across different GNN accelerators, depending on their dataflows. This variability in performance due to differing dataflows and graph properties remains largely unexplored, limiting the adaptability of GNN accelerators. To address this, we propose a data-driven framework for dataflow-aware latency prediction in GNN inference. Our approach involves training regressors to predict the latency of executing specific graphs on particular dataflows, using simulations on synthetic graphs. Experimental results indicate that our regressors can predict the optimal dataflow for a given graph with up to 91.28% accuracy and a Mean Absolute Percentage Error (MAPE) of 3.78%. Additionally, we introduce an online scheduling algorithm that uses these regressors to enhance scheduling decisions. Our experiments demonstrate that this algorithm achieves up to 3.17× speedup in mean completion time and 6.26× speedup in mean execution time compared to the best feasible baseline across all datasets.
Pol Puigdemont, Enrico Russo 0002, Axel Wassington, Abhijit Das 0002, Sergi Abadal, Maurizio Palesi
ASP-DAC6
2025 Optimizing Qubit Assignment in Modular Quantum Systems via Attention-Based Deep Reinforcement Learning
abstract
Modular, distributed, and multi-core architectures are considered a promising solution for scaling quantum computing systems. Optimising communication is crucial to preserve quantum coherence. The compilation and mapping of quantum circuits should minimise state transfers while adhering to architec-tural constraints. To address this problem efficiently, we propose a novel approach using Reinforcement Learning (RL) to learn heuristics for a specific multi-core architecture. Our RL agent uses a Transformer encoder and Graph Neural Networks, encoding quantum circuits with self-attention and producing outputs via an attention-based pointer mechanism to match logical qubits with physical cores efficiently. Experimental results show our method outperform the baseline reducing by 28% inter-core communications for random circuits while minimising time-to-solution.
Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania
DATE2
2025 An Anomaly Detection Model for RISC-V in Automotive Applications: A Domain-Specific Accelerator Perspective
abstract
Early anomaly detection in automotive systems is crucial for enhancing user safety and enabling timely corrective actions, thereby minimizing the risks associated with system malfunctions. This paper presents an approach for implementing Artificial Intelligence (AI)-based algorithms for anomaly detection in the automotive domain, leveraging the RISC-V architecture in conjunction with Domain-Specific Accelerators (DSAs). By exploiting the efficiency of DSAs, the proposed system aims to achieve faster anomaly detection compared to traditional processing methods. A detailed comparison is conducted between the performance of executing the AI-based anomaly detection algorithm on the RISC-V core versus offloading it to an optimized hardware accelerator tailored to the specific AI model. The goal of this work is to provide valuable insights into the potential of RISCV and DSAs to enhance AI-driven safety mechanisms, contributing to the development of more reliable automotive systems.
Elio Vinciguerra, Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia
PDP3
2025 A Characteristics-Based Least Common Multiple Algorithm to Optimize Magnetic-Field-Based Indoor Localization
abstract
Clustering is an unsupervised learning technique that groups data based on similarity criteria. Traditional methods like K-Means and agglomerative clustering often require predefined parameters, struggle with irregular cluster shapes, and fail to classify subcluster points in magnetic fingerprint-based indoor localization. This study proposes the characteristics-based least common multiple (LCM) algorithm to address these challenges. This novel approach autonomously determines cluster number and shape while accurately classifying misclassified points based on characteristic similarities using LCM. We evaluated the proposed technique using state-of-the-art metrics and tested it in magnetic-field-based indoor localization scenarios. Comparisons were made with real-time and benchmark datasets, alongside traditional clustering methods. Results demonstrate that LCM significantly enhances localization accuracy, achieving a mean absolute error rate of 0.1 m.
Hamaad Rafique, Davide Patti, Maurizio Palesi, Gaetano Carmelo La Delfa
IEEE Internet Things J.3
2024 A Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
abstract
Currently, there is a growing trend of outsourcing the execution of DNNs to cloud services. For service providers, managing multitenancy and ensuring high-quality service delivery, particularly in meeting stringent execution time constraints, assumes paramount importance, all while endeavoring to maintain cost-effectiveness. In this context, the utilization of heterogeneous multi-accelerator systems becomes increasingly relevant. This paper presents RELMAS, a low-overhead deep reinforcement learning algorithm designed for the online scheduling of DNNs in multi-tenant environments, taking into account the dataflow heterogeneity of accelerators and memory bandwidths contentions. By doing so, service providers can employ the most efficient scheduling policy for user requests, optimizing Service-Level-Agreement (SLA) satisfaction rates and enhancing hardware utilization. The application of RELMAS to a heterogeneous multi-accelerator system composed of various instances of Simba and Eyeriss sub-accelerators resulted in up to a 173% improvement in SLA satisfaction rate compared to state-of-the-art scheduling techniques across different workload scenarios, with less than a 1.5% energy overhead.
Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania
DAC3
2024 Abstracting Bitcoin Lightning Network Complexity with Ultraviolet
abstract
With this work, we introduce the concept of Timechain-level model, along with an open-source implementation (Ultraviolet), to abstract the complexity of the Lightning Network (LN) while still providing a vision of base layer events and protocol internals. After depicting how each element of the model is mapped into the LN architectural stack, we show a case study to demonstrate its usage in investigating large-scale scenarios for research, development and educational purposes. Finally, we present a comparison to properly contextualize our contribution to the current state of the art of LN modelling, highlighting the advancements introduced by a Timechain-level and future directions of research it opens.
Davide Patti, Salvatore Monteleone, Enrico Russo 0002, Maurizio Palesi, Vincenzo Catania
ICBC4
2024 Characteristics-Based Least Common Multiple: A Novel Clustering Algorithm to Optimize Indoor Positioning
Hamaad Rafique, Davide Patti, Maurizio Palesi, Gaetano Carmelo La Delfa
ICINCO (1)3
2024 Fusing Visuals with Magnetic Signals to Improve Indoor Localization Using Vision Transformer
abstract
Sensor fusion-based indoor localization is an evolving application that uses fused information to determine the location of smartphone users. However, the heterogeneity of sensors across various smartphones has significantly compromised the accuracy of localization algorithms. Therefore, this paper introduces MH-ViL, an infrastructure-free and calibration-free frame-work built on top of the Vision Transformer neural network. MH-ViL seamlessly integrates magnetic field signals (MFS) and visual images for localization tasks. A novel magnetic feature projection (MFP) model is proposed to effectively map MFS onto visual image features, enhancing positional accuracy within the self-attention mechanism. Real-time experiments demonstrate that MH-ViL surpasses alternative models, with an impressive 92% accuracy. We provide a 95% confidence interval with an error below 0.5 meters. Code: https://github.com/Hamaad1/MH-ViL.git
Hamaad Rafique, Davide Patti, Maurizio Palesi, Gaetano Carmelo La Delfa
IPIN3
2024 Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
abstract
This paper addresses the critical challenge of managing Quality of Service (QoS) in cloud services, focusing on the nuances of individual tenant expectations and varying Service Level Indicators (SLIs). It introduces a novel approach utilizing Deep Reinforcement Learning for tenant-specific QoS management in multi-tenant, multi-accelerator cloud environments. The chosen SLI, deadline hit rate, allows clients to tailor QoS for each service request. A novel online scheduling algorithm for Deep Neural Networks in multi-accelerator systems is proposed, with a focus on guaranteeing tenant-wise, model-specific QoS levels while considering real-time constraints.
Enrico Russo 0002, Francesco Giulio Blanco, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Vincenzo Catania
ISCAS3
2024 Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-Based Accelerators
abstract
The need to efficiently execute different Deep Neural Networks (DNNs) on the same computing platform, coupled with the requirement for easy scalability, makes Multi-Chip Module (MCM)-based accelerators a preferred design choice. Such an accelerator brings together heterogeneous sub-accelerators in the form of chiplets, interconnected by a Network-on-Package (NoP). This paper addresses the challenge of selecting the most suitable sub-accelerators, configuring them, determining their optimal placement in the NoP, and mapping the layers of a predetermined set of DNNs spatially and temporally. The objective is to minimise execution time and energy consumption during parallel execution while also minimising the overall cost, specifically the silicon area, of the accelerator.This paper presents MOHaM, a framework for multi-objective hardware-mapping co-optimisation for multi-DNN workloads on chiplet-based accelerators. MOHaM exploits a multi-objective evolutionary algorithm that has been specialised for the given problem by incorporating several customised genetic operators. MOHaM is evaluated against state-of-the-art Design Space Exploration (DSE) frameworks on different multi-DNN workload scenarios. The solutions discovered by MOHaM are Pareto optimal compared to those by the state-of-the-art. Specifically, MOHaM-generated accelerator designs can reduce latency by up to 96% and energy by up to 96.12%.
Abhijit Das 0002, Enrico Russo 0002, Maurizio Palesi
IEEE Trans. Computers3
2024 Correction to: The position-based compression techniques for DNN model
Minghua Tang, Enrico Russo 0002, Maurizio Palesi
J. Supercomput.3
2023 Memory-Aware DNN Algorithm-Hardware Mapping via Integer Linear Programming
abstract
Mapping a deep neural network (DNN) layer onto domain-specific accelerators can require an intractable number of choices regarding loop factorization, ordering, and spatial unrolling. Determining the optimal mapping that achieves the best figures in terms of latency and energy efficiency can be difficult due to the vast number of possible candidates that need to be exhaustively evaluated. Many techniques have been recently proposed for fast and efficient mapping space exploration; some of them adopt a black-box optimization approach, others make assumptions on the underlying accelerator memory hierarchy or require time-consuming model retraining. We propose an integer linear programming (ILP) approach and formulate a mathematical model, namely LEMON, that takes into account number of accesses to each buffer, energy costs and buffer bandwidths in the accelerator and is flexible enough to work with different memory hierarchies. Compared with state-of-the-art techniques, LEMON achieves up to 83% energy-delay product reduction when compared to another ILP-based approach (CoSA) and 27% when compared to a genetic algorithm approach (GAMMA).
Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Salvatore Monteleone, Vincenzo Catania
CF2
2023 Explainable AI-Based Clinical Decision Support System for Obesity Comorbidity Analysis
abstract
This paper presents a novel Clinical Decision Support System based on eXplainable Artificial Intelligence (XAI-CDSS) as a comprehensive structured tool consisting of three main parts: predictive models, XAI interpretation, and a graph-based visualization of non-communicable pathologies. Machine learning models are proposed to predict the risk factors related to the direct association between obesity and comorbidities such as cardiovascular, heart disease, and diabetes. Multilayer perceptron and extreme gradient boosting are chosen among different machine learning algorithms as the best performing for the risk factors prediction of the selected comorbidities. They perform prediction with an accuracy of 0.72 for diabetes, and 0.73 for cardiovascular and heart disease. The intuitive XAI interface gives the end-user insight into the machine learning decision process, while the graph visualization links such co-occurrent pathologies to many other non-communicable diseases and provides a global view to healthcare professionals, usable for obesity and associated pathologies prevention and long-term treatment and care.
Grazia Veronica Aiosa, Maurizio Palesi, Francesca Sapuppo, Maria Gabriella Xibilia
e-Science2
2023 Scalable multi-chip quantum architectures enabled by cryogenic hybrid wireless/quantum-coherent network-in-package
abstract
The grand challenge of scaling up quantum computers requires a full-stack architectural standpoint. In this position paper, we will present the vision of a new generation of scalable quantum computing architectures featuring distributed quantum cores (Qcores) interconnected via quantum-coherent qubit state transfer links and orchestrated via an integrated wireless interconnect.
Eduard Alarcón, Sergi Abadal, Fabio Sebastiano, Masoud Babaie, Edoardo Charbon, Peter Haring Bolívar, Maurizio Palesi, Elena Blokhina, Dirk Leipold, Robert Bogdan Staszewski, Artur García-Sáez, Carmen G. Almudéver
ISCAS7
2023 Multiobjective End-to-End Design Space Exploration of Parameterized DNN Accelerators
abstract
Deep neural network (DNN) hardware accelerators enable the execution of complex DNN inferences on resource-constrained IoT devices. Inference performance and energy figures depend on how the DNN layers are mapped into the accelerator and how the architecture of the accelerator fits the variety of layers’ shapes of the actual DNN. The mapping determines the execution order of the operations, both temporally and spatially. Thus, selecting the best mapping that allows fitting the DNN model to the specific accelerator is of paramount importance to meet the strong constraints imposed by resource-scarce IoT platforms. Although several mapping space exploration techniques have been proposed in the literature, they are focused on determining the best mapping for a given layer, for a given architecture, and for optimizing a single objective. This article largely extends the scope of the exploration by considering the huge design space spanned by mapping related and architectural parameters, considering all the layers of the DNN, and optimizing multiple objectives simultaneously. We present EPOCA, end-to-end Pareto optimization of DNN accelerators, whose goal is to determine the accelerator’s architecture and the mapping for each layer that optimizes end-to-end and in a multiobjective fashion a set of conflicting design criteria. We assess EPOCA on different DNN models on a parameterized hardware accelerator designed for IoT applications and compare them with a state-of-the-art mapping space explorer, considering the area, inference latency, and inference energy as optimization metrics. We show that the set of Pareto solutions found by EPOCA provides the designer with a range of choices from which to select the best tradeoff with respect to the specific application.
Enrico Russo 0002, Maurizio Palesi, Davide Patti, Salvatore Monteleone, Giuseppe Ascia, Vincenzo Catania
IEEE Internet Things J.2
2023 The position-based compression techniques for DNN model
Minghua Tang, Enrico Russo 0002, Maurizio Palesi
J. Supercomput.3
2022 MEDEA: A Multi-objective Evolutionary Approach to DNN Hardware Mapping
abstract
Deep Neural Networks (DNNs) embedded domain-specific accelerators enable inference on resource-constrained devices. Making optimal design choices and efficiently scheduling neural network algorithms on these specialized architectures is challenging. Many choices can be made to schedule computation spatially and temporally on the accelerator. Each choice influences the access pattern to the buffers of the architectural hierarchy, affecting the energy and latency of the inference. Each mapping also requires specific buffer capacities and a number of spatial components instances that translate in different chip area occupation. The space of possible combinations, the mapping space, is so large that automatic tools are needed for its rapid ex-ploration and simulation. This work presents MEDEA, an open-source multi-objective evolutionary algorithm based approach to DNNs accelerator mapping space exploration. MEDEA leverages the Timeloop analytical cost model. Differently from the other schedulers that optimize towards a single objective, MEDEA allows deriving the Pareto set of mappings to optimize towards multiple, sometimes conflicting, objectives simultaneously. We found that solutions found by MEDEA dominates in most cases those found by state-of-the-art mappers.
Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Vincenzo Catania
DATE2
2022 DNN Model Compression for IoT Domain-Specific Hardware Accelerators
abstract
Machine learning techniques, particularly those based on neural networks, are always more often used at the edge of the network by Internet of Things (IoT) nodes. Unfortunately, the computation capabilities demanded by those applications, together with their energy efficiency-related constraints, exceed those exposed by embedded general-purpose processors. For this reason, the use of domain-specific hardware accelerators (DSAs) is considered the most viable solution to the unsustainable “Turing tariff” of general-purpose hardware. Starting from the observation that memory and communication traffic account for a large fraction of the overall latency and energy in deep neural network (DNN) inferences, this article proposes a new compression technique aimed at: 1) reducing the memory footprint for storing the model parameters of a DNN and 2) improving DNN inference latency and energy on resource-constrained IoT devices. The proposed compression technique, namely, LineCompress, is applied on a set of representative convolutional neural networks (CNNs) for object recognition mapped on a state-of-the-art DSA targeted for resource-constrained IoT devices. We show that on average,$7.4\times $memory footprint reduction can be obtained, thus reducing the memory and communication traffic that result to 77% and 87% inference latency and energy reduction, respectively, trading-off efficiency versus accuracy.
Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Andrea Mineo, Giuseppe Ascia, Vincenzo Catania
IEEE Internet Things J.2
2021 Power density aware application mapping in mesh-based network-on-chip architecture: An evolutionary multi-objective approach
Nizar Dahir, Ammar Karkar, Maurizio Palesi, Terrence S. T. Mak, Alexandre Yakovlev
Integr.3
2021 Opportunistic Caching in NoC: Exploring Ways to Reduce Miss Penalty
abstract
Due to limited on-chip caching, data-driven applications with large memory footprint encounter frequent cache misses. Such applications suffer from recurring miss penalty when they re-reference recently evicted cache blocks. To meet the worst-case performance requirements, Network-on-Chip (NoC) routers are provisioned with input port buffers. However, recent studies reveal that these buffers remain underutilised except during network congestion. Trace buffers are Design-for-Debug (DfD) hardware employed in NoC routers for post-silicon debug and validation. Nevertheless, they become non-functional once a design goes into production and remain in the routers left unused. In this article, we exploit the underutilised NoC router buffers and the unused trace buffers to store recently evicted cache blocks. While these blocks are stored in the buffers, future re-reference to these blocks can be replied from the NoC router. Such an opportunistic caching of evicted blocks in NoC routers significantly reduce the miss penalty. Experimental analysis shows that the proposed architectures can achieve up to 21 percent (16 percent on average) reduction in miss penalty and 19 percent (14 percent on average) improvement in overall system performance. While we have a negligible area and leakage power overhead of 2.58 and 3.94 percent, respectively, dynamic power reduces by 6.12 percent due to the improvement in performance.
Abhijit Das 0002, John Jose, Maurizio Palesi
IEEE Trans. Computers4
2021 On Performance Optimization and Quality Control for Approximate-Communication-Enabled Networks-on-Chip
abstract
For many applications showing error forgiveness, approximate computing is a new design paradigm that trades application output accuracy for mitigating computation/communication effort, which results in performance/energy benefit. Since networks-on-chip (NoCs) are one of the major contributors to system performance and power consumption, the underlying communication is approximated to achieve time/energy improvement. However, performing approximation blindly causes unacceptable quality loss. In this article, first, an optimization problem to maximize NoC performance is formulated with the constraint of application quality requirement, and the application quality loss is studied. Second, a congestion-aware quality control method is proposed to improve system performance by aggressively dropping network data, which is based on flow prediction and a lightweight heuristic. In the experiments, two recent approximation methods for NoCs are augmented with our proposed control method to compare with their original ones. Experimental results show that our proposed method can speed up execution by as much as 29.42% over the two state-of-the-art works.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Liang Wang 0020, Terrence S. T. Mak
IEEE Trans. Computers3
2021 COPE: Reducing Cache Pollution and Network Contention by Inter-tile Coordinated Prefetching in NoC-based MPSoCs
abstract
Prefetching helps in reducing the memory access latency in multi-banked NUCA architecture, where the Last Level Cache (LLC) is shared. In such systems, an application running on core generates significant traffic on the shared resources, the underlying network and LLC. While prefetching helps to increase application performance, but an inaccurate prefetcher can cause harm by generating unwanted traffic that additionally increases network and LLC contention. Increased network contention results in untimely prefetching of cache blocks, thereby reducing the effectiveness of a prefetcher. Prefetch accuracy is extensively used to reduce unwanted prefetches that can mitigate the prefetcher caused contention. However, the conventional prefetch accuracy parameter has major limitations in NUCA architectures. The article exposes that prefetch accuracy can create two major false-positive cases of prefetching, Under-estimation and Over-estimation problems, and false feedback loop that can mislead a prefetcher in generating more unwanted traffic. We propose a novel technique, Coordinated Prefetching for Efficient (COPE), which addresses these issues by redefining prefetch accuracy for such architectures and identifies additional parameters that can avoid generating unwanted prefetch requests. Experiment conducted using PARSEC benchmark on a 64-core system shows that COPE achieve 3% reduction in L1 cache miss rate, 12.64% improvement in IPC, 23.2% reduction in average packet latency and 18.56% reduction in dynamic power consumption of the underlying network.
Dipika Deb, John Jose, Maurizio Palesi
ACM Trans. Design Autom. Electr. Syst.3
2020 Implementing On-Chip Wireless Communication in Multi-stage Interconnection NoCs
Sirine Mnejja, Yassine Aydi, Mohamed Abid, Salvatore Monteleone, Maurizio Palesi, Davide Patti
AINA5
2020 DNNZip: Selective Layers Compression Technique in Deep Neural Network Accelerators
abstract
In Deep Neural Network (DNN) accelerators, the on-chip traffic and memory traffic accounts for a relevant fraction of the inference latency and energy consumption. A major component of such traffic is due to the moving of the DNN model parameters from the main memory to the memory interface and from the latter to the processing elements (PEs) of the accelerator. In this paper, we present DNNZip, a technique aimed at compressing the model parameters of a DNN, thus resulting in significant energy and performance improvement. DNNZip implements a lossy compression whose compression ratio is tuned based on the maximum tolerated error on the model parameters provided by the user. DNNZip is assessed on several convolutional NNs and the trade-off inference energy saving vs. inference latency reduction vs. network accuracy degradation is discussed. We found that up to 64% energy saving, and up to 67% latency reduction can be obtained with a limited impact on the accuracy of the network.
Habiba Lahdhiri, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Jordane Lorandel, Emmanuelle Bourdel, Vincenzo Catania
DSD2
2020 Efficient Compression Technique for NoC-based Deep Neural Network Accelerators
abstract
Deep Neural Networks (DNNs) are very powerful neural networks, widely used in many applications. On the other hand, such networks are computation and memory intensive, which makes their implementation difficult onto hardware-constrained systems, that could use network-on-chip as interconnect infrastructure. A way to reduce the traffic generated among memory and the processing elements is to compress the information before their exchange inside the network. In particular, our work focuses on reducing the huge number of DNN parameters, i.e., weights. In this paper, we propose a flexible and low-complexity compression technique which preserves the DNN performance, allowing to reduce the memory footprint and the volume of data to be exchanged while necessitating few hardware resources. The technique is evaluated on several DNN models, achieving a compression rate close to 80% without significant loss in accuracy on AlexNet, ResNet, or LeNet-5.
Jordane Lorandel, Habiba Lahdhiri, Emmanuelle Bourdel, Salvatore Monteleone, Maurizio Palesi
DSD5
2020 Improving Inference Latency and Energy of DNNs through Wireless Enabled Multi-Chip-Module-based Architectures and Model Parameters Compression
abstract
Performance and energy figures of Deep Neural Network (DNN) accelerators are profoundly affected by the communication and memory sub-system. In this paper, we make the case of a state-of-the-art multi-chip-module-based architecture for DNN inference acceleration. We propose a hybrid wired/wireless network-in-package interconnection fabric and a compression technique for drastically improving the communication efficiency and reducing the memory and communication traffic with a consequent improvement of performance and energy metrics. We assess the inference performance and energy improvement vs. accuracy degradation for different CNNs showing that up to 77% and 68% of inference latency reduction and inference energy reduction, respectively, can be obtained while keeping the accuracy degradation below 5% as respect to the original uncompressed CNN.
Giuseppe Ascia, Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti
NOCS5
2020 Exploiting Data Resilience in Wireless Network-on-chip Architectures
abstract
The emerging wireless Network-on-Chip (WiNoC) architectures are a viable solution for addressing the scalability limitations of manycore architectures in which multi-hop long-range communications strongly impact both the performance and energy figures of the system. The energy consumption of wired links as well as that of radio communications account for a relevant fraction of the overall energy budget. In this article, we extend the approximate computing paradigm to the case of the on-chip communication system in manycore architectures. We present techniques, circuitries, and programming interfaces aimed at reducing the energy consumption of a WiNoC by exploiting the trade-off energy saving vs. application output degradation. The proposed platform—namely, xWiNoC—uses variable voltage swing links and tunable transmitting power wireless interfaces along with a programming interface that allows the programmer to specify those data structures that are error-resilient. Thus, communications induced by the access to such error-resilient data structures are carried out by using links and radio channels that are configured to work in a low energy mode, albeit by exposing a higher bit error rate. xWiNoC is assessed on a set of applications belonging to different domains in which the trade-off energy vs. performance vs. application result quality is discussed. We found that up to 50% of communication energy saving can be obtained with a negligible impact on the application output quality and 3% in application performance degradation.
Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose, Valerio Mario Salerno
ACM J. Emerg. Technol. Comput. Syst.4
2020 Special issue on energy-efficient many-core embedded systems and architectures (SI: NoCArc18)
Maurizio Palesi, Kun-Chih Chen, Midia Reshadi
J. Syst. Archit.1
2019 ACDC: An Accuracy- and Congestion-aware Dynamic Traffic Control Method for Networks-on-Chip
abstract
Many applications exhibit error forgiving features. For these applications, approximate computing provides the opportunity of accelerating the execution time or reducing power consumption, by mitigating computation effort to get an approximate result. Among the components on a chip, network-on-chip (NoC) contributes a large portion to system power and performance. In this paper, we exploit the opportunity of aggressively reducing network congestion and latency by selectively dropping data. Essentially, the importance of the dropped data is measured based on a quality model. An optimization problem is formulated to minimize the network congestion with constraint of the result quality. A lightweight online algorithm is proposed to solve this problem. Experiments show that on average, our proposed method can reduce the execution time by as much as 12.87% and energy consumption by 12.42% under strict quality requirement, speed up execution by 19.59% and reduce energy consumption by 21.20% under relaxed requirement, compared to a recent work on approximate computing approach for NoCs.
Siyuan Xiao, Xiaohang Wang 0001, Maurizio Palesi, Amit Kumar Singh 0002, Terrence S. T. Mak
DATE3
2019 Analyzing networks-on-chip based deep neural networks
abstract
One of the most promising architectures for performing deep neural network inferences on resource-constrained embedded devices is based on massive parallel and specialized cores interconnected by means of a Network-on-Chip (NoC). In this paper, we extensively evaluate NoC-based deep neural network accelerators by exploring the design space spanned by several architectural parameters. We show how latency is mainly dominated by the on-chip communication whereas energy consumption is mainly accounted by memory (both on-chip and off-chip).
Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose
NOCS4
2018 Critical Packet Prioritisation by Slack-Aware Re-Routing in On-Chip Networks
abstract
Packet based Network-on-Chip (NoC) connect tens to hundreds of components in a multi-core system. The routing and arbitration policies employed in traditional NoCs treat all application packets equally. However, some packets are critical as they stall application execution whereas others are not. We differentiate packets based on a metric called slack that captures a packet's criticality. We observe that majority of NoC packets generated by standard application based benchmarks do not have slack and hence are critical. Prioritising these critical packets during routing and arbitration will reduce application stall and improve performance. We study the diversity and interference of packets to propose a policy that prioritises critical packets in NoC. This paper presents a slack-aware re-routing (SAR) technique that prioritises lower slack packets over higher slack packets and explores alternate minimal path when two no-slack packets compete for same output port. Experimental evaluation on a 64-core Tiled Chip Multi-Processor (TCMP) with 8×8 2D mesh NoC using both multiprogrammed and multithreaded workloads show that our proposed policy reduces application stall time by upto 22% over traditional round-robin policy and 18% over state-of-the-art slack-aware policy.
Abhijit Das 0002, Sarath Babu 0003, John Jose, Sangeetha Jose, Maurizio Palesi
NOCS5
2018 Traffic Aware Deflection Rerouting Mechanism for Mesh Network on Chip
abstract
In two dimensional mesh Network on Chips (NoC), efficient routing algorithms route majority of the flits through the central routers of the network, whereas routers at the edges and corners experience relatively lesser flit flow. This in turn leads to higher traffic towards central routers than to edge and corner routers. Such uneven traffic distribution causes thermal hot-spots at the center of the chip where the load is high, and reduces the average life-time of the chip. In existing buffer-less deflection routing techniques, load balanced traffic distribution is not considered as a factor during assignment of links to mis-routed flits. Devising deflection routing techniques with greater load balancing capability is a major challenge for efficient thermal management of the chip. This paper proposes an adaptive routing mechanism that can provide a more balanced traffic profile in a deflection router based mesh NoC. Significant number of deflected flits are rerouted towards the edges/corners of the mesh, thereby reducing the load on the central routers. From evaluations, it is seen that the proposed technique reduces traffic variance compared to NoCs using baseline deflection routers. Transient temperature variation studies using Hotspot tool substantiate our findings.
Simi Zerine Sleeba, John Jose, Maurizio Palesi, Rekha K. James, Maniyelil Govindankutty Mini
VLSI-SoC3
2018 An optimized hybrid algorithm in term of energy and performance for mapping real time workloads on 2d based on-chip networks
Sarzamin Khan, Sheraz Anjum, Usman Ali Gulzari, Farruh Ishmanov, Maurizio Palesi, Muhammad Khalil Afzal
Appl. Intell.5
2018 Improving Energy Efficiency in Wireless Network-on-Chip Architectures
abstract
Wireless Network-on-Chip (WiNoC) represents a promising emerging communication technology for addressing the scalability limitations of future manycore architectures. In a WiNoC, high-latency and power-hungry long-range multi-hop communications can be realized by performance- and energy-efficient single-hop wireless communications. However, the energy contribution of such wireless communication accounts for a significant fraction of the overall communication energy budget. This article presents a novel energy managing technique for WiNoC architectures aimed at improving the energy efficiency of the main elements of the wireless infrastructure, namely, radio-hubs. The rationale behind the proposed technique is based on selectively turning off, for the appropriate number of cycles, all the radio-hubs that are not involved in the current wireless communication. The proposed energy managing technique is assessed on several network configurations under different traffic scenarios both synthetic and extracted from the execution of real applications. The obtained results show that the application of the proposed technique allows up to 25% total communication energy saving without any impact on performance and with a negligible impact on the silicon area of the radio-hub.
Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti
ACM J. Emerg. Technol. Comput. Syst.4
2018 The Suboptimal Routing Algorithm for 2D Mesh Network
abstract
Due to the huge routing algorithm search space for 2D mesh based Network-on-Chip (NoC), Divide-Conquer method is presented to effectively explore the search space. When using Divide-Conquer method, a large number of routing algorithms will be created. In order to get the final results in an acceptable time, a precise metric is needed to measure routing performance and discard the poor performance routings. In this paper, we propose a new routing performance metric, namely, network pressure. Network pressure has the following three advantages: (1) it could measure the whole network congestion state; (2) network pressure of a network and that of its partial component is highly related, under the same routing; (3) it is closely related with routing performance. Based on network pressure and Divide-Conquer method, high performance routing could be achieved. The obtained routing is called suboptimal routing due to the following two reasons: (1) there is only a little gap between its performance and that of the fully adaptive routings under both transpose1 and transpose2 traffics. (2) the search space of routing algorithms is systematically and widely exploited.
Minghua Tang, Maurizio Palesi
IEEE Trans. Computers3
2017 The Repetitive Turn Model for Adaptive Routing
abstract
For 2D mesh based Network-on-Chip (NoC), the prohibited turns of routing algorithms should be repetitively distributed in order for the routing algorithms to be implemented by logic-based circuit. In this paper, we aim to exploit the designing space for logicbased routing algorithms, and propose new logic-based routing algorithms that outperform the state-of-the-art counterparts. Toward this direction, we firstly construct all routing algorithms for 5 x 5 2D mesh topology. Then we select those routing algorithms which have repetitive prohibited turns across both the network rows and columns. In addition, we chose those routing algorithms that have smaller routing pressures than Odd-Even routing algorithm. Then the routing algorithms for 2D mesh topology ranging from 6 x 6 to 15 x 15 are respectively constructed according to the prohibited turns distribution of the selected routing algorithms. Two routing algorithms that have smaller routing pressures than Odd-Even routing algorithm are obtained for all considered networks. The obtained logic-based routing algorithms are called as Repetitive Turn Model (RTM). Simulation results show that RTM could achieve up to 51% performance improvement as compared to Odd-Even routing algorithm.
Minghua Tang, Xiaola Lin, Maurizio Palesi
IEEE Trans. Computers3
2016 Improving the energy efficiency of wireless Network on Chip architectures through online selective buffers and receivers shutdown
abstract
The wireless Network-on-Chip (WiNoC) design paradigm represents an emergent and viable solution for addressing the scalability limitations of future manycores architectures. Unfortunately, components such as the buffers and the transceiver of the radio-hubs in a WiNoC, account for a significant fraction of the total communication energy budget. In this paper, we present WIRXSleep, a mechanism aimed at improving the energy efficiency of radio-hubs in WiNoC architectures. WIRXSleep selectively and dynamically disables receiver modules and buffers of those radio-hubs that will be not involved in any communication during the next forthcoming clock cycles. Its application on different WiNoC topologies, with different configurations, and under different traffic scenarios has resulted interesting energy savings (up to 25%) without any impact on performance and with a negligible impact on cost metrics.
Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti
CCNC4
2016 Energy efficient transceiver in wireless Network on Chip architectures
Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti
DATE4
2016 Efficient Congestion-Aware Scheme for Wireless on-Chip Networks
abstract
Wireless NoC is becoming popular to be a promising future on-chip interconnection network as a result of high bandwidth, low latency and flexible topology configurations provided by this emerging technology. Nonetheless, congestion occurrence in wireless routers negatively affects the usability of high speed wireless links and considerably increases the network latency, therefore, in this paper, a congestion-aware platform (CAP-W) is introduced for wireless NoCs in order to reduce both internal and external congestions. The whole platform of CAP-W consists of an adaptive routing algorithm that balances utilization of wired and wireless networks, a dynamic task mapping approach that tries to minimize congestion probability, and a task migration strategy that considers dynamic variation of application behaviors. Simulation results show significant gain in congestion control over PEs of wireless NoC, compared to state-of-the-art works.
Amin Rezaei 0001, Masoud Daneshtalab, Maurizio Palesi, Danella Zhao
PDP3
2016 Special issue on energy efficient methods and systems in the emerging cloud era
Maurizio Palesi, Mario Collotta, Masoud Daneshtalab, Pradip Bose
J. Comput. Syst. Sci.1
2016 On-Chip Communication Energy Reduction Through Reliability Aware Adaptive Voltage Swing Scaling
abstract
In a multi/many-core system, the network-on-chip (NoC)-based communication backbone is responsible for a relevant fraction of the overall energy budget. Reducing the voltage swing for signaling in crossbars and links results in significant energy saving. Unfortunately, as voltage swing reduces, the bit error rate increases, that in turn compromises the communication reliability. Starting from the assumption that not all the communications need same level of reliability, in this paper we propose techniques and architectures for run-time tuning of the voltage swing of the crossbars and interrouter links. The proposed technique is compared with the state of the art in link energy reduction through data encoding under both synthetic and real traffic scenarios. We found that the proposed techniques allow to significantly reduce the energy consumption of the NoC fabric without degrading the performance metrics. Energy savings ranging from 20% to 43% have been observed without any relevant impact on the performance metrics.
Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Partha Pratim Pande, Vincenzo Catania
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Local Congestion Avoidance in Network-on-Chip
abstract
Network-on-Chip (NoC) has been made the communication infrastructure for many-core architecture. NoC are subject to congestion, which is claimed to be avoided by many researchers. However, there is no completely understanding of congestion in literature, which hinders its solution. Toward this direction, we firstly carry out study on congestion in this paper. We find that congestion usually occurs at a portion of nodes in a local network region. Moreover, local congestion will significantly decrease system performance and mostly impact some particular communication pairs. Then we attempt to solve local congestion by addressing different local region size, based on Divide-Conquer approach and routing pressure. It avoids congestion in every local region by keeping routing pressure of every local region minimum. Using different local region size will create different routings. Our study shows that the local region size is closely related with the routing performance. When local region size is 5 × 5 the optimal routing performance of large size network could be achieved.
Minghua Tang, Xiaola Lin, Maurizio Palesi
IEEE Trans. Parallel Distributed Syst.3
2016 Runtime Tunable Transmitting Power Technique in mm-Wave WiNoC Architectures
abstract
Emerging on-chip communication technologies, like wireless networks-on-chip (WiNoCs), have been recently proposed as candidate solutions for addressing the scalability limitations of conventional multihop network on chip (NoC) architectures. In a WiNoC, a subset of network nodes, namely, radio hubs, are equipped with a wireless interface that allows them to wirelessly communicate with other radio hubs. Thus, long-range communications, which would involve multiple hops in a conventional wireline NoC, can be realized by a single hop through the radio medium. Unfortunately, the energy consumed by the RF transceiver into the radio hub (i.e., the main building block in a WiNoC), and in particular by its transmitter, accounts for a significant fraction of the overall communication energy. In order to alleviate such contribution, this paper presents a runtime tunable transmitting power technique for improving the energy efficiency of the transceiver in WiNoC architectures. The basic idea is tuning the transmitting power based on the physical location of the recipient of the current communication. Specifically, based on the destination address of the incoming packet, the radio hub tunes its transmitting power to a minimum level, but high enough to reach the destination antenna without exceeding a certain bit error ratio. The proposed technique is general and can be applied to any WiNoC architecture. Its application on different representative WiNoC architectures results in an average energy reduction up to 50% without any impact on performance and with a negligible overhead in terms of silicon area.
Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Noxim: An open, extensible and cycle-accurate network on chip simulator
abstract
Emerging on-chip communication technologies like wireless Networks-on-Chip (WiNoCs) have been proposed as candidate solutions for addressing the scalability limitations of conventional multi-hop NoC architectures. In a WiNoC, a subset of network nodes are equipped with a wireless interface which allows them long-range communication in a single hop. This paper presents Noxim, an open, configurable, extendible, cycle-accurate NoC simulator developed in SystemC which allows to analyze the performance and power figures of both conventional wired NoC and emerging WiNoC architectures.
Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti
ASAP4
2015 A closed loop transmitting power self-calibration scheme for energy efficient WiNoC architectures
Andrea Mineo, Mohd Shahrizal Rusli, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania, Muhammad N. Marsono
DATE3
2015 Low Energy yet Reliable Data Communication Scheme for Network-on-Chip
abstract
In this paper, a low energy yet reliable communication scheme for network-on-chip is suggested. To reduce the communication energy consumption, we invoke low-swing signals for transmitting data, as well as data encoding techniques, for minimizing both self and coupling switching capacitance activity factors. To maintain the communication reliability of communication at low-voltage swing, an error control coding (ECC) technique is exploited. The decision about end-to-end or hop-to-hop ECC schemes and the proper number of detectable errors are determined through high-level mathematical analysis on the energy and reliability characteristics of the techniques. Based on the analysis, the extended single error correction double-error detecting end-to-end coding technique with three bits of error detection is used in the network layer. For minimization of the self and coupling switching capacitance activity factors, the odd, even, full invert scheme is employed in the data link layer. This coding has an inherent error detection probability for the flits, which is exploited in the suggested technique. The efficiency of the scheme is studied by using both synthetic and real traffic scenarios. The study reveals savings of up to 43% and 58%, for power dissipation and energy consumption, respectively, without any significant performance degradation and overhead in the network interface.
Nima Jafarzadeh, Maurizio Palesi, Saeedeh Eskandari, Shaahin Hessabi, Ali Afzali-Kusha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 An Offline Method for Designing Adaptive Routing Based on Pressure Model
abstract
As a scalable substitute of on-chip bus, network-on-chip (NoC) is proposed as the communication infrastructure in modern multi/many-core system-on-chip (SoC). Efficient communication in NoC is critical to the overall SoC performance. Although local congestion has an important impact on communication delay, it is barely taken into account when designing routing algorithms. In this paper, we propose an offline methodology of designing routing algorithm based on channel pressure model to address the local congestion issue. Specifically, the proposed methodology uses divide-conquer with the aim of generating high performance routing algorithms, which are able to balance the load over the network with a consequent reduction of local congestion. By using the proposed methodology, the obtained routing could achieve up to 37% performance improvement (in terms of average communication delay) as compared to the well-known odd-even routing algorithm for 15 × 15 network.
Minghua Tang, Xiaola Lin, Maurizio Palesi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2015 Routing Pressure: A Channel-Related and Traffic-Aware Metric of Routing Algorithm
abstract
How to precisely measure performance of routing algorithm is an important issue when studying routing algorithm of network-on-chip (NoC). The degree of adaptiveness is the most widely used metric in the literature. However, our study shows that the degree of adaptiveness cannot precisely measure performance of routing algorithm. It cannot account for why routing algorithm with high degree of adaptiveness may have poor performance. Simulation has to be carried out to evaluate performance of routing algorithm. In this paper, we propose a new metric of routing pressure for measuring performance of routing algorithm. It has higher precision of measuring routing algorithm performance than the degree of adaptiveness. Performance of routing algorithm can be evaluated through routing pressure without simulation. It can explain why congestion takes place in network. In addition, where and when congestion takes place can be pointed out without simulation.
Minghua Tang, Xiaola Lin, Maurizio Palesi
IEEE Trans. Parallel Distributed Syst.3
2014 SHiFA: System-Level Hierarchy in Run-Time Fault-Aware Management of Many-Core Systems
abstract
A system-level approach to fault-aware resource management of many-core systems is proposed. The proposed approach, called SHiFA, is able to tolerate run-time faults at system level without any hardware overhead. In contrast to the existing system-level methods, network resources are also considered to be potentially faulty. Accordingly, applications are mapped onto healthy nodes of the system at run-time such that their interaction will not require the use of faulty elements. By utilizing the simple routing approach, results show 100% utilizability of PEs and 99.41% of successful mapping when up to 8 links are broken. SHiFA design is based on distributed operating systems, such that it is kept scalable for future many-core systems. A significant improvement in scalability properties is observed compared to the state-of-the-art distributed approaches.
Mohammad Fattah, Maurizio Palesi, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DAC2
2014 An adaptive transmitting power technique for energy efficient mm-wave wireless NoCs
abstract
Several emerging techniques have been recently proposed for alleviating the communication latency and the energy consumption issues in multi/many-core architectures. One of such emerging communication techniques, namely, WiNoC replaces the traditional wired links with the use of wireless medium. Unfortunately, the energy consumed by the RF transceiver (i.e., the main building block of a WiNoC), and in particular by its transmitter, accounts for a significant fraction of the overall communication energy. In this paper we propose a runtime tunable transmitting power technique for improving the energy efficiency of the transceiver in wireless NoC architectures. The basic idea is tuning the transmitting power based on the location of the recipient of the current communication. The integration of the proposed technique into two known WiNoC architectures, namely, iWise64 and McWiNoC resulted in an energy reduction of 43% and 60%, respectively.
Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania
DATE2
2014 Adaptive power allocation for many-core systems inspired from multiagent auction model
abstract
Scaling of future many-core chips is hindered by the challenge imposed by ever-escalating power consumption. At its worst, an increasing fraction of the chips will have to be shut down, as power supply is inadequate to simultaneously switch all the transistors. This so-called dark silicon problem brings up a critical issue regarding how to achieve the maximum performance within a given limited power budget. This issue is further complicated by two facts. First, high variation in power budget calls for wide range power control capability, whereas most current frequency/voltage scaling techniques cannot effectively adjust power over such a wide range. Second, as the applications' behavior becomes more complicated, there is a pressing need for scalability and global coordination, rendering heuristic-based centralized or fully distributed control schemes inefficient. To address the aforementioned problems, in this paper, a power allocation method employing multiagent auction models is proposed, referred as Hierarchal MultiAgent based Power allocation (HiMAP). Tiles act the role of consumers to bid for power budget and the whole process is modeled by a combinatorial auction, whereas HiMAP finds the Walrasian equilibria. Experimental results have confirmed that HiMAP can reduce the execution time by as much as 45% compared to three competing methods. The runtime overhead and cost of HiMAP are also small, which makes it suitable for adaptive power allocation in many-core systems.
Xiaohang Wang 0001, Baoxin Zhao, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi
DATE7
2014 Editorial: Special issue on design challenges for many-core processors
abstract
No abstract available.
Masoud Daneshtalab, Maurizio Palesi, Juha Plosila
ACM Trans. Embed. Comput. Syst.2
2014 Editorial: Special Section on ESTIMedia'13
abstract
No abstract available.
Maurizio Palesi, Todor P. Stefanov
ACM Trans. Embed. Comput. Syst.1
2014 On self-tuning networks-on-chip for dynamic network-flow dominance adaptation
abstract
Modern network-on-chip (NoC) systems are required to handle complex runtime traffic patterns and unprecedented applications. Data traffics of these applications are difficult to fully comprehend at design time so as to optimize the network design. However, it has been discovered that the majority of dataflows in a network are dominated by less than 10% of the specific pathways. In this article, we introduce a method that is capable of identifying critical pathways in a network at runtime and can then dynamically reconfigure the network to optimize for network performance subject to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerging dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance dataflows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulations. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications.
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Peng Liu 0016, Masoud Daneshtalab, Maurizio Palesi, Terrence S. T. Mak
ACM Trans. Embed. Comput. Syst.6
2014 Data Encoding Techniques for Reducing Energy Consumption in Network-on-Chip
abstract
As technology shrinks, the power dissipated by the links of a network-on-chip (NoC) starts to compete with the power dissipated by the other elements of the communication subsystem, namely, the routers and the network interfaces (NIs). In this paper, we present a set of data encoding schemes aimed at reducing the power dissipated by the links of an NoC. The proposed schemes are general and transparent with respect to the underlying NoC fabric (i.e., their application does not require any modification of the routers and link architecture). Experiments carried out on both synthetic and real traffic scenarios show the effectiveness of the proposed schemes, which allow to save up to 51% of power dissipation and 14% of energy consumption without any significant performance degradation and with less than 15% area overhead in the NI.
Nima Jafarzadeh, Maurizio Palesi, Ahmad Khademzadeh, Ali Afzali-Kusha
IEEE Trans. Very Large Scale Integr. Syst.2
2013 An Adaptive Output Selection Function Based on a Fuzzy Rule Base System for Network on Chip
abstract
In Network-on-chip design two of the most important performance indices are the average delay of packets and power dissipation. The first depends on the level of congestion of the communication system, the second is strongly influenced by the power dissipated by the links of a network-on-chip (NoC) which accounts for a significant fraction of the overall power dissipated by the on-chip communication fabric. Such fraction becomes more and more relevant as technology shrinks. The selection policy of the output port in the NOC router with adaptive routing function affects both the performance and power dissipation. In fact, generally a NOC with a lower average delay has a lower power dissipation on the links for transmission. The selection policies proposed in the literature have the object or the minimization of the average delay or in some cases the reduction of power dissipation. In this paper we propose a selection policy aimed at minimizing both the average delay and power consumption of the communication system of NoC based architectures. The proposed policy uses a cost function obtained using a Fuzzy Rule Base System (FRBS) with two inputs an one output. Inputs are the level of use of the buffers of the downstream routers, and the link power dissipation of the selected output. The output of FRBS is the cost function. The experimental results, obtained using a cycle accurate NoC simulator for different traffic scenarios, shows that the proposed selection policy outperforms other selection approaches both in terms of average delay and in terms on power dissipation and energy consumption.
Giuseppe Ascia, Maurizio Palesi, Vincenzo Catania
DSD2
2013 Runtime Online Links Voltage Scaling for Low Energy Networks on Chip
abstract
The power dissipated by the links of a network on chip (NoC) accounts for a significant fraction of the overall power budget. Reducing the operating voltage of the network links, results in a square reduction of their power contribution. Unfortunately, the voltage reduction has a negative impact on communication reliability in terms of bit error rate. Starting from the assumption that not all the communications require the same reliability level, we present a technique for dynamically changing the voltage of the links based on the communication reliability requirements. The experiments, carried out under both synthetic and real traffic scenarios, show the effectiveness of the proposed technique which allows to save up of 55% of link energy with a total energy saving of 25% of the entire NoC.
Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania
DSD2
2013 On self-tuning networks-on-chip for dynamic network-flow dominance adaptation
abstract
Modern networks-on-chip (NoC) systems are required to handle complex run-time traffic patterns and unprecedented applications. Data traffics of these applications are difficult to be fully comprehended at design-time so as to optimize the network design. However, it has been discovered that the majority data flows in a network are dominated by less than 10% of the specific pathways. In this paper, we introduce a method that is capable of identifying critical pathways in a network at run-time and, then, can dynamically reconfigure the network to optimize for the network performance subjected to the identified dominated flows. An online learning and analysis scheme is employed to quickly discover the emerged dominated traffic flows and provides a statistical traffic prediction using regression analysis. The architecture of a self-tuning network is also discussed which can be reconfigured by setting up the identified point-to-point paths for the dominance data flows in large traffic volumes. The merits of this new approach are experimentally demonstrated using comprehensive NoC simulators. Compared to the conventional network architectures over a range of realistic applications, the proposed self-tuning network approach can effectively reduce the latency and power consumption by as much as 25% and 24%, respectively. We also evaluated the configuration time and additional hardware cost. This new approach demonstrates the capability of an adaptive NoC to handle more complex and dynamic applications.
Xiaohang Wang 0001, Terrence S. T. Mak, Mei Yang 0001, Yingtao Jiang, Masoud Daneshtalab, Maurizio Palesi
NOCS6
2013 Energy Efficient Run-Time Incremental Mapping for 3-D Networks-on-Chip
Xiaohang Wang 0001, Peng Liu 0016, Mei Yang 0001, Maurizio Palesi, Yingtao Jiang, Michael C. Huang 0001
J. Comput. Sci. Technol.4
2013 Efficient multicast schemes for 3-D Networks-on-Chip
Xiaohang Wang 0001, Mei Yang 0001, Yingtao Jiang, Maurizio Palesi, Peng Liu 0016, Terrence S. T. Mak, Nader Bagherzadeh
J. Syst. Archit.4
2013 Guest Editors' Introduction to the Special Issue on "Novel On-Chip Parallel Architectures and Software Support"
Fangyang Shen, Mei Yang 0001, Maurizio Palesi
Parallel Comput.3
2013 Introduction to the special section on on-chip and off-chip network architectures
abstract
No abstract available.
José Flich, Maurizio Palesi
ACM Trans. Embed. Comput. Syst.2
2013 Introduction to the special section on ESTIMedia'12
abstract
No abstract available.
Jian-Jia Chen, Maurizio Palesi
ACM Trans. Embed. Comput. Syst.2
2012 HARAQ: Congestion-Aware Learning Model for Highly Adaptive Routing Algorithm in On-Chip Networks
abstract
The occurrence of congestion in on-chip networks can severely degrade the performance due to increased message latency. In mesh topology, minimal methods can propagate messages over two directions at each switch. When shortest paths are congested, sending more messages through them can deteriorate the congestion condition considerably. In this paper, we present an adaptive routing algorithm for on-chip networks that provide a wide range of alternative paths between each pair of source and destination switches. Initially, the algorithm determines all permitted turns in the network including 180-degree turns on a single channel without creating cycles. The implementation of the algorithm provides the best usage of all allowable turns to route messages more adaptively in the network. On top of that, for selecting a less congested path, an optimized and scalable learning method is utilized. The learning method is based on local and global congestion information and can estimate the latency from each output channel to the destination region.
Masoumeh Ebrahimi, Masoud Daneshtalab, Fahimeh Farahnakian, Juha Plosila, Pasi Liljeberg, Maurizio Palesi, Hannu Tenhunen
NOCS6
2012 Embedded Transitive Closure Network for Runtime Deadlock Detection in Networks-on-Chip
abstract
Interconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at runtime is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes runtime transitive closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoCs). This detection scheme guarantees the discovery of all true deadlocks without false alarms in contrast with state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture, which couples with the NoC infrastructure, is also presented to realize the detection mechanism efficiently. Detailed hardware realization architectures and schematics are also discussed. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the proposed method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and, thus, reducing energy wastage in retransmission for various traffic scenarios including real-world application. We found that timing-based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant.
Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi
IEEE Trans. Parallel Distributed Syst.5
2011 Run-time deadlock detection in networks-on-chip using coupled transitive closure networks
abstract
Interconnection networks with adaptive routing are susceptible to deadlock, which could lead to performance degradation or system failure. Detecting deadlocks at run-time is challenging because of their highly distributed characteristics. In this paper, we present a deadlock detection method that utilizes run-time Transitive Closure (TC) computation to discover the existence of deadlock-equivalence sets, which imply loops of requests in networks-on-chip (NoC). This detection scheme guarantees the discovery of all true deadlocks without false alarms unlike state-of-the-art approximation and heuristic approaches. A distributed TC-network architecture which couples with the NoC architecture is also presented to realize the detection mechanism efficiently. Our results based on a cycle-accurate simulator demonstrate the effectiveness of the TC-network method. It drastically outperforms timing-based deadlock detection mechanisms by eliminating false detections and thus reducing energy dissipation in various traffic scenarios. For example, timing based methods may produce two orders of magnitude more deadlock alarms than the TC-network method. Moreover, the implementations presented in this paper demonstrate that the hardware overhead of TC-networks is insignificant.
Ra'ed Al-Dujaily, Terrence S. T. Mak, Fei Xia 0001, Alexandre Yakovlev, Maurizio Palesi
DATE5
2011 Power-Aware Run-Time Incremental Mapping for 3-D Networks-on-Chip
Xiaohang Wang 0001, Maurizio Palesi, Mei Yang 0001, Yingtao Jiang, Michael C. Huang 0001, Peng Liu 0016
NPC2
2011 Low latency and energy efficient multicasting schemes for 3D NoC-based SoCs
abstract
In this paper, two topology oriented multicast routing algorithms, MXYZ and AL+XYZ, are proposed to support multicasting in 3D Networks on Chips (NoCs). In specific, MXYZ is a dimension order multicast routing algorithm that targets 3D NoC systems built upon regular topologies, while AL+XYZ is applicable to NoCs with irregular topologies. If the output channel found by MXYZ is not available (i.e. in the same region), an alternative output channel is used to forward/replicate the packets in AL+XYZ. MXYZ is evaluated against a path based regular topology oriented multicast routing and AL+XYZ against an irregular region oriented multiple unicast routing algorithm. Our experimental results have demonstrated that the proposed MXYZ and AL+XYZ schemes have lower latency and energy consumption than the conventional path based multicast routing and the multiple unicast routing algorithms, meriting them to be more suitable for supporting multicasting in 3D NoC systems.
Xiaohang Wang 0001, Maurizio Palesi, Mei Yang 0001, Yingtao Jiang, Michael C. Huang 0001, Peng Liu 0016
VLSI-SoC2
2011 Data Encoding Schemes in Networks on Chip
abstract
An ever more significant fraction of the overall power dissipation of a network-on-chip (NoC) based system-on-chip (SoC) is due to the interconnection system. In fact, as technology shrinks, the power contribute of NoC links starts to compete with that of NoC routers. In this paper, we propose the use of data encoding techniques as a viable way to reduce both power dissipation and energy consumption of NoC links. The proposed encoding scheme exploits the wormhole switching techniques and works on an end-to-end basis. That is, flits are encoded by the network interface (NI) before they are injected in the network and are decoded by the destination NI. This makes the scheme transparent to the underlying network since the encoder and decoder logic is integrated in the NI and no modification of the routers architecture is required. We assess the proposed encoding scheme on a set of representative data streams (both synthetic and extracted from real applications) showing that it is possible to reduce the power contribution of both the self-switching activity and the coupling switching activity in inter-routers links. As results, we obtain a reduction in total power dissipation and energy consumption up to 37% and 18%, respectively, without any significant degradation in terms of both performance and silicon area.
Maurizio Palesi, Giuseppe Ascia, Fabrizio Fazzino, Vincenzo Catania
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2010 An Efficient Technique for In-order Packet Delivery with Adaptive Routing Algorithms in Networks on Chip
abstract
Although adaptive routing algorithms promise higher communication performance, as compared to deterministic routing algorithms, they suffer from the out-of-order packet delivery problem. In the context of Network on Chip, the area and computational overhead of ordering packets at the destination is high and may reverse any gain achieved through the use of adaptivity of the routing algorithm. In this paper, we describe a novel scheme for ensuring in-order packet delivery while retaining the performance advantages of adaptive routing. The hardware architecture of a router that supports the proposed scheme is described. Although the basic idea in our proposal is topology independent we evaluate and compare the performance of our scheme with both deterministic as well as adaptive routing algorithms for 2D mesh NoC. As compared to the XY routing algorithm, our technique significantly reduces the packet delay and improves the saturation point. The impact on router area and power dissipation is also discussed. Although the power consumption of routers increase, the energy consumption per flit increases less than 2% on average, since the higher performance allows for draining more traffic during a certain time window.
Maurizio Palesi, Rickard Holsmark, Xiaohang Wang 0001, Shashi Kumar, Mei Yang 0001, Yingtao Jiang, Vincenzo Catania
DSD1
2010 Leveraging Partially Faulty Links Usage for Enhancing Yield and Performance in Networks-on-Chip
abstract
The communication infrastructure of a complex multicore system-on-a-chip is getting an increasing fraction of the overall chip area. According to the International Technology Roadmap for Semiconductors, killer defect density does not decrease over successive technology generations. For this reason, the probability that a manufacturing defect affects the communication system is predicted to increase. In this paper, we deal with manufacturing defects which affect the links in a network-on-chip-based interconnection system. The goal of this paper is to show that by using effective routing functions, supported by appropriate selection policies and with a limited amount of extra logic in the router, it is easy to exploit partially faulty links to improve the performance of the system. We show that, instead of discarding partially faulty links, they can be used at reduced capacity to improve the distribution of the traffic over the network, yielding performance and power improvements. We couple an application-specific routing function with a set of selection policies which are aware of link fault distribution and evaluate them on both synthetic traffic and a real complex multimedia application. We also present an implementation of the router, augmented with the extra logic, to support both the proposed selection functions and the transmission of messages over partially faulty links. We analyze the router in terms of silicon area, timing, and power dissipation.
Maurizio Palesi, Shashi Kumar, Vincenzo Catania
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2009 An Effective Methodology to Multi-objective Design of Application Domain-specific Embedded Architectures
abstract
Today's computer systems have become unbelievably complex. Nowadays register-level design is an overwhelming task, especially in the embedded system area where the time-to-market is very short. Platform based design shifts the challenge on how to tune parametric platforms to achieve the best performance at the smallest cost. This task, called multi-objective design space exploration, requires accurate strategies because the design space is too vast to be exhaustively evaluated. Even using efficient exploration strategies proposed in the literature, simulation times can become a bottleneck in the design flow. In this work we propose a novel approach to application-domain design space exploration using a multi-objective genetic algorithm and employing HPC to reduce exploration times. The genetic algorithm is preceded by a correlation analysis of the different objectives. The search space is thus reduced by combining highly correlated objectives from different domains. We describe the steps needed to parallelize the exploration on the grid, and present the results of extensive testing of the proposed approach. We obtained over one order of magnitude reduction in exploration times without hampering the quality of the solutions. Shorter simulation times allow more ideas to be explored in less time. This leads to shorter product time-to-market and a more thorough design space exploration. Furthermore the combination of correlated objectives favors the design of modern multi-purpose devices.
Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti, Gianmarco De Francisci Morales
DSD3
2009 Data Encoding for Low-Power in Wormhole-Switched Networks-on-Chip
abstract
As the number of cores in a chip increases, the role played by the communication system becomes more and more central. An on-chip communication infrastructure based on the Network-on-Chip (NoC) paradigm is today recognized as the most effective and scalable solution able to deal with the communication issues that will characterize the next generation of many-cores architectures. An ever more significant fraction of the overall chip area is devoted to support advanced and reliable communication protocols making the energy resources used for communication starting to compete with the ones spent for computation. Amongst the communication resources, as technology shrinks, the power ratio between NoC links and routers increases making the links becoming more power-hungry than routers. In this paper we propose a novel endto-end data encoding scheme which exploits the wormhole technique commonly used in NoC-based system to reduce power dissipated by the NoC links. We assess the proposed encoding scheme on a set of representative data streams showing that it is possible to reduce the power contribution of both the self switching activity and the coupling switching activity in inter-routers links. As results, we obtain a reduction in total power dissipation and energy consumption up to 26% and 9% respectively without any significant degradation in terms of both performance and silicon area. The encoder and decoder logic is integrated in the network interface and is transparent to the underling NoC.
Maurizio Palesi, Fabrizio Fazzino, Giuseppe Ascia, Vincenzo Catania
DSD1
2009 A multi-objective strategy for concurrent mapping and routing in networks on chip
abstract
The design flow of network-on-chip (NoCs) include several key issues. Among other parameters, the decision of where cores have to be topologically mapped and also the routing algorithm represent two highly correlated design problems that must be carefully solved for any given application in order to optimize several different performance metrics. The strong correlation between the different parameters often makes that the optimization of a given performance metric has a negative effect on a different performance metric. In this paper we propose a new strategy that simultaneously refines the mapping and the routing function to determine the Pareto optimal configurations which optimize average delay and routing robustness. The proposed strategy has been applied on both synthetic and real traffic scenarios. The obtained results show how the solutions found by the proposed approach outperforms those provided by other approaches proposed in literature, in terms of both performance and fault tolerance.
Rafael Tornero, Valentino Sterrantino, Maurizio Palesi, Juan M. Orduña
IPDPS3
2009 HiRA: A methodology for deadlock free routing in hierarchical networks on chip
abstract
Complexity of designing large and complex NoCs can be reduced/managed by using the concept of hierarchical networks. In this paper, we propose a methodology for design of deadlock free routing algorithms for hierarchical networks, by combining routing algorithms of component subnets. Specifically, our methodology ensures reachability and deadlock freedom for the complete network if routing algorithms for subnets are deadlock free. We evaluate and compare the performance of hierarchical routing algorithms designed using our methodology with routing algorithms for corresponding flat networks. We show that hierarchical routing, combining best routing algorithm for each subnet, has a potential for providing better performance than using any single routing algorithm. This is observed for both synthetic as well as traffic from real applications. We also demonstrate, by measuring jitter in throughput, that hierarchical routing algorithms leads to smoother flow of network traffic. A router architecture that supports scalable table-based routing is briefly outlined.
Rickard Holsmark, Shashi Kumar, Maurizio Palesi, Andres Mejia
NOCS3
2009 Application Specific Routing Algorithms for Networks on Chip
abstract
In this paper we present a methodology to develop efficient and deadlock free routing algorithms for Network-on-Chip (NoC) platforms which are specialized for an application or a set of concurrent applications. The proposed methodology, called application specific routing algorithm (APSRA), exploits the application specific information regarding pairs of cores which communicate and other pairs which never communicate in the NoC platform to maximize communication adaptivity and performance. The methodology also exploits the known information regarding concurrency/non-concurrency of communication transactions among cores for the same purpose. We demonstrate, through analysis of adaptivity as well as simulation based evaluation of latency and throughput, that algorithms produced by the proposed methodology give significantly higher performance as compared to other deadlock free algorithms for both homogeneous as well as heterogeneous 2D mesh topology NoC systems. For example, for homogeneous mesh NoC, APSRA results in approximately 30% less average delay as compared to odd-even algorithm just below saturation load. Similarly the saturation load point for APSRA is significantly higher as compared to other adaptive routing algorithms for both homogeneous and non-homogeneous mesh networks.
Maurizio Palesi, Rickard Holsmark, Shashi Kumar, Vincenzo Catania
IEEE Trans. Parallel Distributed Syst.1
2009 Region-Based Routing: A Mechanism to Support Efficient Routing Algorithms in NoCs
abstract
An efficient routing algorithm is important for large on-chip networks [network-on-chip (NoC)] to provide the required communication performance to applications. Implementing NoC using table-based switches provide many advantages, including possibility of changing routing algorithms and fault tolerance, due to the option of table reconfigurations. However, table-based switches have been considered unsuitable for NoCs due to their perceived high area and power consumption. In this paper, we describe the region-based routing (RBR) mechanism which groups destinations into network regions allowing an efficient implementation with logic blocks. RBR can also be viewed as a mechanism to reduce the number of entries in routing tables. RBR is general and can be used in conjunction with any adaptive routing algorithm. In particular, we have evaluated the proposed scheme in conjunction with a general routing algorithm, namely segment-based routing (SR) and an application specific routing algorithm (APSRA) using regular and irregular mesh topologies. Our study shows that the number of entries in the table is significantly reduced, especially for large networks. Evaluation results show that RBR requires only four regions to support several routing algorithms in a 2-D mesh with no performance degradation. Considering link failures, our results indicate that RBR combined with SR is able to tolerate up to 7 link failures in an 8times8 mesh. RBR also reduces area and power dissipation of an equivalent table-based implementation by factors of 8 and 10, respectively. Moreover, the degradation in performance of the network is insignificant when using APSRA combined with RBR.
Andres Mejia, Maurizio Palesi, José Flich, Shashi Kumar, Pedro López 0001, Rickard Holsmark, José Duato
IEEE Trans. Very Large Scale Integr. Syst.2
2008 High Performance Computing for Embedded System Design: A Case Study
abstract
In this paper we assess the use of high performance computing in design space exploration of a complex highly parameterized very long instruction word based system-on-a-chip platform. Experiments show that the conventional belief of linear decrease in exploration time as the number of available processors increases is discredited starting from a relatively low number of processors mainly due to communication overhead and I/O bottleneck.
Vincenzo Catania, Gianmarco De Francisci Morales, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti
DSD4
2008 Efficient Application Specific Routing Algorithms for NoC Systems utilizing Partially Faulty Links
abstract
In this paper we propose a series of efficient routing strategies to effectively utilize NoC systems with partially faulty links. These strategies try to use partially faulty links when the load is high and distribute traffic uniformly on links. Evaluation of our strategies for 8times8 mesh with 7% partially faulty links shows that, using our best strategy, it is possible to achieve an average reduction of up to 50% on packet delay when the load is high. We have also worked out complete designs of routers which can tolerate partial link faults and implement our routing strategies. Approximately 25% extra area and 5% extra power consumption is required for the design of the upgraded router incorporating link fault tolerance and the best routing strategy. However, the overall performance improvement counter balances such overhead resulting in an overall saving in energy consumption of up to 20%. The proposed strategies offer a way to increase the effective yield of large and complex NoC systems.
Dario Frazzetta, Giuseppe Dimartino, Maurizio Palesi, Shashi Kumar, Vincenzo Catania
DSD3
2008 A Communication-Aware Topological Mapping Technique for NoCs
Rafael Tornero, Juan M. Orduña, Maurizio Palesi, José Duato
Euro-Par3
2008 Design of Bandwidth Aware and Congestion Avoiding Efficient Routing Algorithms for Networks-on-Chip Platforms
Maurizio Palesi, Giuseppe Longo, Salvatore Signorino, Rickard Holsmark, Shashi Kumar, Vincenzo Catania
NOCS1
2008 Deadlock free routing algorithms for irregular mesh topology NoC systems with rectangular regions
Rickard Holsmark, Maurizio Palesi, Shashi Kumar
J. Syst. Archit.2
2008 Reducing complexity of multiobjective design space exploration in VLIW-based embedded systems
abstract
Architectures based on very-long instruction word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of instruction level parallelism (ILP) with a reasonable trade-off in complexity and silicon cost. Specialization of such architectures involves the configuration of both hardware-related aspects (e.g., register files, functional units, memory subsystem) and software-related issues (e.g., the compilation strategy). The complex interactions between the components of such systems will force a human designer to rely on judgment and experience in designing them, possibly eliminating interesting configurations, and making tuning of the system, for either power, energy, or performance, difficult. In this paper we propose tools and methodologies to efficiently cope with this complexity from a multiobjective perspective. We first analyze the impact of ILP-oriented code transformations using two alternative compilation profiles to quantitatively show the effect of such transformations on typical design objectives like performance, power dissipation, and energy consumption. Next, by means of statistical analysis, we collect useful data to predict the effectiveness of a given compilation profiles for a specific application. Information gathered from such analysis can be exploited to drastically reduce the computational effort needed to perform the design space exploration.
Vincenzo Catania, Maurizio Palesi, Davide Patti
ACM Trans. Archit. Code Optim.2
2008 Implementation and Analysis of a New Selection Strategy for Adaptive Routing in Networks-on-Chip
abstract
Efficient and deadlock-free routing is critical to the performance of networks-on-chip. The effectiveness of any adaptive routing algorithm strongly depends on the underlying selection strategy. A selection function is used to select the output channel where the packet will be forwarded on. In this paper we present a novel selection strategy that can be coupled with any adaptive routing algorithm. The proposed selection strategy is based on the concept of Neighbors-on-Path the aims of which is to exploit the situations of indecision occurring when the routing function returns several admissible output channels. The overall objective is to choose the channel that will allow the packet to be routed to its destination along a path that is as free as possible of congested nodes. Performance evaluation is carried out by using a flit-accurate simulator under traffic scenarios generated by both synthetic and real applications. Results obtained show how the proposed selection strategy applied to the Odd-Even routing algorithm yields an improvement in both average delay and saturation point up to 20% and 30% on average respectively, with a minimal overhead in terms of area occupation. In addition, a positive effect on total energy consumption is also observed under near-congestion packet injection rates.
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti
IEEE Trans. Computers3
2007 Multi-Objective Evolutionary Fuzzy Clustering for High-Dimensional Problems
abstract
This paper deals with the application of unsupervised fuzzy clustering to high dimensional data. Two problems are addressed: groups (clusters) number discovery and feature selection without performance losses. In particular we analyze the potential of a genetic fuzzy system, that is the integration of a multi-objective evolutionary algorithm with a fuzzy clustering algorithm. The main characteristic of the integrated approach is the ability to handle the two problems at the same time, suggesting a Pareto set of trade-off solutions which could have a better chance of matching the real needs. We exhibit the high quality clustering and features selection results by applying our approach to a real-world data set.
Alessandro G. Di Nuovo, Maurizio Palesi, Vincenzo Catania
FUZZ-IEEE2
2007 Exploiting Communication Concurrency for Efficient Deadlock Free Routing in Reconfigurable NoC Platforms
abstract
In this paper we make a case for the use of NoC paradigm to develop future FPGAs in which large computational blocks (cores) are connected to each other through a packet switched communication network. We propose a methodology to develop efficient and deadlock free routing algorithms for such NoC platforms which can be specialized for an application or a set of concurrent applications. Application specific topology of communicating cores as well as information about their communication concurrency over time is exploited to maximize communication adaptivity and performance. We demonstrate, both through analysis of adaptivity as well as simulation based evaluation of latency and throughput, that our algorithm gives significantly higher performance as compared to general purpose deadlock free algorithms like XY and odd-even.
Maurizio Palesi, Shashi Kumar, Rickard Holsmark, Vincenzo Catania
IPDPS1
2007 Efficient design space exploration for application specific systems-on-a-chip
Giuseppe Ascia, Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti
J. Syst. Archit.4
2006 A Multiobjective Genetic Fuzzy Approach for Intelligent System-level Exploration in Parameterized VLIW Processor Design
abstract
The design of a complex embedded system is dominated by the definition of an optimal architecture in relation to certain performance indexes. This activity, known as Design Space Exploration (DSE), is a great challenge for the EDA (Electronic Design Automation) community. The enormous size of the design space, in fact, together with the long simulation time required to evaluate each system configuration during the exploration process, cause DSE to become a bottleneck in the design flow. In this paper we propose a Multiobjective Design Space Exploration methodology based on a Genetic Fuzzy System, the aim of which is to drastically reduce the exploration time while guaranteeing a high level of accuracy. The methodology uses a Genetic Algorithm (GA) for heuristic exploration and a Fuzzy System to evaluate the configurations visited. Although of general application, the methodology is applied to a real case study: optimization of the performance and power consumption of an embedded architecture based on a Very Long Instruction Word (VLIW) microprocessor in a mobile multimedia application domain. The results obtained are compared, in terms of both accuracy and efficiency, with the state of the art in multiobjective DSE strategies, represented by the classical GA approach, demonstrating the scalability and effectiveness of the proposed approach.
Giuseppe Ascia, Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti
IEEE Congress on Evolutionary Computation4
2006 Deadlock Free Routing Algorithms for Mesh Topology NoC Systems with Regions
abstract
Region concept helps to accommodate cores larger than the tile size in mesh topology NoC architectures. In addition, it offers many new opportunities for NoC design, as well as provides new design issues and challenges. The most important among these is the design of a deadlock free routing algorithm. In this paper, we present and compare two routing algorithms for mesh topology NoC with regions. The first algorithm is borrowed from the area of fault tolerant networks and is adapted for the NoC context. We compare this with an algorithm designed using a methodology for design of application specific routing algorithms for communication networks. Our study shows that the application specific routing algorithm not only provides much higher adaptivity, but also superior performance as compared to the other algorithm in all traffic cases
Rickard Holsmark, Maurizio Palesi, Shashi Kumar
DSD2
2005 Exploring Design Space of VLIW Architectures
abstract
Architectures based on very long instruction word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of instruction level parallelism (ILP) with a reasonable tradeoff in complexity and silicon costs. Effective compiler support for predicated execution using the hyperblock, drastically increases the ILP even for control-dominated applications in which the branch instruction frequency is very high. The use of these techniques, however, is known to increase the instruction footprint, consequently putting pressure on the memory hierarchy. In this paper, we evaluate the performance/power trade-off in a system comprising a VLIW processor and a two-level hierarchical memory subsystem. Via simulation, we show that the efficiency of a compiler that is able to exploit predicate execution by hyperblock formation is greatly affected by the configuration of the memory subsystem as well as the configurable processor parameters. The enabling or disabling of hyperblock formation should therefore not be evaluated separately or independently, but seen as a further free parameter to be tuned in a strategy of design space exploration.
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti
ASAP3
2005 A system-level framework for evaluating area/performance/power trade-offs of VLIW-based embedded systems
abstract
Architectures based on Very Long Instruction Word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of Instruction Level Parallelism (ILP) with a reasonable trade-off in complexity and silicon costs. In this case Application Specific Instruction-set Processor (ASIP) specialization may require not only manipulation of the instruction-set but also tuning of the architectural parameters of the processor (e.g. the number and type of functional units, register files, etc.) and the memory subsystem (cache size, associativity, etc.). Setting the parameters so as to optimize certain metrics requires the use of efficient Design Space Exploration (DSE) strategies and also simulation tools (retargetable compilers and simulators) and accurate estimation models operating at a high level of abstraction. In this paper we present a framework for evaluation, in terms of performance, cost and power consumption, of a system based on a parameterized VLIW microprocessor together with the memory hierarchy subsystem following execution of a specific application. The framework, which can be freely downloaded from the Internet, implements a number of multi-objective DSE strategies to obtain Pareto-optimal configurations for the system.
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti
ASP-DAC3
2005 An evolutionary approach to network-on-chip mapping problem
abstract
The paper addresses the problem of topological mapping of intellectual properties (IPs) on the tiles of a mesh-based network on chip (NoC) architecture. The aim is to obtain the Pareto mappings that maximize performance and minimize the amount of power consumption. As the problem is an NP-hard one, we propose a heuristic technique based on evolutionary computing to obtain an optimal approximation of the Pareto-optimal front in an efficient and accurate way. At the same time, two of the most widely-known approaches to mapping in mesh-based NoC architectures are extended in order to explore the mapping space in a multi-criteria mode. The approaches are then evaluated and compared, in terms of both accuracy and efficiency, on a platform based on an event-driven trace-based simulator which makes it possible to take account of important dynamic effects that have a great impact on mapping. The evaluation performed on real applications (an MPEG-4 codec) confirms the efficiency, accuracy and scalability of the proposed approach
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi
Congress on Evolutionary Computation3
2005 A multiobjective genetic approach for system-level exploration in parameterized systems-on-a-chip
abstract
This paper deals with a significant problem affecting embedded system design methods based on parameterized systems on a chip (SOCs). It proposes a strategy for exploration of the configuration space of a parameterized SOC architecture to determine an accurate approximation of the power/performance Pareto-front. The strategy is based on genetic algorithms and is thoroughly evaluated in terms of accuracy, efficiency, and scalability using SOC platforms that differ as regards both architectural model and complexity. The results obtained show that the proposed approach gives an excellent approximation of the Pareto-optimal front in very short exploration times (up to two orders of magnitude shorter than those required by one of the best known and widely referenced approaches in the literature). In addition, our approach possesses a good degree of scalability as performance levels are maintained even when the architectural complexity increases.
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2004 A GA-based design space exploration framework for parameterized system-on-a-chip platforms
abstract
The constant increase in levels of integration and reduction in the time-to-market has led to the definition of new methodologies, which lay emphasis on reuse. One emerging approach in this context is platform-based design. The basic idea is to avoid designing a chip from scratch. Some portions of the chip's architecture are predefined for a specific type of application. This implies that the basic micro-architecture of the implementation is essentially "fixed," i.e., the principal components should remain the same within a certain degree of parameterization. Many researchers predict that platforms will take the lion's share of the integrated circuit market. In this paper, we propose an approach based on genetic algorithms for exploring the design space of parameterized system-on-a-chip (SOC) platforms. Our strategy focuses on exploration of the architectural parameters of the processor, memory subsystem and bus, making up the hardware kernel of a parameterized SOC platform for the design of embedded systems with strict power consumption and performance constraints. The approach has been validated on two different parameterized architectures: one based on a RISC processor and another based on a parameterized very long instruction word architecture. The results obtained on a suite of benchmarks for embedded applications are discussed in terms of both accuracy and efficiency. As far as accuracy is concerned, the approach gives solutions uniformly distributed in a region less than 1% from the Pareto-optimal front. As regards efficiency, the exploration times required by the approach are up to 20 times shorter than those required by one of the most efficient and widely referenced approaches in the literature.
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi
IEEE Trans. Evol. Comput.3
2003 An evolutionary approach for reducing the switching activity in address buses
abstract
In this paper we present two new approaches based on genetic algorithms (GA) to reduce power consumption by communication buses in an embedded system. The first approach makes it possible to obtain the truth table of an encoder that minimizes switching activity on a bus, whereas the second outputs the netlist of the encoder using the lowest possible number of logic gates. Both approaches are static, in the sense that the encoders are generated ad hoc for specific traffic. This is not, however, a limiting hypothesis if the application scenario considered is that of embedded systems. An embedded system, in fact, executes the same application throughout its lifetime and so it is possible to have detailed knowledge of the trace of the patterns transmitted on a bus following execution of a specific application. The approach is compared with the most efficient encoding schemes proposed in the literature on both multiplexed and separate buses. The results obtained demonstrate the validity of the approach, which on average saves up to 50% of the transitions normally required, as well as their practical applicability, even in an on-chip environment.
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Antonio Parlato
IEEE Congress on Evolutionary Computation3
2003 A Genetic Approach To Bus Encoding
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi
VLSI-SOC3