VLDB 2026 Research / reviewers in the wild / expert
Giuseppe Ascia
dblp:23/357
· DBLP profile ↗
47ranked-venue papers
27as first author
10since 2021 · last 2026
0000-0001-7452-5828ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 11 first-authorComputer networks · 6 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI WorkloadsabstractEnergy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
DATE | 11 |
| 2026 | Assessing the Role of Communication in Modular Multi-Core Quantum SystemsabstractThe scalability of quantum computing is constrained by the physical and architectural limitations of monolithic quantum processors. Modular multi-core quantum architectures, which interconnect multiple quantum cores (QCs) via classical and quantum-coherent links, offer a promising alternative to address these challenges. However, transitioning to a modular architecture introduces communication overhead, where classical communication plays a crucial role in executing quantum algorithms by transmitting measurement outcomes and synchronizing operations across QCs. Understanding the impact of classical communication on execution time is therefore essential for optimizing system performance. In this work, we introduce qcomm , an open-source simulator designed to evaluate the role of classical communication in modular quantum computing architectures. qcomm provides a high-level execution and timing model that captures the interplay between quantum gate execution, entanglement distribution, teleportation protocols, and classical communication latency. We conduct an extensive experimental analysis to quantify the impact of classical communication bandwidth, interconnect types, and quantum circuit mapping strategies on overall execution time. Furthermore, we assess classical communication overhead when executing real quantum benchmarks mapped onto a cryogenically-controlled multi-core quantum system. Our results show that, while classical communication is generally not the dominant contributor to execution time, its impact becomes increasingly relevant in optimized scenarios—such as improved quantum technology, large-scale interconnects, or communication-aware circuit mappings. These findings provide useful insights for the design of scalable modular quantum architectures and highlight the importance of evaluating classical communication as a performance-limiting factor in future systems. Maurizio Palesi, Enrico Russo 0002, Giuseppe Ascia, Hamaad Rafique, Davide Patti, Vincenzo Catania, Sergi Abadal, Abhijit Das 0002, Pau Escofet, Eduard Alarcón, Carmen G. Almudéver |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | Optimizing Qubit Assignment in Modular Quantum Systems via Attention-Based Deep Reinforcement LearningabstractModular, distributed, and multi-core architectures are considered a promising solution for scaling quantum computing systems. Optimising communication is crucial to preserve quantum coherence. The compilation and mapping of quantum circuits should minimise state transfers while adhering to architec-tural constraints. To address this problem efficiently, we propose a novel approach using Reinforcement Learning (RL) to learn heuristics for a specific multi-core architecture. Our RL agent uses a Transformer encoder and Graph Neural Networks, encoding quantum circuits with self-attention and producing outputs via an attention-based pointer mechanism to match logical qubits with physical cores efficiently. Experimental results show our method outperform the baseline reducing by 28% inter-core communications for random circuits while minimising time-to-solution. Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania |
DATE | 4 |
| 2025 | An Anomaly Detection Model for RISC-V in Automotive Applications: A Domain-Specific Accelerator PerspectiveabstractEarly anomaly detection in automotive systems is crucial for enhancing user safety and enabling timely corrective actions, thereby minimizing the risks associated with system malfunctions. This paper presents an approach for implementing Artificial Intelligence (AI)-based algorithms for anomaly detection in the automotive domain, leveraging the RISC-V architecture in conjunction with Domain-Specific Accelerators (DSAs). By exploiting the efficiency of DSAs, the proposed system aims to achieve faster anomaly detection compared to traditional processing methods. A detailed comparison is conducted between the performance of executing the AI-based anomaly detection algorithm on the RISC-V core versus offloading it to an optimized hardware accelerator tailored to the specific AI model. The goal of this work is to provide valuable insights into the potential of RISCV and DSAs to enhance AI-driven safety mechanisms, contributing to the development of more reliable automotive systems. Elio Vinciguerra, Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia |
PDP | 4 |
| 2024 | A Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator SystemsabstractCurrently, there is a growing trend of outsourcing the execution of DNNs to cloud services. For service providers, managing multitenancy and ensuring high-quality service delivery, particularly in meeting stringent execution time constraints, assumes paramount importance, all while endeavoring to maintain cost-effectiveness. In this context, the utilization of heterogeneous multi-accelerator systems becomes increasingly relevant. This paper presents RELMAS, a low-overhead deep reinforcement learning algorithm designed for the online scheduling of DNNs in multi-tenant environments, taking into account the dataflow heterogeneity of accelerators and memory bandwidths contentions. By doing so, service providers can employ the most efficient scheduling policy for user requests, optimizing Service-Level-Agreement (SLA) satisfaction rates and enhancing hardware utilization. The application of RELMAS to a heterogeneous multi-accelerator system composed of various instances of Simba and Eyeriss sub-accelerators resulted in up to a 173% improvement in SLA satisfaction rate compared to state-of-the-art scheduling techniques across different workload scenarios, with less than a 1.5% energy overhead. Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania |
DAC | 5 |
| 2024 | Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement LearningabstractThis paper addresses the critical challenge of managing Quality of Service (QoS) in cloud services, focusing on the nuances of individual tenant expectations and varying Service Level Indicators (SLIs). It introduces a novel approach utilizing Deep Reinforcement Learning for tenant-specific QoS management in multi-tenant, multi-accelerator cloud environments. The chosen SLI, deadline hit rate, allows clients to tailor QoS for each service request. A novel online scheduling algorithm for Deep Neural Networks in multi-accelerator systems is proposed, with a focus on guaranteeing tenant-wise, model-specific QoS levels while considering real-time constraints. Enrico Russo 0002, Francesco Giulio Blanco, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Vincenzo Catania |
ISCAS | 4 |
| 2023 | Memory-Aware DNN Algorithm-Hardware Mapping via Integer Linear ProgrammingabstractMapping a deep neural network (DNN) layer onto domain-specific accelerators can require an intractable number of choices regarding loop factorization, ordering, and spatial unrolling. Determining the optimal mapping that achieves the best figures in terms of latency and energy efficiency can be difficult due to the vast number of possible candidates that need to be exhaustively evaluated. Many techniques have been recently proposed for fast and efficient mapping space exploration; some of them adopt a black-box optimization approach, others make assumptions on the underlying accelerator memory hierarchy or require time-consuming model retraining. We propose an integer linear programming (ILP) approach and formulate a mathematical model, namely LEMON, that takes into account number of accesses to each buffer, energy costs and buffer bandwidths in the accelerator and is flexible enough to work with different memory hierarchies. Compared with state-of-the-art techniques, LEMON achieves up to 83% energy-delay product reduction when compared to another ILP-based approach (CoSA) and 27% when compared to a genetic algorithm approach (GAMMA). Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Salvatore Monteleone, Vincenzo Catania |
CF | 3 |
| 2023 | Multiobjective End-to-End Design Space Exploration of Parameterized DNN AcceleratorsabstractDeep neural network (DNN) hardware accelerators enable the execution of complex DNN inferences on resource-constrained IoT devices. Inference performance and energy figures depend on how the DNN layers are mapped into the accelerator and how the architecture of the accelerator fits the variety of layers’ shapes of the actual DNN. The mapping determines the execution order of the operations, both temporally and spatially. Thus, selecting the best mapping that allows fitting the DNN model to the specific accelerator is of paramount importance to meet the strong constraints imposed by resource-scarce IoT platforms. Although several mapping space exploration techniques have been proposed in the literature, they are focused on determining the best mapping for a given layer, for a given architecture, and for optimizing a single objective. This article largely extends the scope of the exploration by considering the huge design space spanned by mapping related and architectural parameters, considering all the layers of the DNN, and optimizing multiple objectives simultaneously. We present EPOCA, end-to-end Pareto optimization of DNN accelerators, whose goal is to determine the accelerator’s architecture and the mapping for each layer that optimizes end-to-end and in a multiobjective fashion a set of conflicting design criteria. We assess EPOCA on different DNN models on a parameterized hardware accelerator designed for IoT applications and compare them with a state-of-the-art mapping space explorer, considering the area, inference latency, and inference energy as optimization metrics. We show that the set of Pareto solutions found by EPOCA provides the designer with a range of choices from which to select the best tradeoff with respect to the specific application. Enrico Russo 0002, Maurizio Palesi, Davide Patti, Salvatore Monteleone, Giuseppe Ascia, Vincenzo Catania |
IEEE Internet Things J. | 5 |
| 2022 | MEDEA: A Multi-objective Evolutionary Approach to DNN Hardware MappingabstractDeep Neural Networks (DNNs) embedded domain-specific accelerators enable inference on resource-constrained devices. Making optimal design choices and efficiently scheduling neural network algorithms on these specialized architectures is challenging. Many choices can be made to schedule computation spatially and temporally on the accelerator. Each choice influences the access pattern to the buffers of the architectural hierarchy, affecting the energy and latency of the inference. Each mapping also requires specific buffer capacities and a number of spatial components instances that translate in different chip area occupation. The space of possible combinations, the mapping space, is so large that automatic tools are needed for its rapid ex-ploration and simulation. This work presents MEDEA, an open-source multi-objective evolutionary algorithm based approach to DNNs accelerator mapping space exploration. MEDEA leverages the Timeloop analytical cost model. Differently from the other schedulers that optimize towards a single objective, MEDEA allows deriving the Pareto set of mappings to optimize towards multiple, sometimes conflicting, objectives simultaneously. We found that solutions found by MEDEA dominates in most cases those found by state-of-the-art mappers. Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Vincenzo Catania |
DATE | 5 |
| 2022 | DNN Model Compression for IoT Domain-Specific Hardware AcceleratorsabstractMachine learning techniques, particularly those based on neural networks, are always more often used at the edge of the network by Internet of Things (IoT) nodes. Unfortunately, the computation capabilities demanded by those applications, together with their energy efficiency-related constraints, exceed those exposed by embedded general-purpose processors. For this reason, the use of domain-specific hardware accelerators (DSAs) is considered the most viable solution to the unsustainable “Turing tariff” of general-purpose hardware. Starting from the observation that memory and communication traffic account for a large fraction of the overall latency and energy in deep neural network (DNN) inferences, this article proposes a new compression technique aimed at: 1) reducing the memory footprint for storing the model parameters of a DNN and 2) improving DNN inference latency and energy on resource-constrained IoT devices. The proposed compression technique, namely, LineCompress, is applied on a set of representative convolutional neural networks (CNNs) for object recognition mapped on a state-of-the-art DSA targeted for resource-constrained IoT devices. We show that on average,$7.4\times $memory footprint reduction can be obtained, thus reducing the memory and communication traffic that result to 77% and 87% inference latency and energy reduction, respectively, trading-off efficiency versus accuracy. Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Andrea Mineo, Giuseppe Ascia, Vincenzo Catania |
IEEE Internet Things J. | 6 |
| 2020 | DNNZip: Selective Layers Compression Technique in Deep Neural Network AcceleratorsabstractIn Deep Neural Network (DNN) accelerators, the on-chip traffic and memory traffic accounts for a relevant fraction of the inference latency and energy consumption. A major component of such traffic is due to the moving of the DNN model parameters from the main memory to the memory interface and from the latter to the processing elements (PEs) of the accelerator. In this paper, we present DNNZip, a technique aimed at compressing the model parameters of a DNN, thus resulting in significant energy and performance improvement. DNNZip implements a lossy compression whose compression ratio is tuned based on the maximum tolerated error on the model parameters provided by the user. DNNZip is assessed on several convolutional NNs and the trade-off inference energy saving vs. inference latency reduction vs. network accuracy degradation is discussed. We found that up to 64% energy saving, and up to 67% latency reduction can be obtained with a limited impact on the accuracy of the network. Habiba Lahdhiri, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Jordane Lorandel, Emmanuelle Bourdel, Vincenzo Catania |
DSD | 5 |
| 2020 | Improving Inference Latency and Energy of DNNs through Wireless Enabled Multi-Chip-Module-based Architectures and Model Parameters CompressionabstractPerformance and energy figures of Deep Neural Network (DNN) accelerators are profoundly affected by the communication and memory sub-system. In this paper, we make the case of a state-of-the-art multi-chip-module-based architecture for DNN inference acceleration. We propose a hybrid wired/wireless network-in-package interconnection fabric and a compression technique for drastically improving the communication efficiency and reducing the memory and communication traffic with a consequent improvement of performance and energy metrics. We assess the inference performance and energy improvement vs. accuracy degradation for different CNNs showing that up to 77% and 68% of inference latency reduction and inference energy reduction, respectively, can be obtained while keeping the accuracy degradation below 5% as respect to the original uncompressed CNN. Giuseppe Ascia, Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
NOCS | 1 |
| 2020 | Exploiting Data Resilience in Wireless Network-on-chip ArchitecturesabstractThe emerging wireless Network-on-Chip (WiNoC) architectures are a viable solution for addressing the scalability limitations of manycore architectures in which multi-hop long-range communications strongly impact both the performance and energy figures of the system. The energy consumption of wired links as well as that of radio communications account for a relevant fraction of the overall energy budget. In this article, we extend the approximate computing paradigm to the case of the on-chip communication system in manycore architectures. We present techniques, circuitries, and programming interfaces aimed at reducing the energy consumption of a WiNoC by exploiting the trade-off energy saving vs. application output degradation. The proposed platform—namely, xWiNoC—uses variable voltage swing links and tunable transmitting power wireless interfaces along with a programming interface that allows the programmer to specify those data structures that are error-resilient. Thus, communications induced by the access to such error-resilient data structures are carried out by using links and radio channels that are configured to work in a low energy mode, albeit by exposing a higher bit error rate. xWiNoC is assessed on a set of applications belonging to different domains in which the trade-off energy vs. performance vs. application result quality is discussed. We found that up to 50% of communication energy saving can be obtained with a negligible impact on the application output quality and 3% in application performance degradation. Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose, Valerio Mario Salerno |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2019 | Analyzing networks-on-chip based deep neural networksabstractOne of the most promising architectures for performing deep neural network inferences on resource-constrained embedded devices is based on massive parallel and specialized cores interconnected by means of a Network-on-Chip (NoC). In this paper, we extensively evaluate NoC-based deep neural network accelerators by exploring the design space spanned by several architectural parameters. We show how latency is mainly dominated by the on-chip communication whereas energy consumption is mainly accounted by memory (both on-chip and off-chip). Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose |
NOCS | 1 |
| 2016 | Making Android Apps Data-Leak-Safe by Data Flow Analysis and Code InjectionabstractSome support is needed in order to shun the possibility that sensitive data handled by applications are sent to improper destinations. Although apps running on Android OS declare the accessed services, once the user accepts, the application receives complete permissions and may use sensitive data improperly. Some tools have emerged to check data access and flow, however such tools are either based on static analysis or dynamic tracking. The former brings no overhead at run-time, but is less precise, the latter can bring a costly overhead during execution, having to monitor any access to sensitive data and all destinations. Our approach is innovative in that it takes advantage of static analysis and then monitors at run-time only data paths that potentially give sensitive data out. The correspondent tool is tailored to Android environment, tool-chain, libraries, and typical requirements that applications have to satisfy. Giuseppe Ascia, Vincenzo Catania, Raffaele Di Natale, Andrea Fornaia, Misael Mongiovì, Salvatore Monteleone, Giuseppe Pappalardo, Emiliano Tramontana |
WETICE | 1 |
| 2016 | On-Chip Communication Energy Reduction Through Reliability Aware Adaptive Voltage Swing ScalingabstractIn a multi/many-core system, the network-on-chip (NoC)-based communication backbone is responsible for a relevant fraction of the overall energy budget. Reducing the voltage swing for signaling in crossbars and links results in significant energy saving. Unfortunately, as voltage swing reduces, the bit error rate increases, that in turn compromises the communication reliability. Starting from the assumption that not all the communications need same level of reliability, in this paper we propose techniques and architectures for run-time tuning of the voltage swing of the crossbars and interrouter links. The proposed technique is compared with the state of the art in link energy reduction through data encoding under both synthetic and real traffic scenarios. We found that the proposed techniques allow to significantly reduce the energy consumption of the NoC fabric without degrading the performance metrics. Energy savings ranging from 20% to 43% have been observed without any relevant impact on the performance metrics. Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Partha Pratim Pande, Vincenzo Catania |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Runtime Tunable Transmitting Power Technique in mm-Wave WiNoC ArchitecturesabstractEmerging on-chip communication technologies, like wireless networks-on-chip (WiNoCs), have been recently proposed as candidate solutions for addressing the scalability limitations of conventional multihop network on chip (NoC) architectures. In a WiNoC, a subset of network nodes, namely, radio hubs, are equipped with a wireless interface that allows them to wirelessly communicate with other radio hubs. Thus, long-range communications, which would involve multiple hops in a conventional wireline NoC, can be realized by a single hop through the radio medium. Unfortunately, the energy consumed by the RF transceiver into the radio hub (i.e., the main building block in a WiNoC), and in particular by its transmitter, accounts for a significant fraction of the overall communication energy. In order to alleviate such contribution, this paper presents a runtime tunable transmitting power technique for improving the energy efficiency of the transceiver in WiNoC architectures. The basic idea is tuning the transmitting power based on the physical location of the recipient of the current communication. Specifically, based on the destination address of the incoming packet, the radio hub tunes its transmitting power to a minimum level, but high enough to reach the destination antenna without exceeding a certain bit error ratio. The proposed technique is general and can be applied to any WiNoC architecture. Its application on different representative WiNoC architectures results in an average energy reduction up to 50% without any impact on performance and with a negligible overhead in terms of silicon area. Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | A closed loop transmitting power self-calibration scheme for energy efficient WiNoC architectures
Andrea Mineo, Mohd Shahrizal Rusli, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania, Muhammad N. Marsono |
DATE | 4 |
| 2014 | An adaptive transmitting power technique for energy efficient mm-wave wireless NoCsabstractSeveral emerging techniques have been recently proposed for alleviating the communication latency and the energy consumption issues in multi/many-core architectures. One of such emerging communication techniques, namely, WiNoC replaces the traditional wired links with the use of wireless medium. Unfortunately, the energy consumed by the RF transceiver (i.e., the main building block of a WiNoC), and in particular by its transmitter, accounts for a significant fraction of the overall communication energy. In this paper we propose a runtime tunable transmitting power technique for improving the energy efficiency of the transceiver in wireless NoC architectures. The basic idea is tuning the transmitting power based on the location of the recipient of the current communication. The integration of the proposed technique into two known WiNoC architectures, namely, iWise64 and McWiNoC resulted in an energy reduction of 43% and 60%, respectively. Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania |
DATE | 3 |
| 2013 | An Adaptive Output Selection Function Based on a Fuzzy Rule Base System for Network on ChipabstractIn Network-on-chip design two of the most important performance indices are the average delay of packets and power dissipation. The first depends on the level of congestion of the communication system, the second is strongly influenced by the power dissipated by the links of a network-on-chip (NoC) which accounts for a significant fraction of the overall power dissipated by the on-chip communication fabric. Such fraction becomes more and more relevant as technology shrinks. The selection policy of the output port in the NOC router with adaptive routing function affects both the performance and power dissipation. In fact, generally a NOC with a lower average delay has a lower power dissipation on the links for transmission. The selection policies proposed in the literature have the object or the minimization of the average delay or in some cases the reduction of power dissipation. In this paper we propose a selection policy aimed at minimizing both the average delay and power consumption of the communication system of NoC based architectures. The proposed policy uses a cost function obtained using a Fuzzy Rule Base System (FRBS) with two inputs an one output. Inputs are the level of use of the buffers of the downstream routers, and the link power dissipation of the selected output. The output of FRBS is the cost function. The experimental results, obtained using a cycle accurate NoC simulator for different traffic scenarios, shows that the proposed selection policy outperforms other selection approaches both in terms of average delay and in terms on power dissipation and energy consumption. Giuseppe Ascia, Maurizio Palesi, Vincenzo Catania |
DSD | 1 |
| 2013 | Runtime Online Links Voltage Scaling for Low Energy Networks on ChipabstractThe power dissipated by the links of a network on chip (NoC) accounts for a significant fraction of the overall power budget. Reducing the operating voltage of the network links, results in a square reduction of their power contribution. Unfortunately, the voltage reduction has a negative impact on communication reliability in terms of bit error rate. Starting from the assumption that not all the communications require the same reliability level, we present a technique for dynamically changing the voltage of the links based on the communication reliability requirements. The experiments, carried out under both synthetic and real traffic scenarios, show the effectiveness of the proposed technique which allows to save up of 55% of link energy with a total energy saving of 25% of the entire NoC. Andrea Mineo, Maurizio Palesi, Giuseppe Ascia, Vincenzo Catania |
DSD | 3 |
| 2012 | A Study on Evolutionary Multi-Objective Optimization with Fuzzy Approximation for Computational Expensive Problems
Alessandro G. Di Nuovo, Giuseppe Ascia, Vincenzo Catania |
PPSN (2) | 2 |
| 2011 | Data Encoding Schemes in Networks on ChipabstractAn ever more significant fraction of the overall power dissipation of a network-on-chip (NoC) based system-on-chip (SoC) is due to the interconnection system. In fact, as technology shrinks, the power contribute of NoC links starts to compete with that of NoC routers. In this paper, we propose the use of data encoding techniques as a viable way to reduce both power dissipation and energy consumption of NoC links. The proposed encoding scheme exploits the wormhole switching techniques and works on an end-to-end basis. That is, flits are encoded by the network interface (NI) before they are injected in the network and are decoded by the destination NI. This makes the scheme transparent to the underlying network since the encoder and decoder logic is integrated in the NI and no modification of the routers architecture is required. We assess the proposed encoding scheme on a set of representative data streams (both synthetic and extracted from real applications) showing that it is possible to reduce the power contribution of both the self-switching activity and the coupling switching activity in inter-routers links. As results, we obtain a reduction in total power dissipation and energy consumption up to 37% and 18%, respectively, without any significant degradation in terms of both performance and silicon area. Maurizio Palesi, Giuseppe Ascia, Fabrizio Fazzino, Vincenzo Catania |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Data Encoding for Low-Power in Wormhole-Switched Networks-on-ChipabstractAs the number of cores in a chip increases, the role played by the communication system becomes more and more central. An on-chip communication infrastructure based on the Network-on-Chip (NoC) paradigm is today recognized as the most effective and scalable solution able to deal with the communication issues that will characterize the next generation of many-cores architectures. An ever more significant fraction of the overall chip area is devoted to support advanced and reliable communication protocols making the energy resources used for communication starting to compete with the ones spent for computation. Amongst the communication resources, as technology shrinks, the power ratio between NoC links and routers increases making the links becoming more power-hungry than routers. In this paper we propose a novel endto-end data encoding scheme which exploits the wormhole technique commonly used in NoC-based system to reduce power dissipated by the NoC links. We assess the proposed encoding scheme on a set of representative data streams showing that it is possible to reduce the power contribution of both the self switching activity and the coupling switching activity in inter-routers links. As results, we obtain a reduction in total power dissipation and energy consumption up to 26% and 9% respectively without any significant degradation in terms of both performance and silicon area. The encoder and decoder logic is integrated in the network interface and is transparent to the underling NoC. Maurizio Palesi, Fabrizio Fazzino, Giuseppe Ascia, Vincenzo Catania |
DSD | 3 |
| 2008 | Implementation and Analysis of a New Selection Strategy for Adaptive Routing in Networks-on-ChipabstractEfficient and deadlock-free routing is critical to the performance of networks-on-chip. The effectiveness of any adaptive routing algorithm strongly depends on the underlying selection strategy. A selection function is used to select the output channel where the packet will be forwarded on. In this paper we present a novel selection strategy that can be coupled with any adaptive routing algorithm. The proposed selection strategy is based on the concept of Neighbors-on-Path the aims of which is to exploit the situations of indecision occurring when the routing function returns several admissible output channels. The overall objective is to choose the channel that will allow the packet to be routed to its destination along a path that is as free as possible of congested nodes. Performance evaluation is carried out by using a flit-accurate simulator under traffic scenarios generated by both synthetic and real applications. Results obtained show how the proposed selection strategy applied to the Odd-Even routing algorithm yields an improvement in both average delay and saturation point up to 20% and 30% on average respectively, with a minimal overhead in terms of area occupation. In addition, a positive effect on total energy consumption is also observed under near-congestion packet injection rates. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti |
IEEE Trans. Computers | 1 |
| 2007 | Efficient design space exploration for application specific systems-on-a-chip
Giuseppe Ascia, Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti |
J. Syst. Archit. | 1 |
| 2006 | A Multiobjective Genetic Fuzzy Approach for Intelligent System-level Exploration in Parameterized VLIW Processor DesignabstractThe design of a complex embedded system is dominated by the definition of an optimal architecture in relation to certain performance indexes. This activity, known as Design Space Exploration (DSE), is a great challenge for the EDA (Electronic Design Automation) community. The enormous size of the design space, in fact, together with the long simulation time required to evaluate each system configuration during the exploration process, cause DSE to become a bottleneck in the design flow. In this paper we propose a Multiobjective Design Space Exploration methodology based on a Genetic Fuzzy System, the aim of which is to drastically reduce the exploration time while guaranteeing a high level of accuracy. The methodology uses a Genetic Algorithm (GA) for heuristic exploration and a Fuzzy System to evaluate the configurations visited. Although of general application, the methodology is applied to a real case study: optimization of the performance and power consumption of an embedded architecture based on a Very Long Instruction Word (VLIW) microprocessor in a mobile multimedia application domain. The results obtained are compared, in terms of both accuracy and efficiency, with the state of the art in multiobjective DSE strategies, represented by the classical GA approach, demonstrating the scalability and effectiveness of the proposed approach. Giuseppe Ascia, Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti |
IEEE Congress on Evolutionary Computation | 1 |
| 2006 | An integrated fuzzy-GA approach for buffer managementabstractThis paper deals with a novel buffer management scheme based on evolutionary computing for shared-memory asynchronous transfer mode (ATM) switches. The philosophy behind it is adaptation of the threshold for each logical output queue to the real traffic conditions by means of a system of fuzzy inferences. The optimal fuzzy system is achieved using a systematic methodology, based on genetic algorithms (GAs), which allows the fuzzy system parameters to be derived for each switch size, offering a high degree of scalability to the fuzzy control system. Its performance is comparable to that of the push-out (PO) mechanism, which can be considered ideal from a performance viewpoint, and at any rate much better than that of threshold schemes based on conventional logic. In addition, the fuzzy threshold (FT) scheme is simple and cost-effective when implemented using VLSI technology Giuseppe Ascia, Vincenzo Catania, Daniela Panno |
IEEE Trans. Fuzzy Syst. | 1 |
| 2005 | Exploring Design Space of VLIW ArchitecturesabstractArchitectures based on very long instruction word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of instruction level parallelism (ILP) with a reasonable tradeoff in complexity and silicon costs. Effective compiler support for predicated execution using the hyperblock, drastically increases the ILP even for control-dominated applications in which the branch instruction frequency is very high. The use of these techniques, however, is known to increase the instruction footprint, consequently putting pressure on the memory hierarchy. In this paper, we evaluate the performance/power trade-off in a system comprising a VLIW processor and a two-level hierarchical memory subsystem. Via simulation, we show that the efficiency of a compiler that is able to exploit predicate execution by hyperblock formation is greatly affected by the configuration of the memory subsystem as well as the configurable processor parameters. The enabling or disabling of hyperblock formation should therefore not be evaluated separately or independently, but seen as a further free parameter to be tuned in a strategy of design space exploration. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti |
ASAP | 1 |
| 2005 | A system-level framework for evaluating area/performance/power trade-offs of VLIW-based embedded systemsabstractArchitectures based on Very Long Instruction Word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of Instruction Level Parallelism (ILP) with a reasonable trade-off in complexity and silicon costs. In this case Application Specific Instruction-set Processor (ASIP) specialization may require not only manipulation of the instruction-set but also tuning of the architectural parameters of the processor (e.g. the number and type of functional units, register files, etc.) and the memory subsystem (cache size, associativity, etc.). Setting the parameters so as to optimize certain metrics requires the use of efficient Design Space Exploration (DSE) strategies and also simulation tools (retargetable compilers and simulators) and accurate estimation models operating at a high level of abstraction. In this paper we present a framework for evaluation, in terms of performance, cost and power consumption, of a system based on a parameterized VLIW microprocessor together with the memory hierarchy subsystem following execution of a specific application. The framework, which can be freely downloaded from the Internet, implements a number of multi-objective DSE strategies to obtain Pareto-optimal configurations for the system. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti |
ASP-DAC | 1 |
| 2005 | An evolutionary approach to network-on-chip mapping problemabstractThe paper addresses the problem of topological mapping of intellectual properties (IPs) on the tiles of a mesh-based network on chip (NoC) architecture. The aim is to obtain the Pareto mappings that maximize performance and minimize the amount of power consumption. As the problem is an NP-hard one, we propose a heuristic technique based on evolutionary computing to obtain an optimal approximation of the Pareto-optimal front in an efficient and accurate way. At the same time, two of the most widely-known approaches to mapping in mesh-based NoC architectures are extended in order to explore the mapping space in a multi-criteria mode. The approaches are then evaluated and compared, in terms of both accuracy and efficiency, on a platform based on an event-driven trace-based simulator which makes it possible to take account of important dynamic effects that have a great impact on mapping. The evaluation performed on real applications (an MPEG-4 codec) confirms the efficiency, accuracy and scalability of the proposed approach Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi |
Congress on Evolutionary Computation | 1 |
| 2005 | A multiobjective genetic approach for system-level exploration in parameterized systems-on-a-chipabstractThis paper deals with a significant problem affecting embedded system design methods based on parameterized systems on a chip (SOCs). It proposes a strategy for exploration of the configuration space of a parameterized SOC architecture to determine an accurate approximation of the power/performance Pareto-front. The strategy is based on genetic algorithms and is thoroughly evaluated in terms of accuracy, efficiency, and scalability using SOC platforms that differ as regards both architectural model and complexity. The results obtained show that the proposed approach gives an excellent approximation of the Pareto-optimal front in very short exploration times (up to two orders of magnitude shorter than those required by one of the best known and widely referenced approaches in the literature). In addition, our approach possesses a good degree of scalability as performance levels are maintained even when the architectural complexity increases. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | An evolutionary management scheme in high-performance packet switchesabstractThis paper deals with a novel buffer management scheme based on the combination of evolutionary computing and fuzzy logic for shared-memory packet switches. The philosophy behind it is adaptation of the threshold for each logical output queue to the real traffic conditions by means of a system of fuzzy inferences. The optimal fuzzy system is achieved using a systematic methodology based on Genetic Algorithms for membership-function selecting and tuning. This methodology approach allows the fuzzy system parameters to be automatically derived when the switch parameters vary, offering a high degree of scalability to the fuzzy control system. Its performance is close to that of the push-out mechanism, which can be considered ideal from a performance viewpoint, and at any rate much better than that of threshold schemes based on conventional logic. In addition, the fuzzy threshold scheme is simple to implement, unlike the push-out mechanism which is not practically feasible in high-speed switches due to the amount of time required for computation, and above all inexpensive when implemented using current standard technology. Giuseppe Ascia, Vincenzo Catania, Daniela Panno |
IEEE/ACM Trans. Netw. | 1 |
| 2004 | A GA-based design space exploration framework for parameterized system-on-a-chip platformsabstractThe constant increase in levels of integration and reduction in the time-to-market has led to the definition of new methodologies, which lay emphasis on reuse. One emerging approach in this context is platform-based design. The basic idea is to avoid designing a chip from scratch. Some portions of the chip's architecture are predefined for a specific type of application. This implies that the basic micro-architecture of the implementation is essentially "fixed," i.e., the principal components should remain the same within a certain degree of parameterization. Many researchers predict that platforms will take the lion's share of the integrated circuit market. In this paper, we propose an approach based on genetic algorithms for exploring the design space of parameterized system-on-a-chip (SOC) platforms. Our strategy focuses on exploration of the architectural parameters of the processor, memory subsystem and bus, making up the hardware kernel of a parameterized SOC platform for the design of embedded systems with strict power consumption and performance constraints. The approach has been validated on two different parameterized architectures: one based on a RISC processor and another based on a parameterized very long instruction word architecture. The results obtained on a suite of benchmarks for embedded applications are discussed in terms of both accuracy and efficiency. As far as accuracy is concerned, the approach gives solutions uniformly distributed in a region less than 1% from the Pareto-optimal front. As regards efficiency, the exploration times required by the approach are up to 20 times shorter than those required by one of the most efficient and widely referenced approaches in the literature. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi |
IEEE Trans. Evol. Comput. | 1 |
| 2003 | An evolutionary approach for reducing the switching activity in address busesabstractIn this paper we present two new approaches based on genetic algorithms (GA) to reduce power consumption by communication buses in an embedded system. The first approach makes it possible to obtain the truth table of an encoder that minimizes switching activity on a bus, whereas the second outputs the netlist of the encoder using the lowest possible number of logic gates. Both approaches are static, in the sense that the encoders are generated ad hoc for specific traffic. This is not, however, a limiting hypothesis if the application scenario considered is that of embedded systems. An embedded system, in fact, executes the same application throughout its lifetime and so it is possible to have detailed knowledge of the trace of the patterns transmitted on a bus following execution of a specific application. The approach is compared with the most efficient encoding schemes proposed in the literature on both multiplexed and separate buses. The results obtained demonstrate the validity of the approach, which on average saves up to 50% of the transitions normally required, as well as their practical applicability, even in an on-chip environment. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Antonio Parlato |
IEEE Congress on Evolutionary Computation | 1 |
| 2003 | A Genetic Approach To Bus Encoding
Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi |
VLSI-SOC | 1 |
| 2002 | An efficient buffer management policy based on an integrated Fuzzy-GA approachabstractThis paper deals with a novel buffer management scheme based on evolutionary computing for shared-memory ATM switches. The philosophy behind it is adaptation of the threshold for each output logical queue to the real traffic conditions by means of a system of fuzzy inferences. The optimal fuzzy system is achieved using a systematic methodology based on genetic algorithms for membership-function selecting and tuning. This methodology approach allows the fuzzy system parameters to be automatically derived when the switch parameters vary, offering a high degree of scalability to the fuzzy control system. Its performance is very close to that of an ideal mechanism like the push-out mechanism, and at any rate much better than that of the threshold schemes based on conventional logic. In addition it is simple to implement and above all inexpensive when implemented using VLSI technology. Giuseppe Ascia, Vincenzo Catania, Daniela Panno |
INFOCOM | 1 |
| 2001 | A General Purpose Processor Oriented Fuzzy ReasoningabstractThe paper presents the design of a RISC processor which is equipped with a set of fuzzy-oriented instructions, thus allowing a significant increase in performance. In addition, a hardware solution is defined that can eliminate the data and control hazards typical of pipelined processor for fuzzy applications. Giuseppe Ascia, Vincenzo Catania |
FUZZ-IEEE | 1 |
| 2001 | A Fuzzy Buffer Management Scheme For ATM and IP NetworksabstractWe address the relevant issue of managing traffic flows with different priorities in packet switched networks, namely ATM and IP networks. We consider a reference model in which two traffic flows with different priorities are multiplexed within the buffer of a cell-based switch. The solution we propose, based on the fuzzy system theory, is able to guarantee the QoS requirements of high-priority traffic flow, allowing at the same time the exploitation of unused buffer resources to accommodate low-priority traffic flow in order to maximize the total throughput. The performance assessment of the fuzzy scheme demonstrates that our solution outperforms other popular mechanisms, based on conventional logic, such as threshold and push out mechanisms. Giuseppe Ascia, Vincenzo Catania, Giuseppe Ficili, Daniela Panno |
INFOCOM | 1 |
| 2001 | An efficient fuzzy system for traffic management in high-speed packet-switched networks
Giuseppe Ascia, Vincenzo Catania, Daniela Panno |
Soft Comput. | 1 |
| 2000 | A pipeline parallel architecture for a fuzzy inference processorabstractThe paper presents the architecture of a VLSI processor for applications based on fuzzy logic. The main features of the architecture are: a pre-computation phase of the positive degree of truth of the antecedent with fuzzy inputs; a detection phase of the rules positive degree of activation, parallelism in some phases of inference which is split into a sequence of pipeline stages. The processing speed is up to 16.7 MFLIPS for inferences with 64 rules with 4 linguistic variables. Giuseppe Ascia, Vincenzo Catania |
FUZZ-IEEE | 1 |
| 1999 | VLSI hardware architecture for complex fuzzy systemsabstractThis paper presents the design of a VLSI fuzzy processor, which is capable of dealing with complex fuzzy inference systems, i.e., fuzzy inferences that include rule chaining. The architecture of the processor is based on a computational model whose main features are: the capability to cope effectively with complex fuzzy inference systems; a detection phase of the rule with a positive degree of activation to reduce the number of rules to be processed per inference; parallel computation of the degree of activation of active rules; and representation of membership functions based on /spl alpha/-level sets. As the fuzzy inference can be divided into different processing phases, the processor is made up of a number of stages which are pipelined. In each stage several inference processing phases are performed parallelly. Its performance is in the order of 2 MFLIPS with 256 rules, eight inputs, two chained variables, and four outputs and 5.2 MFLIPS with 32 rules, three inputs, and one output with a clock frequency of 66 MHz. Giuseppe Ascia, Vincenzo Catania, Marco Russo |
IEEE Trans. Fuzzy Syst. | 1 |
| 1997 | A VLSI fuzzy expert system for real-time traffic control in ATM networksabstractConcerns a fuzzy logic-based system which has been purposely designed to achieve real-time traffic control in high-speed networks using the asynchronous transfer mode (ATM) technique. One of the most critical functions is "policing", which has the task of ensuring that each user source complies with the traffic parameters negotiated in the call setup to avoid network congestion. This function is difficult to implement on account of certain conflicting requirements such as selectivity and responsiveness. This is confirmed by the severe limits affecting the most popular mechanisms proposed so far, based on conventional logic. The capacity to formalize approximate reasoning processes offered by fuzzy logic is exploited to derive rules of behavior for a policer starting from the know-how of an expert. We address two key issues related to the implementation of the fuzzy policer. The first focuses on the possibility of hardware implementation of the mechanism using VLSI technology; we present the design of a VLSI fuzzy processor which exhibits a level of performance of over 3 MFLIPS. The second issue concerns the suitability of applying the fuzzy policer to the policing of several classes of sources to reach high levels of cost effectiveness and scalability. Giuseppe Ascia, Vincenzo Catania, Giuseppe Ficili, Sergio Palazzo, Daniela Panno |
IEEE Trans. Fuzzy Syst. | 1 |
| 1996 | A Reconfigurable Parallel Architecture for a Fuzzy Processor
Giuseppe Ascia, Vincenzo Catania, Antonio Puliafito, Lorenzo Vita |
Inf. Sci. | 1 |
| 1995 | An efficient hardware architecture to support complex fuzzy reasoningabstractThe paper presents the design of a VLSI fuzzy processor which is capable of dealing with complex knowledge systems. The architecture of the processor is based on a appropriate computational model, whose main features are: capacity to cope with rule chaining; pre-processing of inferences to reduce the number of rules to be processed; parallel computation of the degree of activation of the active rules; optimized representation of membership function. The processor performance is in the order of 1.5 MFLIPS (256 rule, 8 Fuzzy inputs, 4 output). Giuseppe Ascia, Vincenzo Catania |
ICTAI | 1 |
| 1995 | Design of a VLSI fuzzy processor for ATM traffic sources managementabstractTraffic management in ATM networks aims for the controlled use of network resources to prevent the network from becoming a bottleneck. This is a very complex task due to the random nature of the traffic arrival pattern and contention for network resources. In this paper we deal with one of the most important aspects of traffic control: "the policing". It has the role of ensuring that each source complies with the traffic parameter negotiated, in order to avoid network congestion. We present a policing mechanism based on Fuzzy Logic. As far as selectivity and dynamic response are concerned, it outperforms traditional mechanism thanks to the softness which characterizes Fuzzy Logic. In the paper we also address the design of a VLSI fuzzy processor to implement the mechanism in hardware. The processor is organized as a cascade of pipeline stages in which fuzzy inferences are processed in parallel. A synthesis on CMOS 0.5 /spl mu/m technology has allowed the processor to reach a very high processing speed occupying only a few mm/sup 2/ of silicon area, thus requiring a very low implementation cost. Giuseppe Ascia, Giuseppe Ficili, Daniela Panno |
LCN | 1 |
| 1995 | A VLSI Parallel Architecture for Fuzzy Expert SystemsabstractIn this paper we present a VLSI fuzzy processor whose main features are a scalable parallel architecture, and the computation of fuzzy inferences based on the α-level set theory, both of which are important in the field of intensive fuzzy computing, as in fuzzy expert systems. A specific analysis is made in the paper, of techniques for the representation of fuzzy sets, in relation to the amount of area occupied and the forms they can assume. From this analysis a solution is extracted and then used for the processor presented in the paper. The architecture of the processor is chosen after the assessment of possible alternatives by analyzing an appropriate probabilistic model. The processor comprises a set of units which work parallelly and asynschronously to process the various rules. The structure is easy to scale up, as an increase in the number of processing units does not produce bottlenecks in performance. The performance obtainable is about 310 KFLIPS, with a clock frequency of 60 Mhz, 8 input variables, either crisp or fuzzy, and an 8-bit resolution. Vincenzo Catania, Giuseppe Ascia |
Int. J. Pattern Recognit. Artif. Intell. | 2 |