EDBT 2026 Demo / reviewers in the wild / expert
Davide Patti
dblp:97/589
· DBLP profile ↗
33ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-0874-7793ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 7 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | POSTER: Reinforcement Learning-based QoS-aware Online Scheduling for Multi-Tenant DNN Inference on Heterogeneous AcceleratorsabstractDeep Neural Networks (DNNs) are increasingly deployed through cloud services such as Inference-as-a-Service, where multiple tenants submit deadline-constrained inference requests to shared hardware infrastructures. Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti |
CF | 4 |
| 2026 | Timechain-level modeling and analysis of the bitcoin lightning network
Davide Patti, Salvatore Monteleone, Enrico Russo 0002, Maurizio Palesi |
Comput. Networks | 1 |
| 2026 | Assessing the Role of Communication in Modular Multi-Core Quantum SystemsabstractThe scalability of quantum computing is constrained by the physical and architectural limitations of monolithic quantum processors. Modular multi-core quantum architectures, which interconnect multiple quantum cores (QCs) via classical and quantum-coherent links, offer a promising alternative to address these challenges. However, transitioning to a modular architecture introduces communication overhead, where classical communication plays a crucial role in executing quantum algorithms by transmitting measurement outcomes and synchronizing operations across QCs. Understanding the impact of classical communication on execution time is therefore essential for optimizing system performance. In this work, we introduce qcomm , an open-source simulator designed to evaluate the role of classical communication in modular quantum computing architectures. qcomm provides a high-level execution and timing model that captures the interplay between quantum gate execution, entanglement distribution, teleportation protocols, and classical communication latency. We conduct an extensive experimental analysis to quantify the impact of classical communication bandwidth, interconnect types, and quantum circuit mapping strategies on overall execution time. Furthermore, we assess classical communication overhead when executing real quantum benchmarks mapped onto a cryogenically-controlled multi-core quantum system. Our results show that, while classical communication is generally not the dominant contributor to execution time, its impact becomes increasingly relevant in optimized scenarios—such as improved quantum technology, large-scale interconnects, or communication-aware circuit mappings. These findings provide useful insights for the design of scalable modular quantum architectures and highlight the importance of evaluating classical communication as a performance-limiting factor in future systems. Maurizio Palesi, Enrico Russo 0002, Giuseppe Ascia, Hamaad Rafique, Davide Patti, Vincenzo Catania, Sergi Abadal, Abhijit Das 0002, Pau Escofet, Eduard Alarcón, Carmen G. Almudéver |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | Optimizing Qubit Assignment in Modular Quantum Systems via Attention-Based Deep Reinforcement LearningabstractModular, distributed, and multi-core architectures are considered a promising solution for scaling quantum computing systems. Optimising communication is crucial to preserve quantum coherence. The compilation and mapping of quantum circuits should minimise state transfers while adhering to architec-tural constraints. To address this problem efficiently, we propose a novel approach using Reinforcement Learning (RL) to learn heuristics for a specific multi-core architecture. Our RL agent uses a Transformer encoder and Graph Neural Networks, encoding quantum circuits with self-attention and producing outputs via an attention-based pointer mechanism to match logical qubits with physical cores efficiently. Experimental results show our method outperform the baseline reducing by 28% inter-core communications for random circuits while minimising time-to-solution. Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania |
DATE | 3 |
| 2025 | A Characteristics-Based Least Common Multiple Algorithm to Optimize Magnetic-Field-Based Indoor LocalizationabstractClustering is an unsupervised learning technique that groups data based on similarity criteria. Traditional methods like K-Means and agglomerative clustering often require predefined parameters, struggle with irregular cluster shapes, and fail to classify subcluster points in magnetic fingerprint-based indoor localization. This study proposes the characteristics-based least common multiple (LCM) algorithm to address these challenges. This novel approach autonomously determines cluster number and shape while accurately classifying misclassified points based on characteristic similarities using LCM. We evaluated the proposed technique using state-of-the-art metrics and tested it in magnetic-field-based indoor localization scenarios. Comparisons were made with real-time and benchmark datasets, alongside traditional clustering methods. Results demonstrate that LCM significantly enhances localization accuracy, achieving a mean absolute error rate of 0.1 m. Hamaad Rafique, Davide Patti, Maurizio Palesi, Gaetano Carmelo La Delfa |
IEEE Internet Things J. | 2 |
| 2024 | A Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator SystemsabstractCurrently, there is a growing trend of outsourcing the execution of DNNs to cloud services. For service providers, managing multitenancy and ensuring high-quality service delivery, particularly in meeting stringent execution time constraints, assumes paramount importance, all while endeavoring to maintain cost-effectiveness. In this context, the utilization of heterogeneous multi-accelerator systems becomes increasingly relevant. This paper presents RELMAS, a low-overhead deep reinforcement learning algorithm designed for the online scheduling of DNNs in multi-tenant environments, taking into account the dataflow heterogeneity of accelerators and memory bandwidths contentions. By doing so, service providers can employ the most efficient scheduling policy for user requests, optimizing Service-Level-Agreement (SLA) satisfaction rates and enhancing hardware utilization. The application of RELMAS to a heterogeneous multi-accelerator system composed of various instances of Simba and Eyeriss sub-accelerators resulted in up to a 173% improvement in SLA satisfaction rate compared to state-of-the-art scheduling techniques across different workload scenarios, with less than a 1.5% energy overhead. Francesco Giulio Blanco, Enrico Russo 0002, Maurizio Palesi, Davide Patti, Giuseppe Ascia, Vincenzo Catania |
DAC | 4 |
| 2024 | Abstracting Bitcoin Lightning Network Complexity with UltravioletabstractWith this work, we introduce the concept of Timechain-level model, along with an open-source implementation (Ultraviolet), to abstract the complexity of the Lightning Network (LN) while still providing a vision of base layer events and protocol internals. After depicting how each element of the model is mapped into the LN architectural stack, we show a case study to demonstrate its usage in investigating large-scale scenarios for research, development and educational purposes. Finally, we present a comparison to properly contextualize our contribution to the current state of the art of LN modelling, highlighting the advancements introduced by a Timechain-level and future directions of research it opens. Davide Patti, Salvatore Monteleone, Enrico Russo 0002, Maurizio Palesi, Vincenzo Catania |
ICBC | 1 |
| 2024 | Characteristics-Based Least Common Multiple: A Novel Clustering Algorithm to Optimize Indoor Positioning
Hamaad Rafique, Davide Patti, Maurizio Palesi, Gaetano Carmelo La Delfa |
ICINCO (1) | 2 |
| 2024 | Fusing Visuals with Magnetic Signals to Improve Indoor Localization Using Vision TransformerabstractSensor fusion-based indoor localization is an evolving application that uses fused information to determine the location of smartphone users. However, the heterogeneity of sensors across various smartphones has significantly compromised the accuracy of localization algorithms. Therefore, this paper introduces MH-ViL, an infrastructure-free and calibration-free frame-work built on top of the Vision Transformer neural network. MH-ViL seamlessly integrates magnetic field signals (MFS) and visual images for localization tasks. A novel magnetic feature projection (MFP) model is proposed to effectively map MFS onto visual image features, enhancing positional accuracy within the self-attention mechanism. Real-time experiments demonstrate that MH-ViL surpasses alternative models, with an impressive 92% accuracy. We provide a 95% confidence interval with an error below 0.5 meters. Code: https://github.com/Hamaad1/MH-ViL.git Hamaad Rafique, Davide Patti, Maurizio Palesi, Gaetano Carmelo La Delfa |
IPIN | 2 |
| 2024 | Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement LearningabstractThis paper addresses the critical challenge of managing Quality of Service (QoS) in cloud services, focusing on the nuances of individual tenant expectations and varying Service Level Indicators (SLIs). It introduces a novel approach utilizing Deep Reinforcement Learning for tenant-specific QoS management in multi-tenant, multi-accelerator cloud environments. The chosen SLI, deadline hit rate, allows clients to tailor QoS for each service request. A novel online scheduling algorithm for Deep Neural Networks in multi-accelerator systems is proposed, with a focus on guaranteeing tenant-wise, model-specific QoS levels while considering real-time constraints. Enrico Russo 0002, Francesco Giulio Blanco, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Vincenzo Catania |
ISCAS | 5 |
| 2023 | Memory-Aware DNN Algorithm-Hardware Mapping via Integer Linear ProgrammingabstractMapping a deep neural network (DNN) layer onto domain-specific accelerators can require an intractable number of choices regarding loop factorization, ordering, and spatial unrolling. Determining the optimal mapping that achieves the best figures in terms of latency and energy efficiency can be difficult due to the vast number of possible candidates that need to be exhaustively evaluated. Many techniques have been recently proposed for fast and efficient mapping space exploration; some of them adopt a black-box optimization approach, others make assumptions on the underlying accelerator memory hierarchy or require time-consuming model retraining. We propose an integer linear programming (ILP) approach and formulate a mathematical model, namely LEMON, that takes into account number of accesses to each buffer, energy costs and buffer bandwidths in the accelerator and is flexible enough to work with different memory hierarchies. Compared with state-of-the-art techniques, LEMON achieves up to 83% energy-delay product reduction when compared to another ILP-based approach (CoSA) and 27% when compared to a genetic algorithm approach (GAMMA). Enrico Russo 0002, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Salvatore Monteleone, Vincenzo Catania |
CF | 4 |
| 2023 | Multiobjective End-to-End Design Space Exploration of Parameterized DNN AcceleratorsabstractDeep neural network (DNN) hardware accelerators enable the execution of complex DNN inferences on resource-constrained IoT devices. Inference performance and energy figures depend on how the DNN layers are mapped into the accelerator and how the architecture of the accelerator fits the variety of layers’ shapes of the actual DNN. The mapping determines the execution order of the operations, both temporally and spatially. Thus, selecting the best mapping that allows fitting the DNN model to the specific accelerator is of paramount importance to meet the strong constraints imposed by resource-scarce IoT platforms. Although several mapping space exploration techniques have been proposed in the literature, they are focused on determining the best mapping for a given layer, for a given architecture, and for optimizing a single objective. This article largely extends the scope of the exploration by considering the huge design space spanned by mapping related and architectural parameters, considering all the layers of the DNN, and optimizing multiple objectives simultaneously. We present EPOCA, end-to-end Pareto optimization of DNN accelerators, whose goal is to determine the accelerator’s architecture and the mapping for each layer that optimizes end-to-end and in a multiobjective fashion a set of conflicting design criteria. We assess EPOCA on different DNN models on a parameterized hardware accelerator designed for IoT applications and compare them with a state-of-the-art mapping space explorer, considering the area, inference latency, and inference energy as optimization metrics. We show that the set of Pareto solutions found by EPOCA provides the designer with a range of choices from which to select the best tradeoff with respect to the specific application. Enrico Russo 0002, Maurizio Palesi, Davide Patti, Salvatore Monteleone, Giuseppe Ascia, Vincenzo Catania |
IEEE Internet Things J. | 3 |
| 2022 | MEDEA: A Multi-objective Evolutionary Approach to DNN Hardware MappingabstractDeep Neural Networks (DNNs) embedded domain-specific accelerators enable inference on resource-constrained devices. Making optimal design choices and efficiently scheduling neural network algorithms on these specialized architectures is challenging. Many choices can be made to schedule computation spatially and temporally on the accelerator. Each choice influences the access pattern to the buffers of the architectural hierarchy, affecting the energy and latency of the inference. Each mapping also requires specific buffer capacities and a number of spatial components instances that translate in different chip area occupation. The space of possible combinations, the mapping space, is so large that automatic tools are needed for its rapid ex-ploration and simulation. This work presents MEDEA, an open-source multi-objective evolutionary algorithm based approach to DNNs accelerator mapping space exploration. MEDEA leverages the Timeloop analytical cost model. Differently from the other schedulers that optimize towards a single objective, MEDEA allows deriving the Pareto set of mappings to optimize towards multiple, sometimes conflicting, objectives simultaneously. We found that solutions found by MEDEA dominates in most cases those found by state-of-the-art mappers. Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Vincenzo Catania |
DATE | 4 |
| 2022 | DNN Model Compression for IoT Domain-Specific Hardware AcceleratorsabstractMachine learning techniques, particularly those based on neural networks, are always more often used at the edge of the network by Internet of Things (IoT) nodes. Unfortunately, the computation capabilities demanded by those applications, together with their energy efficiency-related constraints, exceed those exposed by embedded general-purpose processors. For this reason, the use of domain-specific hardware accelerators (DSAs) is considered the most viable solution to the unsustainable “Turing tariff” of general-purpose hardware. Starting from the observation that memory and communication traffic account for a large fraction of the overall latency and energy in deep neural network (DNN) inferences, this article proposes a new compression technique aimed at: 1) reducing the memory footprint for storing the model parameters of a DNN and 2) improving DNN inference latency and energy on resource-constrained IoT devices. The proposed compression technique, namely, LineCompress, is applied on a set of representative convolutional neural networks (CNNs) for object recognition mapped on a state-of-the-art DSA targeted for resource-constrained IoT devices. We show that on average,$7.4\times $memory footprint reduction can be obtained, thus reducing the memory and communication traffic that result to 77% and 87% inference latency and energy reduction, respectively, trading-off efficiency versus accuracy. Enrico Russo 0002, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Andrea Mineo, Giuseppe Ascia, Vincenzo Catania |
IEEE Internet Things J. | 4 |
| 2020 | Implementing On-Chip Wireless Communication in Multi-stage Interconnection NoCs
Sirine Mnejja, Yassine Aydi, Mohamed Abid, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
AINA | 6 |
| 2020 | DNNZip: Selective Layers Compression Technique in Deep Neural Network AcceleratorsabstractIn Deep Neural Network (DNN) accelerators, the on-chip traffic and memory traffic accounts for a relevant fraction of the inference latency and energy consumption. A major component of such traffic is due to the moving of the DNN model parameters from the main memory to the memory interface and from the latter to the processing elements (PEs) of the accelerator. In this paper, we present DNNZip, a technique aimed at compressing the model parameters of a DNN, thus resulting in significant energy and performance improvement. DNNZip implements a lossy compression whose compression ratio is tuned based on the maximum tolerated error on the model parameters provided by the user. DNNZip is assessed on several convolutional NNs and the trade-off inference energy saving vs. inference latency reduction vs. network accuracy degradation is discussed. We found that up to 64% energy saving, and up to 67% latency reduction can be obtained with a limited impact on the accuracy of the network. Habiba Lahdhiri, Maurizio Palesi, Salvatore Monteleone, Davide Patti, Giuseppe Ascia, Jordane Lorandel, Emmanuelle Bourdel, Vincenzo Catania |
DSD | 4 |
| 2020 | Improving Inference Latency and Energy of DNNs through Wireless Enabled Multi-Chip-Module-based Architectures and Model Parameters CompressionabstractPerformance and energy figures of Deep Neural Network (DNN) accelerators are profoundly affected by the communication and memory sub-system. In this paper, we make the case of a state-of-the-art multi-chip-module-based architecture for DNN inference acceleration. We propose a hybrid wired/wireless network-in-package interconnection fabric and a compression technique for drastically improving the communication efficiency and reducing the memory and communication traffic with a consequent improvement of performance and energy metrics. We assess the inference performance and energy improvement vs. accuracy degradation for different CNNs showing that up to 77% and 68% of inference latency reduction and inference energy reduction, respectively, can be obtained while keeping the accuracy degradation below 5% as respect to the original uncompressed CNN. Giuseppe Ascia, Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
NOCS | 6 |
| 2020 | Exploiting Data Resilience in Wireless Network-on-chip ArchitecturesabstractThe emerging wireless Network-on-Chip (WiNoC) architectures are a viable solution for addressing the scalability limitations of manycore architectures in which multi-hop long-range communications strongly impact both the performance and energy figures of the system. The energy consumption of wired links as well as that of radio communications account for a relevant fraction of the overall energy budget. In this article, we extend the approximate computing paradigm to the case of the on-chip communication system in manycore architectures. We present techniques, circuitries, and programming interfaces aimed at reducing the energy consumption of a WiNoC by exploiting the trade-off energy saving vs. application output degradation. The proposed platform—namely, xWiNoC—uses variable voltage swing links and tunable transmitting power wireless interfaces along with a programming interface that allows the programmer to specify those data structures that are error-resilient. Thus, communications induced by the access to such error-resilient data structures are carried out by using links and radio channels that are configured to work in a low energy mode, albeit by exposing a higher bit error rate. xWiNoC is assessed on a set of applications belonging to different domains in which the trade-off energy vs. performance vs. application result quality is discussed. We found that up to 50% of communication energy saving can be obtained with a negligible impact on the application output quality and 3% in application performance degradation. Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose, Valerio Mario Salerno |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2019 | Analyzing networks-on-chip based deep neural networksabstractOne of the most promising architectures for performing deep neural network inferences on resource-constrained embedded devices is based on massive parallel and specialized cores interconnected by means of a Network-on-Chip (NoC). In this paper, we extensively evaluate NoC-based deep neural network accelerators by exploring the design space spanned by several architectural parameters. We show how latency is mainly dominated by the on-chip communication whereas energy consumption is mainly accounted by memory (both on-chip and off-chip). Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose |
NOCS | 5 |
| 2018 | Improving Energy Efficiency in Wireless Network-on-Chip ArchitecturesabstractWireless Network-on-Chip (WiNoC) represents a promising emerging communication technology for addressing the scalability limitations of future manycore architectures. In a WiNoC, high-latency and power-hungry long-range multi-hop communications can be realized by performance- and energy-efficient single-hop wireless communications. However, the energy contribution of such wireless communication accounts for a significant fraction of the overall communication energy budget. This article presents a novel energy managing technique for WiNoC architectures aimed at improving the energy efficiency of the main elements of the wireless infrastructure, namely, radio-hubs. The rationale behind the proposed technique is based on selectively turning off, for the appropriate number of cycles, all the radio-hubs that are not involved in the current wireless communication. The proposed energy managing technique is assessed on several network configurations under different traffic scenarios both synthetic and extracted from the execution of real applications. The obtained results show that the application of the proposed technique allows up to 25% total communication energy saving without any impact on performance and with a negligible impact on the silicon area of the radio-hub. Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2016 | Improving the energy efficiency of wireless Network on Chip architectures through online selective buffers and receivers shutdownabstractThe wireless Network-on-Chip (WiNoC) design paradigm represents an emergent and viable solution for addressing the scalability limitations of future manycores architectures. Unfortunately, components such as the buffers and the transceiver of the radio-hubs in a WiNoC, account for a significant fraction of the total communication energy budget. In this paper, we present WIRXSleep, a mechanism aimed at improving the energy efficiency of radio-hubs in WiNoC architectures. WIRXSleep selectively and dynamically disables receiver modules and buffers of those radio-hubs that will be not involved in any communication during the next forthcoming clock cycles. Its application on different WiNoC topologies, with different configurations, and under different traffic scenarios has resulted interesting energy savings (up to 25%) without any impact on performance and with a negligible impact on cost metrics. Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
CCNC | 5 |
| 2016 | Energy efficient transceiver in wireless Network on Chip architectures
Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
DATE | 5 |
| 2015 | Noxim: An open, extensible and cycle-accurate network on chip simulatorabstractEmerging on-chip communication technologies like wireless Networks-on-Chip (WiNoCs) have been proposed as candidate solutions for addressing the scalability limitations of conventional multi-hop NoC architectures. In a WiNoC, a subset of network nodes are equipped with a wireless interface which allows them long-range communication in a single hop. This paper presents Noxim, an open, configurable, extendible, cycle-accurate NoC simulator developed in SystemC which allows to analyze the performance and power figures of both conventional wired NoC and emerging WiNoC architectures. Vincenzo Catania, Andrea Mineo, Salvatore Monteleone, Maurizio Palesi, Davide Patti |
ASAP | 5 |
| 2015 | User-Generated services: Policy Management and access control in a cross-domain environmentabstractThe rapid evolution of mobile computing, together with the spread of social networks is increasingly moving the role of users from simple information and services consumers to actual producers. Currently, while most of the critical aspects related to User-Generated Contents (UGC) have been addressed, many issues related to service generation still must be faced and represent the next challenge. In this work, we focus on security issues raised by a particular kind of services: those generated by users. User-Generated Services (UGS) are characterized by a set of features that distinguish them from conventional services. To cope with UGS security problems we introduce three possible policy management models, analyzing benefits and drawbacks of each approach. Finally, we propose a cloud-based solution that enables the composition of multiple UGS and policy models, allowing user's devices to share features and services among them. Vincenzo Catania, Giuseppe La Torre, Salvatore Monteleone, Daniela Panno, Davide Patti |
IWCMC | 5 |
| 2015 | Parameter Space Representation of Pareto Front to Explore Hardware-Software DependenciesabstractEmbedded systems design requires conflicting objectives to be optimized with an appropriate choice of hardware-software parameters. A simulation campaign can guide the design in finding the best trade-offs, but due to the big number of possible configurations, it is often unfeasible to simulate them all. For these reasons, design space exploration algorithms aim at finding near-optimal system configurations by simulating only a subset of them. In this work, we present PS, a new multiobjective optimization algorithm, and evaluate it in the context of the embedded system design. The basic idea is to recognize interesting regions—that is, regions of the configuration space that provide better configurations with respect to other ones. PS evaluates more configurations in the interesting regions while less thoroughly exploring the rest of the configuration space. After a detailed formal description of the algorithm and the underlying concepts, we show a case study involving the hardware/software exploration of a VLIW architecture. Qualitative and quantitative comparisons of PS against a well-known multiobjective genetic approach demonstrate that while not outperforming it in terms of Pareto dominance, the proposed approach can balance the uniformity and granularity qualities of the solutions found, obtaining more extended Pareto fronts that provide a wider view of the potentiality of the designed device. Therefore, PS represents a further valid choice for the designer when objective constrains allow it. Vincenzo Catania, Andrea Araldo, Davide Patti |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2009 | An Effective Methodology to Multi-objective Design of Application Domain-specific Embedded ArchitecturesabstractToday's computer systems have become unbelievably complex. Nowadays register-level design is an overwhelming task, especially in the embedded system area where the time-to-market is very short. Platform based design shifts the challenge on how to tune parametric platforms to achieve the best performance at the smallest cost. This task, called multi-objective design space exploration, requires accurate strategies because the design space is too vast to be exhaustively evaluated. Even using efficient exploration strategies proposed in the literature, simulation times can become a bottleneck in the design flow. In this work we propose a novel approach to application-domain design space exploration using a multi-objective genetic algorithm and employing HPC to reduce exploration times. The genetic algorithm is preceded by a correlation analysis of the different objectives. The search space is thus reduced by combining highly correlated objectives from different domains. We describe the steps needed to parallelize the exploration on the grid, and present the results of extensive testing of the proposed approach. We obtained over one order of magnitude reduction in exploration times without hampering the quality of the solutions. Shorter simulation times allow more ideas to be explored in less time. This leads to shorter product time-to-market and a more thorough design space exploration. Furthermore the combination of correlated objectives favors the design of modern multi-purpose devices. Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti, Gianmarco De Francisci Morales |
DSD | 4 |
| 2008 | High Performance Computing for Embedded System Design: A Case StudyabstractIn this paper we assess the use of high performance computing in design space exploration of a complex highly parameterized very long instruction word based system-on-a-chip platform. Experiments show that the conventional belief of linear decrease in exploration time as the number of available processors increases is discredited starting from a relatively low number of processors mainly due to communication overhead and I/O bottleneck. Vincenzo Catania, Gianmarco De Francisci Morales, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti |
DSD | 5 |
| 2008 | Reducing complexity of multiobjective design space exploration in VLIW-based embedded systemsabstractArchitectures based on very-long instruction word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of instruction level parallelism (ILP) with a reasonable trade-off in complexity and silicon cost. Specialization of such architectures involves the configuration of both hardware-related aspects (e.g., register files, functional units, memory subsystem) and software-related issues (e.g., the compilation strategy). The complex interactions between the components of such systems will force a human designer to rely on judgment and experience in designing them, possibly eliminating interesting configurations, and making tuning of the system, for either power, energy, or performance, difficult. In this paper we propose tools and methodologies to efficiently cope with this complexity from a multiobjective perspective. We first analyze the impact of ILP-oriented code transformations using two alternative compilation profiles to quantitatively show the effect of such transformations on typical design objectives like performance, power dissipation, and energy consumption. Next, by means of statistical analysis, we collect useful data to predict the effectiveness of a given compilation profiles for a specific application. Information gathered from such analysis can be exploited to drastically reduce the computational effort needed to perform the design space exploration. Vincenzo Catania, Maurizio Palesi, Davide Patti |
ACM Trans. Archit. Code Optim. | 3 |
| 2008 | Implementation and Analysis of a New Selection Strategy for Adaptive Routing in Networks-on-ChipabstractEfficient and deadlock-free routing is critical to the performance of networks-on-chip. The effectiveness of any adaptive routing algorithm strongly depends on the underlying selection strategy. A selection function is used to select the output channel where the packet will be forwarded on. In this paper we present a novel selection strategy that can be coupled with any adaptive routing algorithm. The proposed selection strategy is based on the concept of Neighbors-on-Path the aims of which is to exploit the situations of indecision occurring when the routing function returns several admissible output channels. The overall objective is to choose the channel that will allow the packet to be routed to its destination along a path that is as free as possible of congested nodes. Performance evaluation is carried out by using a flit-accurate simulator under traffic scenarios generated by both synthetic and real applications. Results obtained show how the proposed selection strategy applied to the Odd-Even routing algorithm yields an improvement in both average delay and saturation point up to 20% and 30% on average respectively, with a minimal overhead in terms of area occupation. In addition, a positive effect on total energy consumption is also observed under near-congestion packet injection rates. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti |
IEEE Trans. Computers | 4 |
| 2007 | Efficient design space exploration for application specific systems-on-a-chip
Giuseppe Ascia, Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti |
J. Syst. Archit. | 5 |
| 2006 | A Multiobjective Genetic Fuzzy Approach for Intelligent System-level Exploration in Parameterized VLIW Processor DesignabstractThe design of a complex embedded system is dominated by the definition of an optimal architecture in relation to certain performance indexes. This activity, known as Design Space Exploration (DSE), is a great challenge for the EDA (Electronic Design Automation) community. The enormous size of the design space, in fact, together with the long simulation time required to evaluate each system configuration during the exploration process, cause DSE to become a bottleneck in the design flow. In this paper we propose a Multiobjective Design Space Exploration methodology based on a Genetic Fuzzy System, the aim of which is to drastically reduce the exploration time while guaranteeing a high level of accuracy. The methodology uses a Genetic Algorithm (GA) for heuristic exploration and a Fuzzy System to evaluate the configurations visited. Although of general application, the methodology is applied to a real case study: optimization of the performance and power consumption of an embedded architecture based on a Very Long Instruction Word (VLIW) microprocessor in a mobile multimedia application domain. The results obtained are compared, in terms of both accuracy and efficiency, with the state of the art in multiobjective DSE strategies, represented by the classical GA approach, demonstrating the scalability and effectiveness of the proposed approach. Giuseppe Ascia, Vincenzo Catania, Alessandro G. Di Nuovo, Maurizio Palesi, Davide Patti |
IEEE Congress on Evolutionary Computation | 5 |
| 2005 | Exploring Design Space of VLIW ArchitecturesabstractArchitectures based on very long instruction word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of instruction level parallelism (ILP) with a reasonable tradeoff in complexity and silicon costs. Effective compiler support for predicated execution using the hyperblock, drastically increases the ILP even for control-dominated applications in which the branch instruction frequency is very high. The use of these techniques, however, is known to increase the instruction footprint, consequently putting pressure on the memory hierarchy. In this paper, we evaluate the performance/power trade-off in a system comprising a VLIW processor and a two-level hierarchical memory subsystem. Via simulation, we show that the efficiency of a compiler that is able to exploit predicate execution by hyperblock formation is greatly affected by the configuration of the memory subsystem as well as the configurable processor parameters. The enabling or disabling of hyperblock formation should therefore not be evaluated separately or independently, but seen as a further free parameter to be tuned in a strategy of design space exploration. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti |
ASAP | 4 |
| 2005 | A system-level framework for evaluating area/performance/power trade-offs of VLIW-based embedded systemsabstractArchitectures based on Very Long Instruction Word (VLIW) have found fertile ground in multimedia electronic appliances thanks to their ability to exploit high degrees of Instruction Level Parallelism (ILP) with a reasonable trade-off in complexity and silicon costs. In this case Application Specific Instruction-set Processor (ASIP) specialization may require not only manipulation of the instruction-set but also tuning of the architectural parameters of the processor (e.g. the number and type of functional units, register files, etc.) and the memory subsystem (cache size, associativity, etc.). Setting the parameters so as to optimize certain metrics requires the use of efficient Design Space Exploration (DSE) strategies and also simulation tools (retargetable compilers and simulators) and accurate estimation models operating at a high level of abstraction. In this paper we present a framework for evaluation, in terms of performance, cost and power consumption, of a system based on a parameterized VLIW microprocessor together with the memory hierarchy subsystem following execution of a specific application. The framework, which can be freely downloaded from the Internet, implements a number of multi-objective DSE strategies to obtain Pareto-optimal configurations for the system. Giuseppe Ascia, Vincenzo Catania, Maurizio Palesi, Davide Patti |
ASP-DAC | 4 |