EDBT 2026 Demo / reviewers in the wild / expert
Paolo Burgio
dblp:00/8746
· DBLP profile ↗
34ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0003-1954-7201ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 6 first-author · 12 since 2021Software engineering, systems software and programming languages · 9 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring and Understanding Visualization Latency Performance for Smart City ApplicationsabstractEnd-to-end latency is a critical metric in interactive and real-time systems, where even short delays can undermine usability, situational awareness, and trust. While most studies focus on network transmission delays, this overlooks other significant sources of latency, including sensor acquisition, processing, and especially visualization. Rendering pipelines are heavily influenced by hardware, visualization strategies, and data complexity, and in visualization-rich domains such as smart cities, these factors can add considerable overhead. Ignoring them leads to overly optimistic performance assessments and missed opportunities for optimization.In this work, we study the last mile of end-to-end latency, explicitly incorporating the visualization stage and bridging the gap between network-level metrics and user-perceived responsiveness. We introduce the Unreal Smart Cities Visualizer (USCV), a framework we have built to support precise latency assessment in realistic, data-intensive urban scenarios. Our contributions include highlighting the limits of network-centric evaluation, presenting a visualization-aware methodology, and demonstrating USCV as a tool for accurate end-to-end latency measurement.Our results show that visualization delays, particularly in latency-critical environments, are not negligible and cannot be underestimated. Alessio Masola, Paolo Burgio, Carlo Augusto Grazia, Luca Bedogni |
CCNC | 2 |
| 2026 | Multi-Partner Project: Scheduling-Deployment Workflow for Autonomous RoboRacer Driving Stacks in the HAL4SDV ProjectabstractThe European-funded HAL4SDV project aims to advance European solutions in software-defined vehicles by introducing a hardware abstraction layer positioned between executed software and execution units. HAL4SDV includes over 60 partners across 12 countries and receives funding within the Chips Joint Undertaking under Horizon Europe since April 2024 and is coordinated by TTTech Computertechnik. The proposed hardware abstraction layer includes safety-critical scheduling and platform deployment of software tasks, and is motivated by the requirement for abstracted hardware with unified interfaces in centralized automotive architectures.This work presents a correct-by-construction workflow which is developed by academic partners to schedule and deploy periodic software tasks onto diverse execution units. The workflow facilitates the execution of the same task stack on multiple unit architectures and consists of a task model and scheduling algorithm, which is followed by platform deployment for diverse hardware units, ensuring safe execution. In this multi-partner project, a bandwidth regulation unit for hardware accelerators and a RISC-V-based multicore system with tightly coupled memories are used as target platforms.A RoboRacer driving stack is chosen for evaluation, showing the viability of our workflow to schedule autonomous driving functions. To show generalization capability, synthetic task sets are additionally used to validate our deployment workflow. Matthias Stammler, Henrik Scheidt, Tanja Harbaum, Jürgen Becker 0001, Konstantin Dudzik, Victor Pazmino Betancourt, Federico Gavioli, Paolo Burgio, Arvind Easwaran, Andreas Eckel |
DATE | 8 |
| 2026 | Leader Election Algorithms for Vehicle Platooning: A Survey Focusing on Resilience to Member Failures
Chinmay Satish Shrivastav, Francesco Castorini, Alessio Masola, Paolo Burgio, Roberto Cavicchioli |
MDM | 4 |
| 2025 | Enabling the Proxy Computing Paradigm on DPU-based FPGA AccelerationabstractFPGA-based heterogeneous systems are a popular choice for accelerating Deep Neural Networks (DNNs), but efficiently integrating and orchestrating HW and SW tasks remains challenging.FPGA overlay architectures have been proposed to simplify accelerator management, yet state-of-the-art solutions struggle with performance bottlenecks caused by frequent CPU-FPGA interactions.We introduce a novel overlay-based methodology enabling the Proxy Computing paradigm, leveraging a local orchestrator and shared memory to (i) reduce accelerator control overhead and (ii) minimize unnecessary data movements.As a case study, we integrate the AMD/Xilinx Deep Learning Processing Unit (DPU) with additional accelerators for unsupported layers.Experiments show that our approach significantly reduces memory transfers, achieving up to 4× speed up in the proposed case study. Gianluca Brilli, Alessandro Capotondi, Paolo Burgio, Andrea Marongiu |
CF | 3 |
| 2025 | ShapeFuture - Technical Progress After Year 1abstractShapeFuture will drive innovation in fundamental Electronic Components and Systems (ECS) that are essential for robust, powerful, fail-operational and integrated perception, cognition, AI-enabled decision making, resilient automation and computing, as well as communications, for highly automated vehicles. The overarching vision of ShapeFuture is to bring ECS Innovation to the heart of Europe’s Mobility Transformation, thereby elevating sovereignty by perfecting programmable ECS solutions for intelligent, safe, connected, and highly automated vehicles. In this paper, we detail not only the vision and mission of the ShapeFuture project, but we also showcase the results achieved during the first year. Norbert Druml, Martin Gschwandtner, Mayeul Jeannin, Rainer Matischek, Edgars Lielamurs, Maksis Celitans, Kaspars Ozols, Nurullah Demiralay, Besir Tayfur, Ismail Sinan Gulbas, Nadir Kucuk, Isa Kiyat, Yahya Nasolo, Jens U. Brandt, Noah Christoph Pütz, Thomas Bartz-Beielstein, Jose Isola, Nikola Mandic, Francesca Flamigni, Alexander Kuehhas, Gianluca Brilli, Paolo Burgio, Giacomo Paolieri, Jorge Villagra, Álvaro Flores Cueto, José Antonio Sánchez, Jacopo Sini, Massimo Violante, Lorenzo Giraudi, Paolo Santero, Uwe Kölbel, Moritz Schaffenroth, Panu Sjövall, Jarno Vanne, Morten Larsen, Nergis Gizem Yilmaz, Ziya Uygar Yengin, George Dimitrakopoulos 0001 |
DSD | 22 |
| 2025 | Scalable Object Geolocation in Traffic Camera Imagery Using 3D World ModelabstractThe rapid expansion of outdoor traffic camera systems requires efficient methods to accurately estimate the geolocation of objects within their scenes. We present an innovative and scalable framework that completely automates this process by combining easy-to-build 3D world modeling with real-world traffic camera imagery. First, using the Cesium plugin for Unreal Engine, we create detailed and scalable 3D representations of urban environments, leveraging publicly available, highly accurate 3D data. This results in the creation of globally curated 3D content, including terrain, imagery, and photogrammetry. The real-world traffic camera imagery is then matched within our model using state-of-the-art feature matching techniques. By estimating the homography between synthetic images from the 3D model and the real images from traffic cameras, we accurately determine the geolocation of observed objects within the scene. This approach not only enhances geolocation accuracy but also enables seamless scalability across diverse urban settings and camera deployments worldwide. Our method significantly reduces the manual effort required for traffic camera calibration, thus streamlining the deployment of intelligent transportation systems at scale. We demonstrate the high performance of our approach by experimenting in an urban trial site with multiple smart city cameras and publicly available cameras around the world. Additionally, we highlight the adaptability of our framework for a wide range of computer vision-based traffic analytics applications, including its potential for drone-based localization. Chinmay Satish Shrivastav, Alessio Masola, Roberto Cavicchioli, Nicola Capodieci, Paolo Burgio |
SMC | 5 |
| 2025 | Fine-Grained QoS Control via Tightly-Coupled Bandwidth Monitoring and Regulation for FPGA-Based Heterogeneous SoCsabstractCommercial embedded systems increasingly rely on heterogeneous architectures that integrate general-purpose, multi-core processors, and various hardware accelerators on the same chip. This provides the high performance required by modern applications at a low cost and low power consumption, but at the same time poses new challenges. Hardware resource sharing at various levels, and in particular at the main memory controller level, results in slower execution time for the application tasks, ultimately making the system unpredictable from the point of view of timing. To enable the adoption of heterogeneous systems-on-chip (System on Chips (SoCs)) in the domain of timing-critical applications several hardware and software approaches have been proposed, bandwidth regulation based on monitoring and throttling being one of the most widely adopted. Existing solutions, however, are either too coarse-grained, limiting the control over computing engines activities, or strongly platform-dependent, addressing the problem only for specific SoCs. This article proposes an innovative approach that can accurately control main memory bandwidth usage in FPGA-based heterogeneous SoCs. In particular, it controls system bandwidth by connecting a runtime bandwidth regulation component to FPGA-based accelerators. Our solution offers dynamically configurable, fine-grained bandwidth regulation – to adapt to the varying requirements of the application over time – at a very low overhead. Furthermore, it is entirely platform-independent, capable of integration with any FPGA-based accelerator. Developed at the register-transfer level using a reference SoC platform, it is designed for easy compatibility with any FPGA-based SoC. Experimental results conducted on the Xilinx Zynq UltraScale+ platform demonstrate that our approach (i) is more than$100\times$faster than loosely-coupled, software controlled regulators; (ii) is capable of exploiting the system bandwidth 28.7% more efficiently than tightly-coupled hardware regulators (e.g., ARM CoreLink QoS-400, where available); (iii) enables task co-scheduling solutions not feasible with state-of-the-art bandwidth regulation methods. Giacomo Valente, Gianluca Brilli, Tania Di Mascio, Alessandro Capotondi, Paolo Burgio, Paolo Valente, Andrea Marongiu |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | The Degree of Entanglement: Cyber-Physical Awareness in Digital Twin ApplicationsabstractA defining feature of a Digital Twin (DT) is its level of ”entanglement”: the degree of strength to which the twin is interconnected with its physical counterpart. Despite its importance, this characteristic has not been yet fully investigated, and its impact on applications' design is underestimated. In this paper, we define the concept of “Degree of Entanglement” (DoE), which provides an operational model for assessing the strength of the entanglement between a DT and its physical counterpart. We also propose an interoperable representation of DoE within the Web of Things (WoT) framework, which enables DT-driven applications to dynamically adapt to changes in the physical environment. We evaluate our proposal using two realistic use cases, demonstrating the practical utility of DoE in supporting, for instance, context-awareness decisions and adaptiveness. Marco Picone 0001, Stefano Mariani 0001, Roberto Cavicchioli, Paolo Burgio, Arslane Hamza Cherif |
CCNC | 4 |
| 2024 | Adaptive Localization for Autonomous Racing Vehicles with Resource-Constrained Embedded PlatformsabstractModern autonomous vehicles have to cope with the consolidation of multiple critical software modules processing huge amounts of real-time data on power- and resource-constrained embedded MPSoCs. In such a highly-congested and dynamic scenario, it is extremely complex to ensure that all components meet their quality-of-service requirements (e.g., sensor frequencies, accuracy, responsiveness, reliability) under all possible working conditions and within tight power budgets. One promising solution consists of taking advantage of complementary resource usage patterns of software components by implementing dynamic resource provisioning. A key enabler of this paradigm consists of augmenting applications with dynamic reconfiguration capability, thus adaptively modulating quality-of-service based on resource availability or proactively demanding resources based just on the complexity of the input at hand. The goal of this paper is to explore the feasibility of such a dynamic model of computation for the critical localization function of self-driving vehicles, so that it can burden on system resources just for what is needed at any point in time or gracefully degrade accuracy in case of resource shortage. We validate our approach in a harsh scenario, by implementing it in the localization module of an autonomous racing vehicle. Experiments show that we can adapt to variations in operational conditions such as the system workload, and that we can also achieve an overall reduction of platform utilization and power consumption for this computation-greedy software module by up to$1.6\times$and$1.5\times$, respectively, for roughly the same quality of service. Federico Gavioli, Gianluca Brilli, Paolo Burgio, Davide Bertozzi |
DATE | 3 |
| 2024 | GPU implementation of the Frenet Path Planner for embedded autonomous systems: A case study in the F1tenth scenarioabstractAutonomous vehicles are increasingly utilized in safety-critical and time-sensitive settings like urban environments and competitive racing. Planning maneuvers ahead is pivotal in these scenarios, where the onboard compute platform determines the vehicle’s future actions. This paper introduces an optimized implementation of the Frenet Path Planner, a renowned path planning algorithm, accelerated through GPU processing. Unlike existing methods, our approach expedites the entire algorithm, encompassing path generation and collision avoidance. We gauge the execution time of our implementation, showcasing significant enhancements over the CPU baseline (up to 22x of speedup). Furthermore, we assess the influence of different precision types (double, float, half) on trajectory accuracy, probing the balance between completion speed and computational precision. Moreover, we analyzed the impact on the execution time caused by the use of Nvidia Unified Memory and by the interference caused by other processes running on the same system. We also evaluate our implementation using the F1tenth simulator and in a real race scenario. The results position our implementation as a strong candidate for the new state-of-the-art implementation for the Frenet Path Planner algorithm. Filippo Muzzini, Nicola Capodieci, Federico Ramanzin, Paolo Burgio |
J. Syst. Archit. | 4 |
| 2023 | Fine-Grained QoS Control via Tightly-Coupled Bandwidth Monitoring and Regulation for FPGA-based Heterogeneous SoCsabstractEmbedded systems are increasingly adopting heterogeneous templates integrating hardware accelerators and application-specific processors, which poses novel challenges. In particular, it is difficult to have accurate control of task activities in Commercial Off-the-shelf (COTS) System on Chips (SoCs), due to complex main memory sharing mechanisms among different computing engines. To address this problem, bandwidth regulation approaches based on monitoring and throttling are widely adopted. Existing solutions, however, are either too coarse-grained, limiting the control over computing engines activities, or platform-dependent, addressing the problem only for specific SoCs. In this paper we propose an innovative, fine-grained and platform-independent approach that can accurately control main memory bandwidth usage in an FPGA-based Heterogeneous System on Chip (HeSoC). Experimental results conducted on the Xilinx Zynq UltraScale+ platform demonstrate that our approach enables solutions not feasible with state-of-the-art bandwidth regulation methods. Gianluca Brilli, Giacomo Valente, Alessandro Capotondi, Paolo Burgio, T. Di Masciov, Paolo Valente, Andrea Marongiu |
DAC | 4 |
| 2023 | Time-sensitive autonomous architectures
Donato Ferraro, Luca Palazzi, Federico Gavioli, Michele Guzzinati, Andrea Bernardi, Benjamin Rouxel, Paolo Burgio, Marco Solieri |
Real Time Syst. | 7 |
| 2022 | An FPGA Overlay for Efficient Real-Time Localization in 1/10th Scale Autonomous VehiclesabstractHeterogeneous systems-on-chip (HeSoC) based on reconfigurable accelerators, such as Field-Programmable Gate Arrays (FPGA), represent an appealing option to deliver the performance/Watt required by the advanced perception and localization tasks employed in the design of Autonomous Vehicles. Different from software-programmed GPUs, FPGA development involves significant hardware design effort, which in the context of HeSoCs is further complicated by the system-level integration of HW and SW blocks. High-Level Synthesis is increasingly being adopted to ease hardware IP design, allowing engineers to quickly prototype their solutions. However, automated tools still lack the required maturity to efficiently build the complex hard-ware/software interaction between the host CPU and the FPGA accelerator(s). In this paper we present a fully integrated system design where a particle filter for LiDAR-based localization is efficiently deployed as FPGA logic, while the rest of the compute pipeline executes on programmable cores. This design constitutes the heart of a fully-functional 1/10th-scale racing autonomous car. In our design, accelerated IPs are controlled locally to the FPGA via a proxy core. Communication between the two and with the host CPU happens via shared memory banks also implemented as FPGA IPs. This allows for a scalable and easy-to-deploy solution both from the hardware and software viewpoint, while providing better performance and energy efficiency compared to state-of-the-art solutions. Andrea Bernardi, Gianluca Brilli, Alessandro Capotondi, Andrea Marongiu, Paolo Burgio |
DATE | 5 |
| 2022 | Understanding and Mitigating Memory Interference in FPGA-based HeSoCsabstractLike most high-end embedded systems, FPGA-based systems-on-chip (SoC) are increasingly adopting heterogeneous designs, where CPU cores, the configurable logic and other ICs all share interconnect and main memory (DRAM) controller. This paradigm is scalable and reduces production costs and time-to-market, but creates resource contention issues, which ultimately affects the programs' timing. This problem has been widely studied on CPU- and GPU-based systems, along with strategies to mitigate such effects, but little has been done so far to systematically study the problem on FPGA-based SoCs. This work provides an in-depth analysis of memory interference on such systems, tar-geting two state-of-the-art commercial FPGA SoCs. We also discuss architectural support for Controlled Memory Request Injection (CMRI), a technique that has proven effective at reducing the bandwidth under-utilization implied by naive schemes that solve the interference problem by only allowing mutually exclusive access to the shared resources. Our experimental results show that: i) memory interference can slow down CPU tasks by up to 16×in the tested FPGA-based SoCs; ii) CMRI allows to exploit more than 40% of the memory bandwidth avail-able to FPGA accelerators (normally completely unused in PREM-like schemes), keeping the slowdown due to interference below 10%. Gianluca Brilli, Alessandro Capotondi, Paolo Burgio, Andrea Marongiu |
DATE | 3 |
| 2022 | Sentient Spaces: Intelligent Totem Use Case in the ECSEL FRACTAL ProjectabstractThe objective of the FRACTAL project is to create a novel approach to reliable edge computing. The FRACTAL computing node will be the building block of scalable Internet of Things (from Low Computing to High Computing Edge Nodes). The node will also have the capability of learning how to improve its performance against the uncertainty of the environment. In such a context, this paper presents in detail one of the key use cases: an Internet-of-Things solution, represented by intelligent totems for advertisement and wayfinding services, within advanced ICT-based shopping malls conceived as a sentient space. The paper outlines the reference scenario and provides an overview of the architecture and the functionality of the demonstrator, as well as a roadmap for its development and evaluation. Federica Caruso, Tania Di Mascio, Daniele Frigioni, Luigi Pomante, Giacomo Valente, Stefano Delucchi, Paolo Burgio, Manuel Di Frangia, Luca Paganin, Chiara Garibotto, Damiano Vallocchia |
DSD | 7 |
| 2021 | Programmable Systems for Intelligence in Automobiles (PRYSTINE): Final results after Year 3abstractAutonomous driving is disrupting the automotive industry as we know it today. For this, fail-operational behavior is essential in the sense, plan, and act stages of the automation chain in order to handle safety-critical situations on its own, which currently is not reached with state-of-the-art approaches.The European ECSEL research project PRYSTINE realizes Fail-operational Urban Surround perceptION (FUSION) based on robust Radar and LiDAR sensor fusion and control functions in order to enable safe automated driving in urban and rural environments. This paper showcases some of the key exploitable results (e.g., novel Radar sensors, innovative embedded control and E/E architectures, pioneering sensor fusion approaches, AI-controlled vehicle demonstrators) achieved until its final year 3. Norbert Druml, Anna Ryabokon, Rupert Schorn, Jochen Koszescha, Kaspars Ozols, Aleksandrs Levinskis, Rihards Novickis, Ethiopia Nigussie, Jouni Isoaho, Selim Solmaz, Georg Stettinger, Sergio E. Diaz, Mauricio Marcano, Jorge Villagra, Juan Medina, Martina Schwarz, Antonio Artuñedo, Mauro Comi, Rutger Beekelaar, Onur Özçelik, Elif Aksu Tasdelen, Yesim Gürbüz, Jan Saijets, Jukka Kyynäräinen, Dmitry Morits, Björn Debaillie, Maxim Rykunov, Joan Escamilla, Jarno Vanne, Tomi Korhonen, Kalle Holma, Eva-Maria Matzhold, Carlo Novara, Fabio Tango, Paolo Burgio, Giuseppe Carlo Calafiore, Milad Karimshoushtari, Emilie Boulay, Miguel Dhaens, Kylian Praet, Han Zwijnenberg, Henri Palm, David Aledo Ortega, Ercan Kalali, Tuomas Pensala, Arto Kyytinen, Morten Larsen, Omar Veledar, Georg Macher, Michael Lafer, Lorenzo Giraudi, Jakob Reckenzaun, Daniel Hammer, Naveen Mohan, Josef Schmid, Alfred Höß, Shai Ophir, Anand Dubey, Jonas Fuchs, Maximilian Lübke, Andrei Anghel, Nicolae-Catalin Ristea, Martin Törngren, Alua Musralina, Marlene Harter, Joseena Memadathil Jose, George Dimitrakopoulos 0001 |
DSD | 35 |
| 2021 | A Full-Featured, Enhanced Cost Function to Mitigate Motion Sickness in Semi- and Fully-autonomous Vehicles
Isa Moazen, Paolo Burgio |
VEHITS | 2 |
| 2020 | An Automatic Scenario Generator for Validation of Automated Valet Parking SystemsabstractA primary goal of self-driving car manufacturers is to create an autonomous car system that is clearly and demonstrably safer than an average human-controlled car.The real-world tests are expensive, time-consuming and potentially dangerous.The virtual simulation is therefore required.The autonomous driving valet parking is expected to be the first commercially available automated driving function without a human driver at the wheel (SAE Level 4).Although many simulation solutions for the automotive market already exist, none of them features the parking environments.In this paper, we propose a new software virtual scenario generator for the parking sites.The tool populates the synthetics parking maps with objects and actions related to these environments: the cars driving from the drop-off point towards the vacant slots and the randomly placed parked cars, each with a given probability of exiting its slot.The generated scenarios are in the OpenSCENARIO format and are fully simulated in the Virtual Test Drive simulator. Andrea Tagliavini, Donato Ferraro, Tomasz Kloda, Paolo Burgio |
VEHITS | 4 |
| 2019 | PRYSTINE - Technical Progress After Year 1abstractAmong the actual trends that will affect society in the coming years, autonomous driving stands out as having the potential to disruptively change the automotive industry as we know it today. For this, fail-operational behavior is essential in the sense, plan, and act stages of the automation chain in order to handle safety-critical situations by its own, which currently is not reached with state-of-the-art approaches also due to missing reliable environment perception and sensor fusion. PRYSTINE will realize Fail-operational Urban Surround perceptION (FUSION) which is based on robust Radar and LiDAR sensor fusion and control functions in order to enable safe automated driving in urban and rural environments. In this paper, we detail the vision of the PRYSTINE project and we showcase the results achieved during the first year. Norbert Druml, Omar Veledar, Georg Macher, Georg Stettinger, Selim Solmaz, Jakob Reckenzaun, Sergio E. Diaz, Mauricio Marcano, Jorge Villagra, Rutger Beekelaar, Johannes Jany-Luig, Marta Maria Corredoira, Paolo Burgio, Christian Ballato, Björn Debaillie, Lars van Meurs, Andrei Sergeevich Terechko, Fabio Tango, Anna Ryabokon, Andrei Anghel, Oguz Icoglu, Sumeet S. Kumar, George Dimitrakopoulos 0001 |
DSD | 13 |
| 2019 | System Performance Modelling of Heterogeneous HW Platforms: An Automated Driving Case StudyabstractThe push towards automated and connected driving functionalities mandates the use of heterogeneous HW platforms in order to provide the required computational resources. For these platforms, the established methods for performance modelling in industry are no longer effective. In this paper, we propose an initial modelling concept for heterogeneous platforms which can then be fed into appropriate tools to derive effective performance predictions. The approach is demonstrated for a prototypical automated driving application on the Nvidia Tegra X2 platform. Falk Wurst, Dakshina Dasari, Arne Hamann 0001, Dirk Ziegenbein, Ignacio Sanudo Olmedo, Nicola Capodieci, Marko Bertogna, Paolo Burgio |
DSD | 8 |
| 2018 | Artificial Neural Networks: The Missing Link Between Curiosity and Accuracy
Giorgia Franchini, Paolo Burgio, Luca Zanni |
ISDA (2) | 2 |
| 2017 | Adaptive Coordination in Autonomous Driving: Motivations and PerspectivesabstractAs autonomous cars are entering mainstream, new research directions are opening involving several domains, from hardware design to control systems, from energy efficiency to computer vision. An exciting direction of research is represented by the coordination of the different vehicles, moving the focus from the single one to a collective system. In this paper we propose some challenging examples thatshow the motivations for a coordination approach in autonomous driving. Moreover, we present some techniques borrowed from distributed artificial intelligence that can be exploited to tackle the previously mentioned challenges. Marko Bertogna, Paolo Burgio, Giacomo Cabri, Nicola Capodieci |
WETICE | 2 |
| 2016 | A Software Stack for Next-Generation Automotive Systems on Many-Core Heterogeneous PlatformsabstractThe advent of commercial-of-the-shelf (COTS) heterogeneous many-core platforms is opening up a series of opportunities in the embedded computing market. Integrating multiple computing elements running at lower frequencies allows obtaining impressive performance capabilities at a reduced power consumption. These platforms can be successfully adopted to build the next-generation of self-driving vehicles, where Advanced Driver Assistance Systems (ADAS) need to process unprecedently higher computing workloads at low power budgets. Unfortunately, the current methodologies for providing real-time guarantees are uneffective when applied to the complex architectures of modern many-cores. Having impressive average performances with no guaranteed bounds on the response times of the critical computing activities is of little if no use to these applications. Project HERCULES will provide the required technological infrastructure to obtain an order-of-magnitude improvement in the cost and power consumption of next generation automotive systems. This paper presents the integrated software framework of the project, which allows achieving predictable performance on top of cutting-edge heterogeneous COTS platforms. The proposed software stack will let both real-time and non real-time application coexist on next-generation, power-efficient embedded platform, with preserved timing guarantees. Paolo Burgio, Marko Bertogna, Ignacio Sanudo Olmedo, Paolo Gai, Andrea Marongiu, Michal Sojka |
DSD | 1 |
| 2014 | A tightly-coupled hardware controller to improve scalability and programmability of shared-memory heterogeneous clustersabstractModern designs for embedded many-core systems increasingly include application-specific units to accelerate key computational kernels with orders-of-magnitude higher execution speed and energy efficiency compared to software counterparts. A promising architectural template is based on heterogeneous clusters, where simple RISC cores and specialized HW units (HWPU) communicate in a tightly-coupled manner via L1 shared memory. Efficiently integrating processors and a high number of HW Processing Units (HWPUs) in such an system poses two main challenges, namely, architectural scalability and programmability. In this paper we describe an optimized Data Pump (DP) which connects several accelerators to a restricted set of communication ports, and acts as a virtualization layer for programming, exposing FIFO queues to offload “HW tasks” to them through a set of lightweight APIs. In this work, we aim at optimizing both these mechanisms, for respectively reducing modules area and making programming sequence easier and lighter. Paolo Burgio, Robin Danilo, Andrea Marongiu, Philippe Coussy, Luca Benini |
DATE | 1 |
| 2014 | Tightly-coupled hardware support to dynamic parallelism acceleration in embedded shared memory clustersabstractModern designs for embedded systems are increasingly embracing cluster-based architectures, where small sets of cores communicate through tightly-coupled shared memory banks and high-performance interconnections. At the same time, the complexity of modern applications requires new programming abstractions to exploit dynamic and/or irregular parallelism on such platforms. Supporting dynamic parallelism in systems which i) are resource-constrained and ii) run applications with small units of work calls for a runtime environment which has minimal overhead for the scheduling of parallel tasks. In this work, we study the major sources of overhead in the implementation of OpenMP dynamic loops, sections and tasks, and propose a hardware implementation of a generic Scheduling Engine (HWSE) which fits the semantics of the three constructs. The HWSE is designed as a tightly-coupled block to the PEs within a multi-core cluster, communicating through a shared-memory interface. This allows very fast programming and synchronization with the controlling PEs, fundamental to achieving fast dynamic scheduling, and ultimately to enable fine-grained parallelism. We prove the effectiveness of our solutions with real applications and synthetic benchmarks, using a cycle-accurate virtual platform. Paolo Burgio, Giuseppe Tagliavini, Francesco Conti 0001, Andrea Marongiu, Luca Benini |
DATE | 1 |
| 2014 | A HLS-Based Toolflow to Design Next-Generation Heterogeneous Many-Core Platforms with Shared MemoryabstractThis work describes how we use High-Level Synthesis to support design space exploration (DSE) of heterogeneous many-core systems. Modern embedded systems increasingly couple hardware accelerators and processing cores on the same chip, to trade specialization of the platform to an application domain for increased performance and energy efficiency. However, the process of designing such a platform is complex and error-prone, and requires skills on algorithmic aspects, hardware synthesis, and software engineering. DSE can partially be automated, and thus simplified, by coupling the use of HLS tools and virtual prototyping platforms. In this paper we enable the design space exploration of heterogeneous many-cores adopting a shared-memory architecture template, where communication and synchronization between the hardware accelerators and the cores happens through L1 shared memory. This communication infrastructure leverages a "zero-copy" scheme, which simplifies both the design process of the platform and the development of applications on top of it. Moreover, the shared-memory template perfectly fits the semantics of several high-level programming models, such as OpenMP. We provide programmers with simple yet powerful abstractions to exploit accelerators from within an OpenMP application, and propose a low-cost implementation of the necessary runtime support. An HLS-based automatic design flow is set up, to quickly explore the design space using a cycle-accurate virtual platform. Paolo Burgio, Andrea Marongiu, Philippe Coussy, Luca Benini |
EUC | 1 |
| 2013 | Enabling fine-grained OpenMP tasking on tightly-coupled shared memory clustersabstractCluster-based architectures are increasingly being adopted to design embedded many-cores. These platforms can deliver very high peak performance within a contained power envelope, provided that programmers can make effective use the available parallel cores. This is becoming an extremely difficult task, as embedded applications are growing in complexity and exhibit irregular and dynamic parallelism. The OpenMP tasking extensions represent a powerful abstraction to capture this form of parallelism. However, efficiently supporting it on cluster-based embedded SoCs is not easy, because the fine-grained parallel workload present in embedded applications can not tolerate high memory and run-time overheads. In this paper we present our design of the runtime support layer to OpenMP tasking for an embedded shared memory cluster, identifying key aspects to achieving performance and discussing important architectural support to removing major bottlenecks. Paolo Burgio, Giuseppe Tagliavini, Andrea Marongiu, Luca Benini |
DATE | 1 |
| 2013 | Variation-tolerant OpenMP tasking on tightly-coupled processor clustersabstractWe present a variation-tolerant tasking technique for tightly-coupled shared memory processor clusters that relies upon modeling advance across the hardware/software interface. This is implemented as an extension to the OpenMP 3.0 tasking programming model. Using the notion of Task-Level Vulnerability (TLV) proposed here, we capture dynamic variations caused by circuit-level variability as a high-level software knowledge. This is accomplished through a variation-aware hardware/software codesign where: (i) Hardware features variability monitors in conjunction with online per-core characterization of TLV metadata; (ii) Software supports a Task-level Errant Instruction Management (TEIM) technique to utilize TLV metadata in the runtime OpenMP task scheduler. This method greatly reduces the number of recovery cycles compared to the baseline scheduler of OpenMP [22], consequently instruction per cycle (IPC) of a 16-core processor cluster is increased up to 1.51× (1.17× on average). We evaluate the effectiveness of our approach with various number of cores (4,8,12,16), and across a wide temperature range(ΔT=90°C). Abbas Rahimi, Andrea Marongiu, Paolo Burgio, Rajesh K. Gupta 0001, Luca Benini |
DATE | 3 |
| 2012 | Fast and lightweight support for nested parallelism on cluster-based embedded many-coresabstractSeveral recent many-core accelerators have been architected as fabrics of tightly-coupled shared memory clusters. A hierarchical interconnection system is used - with a crossbar-like medium inside each cluster and a network-on-chip (NoC) at the global level - which make memory operations non-uniform (NUMA). Nested parallelism represents a powerful programming abstraction for these architectures, where a first level of parallelism can be used to distribute coarse-grained tasks to clusters, and additional levels of fine-grained parallelism can be distributed to processors within a cluster. This paper presents a lightweight and highly optimized support for nested parallelism on cluster-based embedded many-cores. We assess the costs to enable multi-level parallelization and demonstrate that our techniques allow to extract high degrees of parallelism. Andrea Marongiu, Paolo Burgio, Luca Benini |
DATE | 2 |
| 2012 | OpenMP-based Synergistic Parallelization and HW Acceleration for On-Chip Shared-Memory ClustersabstractModern embedded MPSoC designs increasingly couple hardware accelerators to processing cores to trade between energy efficiency and platform specialization. To assist effective design of such systems there is the need on one hand for clear methodologies to streamline accelerator definition and instantiation, on the other for architectural templates and run-time techniques that minimize processors-to-accelerator communication costs. In this paper we present an architecture featuring tightly-coupled processors and accelerators, with zero-copy communication. Efficient programming is supported by an extended OpenMP programming model, where custom directives allow to specialize code regions for execution on parallel cores, accelerators, or a mix of the two. Our integrated approach enables fast yet accurate exploration of accelerator-based HW and SW architectures. Paolo Burgio, Andrea Marongiu, Dominique Heller, Cyrille Chavet, Philippe Coussy, Luca Benini |
DSD | 1 |
| 2011 | Bus Access Design for Combined Worst and Average Case Execution Time Optimization of Predictable Real-Time Applications on Multiprocessor Systems-on-ChipabstractOptimization techniques for improving the average-case execution time of an application, for which predictability with respect to time is not required, have been investigated for a long time in many different contexts. However, this has traditionally been done without paying attention to the worst-case execution time. For predictable real-time applications, on the other hand, the focus has been solely on worst-case execution time optimization, ignoring how this affects the execution time in the average case. In this paper, we show that having a good average-case delay can be important also for real-time applications for which predictability is required. Furthermore, for real-time applications running on multiprocessor systems-on-chip, we present a technique for optimizing the average case and the worst case simultaneously, allowing for a good average-case execution time while still keeping the worst case as small as possible. Jakob Rosen, Carl-Fredrik Neikter, Petru Eles, Zebo Peng, Paolo Burgio, Luca Benini |
IEEE Real-Time and Embedded Technology and Applications Symposium | 5 |
| 2010 | Vertical stealing: robust, locality-aware do-all workload distribution for 3D MPSoCsabstractIn this paper we address the issue of efficient doall workload distribution on a embedded 3D MPSoC. 3D stacking technology enables low latency and high bandwidth access to multiple, large memory banks in close spatial proximity. In our implementation one silicon layer contains multiple processors, whereas one or more DRAM layers on top host a NUMA memory subsystem. To obtain high locality and balanced workload we consider a two-step approach. First, a compiler pass analyzes memory references in a loop and schedules each iteration to the processor owning the most frequently accessed data. Second, if locality-aware loop parallelization has generated unbalanced workload we allow idle processors to execute part of the remaining work from neighbors by implementing runtime support for work stealing. Andrea Marongiu, Paolo Burgio, Luca Benini |
CASES | 2 |
| 2010 | Evaluating OpenMP Support Costs on MPSoCsabstractThe ever-increasing complexity of MPSoCs is making the production of software the critical path in embedded system development. Several programming models and tools have been proposed in the recent past that aim at facilitating application development for embedded MPSoCs. OpenMP is a mature and easy-to-use standard for shared memory programming, which has recently been successfully adopted in embedded MPSoC programming as well. To achieve performance, however, it is necessary that the implementation of OpenMP constructs efficiently exploits the many peculiarities of MPSoC hardware. In this paper we present an extensive evaluation of the cost associated with supporting OpenMP on such a machine, investigating several implementative variants that efficiently exploit the memory hierarchy. Experimental results on different benchmarks confirm the effectiveness of the optimizations in terms of performance improvements. Andrea Marongiu, Paolo Burgio, Luca Benini |
DSD | 2 |
| 2010 | Adaptive TDMA bus allocation and elastic scheduling: A unified approach for enhancing robustness in multi-core RT systemsabstractNext-generation real-time systems will be increasingly based on heterogeneous MPSoC design paradigms, where predictability and performance will be key issues to deal with. Such issues can be tackled both at the hardware level, by embedding technologies such as TDMA busses, and at the OS level, where suitable scheduling techniques can improve performance and reduce energy consumption. Among these, elastic scheduling has been proved to provide satisfactory results by dynamically reducing task periods at run-time to ensure the highest utilization possible of the processors. On the other hand, elastic scheduling lowers the degree of predictability and increases the complexity of the analysis at the system level. This reduces the benefits given by the TDMA bus, which relies on the high level task analysis for a robust and efficient slot allocation. Starting from this consideration, we propose a system where the elastic scheduling and the TDMA bus work synergistically. We introduce a QoS-aware adaptive bus service which takes the best of both techniques, mitigating their drawbacks at the same time. We show how the overhead introduced by coordination action is small, and it is however dominated by the benefits of the overall strategy in terms of performance and predictability guarantees. Paolo Burgio, Martino Ruggiero, Francesco Esposito, Mauro Marinoni, Giorgio C. Buttazzo, Luca Benini |
ICCD | 1 |