VLDB 2026 Research / reviewers in the wild / expert
Roberto Cavicchioli
dblp:136/4964
· DBLP profile ↗
35ranked-venue papers
7as first author
26since 2021 · last 2026
0000-0003-0166-0898ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Certification of Industrial Digital Twins via Blockchain for Trusted Lifecycle ManagementabstractThe integration of Digital Twins (DTs) and Blockchain technologies represents a promising direction for building trustworthy, auditable, and interoperable industrial systems. Yet, most existing approaches focus on static identity anchoring rather than on the continuous certification of DT state evolution. This paper proposes a novel framework for the dynamic certification of DTs in Industrial Internet of Things (IIoT) environments, combining a lightweight, permissioned blockchain with adaptive batching and ordering mechanisms. The proposed architecture connects physical assets, DT models and a blockchain-based certification layer through three coordinated components: a DT Instance Manager, a Smart Contract for state hashing and metadata storage, and Verifier Nodes for crosspeer consistency checking. A complete experimental campaign evaluates certification latency, drop rate, commit ratio, and energy overhead under realistic IIoT network conditions. Results demonstrate sub–30ms end-to-end latency for full IIoT emulation and up to 60% energy savings with micro-batching, confirming the feasibility of scalable and energy-aware DT certification across the edge–cloud continuum. Marcello Pietri, Matteo Martinelli 0001, Fabio Turazza, Roberto Cavicchioli, Marco Picone 0001, Marco Mamei |
CCNC | 4 |
| 2026 | Seeing is Believing? Understanding Truth in Synthetic MediaabstractThis paper presents preliminary findings from an ongoing investigation into user perceptions of truthfulness in multimedia news, with a focus on AI-generated content. A 26-item questionnaire was designed and administered to 1,804 volunteers to examine how individuals evaluate the credibility of news items, combining real and fabricated events illustrated with authentic or AI/Synthetic images. Descriptive results reveal systematic age-related differences: older participants were more likely to misinterpret fabricated events as authentic and reported lower confidence in distinguishing AI-generated material, whereas younger groups displayed higher exposure to AI tools but limited training. Clustering analysis further identified heterogeneous profiles, ranging from distrustful users perceiving high levels of misinformation, to balanced realists, to vigilant monitors. Building on these insights, we propose a perceptual profiling model that supports the design of personalized alerts, moving beyond one-size-fits-all moderation strategies. Although limited by non-random sampling, a predominantly female cohort, and reliance on self-reported measures, the study demonstrates the relevance of user-centered approaches in addressing the challenges of synthetic media. Camilla Totti, Roberto Cavicchioli, Marco Furini, Giacomo Tagliani |
CCNC | 2 |
| 2026 | Multi-Partner Project: dAIEDGE - A Network of Excellence for Distributed, Trustworthy, Efficient and Scalable AI at the EdgeabstractThe dAIEDGE Network of Excellence (NoE) seeks to strengthen and support the development of a dynamic European cutting-edge Artificial intelligence (AI) ecosystem under the umbrella of the European Lighthouse for AI, and to sustain the development of advanced AI. dAIEDGE fosters the exchange of ideas, concepts, and trends on cutting-edge next generation AI, creating links between ecosystem actors to help both the European Commission (EC) and the European Union (EU) and the peripheral AI constituency identify strategies for future developments in Europe. Our main objective is to advance Europe’s innovation and technology base by developing a comprehensive policy and governance approach to AI in order for the EU to become a world leader in innovation in the data economy and its applications. Alain Pagani, Haralampos-G. D. Stratigopoulos, Aysajan Abidin, Mhd Rashed Al Koutayni, Luca Benini, Angelos Bilas, Alessandro Capotondi, Roberto Cavicchioli, Brian Clerkin, Oscar Déniz-Suárez, Margaux Divernois, Baptiste Dupertuis, Dorvan Favre, Giulio Gambardella, Ander García Gangoiti, Carlo Augusto Grazia, Dominik Günzel, Jude Haris, Klodjan K. Hidri, Maïck Huguenin-Vuillemin, Manal Jammal, Paul Kling, Christos Kozanitis, Xavier Lessage, Srikanth Mandapati, Philippe Massonet, Alfio Di Mauro, Varesh Mishra, Juan Odriozola, Javier Parra 0001, Nuria Pazos, Viviane Potocnik, Miguel de Prado, Rohit Prasad, Spyridon Raptis, Gregoire Rebstein, Ignacio Sanudo Olmedo, Mohamed Selim, Chinmay Satish Shrivastav, Noelia Vállez, Giorgos Vasiliadis, Micaela Verrucchi, Enrico Vincenzi, Damian Vizár, Devendra Vyas, Stefan Wiehle |
DATE | 9 |
| 2026 | Leader Election Algorithms for Vehicle Platooning: A Survey Focusing on Resilience to Member Failures
Chinmay Satish Shrivastav, Francesco Castorini, Alessio Masola, Paolo Burgio, Roberto Cavicchioli |
MDM | 5 |
| 2025 | Diffusion is Your Friend in Show, Suggest and Tell
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi |
IEEE Big Data | 2 |
| 2025 | AI-Based Melody GenerationabstractIn the evolving realm of social media, music significantly enhances post appeal and viewer engagement. However, challenges such as copyright and royalties complicate its usage. Artificial intelligence (AI) might be used to generate royalty- free music that complements visual content. This paper presents an AI-based melody generator designed to create music suitable for various applications, particularly for enhancing social media posts. Unlike full songs, which involve complex AI models, our focus on melodies addresses a more specific and manageable aspect of music generation. We developed an algorithm to differentiate between main and background melodies, leveraging an LSTM and Transformer architecture to capture musical dependencies. Training on the Lakn MIDI dataset, which includes 178,000 files, our model achieved 64% accuracy in predicting main melodies and 78% in background melodies. Evaluation by 23 volunteers revealed that AI-generated melodies were as pleasant as human-composed ones and revealed that participants struggled to distinguish whether the melody they heard was human-composed or AI-generated. This indicates that our AI model might offer significant benefits in scenarios where melodies play an important role. Roberto Cavicchioli, Jia-Cheng Hu, Marco Furini |
CCNC | 1 |
| 2025 | From Physical to Digital: Exploring Digital Twins within the Modena Automotive Smart AreaabstractThe Modena Automotive Smart Area (MASA) is a cutting-edge testing environment featuring a variety of dynamic physical assets, including smart cameras, roadside units, and connected vehicles. These assets support numerous digital applications, ranging from real-time safety systems to mobility intelligence and 3D visualization of the MASA area. However, the complexity of the physical environment and the diverse needs of these digital applications necessitate a decoupling strategy to ensure efficient operation. This paper presents the design of the MASA Digital Twin, detailing its hierarchical structure, the associated design challenges, and the technological approaches used in its implementation. The MASA Digital Twin serves as a crucial tool for managing the interplay between physical and digital elements, enabling a more structured and adaptable approach to connected mobility and smart city applications. Marco Picone 0001, Antonello Barbone, Riccardo Morandi, Enrico Rossini, Alessio Masola, Marcello Pietri, Roberto Cavicchioli, Carlo Augusto Grazia, Marco Mamei, Marko Bertogna |
CCNC | 7 |
| 2025 | Listening to Emotions: Inferring User Mood Through Music Consumption PatternsabstractMusic serves as a powerful medium for emotional expression, influencing and reflecting an individual’s mood. This study proposes a non-intrusive method to infer a user’s emotional state based solely on their music listening history. Unlike existing approaches that rely on biometric data or textual sentiment analysis, this method ensures privacy while leveraging analytical features of songs, such as tempo, valence, and energy, to classify emotions. The proposed framework maps songs into an emotional space and tracks long-term mood trends by clustering user listening patterns. The experimental phase involved 30 volunteers whose listening histories were analyzed to reconstruct emotional profiles. The results indicate that our proposal might be used to identify stable moods, seasonal trends, and fluctuations. These insights have broad applications, from improving music recommendation systems to informing mental health interventions and improving user experiences in adaptive environments. In summary, this study highlights music as a valuable tool for emotion recognition, paving the way for innovative, privacypreserving applications in personalized digital services. Roberto Cavicchioli, Marco Furini |
ISCC | 1 |
| 2025 | Scalable Object Geolocation in Traffic Camera Imagery Using 3D World ModelabstractThe rapid expansion of outdoor traffic camera systems requires efficient methods to accurately estimate the geolocation of objects within their scenes. We present an innovative and scalable framework that completely automates this process by combining easy-to-build 3D world modeling with real-world traffic camera imagery. First, using the Cesium plugin for Unreal Engine, we create detailed and scalable 3D representations of urban environments, leveraging publicly available, highly accurate 3D data. This results in the creation of globally curated 3D content, including terrain, imagery, and photogrammetry. The real-world traffic camera imagery is then matched within our model using state-of-the-art feature matching techniques. By estimating the homography between synthetic images from the 3D model and the real images from traffic cameras, we accurately determine the geolocation of observed objects within the scene. This approach not only enhances geolocation accuracy but also enables seamless scalability across diverse urban settings and camera deployments worldwide. Our method significantly reduces the manual effort required for traffic camera calibration, thus streamlining the deployment of intelligent transportation systems at scale. We demonstrate the high performance of our approach by experimenting in an urban trial site with multiple smart city cameras and publicly available cameras around the world. Additionally, we highlight the adaptability of our framework for a wide range of computer vision-based traffic analytics applications, including its potential for drone-based localization. Chinmay Satish Shrivastav, Alessio Masola, Roberto Cavicchioli, Nicola Capodieci, Paolo Burgio |
SMC | 3 |
| 2025 | Embeddings hidden layers learning for neural network compressionabstractSequence modeling neural networks are being deployed over a growing number of applications, a phenomenon partially motivated by the advent of Large Language Models (LLMs), Neural Networks often characterized by billions of parameters. Such a size poses an obstacle to their deployment in constrained devices, which motivates the development of compression methods. In this work, we introduce a new parameter-sharing method that leverages the embedding matrix to learn the model's hidden layers. To demonstrate its effectiveness, we present a new architecture family called ShareBERT, which can preserve up to 95.5% of BERT accuracy performances, using only 5M parameters (21.9× fewer parameters) without the help of Knowledge Distillation. The evaluation of multiple linguistic benchmarks showcases that our compression method does not negatively affect the model's learning capabilities, instead, it can be beneficial for representation learning. The method is robust and flexible across different neural architecture types (such as Recurrent, Convolution, and Transformers), layers (e.g., encoder, decoder, autoregressive, and non-autoregressive modules), and tasks (e.g., translation, captioning, and language modeling). Our proposal pushes the model compression to a new level by enabling the design of near-zero architectures, and on top of that, it is orthogonal to most existing approaches, which can be further applied to ease the deployment in low-powered and embedded devices. Code is available at https://github.com/jchenghu/sharebert. Jia-Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi |
Neural Networks | 2 |
| 2024 | ShareBERT: Embeddings Are Capable of Learning Hidden LayersabstractThe deployment of Pre-trained Language Models in memory-limited devices is hindered by their massive number of parameters, which motivated the interest in developing smaller architectures. Established works in the model compression literature showcased that small models often present a noticeable performance degradation and need to be paired with transfer learning methods, such as Knowledge Distillation. In this work, we propose a parameter-sharing method that consists of sharing parameters between embeddings and the hidden layers, enabling the design of near-zero parameter encoders. To demonstrate its effectiveness, we present an architecture design called ShareBERT, which can preserve up to 95.5% of BERT Base performances, using only 5M parameters (21.9× fewer parameters) without the help of Knowledge Distillation. We demonstrate empirically that our proposal does not negatively affect the model learning capabilities and that it is even beneficial for representation learning. Code will be available at https://github.com/jchenghu/sharebert. Jia-Cheng Hu, Roberto Cavicchioli, Giulia Berardinelli, Alessandro Capotondi |
AAAI | 2 |
| 2024 | On Using Artificial Intelligence to Predict Music Playlist SuccessabstractThe emergence of digital music platforms has fundamentally transformed the way we engage with and organize music. As playlist creation has gained widespread popularity, there is an increasing desire among music aficionados and industry experts to comprehend the factors that drive playlist success. This paper presents a machine learning-based approach designed to predict the success of music playlists. By analyzing various musical characteristics of songs, our model achieves an impressive accuracy of 89.6% in predicting playlist success. Notably, it exhibits a remarkable 92.0% accuracy in forecasting the success of popular playlists, while also effectively identifying unpopular playlists with an accuracy of 89.4%. These findings provide invaluable insights into playlist creation, ultimately enhancing the overall music-listening experience. By harnessing the power of machine learning, our proposed approach unlocks new prospects for optimizing playlist design strategies and delivering personalized music recommendations. This has significant ramifications for music enthusiasts and industry professionals seeking to elevate playlist creation and enrich the music consumption experience. Roberto Cavicchioli, Jia-Cheng Hu, Marco Furini |
CCNC | 1 |
| 2024 | Collaborative Misbehaviour Response System for Improving Road SafetyabstractWrong-way driving (WWD), Driver Monitoring System (DMS), and parking violations pose significant threats to road safety. To address these challenges, we propose a collaborative misbehavior response system (MBR) that generates real-time, context-aware navigation recommendations to the nearest available parking spot. The MBR integrates individual misbehavior detection systems(MBDs) for a holistic approach to road safety and leverages Kafka and Avro for efficient communication under the 5GMETA Platform. Khaled Chikh, Chinmay Satish Shrivastav, Roberto Cavicchioli |
CCNC | 3 |
| 2024 | The Degree of Entanglement: Cyber-Physical Awareness in Digital Twin ApplicationsabstractA defining feature of a Digital Twin (DT) is its level of ”entanglement”: the degree of strength to which the twin is interconnected with its physical counterpart. Despite its importance, this characteristic has not been yet fully investigated, and its impact on applications' design is underestimated. In this paper, we define the concept of “Degree of Entanglement” (DoE), which provides an operational model for assessing the strength of the entanglement between a DT and its physical counterpart. We also propose an interoperable representation of DoE within the Web of Things (WoT) framework, which enables DT-driven applications to dynamically adapt to changes in the physical environment. We evaluate our proposal using two realistic use cases, demonstrating the practical utility of DoE in supporting, for instance, context-awareness decisions and adaptiveness. Marco Picone 0001, Stefano Mariani 0001, Roberto Cavicchioli, Paolo Burgio, Arslane Hamza Cherif |
CCNC | 3 |
| 2024 | Learning from Wrong Predictions in Low-Resource Neural Machine TranslationabstractResource scarcity in Neural Machine Translation is a challenging problem in both industry applications and in the support of less-spoken languages represented, in the worst case, by endangered and low-resource languages. Many Data Augmentation methods rely on additional linguistic sources and software tools but these are often not available in less favoured language. For this reason, we present USKI (Unaligned Sentences Keytokens pre-traIning), a pre-training strategy that leverages the relationships and similarities that exist between unaligned sentences. By doing so, we increase the dataset size of endangered and low-resource languages by the square of the initial quantity, matching the typical size of high-resource language datasets such as WMT14 En-Fr. Results showcase the effectiveness of our approach with an increase on average of 0.9 BLEU across the benchmarks using a small fraction of the entire unaligned corpus, suggesting the importance of the research topic and the potential of a currently under-utilized resource and under-explored approach. Jia-Cheng Hu, Roberto Cavicchioli, Giulia Berardinelli, Alessandro Capotondi |
LREC/COLING | 2 |
| 2024 | High-Performance Feature Extraction for GPU -Accelerated ORB-SLAMxabstractIn the autonomous vehicles field, localization is a crucial aspect. While the ORB-SLAM algorithm is a recognized solution for these tasks, it poses challenges due to its computational intensity. Although accelerated implementation exists, a bottleneck persists in the Point Filtering phase which relies on the Distribute Octree algorithm that is not suitable for GPU processing. In this paper, we introduce a novel GPU-suitable algorithm designed to enhance the Point Filtering step, surpassing Distribute Octree. We conducted a comprehensive comparison with state-of-the-art CPU and GPU implementations, considering both computational time and trajectory accuracy. Our experimental results, demonstrate significant speed-ups up to 3x compared to previous contributions. Filippo Muzzini, Nicola Capodieci, Roberto Cavicchioli, Benjamin Rouxel |
DATE | 3 |
| 2024 | Shifted Window Fourier Transform and Retention for Image Captioning
Jia-Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi |
ICONIP (8) | 2 |
| 2024 | Adaptive Frame-Aware Network for Driver Monitoring SystemsabstractDriver monitoring systems (DMS) are crucial for enhancing road safety by detecting driver behaviors and states such as drowsiness, distraction, and other potentially hazardous actions. In this research, we propose a novel approach using Adaptive Frame Aware Network (AFAN) for sequence classification in driver monitoring. Our model leverages frame embedding, adaptive attention, and sequence classification to accurately classify driver behaviors. We evaluate our model on two distinct datasets, including sequences collected on a private test setup, demonstrating its effectiveness in diverse conditions. The results show significant improvements over existing methods, achieving up to more than twice faster inference time and up to 88% memory size reduction, highlighting the potential of AFAN for real-time driver monitoring applications. Khaled Chikh, Roberto Cavicchioli |
SEC | 2 |
| 2023 | Exploiting Multiple Sequence Lengths in Fast End to End Training for Image CaptioningabstractWe introduce a method called the Expansion mechanism that processes the input unconstrained by the number of elements in the sequence. By doing so, the model can learn more effectively compared to traditional attention-based approaches. To support this claim, we design a novel architecture ExpansionNet v2 that achieved strong results on the MS COCO 2014 Image Captioning challenge and the State of the Art in its respective category, with a score of 143.7 CIDErD in the offline test split, 140.8 CIDErD in the online evaluation server and 72.9 AllCIDEr on the nocaps validation set. Additionally, we introduce an End to End training algorithm up to 2.8 times faster than established alternatives. Jia-Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi |
IEEE Big Data | 2 |
| 2023 | Memory-Aware Latency Prediction Model for Concurrent Kernels in Partitionable GPUs: Simulations and Experiments
Alessio Masola, Nicola Capodieci, Roberto Cavicchioli, Ignacio Sanudo Olmedo, Benjamin Rouxel |
JSSPP | 3 |
| 2023 | 5G MEC Architecture for Vulnerable Road Users Management Through Smart City Data FusionabstractEnhancing the safety of Vulnerable Road Users (VRUs) poses a significant research challenge in the context of connected mobility and a plethora of technological opportunities trying to balance efficiency and widespread applicability. This paper presents a demo focused on applying 5G Multi-Access Edge Computing (MEC) to address this challenge through the combination of commercial mobile devices, public cellular networks, and data fusion between vehicle positioning and city camera infrastructure. The demo showcases the designed system and its experimental evaluation in the Modena Automotive Smart Area (MASA) through the 5G MEC infrastructure of Telecom Italia (TIM) with the aim to build a secure and efficient connected mobility environment. Enrico Rossini, Marcello Pietri, Roberto Cavicchioli, Marco Picone 0001, Marco Mamei, Roberto Querio, Laura Colazzo, Roberto Procopio |
MobiCom | 3 |
| 2023 | Machine Learning Techniques for Understanding and Predicting Memory Interference in CPU-GPU Embedded SystemsabstractNowadays, heterogeneous embedded platforms are extensively used in various low-latency applications, including the automotive industry, real-time IoT systems, and automated factories. These platforms utilize specific components, such as CPUs, GPUs, and neural network accelerators for efficient task processing and to solve specific problems with a lower power consumption compared to more traditional systems. However, since these accelerators share resources such as the global memory, it is crucial to understand how workloads behave under high computational loads to determine how parallel computational engines on modern platforms can interfere and adversely affect the system's predictability and performance. One area that remains unclear is the interference effect on shared memory resources between the CPU and GPU: more specifically, the latency degradation experienced by GPU kernels when memory-intensive CPU applications run concurrently. In this work, we first analyze the metrics that characterize the behavior of different kernels under various board conditions caused by CPU memory-intensive workloads on a Nvidia Jetson Xavier. Then, we exploit various machine learning methodologies aiming to estimate the latency degradation of kernels based on their metrics. As a result of this, we are able to identify the metrics that could potentially have the most significant impact when predicting the kernels completion latency degradation. Alessio Masola, Nicola Capodieci, Benjamin Rouxel, Giorgia Franchini, Roberto Cavicchioli |
RTCSA | 5 |
| 2023 | Brief Announcement: Optimized GPU-accelerated Feature Extraction for ORB-SLAM SystemsabstractReducing the execution time of ORB-SLAM algorithm is a crucial aspect of autonomous vehicles since it is computationally intensive for embedded boards. We propose a parallel GPU-based implementation, able to run on embedded boards, of the Tracking part of the ORB-SLAM2/3 algorithm. Our implementation is not simply a GPU port of the tracking phase. Instead, we propose a novel method to accelerate image Pyramid construction on GPUs. Comparison against state-of-the-art CPU and GPU implementations, considering both computational time and trajectory errors shows improvement on execution time in well-known datasets, such as KITTI and EuRoC. Filippo Muzzini, Nicola Capodieci, Roberto Cavicchioli, Benjamin Rouxel |
SPAA | 3 |
| 2023 | Evaluating Controlled Memory Request Injection for Efficient Bandwidth Utilization and Predictable Execution in Heterogeneous SoCsabstractHigh-performance embedded platforms are increasingly adopting heterogeneous systems-on-chip (HeSoC) that couple multi-core CPUs with accelerators such as GPU, FPGA, or AI engines. Adopting HeSoCs in the context of real-time workloads is not immediately possible, though, as contention on shared resources like the memory hierarchy—and in particular the main memory (DRAM)—causes unpredictable latency increase. To tackle this problem, both the research community and certification authorities mandate (i) that accesses from parallel threads to the shared system resources (typically, main memory) happen in a mutually exclusive manner by design, or (ii) that per-thread bandwidth regulation is enforced. Such arbitration schemes provide timing guarantees, but make poor use of the memory bandwidth available in a modern HeSoC. Controlled Memory Request Injection (CMRI) is a recently-proposed bandwidth limitation concept that builds on top of a mutually-exclusive schedule but still allows the threads currently not entitled to access memory to use as much of the unused bandwidth as possible without losing the timing guarantee. CMRI has been discussed in the context of a multi-core CPU, but the same principle applies also to a more complex system such as an HeSoC. In this article, we introduce two CMRI schemes suitable for HeSoCs: Voluntary Throttling via code refactoring and Bandwidth Regulation via dynamic throttling. We extensively characterize a proof-of-concept incarnation of both schemes on two HeSoCs: an NVIDIA Tegra TX2 and a Xilinx UltraScale+, highlighting the benefits and the costs of CMRI for synthetic workloads that model worst-case DRAM access. We also test the effectiveness of CMRI with real benchmarks, studying the effect of interference among the host CPU and the accelerators. Gianluca Brilli, Roberto Cavicchioli, Marco Solieri, Paolo Valente, Andrea Marongiu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | A Taxonomy of Modern GPGPU Programming Methods: On the Benefits of a Unified SpecificationabstractSeveral Application Programming Interfaces (APIs) and frameworks have been proposed to simplify the development of General-Purpose GPU (GPGPU) applications. GPGPU application development typically involves specific customization for the target operating systems and hardware devices. The effort to port applications from one API to the other (or to develop multi-target applications) is complicated by the availability of a plethora of specifications, which in essence offers very similar underlying functionality. In this work we provide an in-depth study of six state-of-the-art GPGPU APIs. From these we derive a taxonomy of the common semantics and propose a unified specification. We describe a methodology to translate this unified specification into different target APIs. This simplifies cross-platform application development and provides a clean framework for benchmarking. Our proposed unified specification is called GUST (GPGPU Unified Specification and Translation) and it captures common functionality found in compute-only APIs (e.g., CUDA and OpenCL), in the compute pipeline of traditional graphic-oriented APIs (e.g., OpenGL and Direct3D11) and in last-generation bare-metal APIs (e.g., Vulkan and Direct3D12). The proposed translation methodology solves differences between specific APIs in a transparent manner, without hiding available tuning knobs for compute kernel optimizations and fostering best programming practices in a simple manner. Nicola Capodieci, Roberto Cavicchioli, Andrea Marongiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | The HPC-DAG Task Model for Heterogeneous Real-Time SystemsabstractRecent commercial hardware platforms for embedded real-time systems feature heterogeneous processing units and computing accelerators on the same System-on-Chip. When designing complex real-time applications for such architectures, the designer is exposed to a number of difficult choices, like deciding on which compute engine to execute a certain task, or what degree of parallelism to adopt for a given function. To help the designer exploring the wide space of design choices and tune the scheduling parameters, we propose a novel real-time application model, called HPC-DAG (Heterogeneous Parallel Condition Directed Acyclic Graph Model), specifically conceived for heterogeneous platforms. An HPC-DAG allows the system designer to specify alternative implementations of a software component for different processing engines, as well as conditional branches to modelif-then-elsestatements. We also propose a schedulability analysis for the HPC-DAG model and a set of heuristic allocation algorithms aimed at improving schedulability for latency sensitive applications. Our analysis takes into account the cost of preempting a task, which can be non-negligible on certain processors. We show the use of our approach on a realistic case study, and we demonstrate its effectiveness by comparing it with state-of-the-art algorithms previously proposed in literature. Houssam-Eddine Zahaf, Nicola Capodieci, Roberto Cavicchioli, Giuseppe Lipari, Marko Bertogna |
IEEE Trans. Computers | 3 |
| 2020 | A Systematic Assessment of Embedded Neural Networks for Object DetectionabstractObject detection is arguably one of the most important and complex tasks to enable the advent of next-generation autonomous systems. Recent advancements in deep learning techniques allowed a significant improvement in detection accuracy and latency of modern neural networks, allowing their adoption in automotive, avionics and industrial embedded systems, where performances are required to meet size, weight and power constraints.Multiple benchmarks and surveys exist to compare state-of-the-art detection networks, profiling important metrics, like precision, latency and power efficiency on Commercial-off-the-Shelf (COTS) embedded platforms. However, we observed a fundamental lack of fairness in the existing comparisons, with a number of implicit assumptions that may significantly bias the metrics of interest. This includes using heterogeneous settings for the input size, training dataset, threshold confidences, and, most importantly, platform-specific optimizations, that are especially important when assessing latency and energy-related values. The lack of uniform comparisons is mainly due to the significant effort required to re-implement network models, whenever openly available, on the specific platforms, to properly configure the available acceleration engines for optimizing performance, and to re-train the model using a homogeneous dataset.This paper aims at filling this gap, providing a comprehensive and fair comparison of the best-in-class Convolution Neural Networks (CNNs) for real-time embedded systems, detailing the effort made to achieve an unbiased characterization on cutting-edge system-on-chips. Multi-dimensional trade-offs are explored for achieving a proper configuration of the available programmable accelerators for neural inference, adopting the best available software libraries. To stimulate the adoption of fair benchmarking assessments, the framework is released to the public in an open source repository. Micaela Verucchi, Gianluca Brilli, Davide Sapienza, Mattia Verasani, Marco Arena, Francesco Gatti, Alessandro Capotondi, Roberto Cavicchioli, Marko Bertogna, Marco Solieri |
ETFA | 8 |
| 2020 | Evaluating Controlled Memory Request Injection to Counter PREM Memory Underutilization
Roberto Cavicchioli, Nicola Capodieci, Marco Solieri, Marko Bertogna, Paolo Valente, Andrea Marongiu |
JSSPP | 1 |
| 2020 | Contending memory in heterogeneous SoCs: Evolution in NVIDIA Tegra embedded platformsabstractModern embedded platforms are known to be constrained by size, weight and power (SWaP) requirements. In such contexts, achieving the desired performance-per-watt target calls for increasing the number of processors rather than ramping up their voltage and frequency. Hence, generation after generation, modern heterogeneous System on Chips (SoC) present a higher number of cores within their CPU complexes as well as a wider variety of accelerators that leverages massively parallel compute architectures. Previous literature demonstrated that while increasing parallelism is theoretically optimal for improving on average performance, shared memory hierarchies (i.e. caches and system DRAM) act as a bottleneck by exposing the platform processors to severe contention on memory accesses, hence dramatically impacting performance and timing predictability. In this work we characterize how subsequent generations of embedded platforms from the NVIDIA Tegra family balanced the increasing parallelism of each platform's processors with the consequent higher potential on memory interference. We also present an open-source software for generating test scenarios aimed at measuring memory contention in highly heterogeneous SoCs. Nicola Capodieci, Roberto Cavicchioli, Ignacio Sanudo Olmedo, Marco Solieri, Marko Bertogna |
RTCSA | 2 |
| 2019 | Novel Methodologies for Predictable CPU-To-GPU Command OffloadingabstractContainerisation is becoming a cornerstone of modern distributed systems, thanks to their lightweight virtualisation, high portability, and seamless integration with orchestration tools such as Kubernetes. The usage of containers has also gained traction in real-time cyber-physical systems, such as software-defined vehicles, which are characterised by strict timing requirements to ensure safety and performance. Nevertheless, ensuring real-time execution of co-located containers is challenging because of mutual interference due to the sharing of the same processing hardware. Existing parallel computing frameworks such as Ray and its Kubernetes-enabled variant, KubeRay, excel in distributed computation but lack support for scheduling policies that allow guaranteeing real-time timing constraints and CPU resource isolation between containers, such as the SCHED_DEADLINE policy of Linux. To fill this gap, this paper extends Ray to support real-time containers that leverage SCHED_DEADLINE. To this end, we propose KubeDeadline, a novel, modular Kubernetes extension to support SCHED_DEADLINE. We evaluate our approach through extensive experiments, using synthetic workloads and a case study based on the MobileNet and EfficientNet deep neural networks. Our evaluation shows that KubeDeadline ensures deadline compliance in all synthetic workloads, adds minimal deployment overhead (in the order of milliseconds), and achieves lower worst-case response times, up to 4 times lower, than vanilla Kubernetes under background interference. Roberto Cavicchioli, Nicola Capodieci, Marco Solieri, Marko Bertogna |
ECRTS | 1 |
| 2018 | NVIDIA GPU scheduling details in virtualized environments: work-in-progressabstractModern automotive grade embedded platforms feature high performance Graphics Processing Units (GPUs) to support the massively parallel processing power needed for next-generation autonomous driving applications. Hence, a GPU scheduling approach with strong Real-Time guarantees is needed. While previous research efforts focused on reverse engineering the GPU ecosystem in order to understand and control GPU scheduling on NVIDIA platforms, we provide an in depth explanation of the NVIDIA standard approach to GPU application scheduling on a Drive PX platform. Then, we discuss how a privileged scheduling server can be used to enforce arbitrary scheduling policies in a virtualized environment. Nicola Capodieci, Roberto Cavicchioli, Marko Bertogna |
EMSOFT | 2 |
| 2018 | A Perspective on Safety and Real-Time Issues for GPU Accelerated ADASabstractThe current trend in designing Advanced Driving Assistance System (ADAS) is to enhance their computing power by using modern multi/many core accelerators. For many critical applications such as pedestrian detection, line following, and path planning the Graphic Processing Unit (GPU) is the most popular choice for obtaining orders of magnitude increases in performance at modest power consumption. This is made possible by exploiting the general purpose nature of today's GPUs, as such devices are known to express unprecedented performance per watt on generic embarrassingly parallel workloads (as opposed of just graphical rendering, as GPUs where only designed to sustain in previous generations). In this work, we explore novel challenges that system engineers have to face in terms of real-time constraints and functional safety when the GPU is the chosen accelerator. More specifically, we investigate how much of the adopted safety standards currently applied for traditional platforms can be translated to a GPU accelerated platform used in critical scenarios. Ignacio Sanudo Olmedo, Nicola Capodieci, Roberto Cavicchioli |
IECON | 3 |
| 2018 | Deadline-Based Scheduling for GPU with Preemption SupportabstractModern automotive-grade embedded computing platforms feature high-performance Graphics Processing Units (GPUs) to support the massively parallel processing power needed for next-generation autonomous driving applications (e.g., Deep Neural Network (DNN) inference, sensor fusion, path planning, etc). As these workload-intensive activities are pushed to higher criticality levels, there is a stronger need for more predictable scheduling algorithms that are able to guarantee predictability without overly sacrificing GPU utilization. Unfortunately, the real-rime literature on GPU scheduling mostly considered limited (or null) preemption capabilities, while previous efforts in broader domains were often based on programming models and APIs that were not designed to support the real-rime requirements of recurring workloads. In this paper, we present the design of a prototype real-time scheduler for GPU activities on an embedded System on a Chip (SoC) featuring a cutting edge GPU architecture by NVIDIA adopted in the autonomous driving domain. The scheduler runs as a software partition on top of the NVIDIA hypervisor, and it leverages latest generation architectural features, such as pixel-level preemption and threadlevel preemption. Such a design allowed us to implement and test a preemptive Earliest Deadline First (EDF) scheduler for GPU tasks providing bandwidth isolations by means of a Constant Bandwidth Server (CBS). Our work involved investigating alternative programming models for compute APIs, allowing us to characterize CPU-to-GPU command submission with more detailed scheduling information. A detailed experimental characterization is presented to show the significant schedulability improvement of recurring real-time GPU tasks. Nicola Capodieci, Roberto Cavicchioli, Marko Bertogna, Aingara Paramakuru |
RTSS | 2 |
| 2017 | Memory interference characterization between CPU cores and integrated GPUs in mixed-criticality platformsabstractMost of today's mixed criticality platforms feature Systems on Chip (SoC) where a multi-core CPU complex (the host) competes with an integrated Graphic Processor Unit (iGPU, the device) for accessing central memory. The multi-core host and the iGPU share the same memory controller, which has to arbitrate data access to both clients through often undisclosed or non-priority driven mechanisms. Such aspect becomes critical when the iGPU is a high performance massively parallel computing complex potentially able to saturate the available DRAM bandwidth of the considered SoC. The contribution of this paper is to qualitatively analyze and characterize the conflicts due to parallel accesses to main memory by both CPU cores and iGPU, so to motivate the need of novel paradigms for memory centric scheduling mechanisms. We analyzed different well known and commercially available platforms in order to estimate variations in throughput and latencies within various memory access patterns, both at host and device side. Roberto Cavicchioli, Nicola Capodieci, Marko Bertogna |
ETFA | 1 |
| 2013 | ML estimation of wavelet regularization hyperparameters in inverse problemsabstractIn this paper we are interested in regularizing hyperparameter estimation by maximum likelihood in inverse problems with wavelet regularization. One parameter per subband will be estimated by gradient ascent algorithm. We have to face with two main difficulties: i) sampling the a posteriori image distribution to compute the gradient; ii) choosing a suited step-size to ensure good convergence properties. We first show that introducing an auxiliary variable makes the sampling feasible using classical Metropolis-Hastings algorithm and Gibbs sampler. Secondly, we propose an adaptive step-size selection and a line-search strategy to improve the gradient-based method. Good performances of the proposed approach are demonstrated on both synthetic and real data. Roberto Cavicchioli, Caroline Chaux, Laure Blanc-Féraud, Luca Zanni |
ICASSP | 1 |