EDBT 2026 Demo / reviewers in the wild / expert
Nicola Bombieri
dblp:58/1093
· DBLP profile ↗
89ranked-venue papers
33as first author
28since 2021 · last 2026
0000-0003-3256-5885ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 61 · 28 first-author · 17 since 2021Software engineering, systems software and programming languages · 25 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Theory of computation · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COMETH: Convex optimization for multiview estimation and tracking of humansabstractIn the era of Industry 5.0, monitoring human activity is essential for ensuring both ergonomic safety and overall well-being. While multi-camera centralized setups improve pose estimation accuracy, they often suffer from high computational costs and bandwidth requirements, limiting scalability and real-time applicability. Distributing processing across edge devices can reduce network bandwidth and computational load. On the other hand, the constrained resources of edge devices lead to accuracy degradation, and the distribution of computation leads to temporal and spatial inconsistencies. We address this challenge by proposing COMETH (Convex Optimization for Multiview Estimation and Tracking of Humans), a lightweight algorithm for real-time multi-view human pose fusion that relies on three concepts: it integrates kinematic and biomechanical constraints to increase the joint positioning accuracy; it employs convex optimization-based inverse kinematics for spatial fusion; and it implements a state observer to improve temporal consistency. We evaluate COMETH on both public and industrial datasets, where it outperforms state-of-the-art methods in localization, detection, and tracking accuracy. The proposed fusion pipeline enables accurate and scalable human motion tracking, making it well-suited for industrial and safety-critical applications. The code is publicly available at https://github.com/PARCO-LAB/COMETH . Enrico Martini, Ho Jin Choi, Nadia Figueroa, Nicola Bombieri |
Expert Syst. Appl. | 4 |
| 2026 | Edge-Cloud Orchestration of Assertion-Based Monitors for Robotic ApplicationsabstractThe runtime verification of multi-domain software applications implementing the behaviors of modern robots is a challenging task. On the one hand, assertion-based verification (ABV) has shown great potential to check the correctness of complex systems at runtime. On the other hand, the computational overhead introduced by runtime ABV can be substantial, variable and non-deterministic. As a consequence, applying accurate ABV at runtime to autonomous robots, which are often characterized by resource-constrained computing architectures, can lead to severe slowdowns of the software execution and failures of temporal constraints, thus compromising the overall system’s correctness. We address this challenge by proposing a platform for runtime ABV that implements monitor synthesis from signal temporal logic assertions and dynamic monitor migration across edge devices and the cloud. The synthesized monitors are wrapped into ROS-compliant nodes and connected to the system under verification. The overall ABV framework and the related migration mechanism are then containerized with Docker for both edge and cloud computing. To evaluate the proposed platform, we present the results obtained with a set of synthetic benchmarks and with an industrial case study, which implements the mission of a Robotnik RB-Kairos mobile robot in a smart manufacturing production line. Note to Practitioners . This article was motivated by the need for accurate and runtime verification of robotic systems software. Verification and validation of intelligent systems are often incomplete, as they cannot anticipate all potential scenarios, including errors or unexpected events. On top of this, assertion-based verification can also be resource-intensive; therefore, careful use of resources is required to avoid overloading the robot’s computational resources with the monitors. To achieve this, we used signal temporal logic, a widely accepted solution to monitor robotic and distributed applications. The main contribution of this work is a framework that can automatically synthesize the monitors that interface with the Robot Operating System (ROS) and also the capability of optimizing the end-to-end latency of verification at runtime by exploiting a distributed computing architecture (i.e., edge-cloud). In future work, we will address not only the minimization of end-to-end latency but also the timing upper bound of monitors to achieve runtime deterministic verification. Nicola Bombieri, Samuele Germiniani, Francesco Lumpp, Graziano Pravadelli |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2025 | OLORAS: Online LOng Range Action Segmentation for Edge DevicesabstractTemporal action segmentation (TAS) is essential for identifying when actions are performed by a subject, with applications ranging from healthcare to Industry 5.0. In such contexts, the need for real-time, low-latency responses and privacy-aware data handling often requires the use of edge devices, despite their limited memory, power, and computational resources. This paper presents OLORAS, a novel TAS model designed for real-time performance on edge devices. By leveraging human pose data instead of video frames and employing linear recurrent units (LRUs), OLORAS efficiently processes long sequences while minimizing memory usage. Tested on the standard Assembly101 dataset, the model outperforms state-of-the-art TAS methods in accuracy with 10x memory footprint reduction, making it well-suited for deployment on resource-constrained devices. Filippo Ziche, Nicola Bombieri |
DATE | 2 |
| 2025 | A Deep Learning-Based Emotion Recognition Pipeline for Public Speaking Anxiety Detection in Social RoboticsabstractSocial robots are increasingly employed as personalized coaches in educational settings, offering new opportunities for applications such as public speaking training. In this domain, emotional self-regulation plays a crucial role, especially for students presenting in a non-native language. This study proposes a novel pipeline for detecting public speaking anxiety (PSA) using multimodal emotion recognition. Unlike traditional datasets that typically rely on acted emotions, we consider spontaneous data from students interacting naturally with a social robot coach. Emotional labels are generated through knowledge distillation, enabling the creation of soft labels that reflect the emotional valence of each presentation. We introduce a lightweight multimodal model that integrates speech prosody and body posture to classify speakers by anxiety level, without relying on linguistic content. Evaluated on a collected dataset of student presentations, the system achieves 74.67% accuracy and an F1-score of 0.64. The model can operate completely disconnected from the transmission network on an NVIDIA Jetson board, safeguarding data privacy and demonstrating its feasibility for real-world deployment. Michele Boldo, Delara Forghani, Nicola Bombieri, Kerstin Dautenhahn, Chrystopher L. Nehaniv |
RO-MAN | 3 |
| 2025 | An efficient solution for GPUs to the ST-connectivity problem on dynamic graphsabstractST-connectivity poses a decision problem, determining whether, for vertices s and t within a graph, t is reachable from s . The challenge arises in the context of dynamic real-world graphs that undergo rapid evolution over time. In these scenarios, repeatedly solving the s-t connectivity problem from the beginning after each graph modification becomes impractical. Although parallel solutions, especially designed for GPUs, have been introduced to tackle the size complexity of static graphs, none have specifically addressed the concern of work efficiency in dynamic graphs. We propose an efficient solution for GPUs to the st-connectivity problem that can handle concurrent processing of batches of graph updates. We use batch information strategically to reduce the overall workload needed for updating the connectivity result. We provide experimental results based on standard datasets and with graphs of different characteristics and batch sizes to evaluate the proposed solutions efficiency. • The article presents an STCON solution for GPU architectures. • The solution targets dynamic graphs. • It shows the speedup and improvements wrt the static solution. • The results have been conducted on two large and standard datasets. Leonardo Fraccaroli, Federico Busato, Rosalba Giugno, Nicola Bombieri |
Pattern Recognit. Lett. | 4 |
| 2024 | Automating FinOps in Cloud Computing: An Integrated Solution for Efficient Data Collection with Dynamic Scraper GenerationabstractThis paper introduces a framework to integrate Financial Operations (FinOps) practices in Kubernetes, addressing the challenge of managing cloud services' costs across multicloud environments. The framework automates the collection of service providers' costs by deploying Prometheus exporters and scrapers, then standardizes cost data according to the FinOps Cost and Usage Specification (FOCUS) through Large Language Models (LLMs). It enables automatic data analysis, implemented by a metrics aggregator, to compute insightful cost and quality of service optimizations. We present the experimental results on three standard cloud service providers, achieving up to 91% accuracy in automatic data standardization. The results corroborate that the framework helps simplify cloud cost management and promote FinOps principles. Francesco Lumpp, Diego Braga, Franco Fummi, Nicola Bombieri |
CloudCom | 4 |
| 2024 | Late Breaking Results: A real-time diffusion-based filter for human pose estimation on edge devicesabstractHuman Pose Estimation (HPE) is increasingly being adopted in a wide range of applications, from healthcare to Industry 5.0. To address the intrinsic inaccuracy of such CNN-based software, the current trend involves applying filtering models to refine and improve the inference results. However, state-of-the-art filtering models are computationally intensive, limiting their use in resource-constrained devices. To overcome this limitation, we propose a real-time filtering technique based on diffusion models designed specifically for edge devices. Through a micro-benchmarking phase, we analyze how the model responds to various levels of noise and select the optimal setup for specific application scenarios. Using a widely available edge device, we evaluated the model's performance on both synthetic and real noise generated by a state-of-the-art HPE system. Preliminary results demonstrate a significant improvement in real-time filtering performance with minimal computational overhead. Chiara Bozzini, Michele Boldo, Enrico Martini, Nicola Bombieri |
DAC | 4 |
| 2024 | Late Breaking Results: Evaluation of Human Action Quality with Linear Recurrent Units and Graph Attention Networks on Embedded SystemsabstractRecent evolutions of recurrent neural networks (RNN) such as S4, S4D, and LRU, have shown remarkable potential for very long-range sequence modeling tasks for vision, language, and audio. They have shown a capacity to capture dependencies over tens of thousands of steps. Unlike transformers, which face significant memory consumption challenges with large context sizes, they are a promising alternative with their ability to operate effectively on embedded systems. While they have been evaluated for classification and segmentation tasks, no work in the literature has applied them in the context of human pose estimation. In this work we propose an architecture that combines such state space models (SSM) to graph attention networks (GAT) to enable their application to evaluate human action tasks on embedded systems. Filippo Ziche, Nicola Bombieri |
DAC | 2 |
| 2024 | Orchestration-Aware Optimization of ROS2 Communication ProtocolsabstractThe robot operating system (ROS) standard has been extended with different communication mechanisms to address real-time and scalability requirements. On the other hand, containerization and orchestration platforms like Docker and Kubernetes are increasingly being adopted to strengthen platform-independent development and automatic software deployment. In this paper, we quantitatively analyze the impact of topology, containerization, and edge-cloud distribution of ROS nodes on the efficiency of the ROS2 communication protocols. We then present a framework that automatically binds the most efficient ROS protocol for each node-to-node communication by considering the architectural characteristics of both software and edge-cloud computing platforms. The framework is available at https://github.com/PARCO-LAB/ros4k. Mirco De Marchi, Nicola Bombieri |
DATE | 2 |
| 2024 | Real-Time Multi-Person Identification and Tracking via HPE and IMU Data FusionabstractIn the context of smart environments, crafting remote monitoring systems that are efficient, cost-effective, user-friendly, and respectful of privacy is crucial for many scenarios. Recognizing and tracing individuals via markerless motion capture systems in multi-person settings poses challenges due to obstructions, varying light conditions, and intricate interactions among subjects. In contrast, methods based on data gathered by Inertial Measurement Units (IMUs) located in wearables grapple with other issues, including the precision of the sensors and their optimal placement on the body. We claim that more accurate results can be achieved by mixing Human Pose Estimation (HPE) techniques with information collected by wearables. To do that, we introduce a real-time platform that fuses HPE and IMU data to track and identify people. It exploits a matching model that consists of two synergistic components: the first employs a geometric approach, correlating orientation, acceleration, and velocity readings from the input sources. The second utilizes a Convolutional Neural Network (CNN) to yield a correlation coefficient for each HPE and IMU data pair. The proposed platform achieves promising results in identification and tracking, with an accuracy rate of 96.9%. Mirco De Marchi, Cristian Turetta, Graziano Pravadelli, Nicola Bombieri |
DATE | 4 |
| 2024 | GPU-Accelerated BFS for Dynamic Networks
Filippo Ziche, Nicola Bombieri, Federico Busato, Rosalba Giugno |
Euro-Par (3) | 2 |
| 2024 | A Real-time Filter for Human Pose Estimation based on Denoising Diffusion Models for Edge DevicesabstractHuman Pose Estimation (HPE) is increasingly utilized across various sectors, from healthcare to Industry 5.0. To address the inherent inaccuracies in CNN-based HPE systems, filtering models are commonly employed to refine and improve inference results. However, state-of-the-art filtering models often require substantial computational resources, limiting their applicability in resource-constrained environments. To overcome this limitation, we propose a real-time filtering approach based on denoising diffusion models (DM) specifically optimized for edge devices. Through a micro-benchmarking process, we analyze the DM adaptability to different types and levels of noise and determine the optimal setup for specific application scenarios. We present a real-time filter that takes advantage of the DM setup with two configurations to address different application scenarios. Using a widespread edge device, we evaluate the model’s effectiveness in handling both synthetic and real noise generated by state-of-the-art HPE systems. The results demonstrate a significant improvement in real-time filtering performance with minimal computational overhead. The code is available on github.com/PARCO-LAB/LUT-DM-filters. Chiara Bozzini, Michele Boldo, Enrico Martini, Nicola Bombieri |
IROS | 4 |
| 2024 | Optimizing Kubernetes Deployment of Robotic Applications with HEFT-based Container OrchestrationabstractThis study addresses the challenge of deploying robotic software with Quality of Service (QoS) constraints in Edge-Cloud computing clusters. The paper introduces HEFT4K, an event-driven scheduling method tailored for Kubernetes-managed systems based on the Heterogeneous Early Finish Time (HEFT) algorithm. This algorithm reduces software execution time (makespan) and facilitates re-mapping in case of node failures, involving only essential containers to maintain uninterrupted robot functionality. Experimental results, conducted on a real-world robot and synthetic benchmarks, show a 75% speedup in makespan compared to the standard Kubernetes scheduler, enhancing the efficiency of QoS-focused scheduling for robotic applications in distributed systems. Francesco Lumpp, Franco Fummi, Nicola Bombieri |
IROS | 3 |
| 2024 | A Robust Filter for Marker-less Multi-person Tracking in Human-Robot Interaction ScenariosabstractPursuing natural and marker-less human-robot interaction (HRI) has been a long-standing robotics research focus, driven by the vision of seamless collaboration without physical markers. Marker-less approaches promise an improved user experience, but state-of-the-art struggles with the challenges posed by intrinsic errors in human pose estimation (HPE) and depth cameras. These errors can lead to issues such as robot jittering, which can significantly impact the trust users have in collaborative systems. We propose a filtering pipeline that refines incomplete 3D human poses from an HPE backbone and a single RGB-D camera to address these challenges, solving for occlusions that can degrade the interaction. Experimental results show that using the proposed filter leads to more consistent and noise-free motion representation, reducing unexpected robot movements and enabling smoother interaction. Enrico Martini, Harshil Parekh, Shaoting Peng, Nicola Bombieri, Nadia Figueroa |
RO-MAN | 4 |
| 2024 | Real-time multi-camera 3D human pose estimation at the edge for industrial applicationsabstractThere is an increasing interest in exploiting human pose estimation (HPE) software in human-machine interaction systems. Nevertheless, adopting such a computer vision application in real industrial scenarios is challenging. To overcome occlusion limitations, it requires multiple cameras, which in turn require multiple, distributed, and synchronized HPE software nodes running on resource-constrained edge devices. We address this challenge by presenting a real-time distributed 3D HPE platform, which consists of a set of 3D HPE software nodes on edge devices (i.e., one per camera) to redundantly extrapolate the human pose from different points of view. A centralized aggregator collects the pose information through a shared communication network and merges them, in real time, through a pipeline of filtering, clustering and association algorithms. It addresses network communication issues (e.g., delay and bandwidth variability) through a two-levels synchronization, and supports both single and multi-person pose estimation. We present the evaluation results with a real case of study (i.e., HPE for human-machine interaction in an intelligent manufacturing line), in which the platform accuracy and scalability are compared with state-of-the-art approaches and with a marker-based infra-red motion capture system. Michele Boldo, Mirco De Marchi, Enrico Martini, Stefano Aldegheri, Davide Quaglia, Franco Fummi, Nicola Bombieri |
Expert Syst. Appl. | 7 |
| 2024 | FLK: A filter with learned kinematics for real-time 3D human pose estimationabstractThere is a growing interest in adopting 3D human pose estimation in safety-critical systems, from healthcare to Industry 5.0. Nevertheless, when applied in such settings, these neural networks may suffer from estimation inaccuracy. Besides imprecise or inconsistent annotations in the training dataset, the inaccuracy is caused by poor image quality, rare poses, dropped frames, or heavy occlusions in the scene. In addition, these scenarios often require the software results to have temporal constraints, such as real-time and zero- or low-latency, which make many of the filtering solutions proposed in the literature inapplicable. This paper proposes FLK, a Filter with Learned Kinematics, to refine 3D human motion data in real-time and at zero/low latency. The temporal core combines a Kalman filter and a low-pass filter, which learns the motion model through a recurrent neural network. The spatial core takes advantage of the biomechanical constraints of the human body to provide spatial coherency between keypoints. The combination of the cores allows the filter to adequately address different types of noise, from jittering to dropped frames. We test the filter on motion data from multiple datasets and seven 3D human pose estimation backbones, improving accuracy up to 140 mm with non-Gaussian noise and 53 mm with missing information. Enrico Martini, Michele Boldo, Nicola Bombieri |
Signal Process. | 3 |
| 2024 | Domain-Adaptive Online Active Learning for Real-Time Intelligent Video Analytics on Edge DevicesabstractDeep learning (DL) for intelligent video analytics is increasingly pervasive in various application domains, ranging from Healthcare to Industry 5.0. A significant trend involves deploying DL models on edge devices with limited resources. Techniques, such as pruning, quantization, and early exit, have demonstrated the feasibility of real-time inference at the edge by compressing and optimizing deep neural networks (DNNs). However, adapting pretrained models to new and dynamic scenarios remains a significant challenge. While solutions like domain adaptation, active learning (AL), and teacher-student knowledge distillation (KD) contribute to addressing this challenge, they often rely on cloud or well-equipped computing platforms for fine tuning. In this study, we propose a framework for domain-adaptive online AL of DNN models tailored for intelligent video analytics on resource-constrained devices. Our framework employs a KD approach where both teacher and student models are deployed on the edge device. To determine when to retrain the student DNN model without ground-truth or cloud-based teacher inference, our model utilizes singular value decomposition of input data. It implements the identification of key data frames and efficient retraining of the student through the teacher execution at the edge, aiming to prevent model overfitting. We evaluate the framework through two case studies: 1) human pose estimation and 2) car object detection, both implemented on an NVIDIA Jetson NX device. Michele Boldo, Mirco De Marchi, Enrico Martini, Stefano Aldegheri, Nicola Bombieri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | A Design Flow Based on Docker and Kubernetes for ROS-based Robotic Software ApplicationsabstractHuman-centered robotic applications are becoming pervasive in the context of robotics and smart manufacturing, and such a pervasiveness is even more expected with the shift to Industry 5.0. The always increasing level of autonomy of modern robotic platforms requires the integration of software applications from different domains to implement artificial intelligence, cognition, and human-robot/robot-robot interaction. Developing and (re)configuring such a multi-domain software to meet functional constraints is a challenging task. Even more challenging is customizing the software to satisfy non-functional requirements such as real-time, reliability, and energy efficiency. In this context, the concept of Edge-Cloud continuum is gaining consensus as a solution to address functional and non-functional constraints in a seamless way. Containerization and orchestration are becoming a standard practice, as they allow for better information flow among different network levels as well as increased modularity in the use of multi-domain software components. Nevertheless, the adoption of such a practice along the design flow, from simulation to the deployment of complex robotic applications by addressing the de facto development standards (e.g., ROS - Robotic Operating System) is still an open problem. We present a design methodology based on Docker and Kubernetes that enables containerization and orchestration of ROS-based robotic SW applications for heterogeneous and hierarchical HW architectures. The methodology aims at (i) integrating and verifying multi-domain components since early in the design flow, (ii) mapping software tasks to containers to minimize the performance and memory footprint overhead, (iii) clustering containers to efficiently distribute load across the edge-cloud architecture by minimizing resource utilization, and (iv) enabling multi-domain verification of functional and non-functional constraints before deployment. The article presents the results obtained with a real case of study, in which the design methodology has been applied to program the mission of a Robotnik RB-Kairos mobile robot in an industrial agile production chain. We have obtained reduced load on the robot’s HW with minimal performance and network overhead, thanks to the optimized distributed system. Francesco Lumpp, Marco Panato, Nicola Bombieri, Franco Fummi |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Enabling Kubernetes Orchestration of Mixed-Criticality Software for Autonomous Mobile RobotsabstractContainerization and orchestration have become two key requirements in software development best practices. Containerization allows for better resource utilization, platform-independent development, and secure deployment of software. Orchestration automates the deployment, networking, scaling, and availability of containerized workloads and services. While containerization is increasingly being adopted in the robotic community, the use of task orchestration platforms (e.g., Kubernetes) is still an open challenge. The biggest limitation is due to the fact that state-of-the-art orchestrators do not support real-time containers, while advanced robotic software often consists of a mix of heterogeneous tasks (i.e., ROS nodes) with different levels of temporal constraints (i.e., mixed-criticality systems). This work addresses this challenge by presenting RT-Kube, a platform that extends the de-facto reference standard for container orchestration, Kubernetes, to schedule tasks with mixed-criticality requirements. It implements monitoring of tasks and detects missed deadlines for those with real-time constraints. It selects low-priority tasks to be migrated at runtime to different units of the computing cluster to free resources and recover from temporal violations. We present quantitative experimental results on the software implementing the mission of a Robotnik RB-Kairos mobile robot to demonstrate the effectiveness of the proposed approach. The source code is publicly available on GitHub. Francesco Lumpp, Franco Fummi, Hiren D. Patel, Nicola Bombieri |
IEEE Trans. Robotics | 4 |
| 2022 | Real-time Human Pose Estimation at the Edge for Gait Analysis at a DistanceabstractHealth telematics is a major improvement on patient lives and has shown to be a key practice to deliver healthcare services, overcoming geographical, temporal, and even organizational barriers. One of the main challenges is to perform gait analysis at a distance through camera-based platforms, which requires the system to satisfy, beside accuracy and real-time, also portability and privacy compliance at the same time. We address this challenge by proposing a portable and low-cost platform that implements real-time and accurate 3D human pose estimation through an embedded software on a low-power off-the-shelf computing device that guarantees privacy by default and by design. We evaluated both accuracy and performance of the proposed solution through an infra-red marker-based motion capture system as ground truth to understand if and how such a portable technology can be used for gait analysis at a distance without leading to different clinical interpretations. Enrico Martini, Michele Boldo, Stefano Aldegheri, Mirco De Marchi, Nicola Valè, Mirko Filippetti, Nicola Smania, Matteo Bertucco, Alessandro Picelli, Nicola Bombieri |
DCOSS | 10 |
| 2022 | Process-driven Collision Prediction in Human-Robot Work EnvironmentsabstractIn mixed human-robot work cells the emphasis is traditionally on collision avoidance to circumvent injuries and production down times. In this paper we discuss how long in advance a collision can be predicted given the behavior of a robotic arm and the current occupancy of both the robot and the human. Assuming that the behavior of the robot is a combination of a set of predefined operations, we propose an approach to learn this behavior and use it to estimate the time before a collision. The pose of the human is estimated by a multicamera inference application based on neural networks at the edge to preserve privacy and enforce scalability. The occupancy of the manipulator and of the human are modeled through the composition of segments which overcomes the traditional "virtual cage" and can be adapted to different human beings and robots. The system has been implemented in a real factory scenario to demonstrate its readiness regarding both industrial constraints and computational complexity. Luca Geretti, Stefano Centomo, Michele Boldo, Enrico Martini, Nicola Bombieri, Davide Quaglia, Tiziano Villa |
ETFA | 5 |
| 2022 | Containerization and Orchestration of Software for Autonomous Mobile Robots: a Case Study of Mixed-Criticality Tasks across Edge-Cloud Computing PlatformsabstractContainerization promises to strengthen platform-independent development, better resource utilization, and secure deployment of software. As these benefits come with negligible overhead in CPU and memory utilization, containerization is increasingly being adopted in mobile robotic applications. An open challenge is supporting software tasks that have mixed-criticality requirements. Even more challenging is the combination of real-time containers with orchestration, which is an emerging paradigm to automate the deployment, networking, scaling, and availability of containerized workloads and services. This paper addresses this challenge by presenting a framework that extends the de-facto reference standard for container orchestration, Kubernetes, to schedule tasks with mixed-criticality requirements. Quantitative experimental results on the software implementing the mission of a Robotnik RB-Kairos mobile robot demonstrate the effectiveness of the proposed approach. The source code is publicly available on GitHub. Francesco Lumpp, Franco Fummi, Hiren D. Patel, Nicola Bombieri |
IROS | 4 |
| 2022 | Integrating Wearable and Camera Based Monitoring in the Digital Twin for Safety Assessment in the Industry 4.0 Era
Michele Boldo, Nicola Bombieri, Stefano Centomo, Mirco De Marchi, Florenc Demrozi, Graziano Pravadelli, Davide Quaglia, Cristian Turetta |
ISoLA (4) | 2 |
| 2021 | A Framework for Optimizing CPU-iGPU Communication on Embedded PlatformsabstractMany modern programmable embedded devices contain CPUs and a GPU that share the same system memory on a single die. Such a unified memory architecture allows the explicit data copying between CPU and integrated GPU (iGPU) to be eliminated with the benefit of significantly improving performance and energy savings. However, to enable such a “zero-copy” communication model, many devices either implement intricate cache coherence protocols or they may disable the last level caches. This often leads to strong performance degradation of cache-dependent applications, for which CPU-iGPU data transfer based on standard copy remains the best solution. This paper presents a framework based on a performance model, a set of micro-benchmarks, and a novel zero-copy communication pattern to accurately estimate the potential speedup a CPU-iGPU application may have by considering different communication models (i.e., standard copy, unified memory, or pinned “zerocopy”). It shows how the framework can be combined with standard profiler information to efficiently drive the application tuning for a given programmable embedded device. Francesco Lumpp, Hiren D. Patel, Nicola Bombieri |
DAC | 3 |
| 2021 | A containerized ROS-compliant verification environment for robotic systemsabstractThis paper proposes an architecture and a related automatic flow to generate, orchestrate and deploy a ROS-compliant verification environment for robotic systems. The architecture enables assertion-based verification by exploiting monitors automatically synthesized from LTL assertions. The monitors are encapsulated in plug-and-play ROS nodes that do not require any modification to the system under verification (SUV). To guarantee both verification accuracy and real-time constraints of the system in a resource-constrained environment even after the monitor integration, we define a novel approach to move the monitor evaluation across the different layers of an edge-to-cloud computing platform. The verification environment is containerized for both cloud and edge computing using Docker to enable system portability and to handle, at run-time, the resources allocated for verification. The effectiveness and efficiency of the proposed architecture have been evaluated on a complex distributed system implementing a mobile robot path planner based on 3D simultaneous localization and mapping. Stefano Aldegheri, Nicola Bombieri, Samuele Germiniani, Federico Moschin, Graziano Pravadelli |
DATE | 2 |
| 2021 | A Container-based Design Methodology for Robotic Applications on Kubernetes Edge-Cloud architecturesabstractProgramming modern Robots' missions and behavior has become a very challenging task. The always increasing level of autonomy of such platforms requires the integration of multi-domain software applications to implement artificial intelligence, cognition, and human-robot/robot-robot interaction applications. In addition, to satisfy both functional and nonfunctional requirements such as reliability and energy efficiency, robotic SW applications have to be properly developed to take advantage of heterogeneous (Edge-Fog-Cloud) architectures. In this context, containerization and orchestration are becoming a standard practice as they allow for better information flow among different network levels as well as increased modularity in the use of software components. Nevertheless, the adoption of such a practice along the design flow, from simulation to the deployment of complex robotic applications by addressing the de-facto development standards (i.e., robotic operating system - ROS - compliancy for robotic applications) is still an open problem. We present a design methodology based on Docker and Kubernetes that enables containerization and orchestration of ROS-based robotic SW applications for heterogeneous and hierarchical HW architectures. The design methodology allows for (i) integration and verification of multi-domain components since early in the design flow, (ii) task-to-container mapping techniques to guarantee minimum overhead in terms of performance and memory footprint, and (iii) multi-domain verification of functional and non-functional constraints before deployment. We present the results obtained in a real case of study, in which the design methodology has been applied to program the mission of a Robotnik RB-Kairos mobile robot in an industrial agile production chain. The source code of the mobile robot is publicly available on GitHub. Francesco Lumpp, Marco Panato, Franco Fummi, Nicola Bombieri |
FDL | 4 |
| 2021 | Task Mapping and Scheduling for OpenVX Applications on Heterogeneous Multi/Many-Core ArchitecturesabstractComputer vision applications have stringent performance constraints that must be satisfied when they are run at the edge on programmable low-power embedded devices. OpenVX has emerged as the de-facto reference standard to develop such applications. OpenVX uses a primitive-based programming model that results in a directed-acyclic graph (DAG) representation of the application, which can then be used for automatic system-level optimizations and synthesis to heterogeneous multi- and many-core platforms. Although OpenVX has been standardized, its state-of-the-art algorithm for task mapping and scheduling does not deliver the performance necessary for such applications to be deployed on heterogeneous multi-/many-core platforms. This article focuses on addressing this challenge with three main contributions: First, we implemented a static task scheduling and mapping approach for OpenVX using the heterogeneous earliest finish time (HEFT) heuristic. We show that HEFT allows us to improve the system performance up to 70 percent on one of the most widespread smart systems for applying computer vision and intelligent video analytics in general at the edge (i.e., NVIDIA VisionWorks on NVIDIA Jetson TX2). Second, we show that HEFT, in the context of a vision application for edge computing where some primitives may have multiple implementations (e.g., for CPU and GPU), can lead to load imbalance amongst heterogeneous computing elements (CEs), thus suffering from degraded performance. Third, we present an algorithm called exclusive earliest finish time (XEFT) that introduces the notion of exclusive overlap between single implementation primitives to improve the load balancing. We show that XEFT can further improve the system performance up to 33 percent over HEFT, and 82 percent over the native OpenVX scheduler. We present the results on a large set of benchmarks, including a real-world localization and mapping application (ORB-SLAM) combined with an NVIDIA inference application based on convolutional neural networks (CNNs) for object detection. Francesco Lumpp, Stefano Aldegheri, Hiren D. Patel, Nicola Bombieri |
IEEE Trans. Computers | 4 |
| 2021 | SystemC Implementation of Stochastic Petri Nets for Simulation and Parameterization of Biological Networks
Nicola Bombieri, Silvia Scaffeo, Antonio Mastrandrea, Simone Caligola, Tommaso Carlucci, Franco Fummi, Carlo Laudanna, Gabriela Constantin, Rosalba Giugno |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2020 | Late Breaking Results: Enabling Containerized Computing and Orchestration of ROS-based Robotic SW Applications on Cloud-Server-Edge ArchitecturesabstractWe present a toolehain based on Docker and KubeEdge that enables containerization and orchestration of ROS-based robotic SW applications on heterogeneous and hierarchical HW architectures. The toolehain allows for verification of functional and real-time constraints through HW-in-the-loop simulation, and for automatic mapping exploration of the SW across Cloud-Server-Edge architectures. We present the results obtained for the deployment of a real case of study composed by an ORB-SLAM application combined to local/global planners with obstacle avoidance for a mobile robot navigation. Stefano Aldegheri, Nicola Bombieri, Franco Fummi, Simone Girardi, Riccardo Muradore, Nicola Piccinelli |
DAC | 2 |
| 2020 | On the Task Mapping and Scheduling for DAG-based Embedded Vision Applications on Heterogeneous Multi/Many-core ArchitecturesabstractIn this work, we show that applying the heterogeneous earliest finish time (HEFT) heuristic for the task scheduling of embedded vision applications can improve the system performance up to 70% w.r.t. the scheduling solutions at the state of the art. We propose an algorithm called exclusive earliest finish time (XEFT) that introduces the notion of exclusive overlap between application primitives to improve the load balancing. We show that XEFT can improve the system performance up to 33% over HEFT, and 82% over the state of the art approaches. We present the results on different benchmarks, including a real-world localization and mapping application (ORB-SLAM) combined with the NVIDIA object detection application based on deep-learning. Stefano Aldegheri, Nicola Bombieri, Hiren D. Patel |
DATE | 2 |
| 2020 | CRISPRitz: rapid, high-throughput and variant-aware in silico off-target site identification for CRISPR genome editingabstractMOTIVATION: Clustered regularly interspaced short palindromic repeats (CRISPR) technologies allow for facile genomic modification in a site-specific manner. A key step in this process is the in silico design of single guide RNAs to efficiently and specifically target a site of interest. To this end, it is necessary to enumerate all potential off-target sites within a given genome that could be inadvertently altered by nuclease-mediated cleavage. Currently available software for this task is limited by computational efficiency, variant support or annotation, and assessment of the functional impact of potential off-target effects. RESULTS: To overcome these limitations, we have developed CRISPRitz, a suite of software tools to support the design and analysis of CRISPR/CRISPR-associated (Cas) experiments. Using efficient data structures combined with parallel computation, we offer a rapid, reliable, and exhaustive search mechanism to enumerate a comprehensive list of putative off-target sites. As proof-of-principle, we performed a head-to-head comparison with other available tools on several datasets. This analysis highlighted the unique features and superior computational performance of CRISPRitz including support for genomic searching with DNA/RNA bulges and mismatches of arbitrary size as specified by the user as well as consideration of genetic variants (variant-aware). In addition, graphical reports are offered for coding and non-coding regions that annotate the potential impact of putative off-target sites that lie within regions of functional genomic annotation (e.g. insulator and chromatin accessible sites from the ENCyclopedia Of DNA Elements [ENCODE] project). AVAILABILITY AND IMPLEMENTATION: The software is freely available at: https://github.com/pinellolab/CRISPRitzhttps://github.com/InfOmics/CRISPRitz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Samuele Cancellieri, Matthew C. Canver, Nicola Bombieri, Rosalba Giugno, Luca Pinello |
Bioinform. | 3 |
| 2020 | Mangrove: An Inference-Based Dynamic Invariant Mining for GPU ArchitecturesabstractLikely invariants model properties that hold in operating conditions of a computing system. Dynamic mining of invariants aims at extracting logic formulas representing such properties from the system execution traces, and it is widely used for verification of intellectual property (IP) blocks. Although the extracted formulas represent likely invariants that hold in the considered traces, there is no guarantee that they are true in general for the system under verification. As a consequence, to increase the probability that the mined invariants are true in general, dynamic mining has to be performed to large sets of representative execution traces. This makes the execution-based mining process of actual IP blocks very time-consuming due to the trace lengths and to the large sets of monitored signals. This article presents Mangrove, an efficient implementation of a dynamic invariant mining algorithm for GPU architectures. Mangrove exploits inference rules, which are applied at run time to filter invariants from the execution traces and, thus, to sensibly reduce the problem complexity. Mangrove allows users to define invariant templates and, from these templates, it automatically generates kernels for parallel and efficient mining on GPU architectures. The article presents the tool, the analysis of its performance, and its comparison with the best sequential and parallel implementations at the state of the art. Nicola Bombieri, Federico Busato, Alessandro Danese, Luca Piccolboni, Graziano Pravadelli |
IEEE Trans. Computers | 1 |
| 2019 | Automatic Parameterization of the Purine Metabolism Pathway through Discrete Event-based SimulationabstractStochastic Petri Nets (SPN)are recognized as one of the standard formalisms to model metabolic networks. They allow incorporating randomness in the model and taking into account possible fluctuations and noise due to molecule interactions in the environment. Even though some frameworks have been proposed to implement and simulate SPN (e.g., Snoopy, Monalisa), they do not allow for automatic model parameterization, which is a crucial task to identify the network configurations that lead the model to satisfy certain biological properties. We present a framework to synthesize the SPN model of a metabolic network into executable code that can be simulated through a discrete event-based simulator. The framework allows the user to formally define the network properties to be observed and to automatically extrapolate, through Assertion-based Verification (ABV), the parameter configurations that lead the network to satisfy such properties. We applied the framework to model the purine metabolism and to reproduce the metabolomics data obtained from naive lymphocytes and autoreactive T cells implicated in the induction of experimental autoimmune disorders. We show system parameterization extrapolated by the framework to reproduce the experimental results and to simulate the model under different conditions. Simone Caligola, Tommaso Carlucci, Franco Fummi, Carlo Laudanna, Gabriela Constantin, Nicola Bombieri, Rosalba Giugno |
CIBCB | 6 |
| 2019 | Efficient Simulation and Parametrization of Stochastic Petri Nets in SystemC: A Case study from Systems BiologyabstractStochastic Petri nets (SPN) are a form of Petri net where the transitions fire after a probabilistic and randomly determined delay. They are adopted in a wide range of applications thanks to their capability of incorporating randomness in the models and taking into account possible fluctuations and environmental noise. In Systems Biology, they are becoming a reference formalism to model metabolic networks, in which the noise due to molecule interactions in the environment plays a crucial role. Some frameworks have been proposed to implement and dynamically simulate SPN. Nevertheless, they do not allow for automatic model parametrization, which is a crucial task to identify the network configurations that lead the model to satisfy temporal properties of the model. This paper presents a framework that synthesizes the SPN models into SystemC code. The framework allows the user to formally define the network properties to be observed and to automatically extrapolate, through Assertion-based Verification (ABV), the parameter configurations that lead the network to satisfy such properties. We applied the framework to implement and simulate a complex biological network, i.e., the purine metabolism, with the aim of reproducing the metabolomics data obtained in-vitro from naive lymphocytes and autoreactive T cells implicated in the induction of experimental autoimmune disorders. Simone Caligola, Tommaso Carlucci, Franco Fummi, Carlo Laudanna, Gabriela Constantin, Nicola Bombieri, Rosalba Giugno |
FDL | 6 |
| 2019 | RTL Assertion Mining with Automated RTL-to-TLM AbstractionabstractWe present a three-step flow to improve Assertion-based Verification methodology with integrated RTL-to-TLM abstraction: First, an automatic assertion miner generates a large set of possible assertions from an RTL design. Second, automatic assertion qualification identifies the most interesting assertions from this set. Third, the assertions are abstracted to the transaction level, such that they can be re-used in TLM verification. We show that the proposed flow automatically chooses the best assertions among the ones generated to verify the design components when abstracted from RTL to TLM. Our experimental results indicate that the proposed methodology allows us to re-use the most interesting set at TLM without relying on any time consuming or error-prone manual transformations with a considerable amount of speed up and considerable reduction in the execution time. Tara Ghasempouri, Alessandro Danese, Graziano Pravadelli, Nicola Bombieri, Jaan Raik |
FDL | 4 |
| 2019 | Data Flow ORB-SLAM for Real-time Performance on Embedded GPU BoardsabstractThe use of embedded boards on robots, including unmanned aerial and ground vehicles, is increasing thanks to the availability of GPU equipped low-cost embedded boards in the market. Porting algorithms originally designed for desktop CPUs on those boards is not straightforward due to hardware limitations. In this paper, we present how we modified and customized the open source SLAM algorithm ORB-SLAM2 to run in real-time on the NVIDIA Jetson TX2. We adopted a data flow paradigm to process the images, obtaining an efficient CPU/GPU load distribution that results in a processing speed of about 30 frames per second. Quantitative experimental results on four different sequences of the KITTI datasets demonstrate the effectiveness of the proposed approach. The source code of our data flow ORB-SLAM2 algorithm is publicly available on GitHub. Stefano Aldegheri, Nicola Bombieri, Domenico Daniele Bloisi, Alessandro Farinelli |
IROS | 2 |
| 2019 | Parallel Searching on Biological NetworksabstractSoftware applications for biological networks analysis rely on graphs to model the structure interactions. A great part of them requires searching for subgraphs in a target graph or in collections of graphs. Even though very efficient algorithms have been defined to solve such a subgraph isomorphisms problem, the complexity of current real biological networks make their sequential execution time prohibitive. On the other hand, parallel architectures, from multi-core to manycore, have become pervasive to deal with the problem of the data size. Nevertheless, the sequential nature of the graph searching algorithms makes their implementation for parallel architectures very challenging. This paper presents three different parallel solutions for the graph searching problem. The first two target the exact search for multi-core CPUs and manycore GPUs, respectively. The third one targets the approximate search for GPUs, which handles node, edge, and node label mismatches. The paper shows how different techniques have been developed in all the solutions to reduce the search space complexity. The paper shows the performance of the proposed solutions on representative biological networks containing antiviral chemical compounds and protein interactions networks. Nicola Bombieri, Vincenzo Bonnici, Rosalba Giugno |
PDP | 1 |
| 2019 | A Cross-level Verification Methodology for Digital IPs Augmented with Embedded Timing MonitorsabstractSmart systems are characterized by the integration in a single device of multi-domain subsystems of different technological domains, namely, analog, digital, discrete and power devices, MEMS, and power sources. Such challenges, emerging from the heterogeneous nature of the whole system, combined with the traditional challenges of digital design, directly impact on performance and on propagation delay of digital components. This article proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems with the automatic insertion of delay sensors, which can detect and correct timing failures. The article then proposes a methodology to verify such added features at system level. The augmented model is abstracted to SystemC TLM, which is automatically injected with mutants (i.e., code mutations) to emulate delays and timing failures. The resulting TLM model is finally simulated to identify timing failures and to verify the correctness of the inserted delay monitors. Experimental results demonstrate the applicability of the proposed design and verification methodology, thanks to an efficient sensor-aware abstraction methodology, by applying the flow to three complex case studies. Sara Vinco, Nicola Bombieri, Daniele Jahier Pagliari, Franco Fummi, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | An Efficient Implementation of a Subgraph Isomorphism Algorithm for GPUs
Vincenzo Bonnici, Rosalba Giugno, Nicola Bombieri |
BIBM | 3 |
| 2018 | Efficient Load Balancing Techniques for Graph Traversal Applications on GPUs
Federico Busato, Nicola Bombieri |
Euro-Par | 2 |
| 2018 | A Framework for the Design and Simulation of Embedded Vision Applications Based on OpenVX and ROSabstractCustomizing computer vision applications for embedded systems is a common and widespread problem in the cyber-physical systems community. Such a customization means parametrizing the algorithm by considering the external environment and mapping the Software application to the heterogeneous Hardware resources by satisfying non-functional constraints like performance, power, and energy consumption. This work presents a framework for the design and simulation of embedded vision applications that integrates the OpenVX standard platform with the Robot Operating System (ROS). The paper shows how the framework has been applied to tune the ORB-SLAM application for an NVIDIA Jetson TX2 board by considering different environment contexts and different design constraints. Stefano Aldegheri, Nicola Bombieri, Nicola Dall'Ora, Franco Fummi, Simone Girardi, Marco Panato |
ISCAS | 2 |
| 2018 | Rapid Prototyping of Embedded Vision Systems: Embedding Computer Vision Applications into Low-Power Heterogeneous ArchitecturesabstractEmbedded vision is a disruptive new technology in the vision industry. It is a revolutionary concept with far reaching implications, and it is opening up new applications and shaping the future of entire industries. It is applied in self-driving cars, autonomous vehicles in agriculture, digital dermascopes that help specialists make more accurate diagnoses, among many other unique and cutting-edge applications. The design of such systems gives rise to new challenges for embedded Software developers. Embedded vision applications are characterized by stringent performance constraints to guarantee real-time behaviours and, at the same time, energy constraints to save battery on the mobile platforms. In this paper, we address such challenges by proposing an overall view of the problem and by analysing current solutions. We present our last results on embedded vision design automation over two main aspects: the adoption of the model-based paradigm for the embedded vision rapid prototyping, and the application of heterogeneous programming languages to improve the system performance. The paper presents our recent results on the design of a localization and mapping application combined with image recognition based on deep learning optimized for an NVIDIA Jetson TX2. Stefano Aldegheri, Nicola Bombieri |
RSP | 2 |
| 2018 | Enhancing Performance of Computer Vision Applications on Low-Power Embedded Systems Through Heterogeneous Parallel ProgrammingabstractEnabling computer vision applications on low-power embedded systems gives rise to new challenges for embedded SW developers. Such applications implement different functionalities, like image recognition based on deep learning, simultaneous localization and mapping tasks. They are characterized by stringent performance constraints to guarantee real-time behaviors and, at the same time, energy constraints to save battery on the mobile platform. Even though heterogeneous embedded boards are getting pervasive for their high computational power at low power costs, they need a time consuming customization of the whole application (i.e., mapping of application blocks to CPU-GPU processing elements and their synchronization) to efficiently exploit their potentiality. Different languages and environments have been proposed for such an embedded SW customization. Nevertheless, they often find limitations on complex real cases, as their application is mutual exclusive. This paper presents a comprehensive framework that relies on a heterogeneous parallel programming model, which combines OpenMP, PThreads, OpenVX, OpenCV, and CUDA to best exploit different levels of parallelism while guaranteeing a semi-automatic customization. The paper shows how such languages and API platforms have been interfaced, synchronized, and applied to customize an ORB-SLAM application for an NVIDIA Jetson TX2 board. Stefano Aldegheri, Silvia Manzato, Nicola Bombieri |
VLSI-SoC | 3 |
| 2018 | cuRnet: an R package for graph traversing on GPUabstractBACKGROUND: R has become the de-facto reference analysis environment in Bioinformatics. Plenty of tools are available as packages that extend the R functionality, and many of them target the analysis of biological networks. Several algorithms for graphs, which are the most adopted mathematical representation of networks, are well-known examples of applications that require high-performance computing, and for which classic sequential implementations are becoming inappropriate. In this context, parallel approaches targeting GPU architectures are becoming pervasive to deal with the execution time constraints. Although R packages for parallel execution on GPUs are already available, none of them provides graph algorithms. RESULTS: This work presents cuRnet, a R package that provides a parallel implementation for GPUs of the breath-first search (BFS), the single-source shortest paths (SSSP), and the strongly connected components (SCC) algorithms. The package allows offloading computing intensive applications to GPU devices for massively parallel computation and to speed up the runtime up to one order of magnitude with respect to the standard sequential computations on CPU. We have tested cuRnet on a benchmark of large protein interaction networks and for the interpretation of high-throughput omics data thought network analysis. CONCLUSIONS: cuRnet is a R package to speed up graph traversal and analysis through parallel computation on GPUs. We show the efficiency of cuRnet applied both to biological network analysis, which requires basic graph algorithms, and to complex existing procedures built upon such algorithms. Vincenzo Bonnici, Federico Busato, Stefano Aldegheri, Murodzhon Akhmedov, Luciano Cascione, Alberto Arribas Carmena, Francesco Bertoni, Nicola Bombieri, Ivo Kwee, Rosalba Giugno |
BMC Bioinform. | 8 |
| 2018 | Correction to: cuRnet: an R package for graph traversing on GPUabstractAfter publication of this supplement article [1], it was brought to our attention that reference 10 and reference 12 in the article are incorrect. Vincenzo Bonnici, Federico Busato, Stefano Aldegheri, Murodzhon Akhmedov, Luciano Cascione, Alberto Arribas Carmena, Francesco Bertoni, Nicola Bombieri, Ivo Kwee, Rosalba Giugno |
BMC Bioinform. | 8 |
| 2018 | Arena-Idb: a platform to build human non-coding RNA interaction networksabstractBACKGROUND: High throughput technologies have provided the scientific community an unprecedented opportunity for large-scale analysis of genomes. Non-coding RNAs (ncRNAs), for a long time believed to be non-functional, are emerging as one of the most important and large family of gene regulators and key elements for genome maintenance. Functional studies have been able to assign to ncRNAs a wide spectrum of functions in primary biological processes, and for this reason they are assuming a growing importance as a potential new family of cancer therapeutic targets. Nevertheless, the number of functionally characterized ncRNAs is still too poor if compared to the number of new discovered ncRNAs. Thus platforms able to merge information from available resources addressing data integration issues are necessary and still insufficient to elucidate ncRNAs biological roles. RESULTS: In this paper, we describe a platform called Arena-Idb for the retrieval of comprehensive and non-redundant annotated ncRNAs interactions. Arena-Idb provides a framework for network reconstruction of ncRNA heterogeneous interactions (i.e., with other type of molecules) and relationships with human diseases which guide the integration of data, extracted from different sources, via mapping of entities and minimization of ambiguity. CONCLUSIONS: Arena-Idb provides a schema and a visualization system to integrate ncRNA interactions that assists in discovering ncRNA functions through the extraction of heterogeneous interaction networks. The Arena-Idb is available at http://arenaidb.ba.itb.cnr.it. Vincenzo Bonnici, Giorgio De Caro, Giorgio Constantino, Sabino Liuni, Domenica D'Elia, Nicola Bombieri, Flavio Licciulli, Rosalba Giugno |
BMC Bioinform. | 6 |
| 2017 | Power-aware Performance Tuning of GPU Applications Through MicrobenchmarkingabstractTuning GPU applications is a very challenging task as any source-code optimization can sensibly impact performance, power, and energy consumption of the GPU device. Such an impact also depends on the GPU on which the application is run. This paper presents a suite of microbenchmarks that provides the actual characteristics of specific GPU device components (e.g., arithmetic instruction units, memories, etc.) in terms of throughput, power, and energy consumption. It shows how the suite can be combined to standard profiler information to efficiently drive the application tuning by considering the three design constraints (power, performance, energy consumption) and the characteristics of the target GPU device. Nicola Bombieri, Federico Busato, Franco Fummi |
DAC | 1 |
| 2017 | Extending OpenVX for model-based design of embedded vision applicationsabstractDeveloping computer vision applications for low-power heterogeneous systems is increasingly gaining interest in the embedded systems community. Even more interesting is the tuning of such embedded software for the target architecture when this is driven by multiple constraints (e.g., performance, peak power, energy consumption). Indeed, developers frequently run into system-level inefficiencies and bottlenecks that can not be quickly addressed by traditional methods. In this context OpenVX has been proposed as the standard platform to develop portable, optimized and power-efficient applications for vision algorithms targeting embedded systems. Nevertheless, adopting OpenVX for rapid prototyping, early algorithm parametrization and validation of complex embedded applications is a very challenging task. This paper presents a methodology to integrate a model-based design environment to OpenVX. The methodology allows applying Matlab/Simulink for the model-based design, parametrization, and validation of computer vision applications. Then, it allows for the automatic synthesis of the application model into an OpenVX description for the hardware and constraints-aware application tuning. Experimental results have been conducted with an application for digital image stabilization developed through Simulink and, then, automatically synthesized into OpenVX-VisionWorks code for an NVIDIA Jetson TX1 board. Stefano Aldegheri, Nicola Bombieri |
VLSI-SoC | 2 |
| 2017 | An Efficient Approach for Accelerating Bucket Elimination on GPUsabstractBucket elimination (BE) is a framework that encompasses several algorithms, including belief propagation (BP) and variable elimination for constraint optimization problems (COPs). BE has significant computational requirements that can be addressed by using graphics processing units (GPUs) to parallelize its fundamental operations, i.e., composition and marginalization, which operate on functions represented by large tables. We propose a novel approach to parallelize these operations with GPUs, which optimizes the table layout so to achieve better performance in terms of increased speedup and scalability. Our approach allows us to process incomplete tables (i.e., tables with some missing variables assignments), which often occur in several practical applications (such as the ones we consider in our dataset). Finally, we can process tables that are larger than the GPU memory. Our approach outperforms the state-of-the-art technique to parallelize BP on GPUs, achieving better speedups (up to +466% with respect to such parallel technique). We test our method on a publicly available COP dataset, measuring a speedup up to with respect to the sequential version. The ability of our technique to process large tables is crucial in this scenario, in which most of the instances generate tables larger than the GPU memory, and hence they cannot be solved with previous GPU techniques related to BE. Filippo Bistaffa, Nicola Bombieri, Alessandro Farinelli |
IEEE Trans. Cybern. | 2 |
| 2017 | A Dynamic Approach for Workload Partitioning on GPU ArchitecturesabstractWorkload partitioning and the subsequent work item-to-thread mapping are key aspects to face when implementing any efficient GPU application. Different techniques have been proposed to deal with such issues, ranging from the computationally simplest static to the most complex dynamic ones. Each of them finds the best use depending on the workload characteristics (static for more regular workloads, dynamic for irregular workloads). Nevertheless, no one of them provides a sound tradeoff when applied in both cases. Static approaches lead to load unbalancing with irregular problems, while the computational overhead introduced by the dynamic or semi-dynamic approaches often worsens the overall application performance when run on regular problems. This article presents an efficient dynamic technique for workload partitioning and work item-to-thread mapping whose complexity is significantly reduced with respect to the other dynamic approaches in literature. The article shows how the partitioning and mapping algorithm has been implemented by fully taking advantage of the GPU device characteristics with the aim of minimizing the involved computational overhead. The article shows, compares, and analyses the experimental results obtained by applying the proposed approach and several static, dynamic, and semi-dynamic techniques at the state of the art to different benchmarks and over different GPU technologies (i.e., NVIDIA Fermi, Kepler, and Maxwell) to understand when and how each technique best applies. Federico Busato, Nicola Bombieri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | A fine-grained performance model for GPU architectures
Nicola Bombieri, Federico Busato, Franco Fummi |
DATE | 1 |
| 2016 | CUBE: A CUDA Approach for Bucket Elimination on GPUsabstractWe consider Bucket Elimination (BE), a popular algorithmic framework to solve Constraint Optimisation Problems (COPs). We focus on the parallelisation of the most computationally intensive operations of BE, i.e., join sum and maximisation, which are key ingredients in several close variants of the BE framework (including Belief Propagation on Junction Trees and Distributed COP techniques such as ActionGDL and DPOP). In particular, we propose CUBE, a highly-parallel GPU implementation of such operations, which adopts an efficient memory layout allowing all threads to independently locate their input and output addresses in memory, hence achieving a high computational throughput. We compare CUBE with the most recent GPU implementation of BE. Our results show that CUBE achieves significant speed-ups (up to two orders of magnitude) w.r.t. the counterpart approach, showing a dramatic decrease of the runtime w.r.t. the serial version (i.e., up to 652× faster). More important, such speed-ups increase when the complexity of the problem grows, showing that CUBE correctly exploits the additional degree of parallelism inherent in the problem. Filippo Bistaffa, Nicola Bombieri, Alessandro Farinelli |
ECAI | 2 |
| 2016 | APPAGATO: an APproximate PArallel and stochastic GrAph querying TOol for biological networksabstractMOTIVATION: Biological network querying is a problem requiring a considerable computational effort to be solved. Given a target and a query network, it aims to find occurrences of the query in the target by considering topological and node similarities (i.e. mismatches between nodes, edges, or node labels). Querying tools that deal with similarities are crucial in biological network analysis because they provide meaningful results also in case of noisy data. In addition, as the size of available networks increases steadily, existing algorithms and tools are becoming unsuitable. This is rising new challenges for the design of more efficient and accurate solutions. RESULTS: This paper presents APPAGATO, a stochastic and parallel algorithm to find approximate occurrences of a query network in biological networks. APPAGATO handles node, edge and node label mismatches. Thanks to its randomic and parallel nature, it applies to large networks and, compared with existing tools, it provides higher performance as well as statistically significant more accurate results. Tests have been performed on protein-protein interaction networks annotated with synthetic and real gene ontology terms. Case studies have been done by querying protein complexes among different species and tissues. AVAILABILITY AND IMPLEMENTATION: APPAGATO has been developed on top of CUDA-C ++ Toolkit 7.0 framework. The software is available online http://profs.sci.univr.it/∼bombieri/APPAGATO CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vincenzo Bonnici, Federico Busato, Giovanni Micale, Nicola Bombieri, Alfredo Pulvirenti, Rosalba Giugno |
Bioinform. | 4 |
| 2016 | An Efficient Implementation of the Bellman-Ford Algorithm for Kepler GPU ArchitecturesabstractFinding the shortest paths from a single source to all other vertices is a common problem in graph analysis. The Bellman-Ford's algorithm is the solution that solves such a single-source shortest path (SSSP) problem and better applies to be parallelized for many-core architectures. Nevertheless, the high degree of parallelism is guaranteed at the cost of low work efficiency, which, compared to similar algorithms in literature (e.g., Dijkstra's) involves much more redundant work and a consequent waste of power consumption. This article presents a parallel implementation of the Bellman-Ford algorithm that exploits the architectural characteristics of recent GPU architectures (i.e., NVIDIA Kepler, Maxwell) to improve both performance and work efficiency. The article presents different optimizations to the implementation, which are oriented both to the algorithm and to the architecture. The experimental results show that the proposed implementation provides an average speedup of$5 \times$higher than the existing most efficient parallel implementations for SSSP, that it works on graphs where those implementations cannot work or are inefficient (e.g., graphs with negative weight edges, sparse graphs), and that it sensibly reduces the redundant work caused by the parallelization process. Federico Busato, Nicola Bombieri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | RTL property abstraction for TLM assertion-based verification
Nicola Bombieri, Riccardo Filippozzi, Graziano Pravadelli, Francesco Stefanni |
DATE | 1 |
| 2015 | A SystemC Platform for Signal Transduction Modelling and Simulation in Systems BiologyabstractSignal transduction is a class of cell's biological processes, which are commonly represented as highly concurrent reactive systems. In the Systems Biology community, modelling and simulation of signal transduction require overcoming issues like discrete event-based execution of complex systems, description from building blocks through composition and encapsulation, description at different levels of granularity, methods for abstraction and refinement. This paper presents a signal transduction modelling and simulation platform based on SystemC, and shows how the platform allows handling the system complexity by modelling it at different abstraction levels. The paper reports the results obtained by applying the platform to model the intracellular signalling network controlling integrin activation mediating leukocyte recruitment from the blood into the tissues. Rosario Distefano, Franco Fummi, Carlo Laudanna, Nicola Bombieri, Rosalba Giugno |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | Exploiting GPU architectures for dynamic invariant miningabstractDynamic mining of invariants is a class of approaches to extract logic formulas from the execution traces of a system under verification (SUV), with the purpose of expressing stable conditions in the behaviour of the SUV. The mined formulas represent likely invariants for the SUV, which certainly hold on the considered traces, but there is no guarantee that they are true in general. A large set of representative execution traces must be analysed to increase the probability that mined invariants are generally true. However, this becomes extremely time-consuming for current sequential approaches when long execution traces and large set of SUV variables are considered. To overcome this limitation, the paper presents a parallel approach for invariant mining that exploits GPU architectures for processing an execution trace composed of millions of clock cycles in few seconds. Nicola Bombieri, Federico Busato, Alessandro Danese, Luca Piccolboni, Graziano Pravadelli |
ICCD | 1 |
| 2015 | Reusing RTL Assertion Checkers for Verification of SystemC TLM Models
Nicola Bombieri, Franco Fummi, Valerio Guarnieri, Graziano Pravadelli, Francesco Stefanni, Tara Ghasempouri, Michele Lora, Giovanni Auditore, Mirella Negro Marcigaglia |
J. Electron. Test. | 1 |
| 2015 | A Methodology to Recover RTL IP Functionality for Automatic Generation of SW ApplicationsabstractWith the advent of heterogeneous multiprocessor system-on-chips (MPSoCs), hardware/software partitioning is again on the rise both in research and in product development. In this new scenario, implementing intellectual-property (IP) blocks as SW applications rather than dedicated HW is an increasing trend to fully exploit the computation power provided by the MPSoC CPUs. On the other hand, whole libraries of IP blocks are available as RTL descriptions, most of them without a corresponding high-level SW implementation. In this context, this article presents a methodology to automatically generate SW applications in C++, by starting from existing RTL IPs implemented in hardware description language (HDL). The methodology exploits an abstraction algorithm to eliminate implementation details typical of HW descriptions (such as cycle-accurate functionality and data types) to guarantee relevant performance of the generated code. The experimental results show that, in many cases, the C++ code automatically generated in a few seconds with the proposed methodology is as efficient as the corresponding code manually implemented from scratch. Nicola Bombieri, Franco Fummi, Sara Vinco |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2015 | BFS-4K: An Efficient Implementation of BFS for Kepler GPU ArchitecturesabstractBreadth-first search (BFS) is one of the most common graph traversal algorithms and the building block for a wide range of graph applications. With the advent of graphics processing units (GPUs), several works have been proposed to accelerate graph algorithms and, in particular, BFS on such many-core architectures. Nevertheless, BFS has proven to be an algorithm for which it is hard to obtain better performance from parallelization. Indeed, the proposed solutions take advantage of the massively parallelism of GPUs but they are often asymptotically less efficient than the fastest CPU implementations. This paper presents BFS-4K, a parallel implementation of BFS for GPUs that exploits the more advanced features of GPU-based platforms (i.e., NVIDIA Kepler) and that achieves an asymptotically optimal work complexity. The paper presents different strategies implemented in BFS-4K to deal with the potential workload imbalance and thread divergence caused by any actual graph non-homogeneity. The paper presents the experimental results conducted on several graphs of different size and characteristics to understand how the proposed techniques are applied and combined to obtain the best performance from the parallel BFS visits. Finally, an analysis of the most representative BFS implementations for GPUs at the state of the art and their comparison with BFS-4K are reported to underline the efficiency of the proposed solution. Federico Busato, Nicola Bombieri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | A cross-level verification methodology for digital IPs augmented with embedded timing monitorsabstractSmart systems implement the leading technology advances in the context of embedded devices. Current design methodologies are not suitable to deal with tightly interacting subsystems of different technological domains, namely analog, digital, discrete and power devices, MEMS and power sources. The effects of interaction between components and with the environment must be modeled and simulated at system level to achieve high performance. Focusing on the digital domain, additional design constraints have to be considered as a result of the integration of multi-domain subsystems in a single device. The main digital design challenges, combined with those emerging from the heterogeneous nature of the whole system, directly impact on performance and on propagation delay of the digital component. This paper proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems, and a methodology to verify the added features at system-level. The design approach consists of augmenting the RTL model through the automatic insertion of delay sensors, which can detect and correct timing failures. The augmented model is abstracted to SystemC TLM and, then, mutants (i.e., code mutations for emulating timing failures) are automatically injected into the model. Experimental results demonstrate the applicability of the proposed design and verification methodology and the effectiveness of the simulation performance. Valerio Guarnieri, Massimo Petricca, Alessandro Sassone, Sara Vinco, Nicola Bombieri, Franco Fummi, Enrico Macii, Massimo Poncino |
DATE | 5 |
| 2014 | Optimising memory management for Belief Propagation in Junction Trees using GPGPUsabstractBelief Propagation (BP) in Junction Trees (JT) is one of the most popular approaches to compute posteriors in Bayesian Networks (BN). Such approach has significant computational requirements that can be addressed by using highly parallel architectures (i.e., General Purpose Graphic Processing Units) to parallelise the message update phases of BP. In this paper, we propose a novel approach to parallelise BP with GPGPUs, which focuses on optimising the memory layout of the BN tables so to achieve better performance in terms of increased speedup, reduced data transfers between the host and the GPGPU, and scalability. Our empirical comparison with the state of the art approach on standard datasets confirms significant improvements in speedups (up to +594%), and scalability (as our method can operate on networks whose potential tables exceed the global memory of the GPGPU). Filippo Bistaffa, Alessandro Farinelli, Nicola Bombieri |
ICPADS | 3 |
| 2014 | Testbench Qualification of SystemC TLM Protocols through Mutation AnalysisabstractTransaction-level modeling (TLM) has become the de-facto reference modeling style for system-level design and verification of embedded systems. It allows designers to implement high-level communication protocols for simulations up to 1000 × faster than at register-transfer level (RTL). To guarantee interoperability between TLM IP suppliers and users, designers implement the TLM communication protocols by relying on a reference standard, such as the standard OSCI for SystemC TLM. Functional correctness of such protocols as well as their compliance to the reference TLM standard are usually verified through user-defined testbenches, whose high quality and completeness play a key role for an efficient TLM design and verification flow. This article presents a methodology to apply mutation analysis, a technique applied in literature for SW testing, for measuring the testbench quality in verifying TLM protocols. In particular, the methodology aims at (i) qualifying the testbenches by considering both the TLM protocol correctness and their compliance to a defined standard (i.e., OSCI TLM), (ii) optimizing the simulation time during mutation analysis by avoiding mutation redundancies, and (iii) driving the designers in the testbench improvement. Experimental results on benchmarks of different complexity and architectural characteristics are reported to analyze the methodology applicability. Nicola Bombieri, Franco Fummi, Valerio Guarnieri, Graziano Pravadelli |
IEEE Trans. Computers | 1 |
| 2013 | A method to abstract RTL IP blocks into C++ code and enable high-level synthesisabstractWe present a method to automatically generate a synthesizable C++ specification from the given RTL design of an IP block, by abstracting away most of its micro-architectural characteristics while preserving its functionality. The goal is twofold: recover the IP block specification for system-level design, and enable the derivation of more optimized implementations through high-level synthesis. The C++ specification can be generated with different interfaces thus allowing the IP model to be reused across different system platforms. Experimental results show that the proposed approach not only enhances the reusability of the recovered IP block but also unveils a richer design space to explore. Nicola Bombieri, Hung-Yi Liu, Franco Fummi, Luca P. Carloni |
DAC | 1 |
| 2013 | On the use of GP-GPUs for accelerating compute-intensive EDA applicationsabstractGeneral purpose graphics processing units (GP-GPUs) have recently been explored as a new computing paradigm for accelerating compute-intensive EDA applications. Such massively parallel architectures have been applied in accelerating the simulation of digital designs during several phases of their development - corresponding to different abstraction levels, specifically: (i) gate-level netlist descriptions, (ii) register-transfer level and (iii) transaction-level descriptions. This embedded tutorial presents a comprehensive analysis of the best results obtained by adopting GP-GPUs in all these EDA applications. Valeria Bertacco, Debapriya Chatterjee, Nicola Bombieri, Franco Fummi, Sara Vinco, Anirudh M. Kaushik, Hiren D. Patel |
DATE | 3 |
| 2013 | SMAC: Smart Systems Co-designabstractIn this paper we present the concepts and the organization of the FP7 Project SMAC (Smart systems Co-design), an Integrated Project (IP) of the 7th ICT Call under the Objective 3.2 "Smart components and Smart Systems integration". We describe in particular the project objectives and its organization, and how it addresses the challenges of the integration of heterogeneous and conflicting domains that emerge in the design of smart systems. The main outcome of the SMAC project is the development of flexible software platform (the SMAC platform) for smart subsystems/components design include methodologies and EDA tools enabling multi-disciplinary and multi-scale modeling and design, simulation of multi-domain systems, subsystems and components at all levels of abstraction, system integration and exploration for optimization of functional and non-functional metrics. Nicola Bombieri, Giuliana Drogoudis, Giuliana Gangemi, Renaud Gillon, Enrico Macii, Massimo Poncino, Salvatore Rinaudo, Francesco Stefanni, Dimitrios Trachanis, Mark van Helvoort |
DSD | 1 |
| 2013 | On the Reuse of Heterogeneous IPs into SysML Models for Integration Validation
Nicola Bombieri, Emad Samuel Malki Ebeid, Franco Fummi, Michele Lora |
J. Electron. Test. | 1 |
| 2013 | Semi-Automatic Generation of Device Drivers for Rapid Embedded Platform DevelopmentabstractIP core integration into an embedded platform implies the implementation of a customized device driver complying with both the IP communication protocol and the CPU organization (single processor, SMP, AMP). Such a close dependence between driver and platform organization makes reuse of already existing device drivers very hard. Designers are forced to manually customize the driver code to any different organization of the target platform. This results in a very time-consuming and error-prone task. In this paper, we propose a methodology to semi-automatically generate customized device drivers, thus allowing a more rapid embedded platform development. The methodology exploits the testbench provided with the RTL IP module for extracting the formal model of the IP communication protocol. Then, a taxonomy of device drivers based on the CPU organization allows the system to determine the characteristics of the target platform and to obtain a template of the device driver code. This requires some manual support to identify the target architecture and to generate the desired device driver functionality. The template is used then to automatically generate drivers compliant with 1) the CPU organization, 2) the use in a simulated or in a real platform, 3) the interrupt support, 4) the operating system, 5) the I/O architecture, and 6) possible parallel execution. The proposed methodology has been successfully tested on a family of embedded platforms with different CPU organizations. Andrea Acquaviva, Nicola Bombieri, Franco Fummi, Sara Vinco |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | FAST-GP: An RTL functional verification framework based on fault simulation on GP-GPUsabstractThis paper presents FAST-GP, a framework for functional verification of RTL designs, which is based on fault injection and parallel simulation on GP-GPUs. Given a fault model, the framework translates the RTL code into an injected C code targeting NVIDIA GPUs, thus allowing a very fast parallel automatic test pattern generation and fault simulation. The paper compares different configurations of the framework to better exploit the architectural characteristics of such GP-GPUs (such as thread synchronization, branch divergence, etc.) by considering the architectural characteristics of the RTL design under verification (i.e., complexity, size, number of injected faults, etc.). Experimental results have been conducted by applying the framework to different designs, in order to prove the methodology effectiveness. Nicola Bombieri, Franco Fummi, Valerio Guarnieri |
DATE | 1 |
| 2012 | On the automatic synthesis of parallel SW from RTL models of hardware IPsabstractHeterogeneous multicore system-on-chips (MPSoCs) provide many degrees of freedom to map functionalities on either SW and HW components. In this scenario, enabling the remapping of HW IPs as SW routines allows to fully exploit the computation power and flexibility provided by heterogeneous MPSoCs. On the other hand, reuse of existent IP cores is the key strategy to explore this large design space in a reasonable amount of time and to reduce the error risk during the MPSoC design flow. A methodology for automatic generation of parallel SW code taking into account these aspects is currently missing. This paper aims at overcoming this limitation, by presenting a methodology to automatically generate parallel SW IPs starting from existent RTL IP models. Andrea Acquaviva, Nicola Bombieri, Franco Fummi, Sara Vinco |
ACM Great Lakes Symposium on VLSI | 2 |
| 2012 | FAST: An RTL Fault Simulation Framework based on RTL-to-TLM Abstraction
Nicola Bombieri, Franco Fummi, Valerio Guarnieri |
J. Electron. Test. | 1 |
| 2012 | On the Reuse of TLM Mutation Analysis at RTL
Valerio Guarnieri, Giuseppe Di Guglielmo, Nicola Bombieri, Graziano Pravadelli, Franco Fummi, Hanno Hantson, Jaan Raik, Maksim Jenihhin, Raimund Ubar |
J. Electron. Test. | 3 |
| 2011 | Automatic Interface Generation for Component Reuse in HW-SW PartitioningabstractHW-SW partitioning is a key problem in HW-SW code sign of embedded systems studied extensively in the past. All proposed approaches are top down flows, that start from a homogeneous formal specification of the system and determine an optimal partitioning. Thus, the proposed techniques do not exploit reuse nor reconsider the HWSW partitioning of an already designed platform. This paper proposes an extension of traditional flows that allows reuse and automatic generation of components and interfaces. The final flow has been applied to a complex industrial platform to prove the effectiveness and the advantages of the proposed approach. Nicola Bombieri, Franco Fummi, Sara Vinco, Davide Quaglia |
DSD | 1 |
| 2011 | Accelerating RTL Fault Simulation through RTL-to-TLM AbstractionabstractDifferent fault injection techniques based on simulation have been proposed in the past for functional verification of register transfer level (RTL) IP models. They allow designers to model any type of fault and provide the quality of test patterns through the fault coverage estimation. Nevertheless, the low speed of such a cycle-accurate RTL simulation involves a trade-off between the simulation time and the achieved fault coverage. On the other hand, Transaction-level modeling (TLM) allows a simulation speed-up up to 1000x with respect to RTL. This paper presents a methodology to accelerate RTL fault simulation through automatic RTL-to-TLM abstraction. The methodology abstracts injected RTL models into equivalent injected TLM models thus allowing a very fast automatic test pattern generation at TLM level. The paper shows how the generated TLM test patterns can be automatically synthesized into RTL test patterns by exploiting the structural information of the RTL model extracted during the abstraction process. Experimental results have been applied to several designs of different size and complexity to show the methodology effectiveness. Nicola Bombieri, Franco Fummi, Valerio Guarnieri |
ETS | 1 |
| 2011 | Efficient implementation and abstraction of systemc data types for fast simulation
Nicola Bombieri, Franco Fummi, Valerio Guarnieri, Francesco Stefanni, Sara Vinco |
FDL | 1 |
| 2011 | Automatic Abstraction of RTL IPs into Equivalent TLM DescriptionsabstractTransaction-level modeling (TLM) is the most promising technique to deal with the increasing complexity of modern embedded systems. However, modeling a complex system completely at transaction level could be inconvenient when IP cores are available on the market, since they are usually modeled at register transfer level (RTL). In this context, modeling and verification methodologies based on transactors allow designers to reuse RTL IPs into TLM-RTL mixed designs, thus guaranteeing a considerable saving of time. Practical advantages of such an approach are evident, but mixed TLM-RTL designs cannot completely provide the well-known effectiveness in terms of simulation speed provided by TLM. This paper presents a methodology to automatically abstract RTL IPs into equivalent TLM descriptions. To do that, the paper first proposes a formal definition of equivalence based on events, showing how such a definition can be applied to prove the correctness of a code manipulation methodology, such as code abstraction. Then, the paper proposes a technique to automatically abstract RTL IPs into TLM descriptions. Finally, the paper shows that the TLM descriptions obtained by applying the proposed technique are correct by construction, relying on the given definition of event-based equivalence. A set of experimental results is reported to confirm the effectiveness of the methodology. Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
IEEE Trans. Computers | 1 |
| 2010 | Abstraction of RTL IPs into embedded softwareabstractHigh performance provided by multi-processor System-on-Chips (MPSoCs) often induces designers to choose customized processors to execute specific functions rather than using dedicated hardware. On the other hand, reuse of pre-designed and pre-verified IP cores is the key strategy to meet time-to-market while at the same time reducing the error risk during the development of MPSoC designs. In this context, it becomes convenient to translate an existent RTL IP description, originally dedicated to implement an HW component, into pure SW code (i.e., C/C++) to be executed by one or more processors of the MPSoC. This work proposes a methodology to automatically generate SW code by abstracting RTL IP models implemented in hardware description language (HDL). The methodology exploits an abstraction algorithm to eliminate many implementation details typical of the HW descriptions, in order to improve the performance of the generated code. Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
DAC | 1 |
| 2010 | Model checking on TLM-2.0 IPs through automatic TLM-to-RTL synthesisabstractTransaction-level modeling (TLM) is the leading design style to deal with the increasing complexity of modern embedded systems. TLM provides designers with high-level interfaces and communication protocols for abstract modeling and efficient simulation of system platforms. The Open SystemC Initiative (OSCI) has recently released the TLM-2.0 standard for facilitating the interchange of models between suppliers and users, and thus encouraging the use of virtual platforms for fast simulation prior to the availability of register-transfer level (RTL) code. On the other hand, verification of TLM IPs still relies on simulation-based techniques since formal verification methodologies (such as model checking) and tools are not mature enough to be applied at TLM level. In this paper, we propose a methodology to apply existing RTL model checkers for verifying TLM IPs. The methodology relies on the automatic synthesis of the TLM IP models into equivalent RTL descriptions, in order to verify the TLM properties through the equivalent RTL model of the IPs. Nicola Bombieri, Franco Fummi, Valerio Guarnieri |
VLSI-SoC | 1 |
| 2010 | System/network design-space exploration based on TLM for networked embedded systemsabstractThis article presents a methodology for the design of Networked Embedded Systems (NESs), which extends Transaction Level Modeling (TLM) to perform system/network design-space exploration. As a result, a new design dimension is added to the traditional TLM refinement process to represent network configuration alternatives. Each network configuration can be used to drive both architecture exploration and system validation after each refinement step. A system/network simulation taxonomy is investigated aiming at precisely identifying the role of cosimulation in system/network design-space exploration. Furthermore, a general criterion to map functionalities to system and network models is presented. As a case study, the proposed methodology is applied to the design of a Voice-over-IP client. Nicola Bombieri, Franco Fummi, Davide Quaglia |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2009 | Functional qualification of TLM verificationabstractThe topic will cover the use of functional qualification for measuring the quality of functional verification of TLM models. Functional qualification is based on the theory of mutation analysis but considers a mutation to have been killed only if a test case fails. A mutation model of TLM behaviors is proposed to qualify a verification environment based on both testcases and assertions. The presentation describes at first the theoretic aspects of this topic and then it focuses on its application to real cases by using actual EDA tools, thus showing advantages and limitations of the application of mutation analysis to TLM. Nicola Bombieri, Franco Fummi, Graziano Pravadelli, Mark Hampton, Florian Letombe |
DATE | 1 |
| 2009 | Correct-by-construction generation of device drivers based on RTL testbenchesabstractThe generation of device drivers is a very time consuming and error prone activity. All the strategies proposed up to now to simplify this operation require a manual, even formal, specification of the device driver functionalities. In the system-level design, IP functionalities are tested by using testbenches, implemented to contain the communication protocols to correctly interact with the device. The aim of this paper is to present a methodology to automatically generate device drivers from the testbench of any RTL IP. The only manual step required is to tag the states corresponding to the different device functionalities. The Extended Finite State Machines (EFSMs) are then used to create a correct-by-construction two-level device driver: the lower level deals with architectural choices, while the higher one is derived from the EFSMs and it implements the communication protocols. The effectiveness of this methodology has been proved by applying it to a platform provided by STMicroelectronics. Nicola Bombieri, Franco Fummi, Graziano Pravadelli, Sara Vinco |
DATE | 1 |
| 2008 | Integrating RTL IPs into TLM Designs Through Automatic Transactor GenerationabstractTransaction Level Modeling (TLM) is an emerging design practice for overcoming increasing design complexity. It aims at simplifying the design flow of embedded systems by designing and verifying a system at different abstraction levels. In this context, transactors play a fundamental role since they allow communication between the system components, implemented at different abstraction levels. Reuse of RTL IPs into TLM systems is a meaningful example of key advantage guaranteed by exploiting transactors. Nevertheless, transactors implementation is still manual, tedious and error-prone, and the effort spent to verify their correctness often overcomes the benefits of the TLM-based design flow. In this paper we present a methodology to automatically generate transactors for RTL IPs. We show how the transactor code can be automatically generated by exploiting the testbench of any RTL IP. Nicola Bombieri, Nicola Deganello, Franco Fummi |
DATE | 1 |
| 2008 | A Mutation Model for the SystemC TLM 2.0 Communication InterfacesabstractMutation analysis is a widely-adopted strategy in software testing with two main purposes: measuring the quality of test suites, and identifying redundant code in programs. Similar approaches are applied in hardware verification and testing too, especially at RTL or gate level, where mutants are generally referred as faults, and mutation analysis is performed by means of fault modeling and fault simulation. However, in modern embedded systems there is a close integration between HW and SW parts, and verification strategies should be applied early in the design flow. This requires the definition of new mutation analysis-based strategies that work at system level, where HW and SW functionalities are not partitioned yet. In this context, the paper proposes a mutation model for perturbing transaction level modeling (TLM) SystemC descriptions. In particular, the main constructs provided by the SystemC TLM 2.0 library have been analyzed, and a set of mutants is proposed to perturb the primitives related to the TLM communication interfaces. Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
DATE | 1 |
| 2008 | Reuse and optimization of testbenches and properties in a TLM-to-RTL design flowabstractIn transaction-level modeling (TLM), verification methodologies based on transactions allow testbenches, properties, and IP cores in mixed TL-RTL designs to be reused. However, no papers in the literature analyze the effectiveness of transaction-based verification (TBV) in comparison to the more traditional RTL approach. The first contribution of this article is the introduction of a functional-fault-model-based methodology for demonstrating the effectiveness of reuse through TBV. A second contribution is the introduction of a similar methodology for efficient property checking which identifies and removes redundant properties prior to assertion-based verification or model checking. Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2007 | Incremental ABV for functional validation of TL-to-RTL design refinementabstractTransaction-level modeling (TLM) has been proposed as the leading strategy to address the always increasing complexity of digital systems. However, its introduction arouses a new challenge for designers and verification engineers, since there are no mature tools to automatically synthesize an RTL implementation from a transaction-level (TL) design, thus manual refinements are mandatory. In this context, the paper presents an incremental assertion-based verification (ABV) methodology to check the correctness of the TL-to-RTL refinement. The methodology relies on reusing assertions and already checked code, and it is guided by an assertion coverage metrics Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
DATE | 1 |
| 2007 | Towards Equivalence Checking Between TLM and RTL ModelsabstractThe always increasing complexity of digital system is overcome in design flows based on transaction level modeling (TLM) by designing and verifying the system at different abstraction levels. The design implementation starts from a TLM high-level description and, following a top- down approach, it is refined towards a corresponding RTL model. However, the bottom-up approach is also adopted in the design flow when already existing RTL IPs are abstracted to be reused into the TLM system. In this context, proving the equivalence between a model and its refined or abstracted version is still an open problem. In fact, traditional equivalence definitions and formal equivalence checking methodologies presented in the literature cannot be applied due to the very different internal characteristics of the models, including structure organization and timing. Targeting this topic, the paper presents a formal definition of equivalence based on events, and then, it shows how such a definition can be used for proving the equivalence in the RTL vs. TLM context, without requiring timing or structural similarities between the modules to be compared. Finally, the paper presents a practical use of the proposed theory, by proving the correctness of a methodology that automatically abstracts RTL IPs towards TLM implementations. Nicola Bombieri, Franco Fummi, Graziano Pravadelli, João Marques-Silva 0001 |
MEMOCODE | 1 |
| 2006 | On the evaluation of transactor-based verification for reusing TLM assertions and testbenches at RTLabstractTransaction level modeling (TLM) is becoming a usual practice for simplifying system-level design and architecture exploration. It allows the designers to focus on the functionality of the design, while abstracting away implementation details that will be added at lower abstraction levels. However, moving from transaction level to RTL requires redefining TLM test benches and assertions. Such a wasteful and error prone conversion can be avoided by adopting transactor-based verification (TBV). Many recent works adopt this strategy to propose verification methodologies that allow: (1) mixing TLM and RTL components; and (2) reusing TLM assertions and test benches at RTL. Even if practical advantages of such an approach are evident, there are no papers in the literature that evaluate the effectiveness of the TBV compared to a more traditional RTL verification strategy. This paper is intended to fill in the gap. It theoretically compares the quality of the TBV towards the rewriting of assertions and test benches at RTL with respect to both fault coverage and assertion coverage Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
DATE | 1 |
| 2006 | A methodology for abstracting RTL designs into TL descriptionsabstractTransaction-level modeling (TLM) has been proposed as the leading strategy to address the always increasing complexity of digital systems. However, modeling a complex system completely at transaction level (TL) could be inconvenient when IP cores are available on the market, usually modeled at RT level. In this context, modeling and verification methodologies based on transactors allow one to reuse RTL IP-cores in TL-RTL mixed designs, thus guaranteeing a considerable saving of time. Even if practical advantages of such an approach are evident, mixed TL-RTL designs cannot completely benefit from the well-known effectiveness provided by TLM. Thus, this paper proposes a methodology to abstract RTL IPs into corresponding TL descriptions. The possibility of automating the methodology is deeply analyzed. In particular, the paper shows which abstraction levels can be reached without requiring human intervention, according to the characteristics of the design to be refined. A set of experimental results are finally reported to confirm the effectiveness of the methodology Nicola Bombieri, Franco Fummi, Graziano Pravadelli |
MEMOCODE | 1 |
| 2005 | Extended abstract: on the property-based verification in SoC design flow founded on transaction level modelingabstractIn this paper we present the properties specification language (PSL) utilization in a system level verification flow for system on chip (SoC) designs, A compositional approach is proposed and two properties-based techniques are described and compared in terms of properties refinement effort and simulation speed delay. Nicola Bombieri, Andrea Fedeli, Franco Fummi |
MEMOCODE | 1 |