EDBT 2026 Demo / reviewers in the wild / expert
Biagio Peccerillo
dblp:206/0091
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-4998-0092ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DeVAS: Decoupled Virtual Address SpacesabstractThe constant growth of workload size in modern applications is making address translation a performance bottleneck. In principle, increasing the virtual page size could be advantageous, as it would allow each cached address translation to cover a larger memory space. Nevertheless, the utilization of larger pages introduces challenges, such as issues related to memory fragmentation and physical page management. In this paper, we present Decoupled Virtual Address Spaces (DeVAS), a virtual memory proposal that enables the decoupling of address translation and memory allocation, by allowing different sizes for virtual and physical pages, aiming to exploit the benefits of both. DeVAS introduces an intermediate virtual address space allocated by the Operating System employing larger pages (e.g., 2MiB), corresponding to the virtual page size seen by the processor, and a memory controller extension devoted to their allocation in physical memory at a smaller granularity (e.g., 4KiB). We show that DeVAS achieves an average 1.13 × performance improvement over a traditional configuration (i.e., with 4KiB pages) with no specific architectural modifications and for memory-intensive benchmarks. Moreover, DeVAS strategy enables architectural modifications for increasing performance, such as simplified/optimized TLB structure and L1-cache design flexibility. When considering these adjustments, DeVAS achieves a speedup of up to 1.20 × compared to the same reference. Furthermore, it matches and even edges the performance of an ideal (i.e., not implementable in practice) virtual memory configuration based on 2MiB pages only. Mirco Mannino, Biagio Peccerillo, Andrea Mondelli, Sandro Bartolini |
SBAC-PAD | 2 |
| 2023 | Energy and Performance Improvements for Convolutional Accelerators Using Lightweight Address Translation SupportabstractThe growing demand for deep learning applications has led to the design and development of several hardware accelerators to increase performance and energy efficiency. In particular, convolutional accelerators are among those receiving the most attention due to their applicability in many fields. Another aspect that is gaining increasing attention is the use of a shared virtual address space between processor and accelerators. It can provide several advantages such as programmability and security. The use of a shared address space relies on a time-consuming IOMMU to satisfy address translation requests. In this work, we analyze convolutional workloads in convolutional accelerators, identifying the sensitivity of performance to IOMMU activity. Additionally, based on the analysis done on convolutional workloads, we propose the use of dedicated accelerator registers (Translation Registers) to reduce costly IOMMU accesses. Translation Registers allow reducing execution time by about 20% and the energy consumption related to address translation up to about 55%. Mirco Mannino, Biagio Peccerillo, Andrea Mondelli, Sandro Bartolini |
CF | 2 |
| 2022 | Applying Intel's oneAPI to a machine learning case studyabstractAbstract Different technologies and approaches exist to work around the performance portability problem. Companies and academia work together to find a way to preserve performance across heterogeneous hardware using a unified language, one language to rule them all. Intel's oneAPI appears with this idea in mind. In this article, we try the new Intel solution to approach heterogeneous programming, choosing machine learning as our case study. More precisely, we choose Caffe, a machine learning framework that was created six years ago. Nevertheless, how would it be to make Caffe again from the beginning, using a fresh new technology like oneAPI? In terms of not only the ease of programming‐because only one source code would be needed to deploy Caffe to CPUs, GPUs, FPGAs, and accelerators (platforms that oneAPI currently supports)‐but also performance, where oneAPI may be capable of taking advantage of specific hardware automatically. Is Intel's oneAPI ready to take the leap? Pablo Antonio Martínez, Biagio Peccerillo, Sandro Bartolini, José M. García 0001, Gregorio Bernabé |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Flexible task-DAG management in PHAST library: Data-parallel tasks and orchestration support for heterogeneous systemsabstractSummary Heterogeneous architectures proved successful in achieving unprecedented performance and energy‐efficiency. However, taking advantage of these diverse processing elements is still hard. Programmers need to code through the different approaches suitable for each target architecture and need to decide the distribution of activities on the different resources. The majority of current frameworks focuses on either performance or productivity. The former mainly provides low‐level target‐specific programming interfaces, and the latter offers high‐level tools that often fail in achieving high‐performance. In both cases, the design is usually data‐parallel, as task‐parallelism is not supported. In this work, we propose a task‐based solution within the data‐parallel heterogeneous single‐source PHAST library. Tasks can be coded in a target‐agnostic fashion, can be compiled and parallelized on multi‐core CPUs and NVIDIA GPUs automatically and support the choice of the execution platform at runtime. We evaluate the capabilities of the proposed task‐directed acyclic graph support in case of an extensive set of randomly generated task‐based applications with different sizes and characteristics. We compare it against a SYCL implementation in terms of performance and complexity metrics, highlighting that PHAST achieves about 1.56× and 2.60× speedup over SYCL for multi‐core CPU and GPU, respectively, while improving also code complexity metrics. Biagio Peccerillo, Sandro Bartolini |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | A survey on hardware accelerators: Taxonomy, trends, challenges, and perspectivesabstractIn recent years, the limits of the multicore approach emerged in the so-called “dark silicon” issue and diminishing returns of an ever-increasing core count. Hardware manufacturers, out of necessity, switched their focus to accelerators, a new paradigm that pursues specialization and heterogeneity over generality and homogeneity. They are special-purpose hardware structures separated from the CPU with aspects that exhibit a high degree of variability. We define a taxonomy based on fourteen of these aspects, grouped in four macro-categories: general aspects, host coupling, architecture, and software aspects. According to it, we categorize around 100 accelerators of the last decade from both industry and academia, and critically analyze emerging trends. We complete our discussion with throughput and efficiency figures. Then, we discuss some prominent open challenges that accelerators are facing, analyzing state-of-the-art solutions, and suggesting prospective research directions for the future. Biagio Peccerillo, Mirco Mannino, Andrea Mondelli, Sandro Bartolini |
J. Syst. Archit. | 1 |
| 2019 | PHAST - A Portable High-Level Modern C++ Programming Library for GPUs and Multi-CoresabstractA decade after the beginning of the many-core era, multi-core CPU and GPU architectures are everywhere, from mobile devices up to high-performance workstations and servers. To this day, programmers willing to harness their power need to express their code via languages and frameworks that often lack of expressivity and high-level abstractions. These solutions, despite allowing users to reach unprecedented performance, can still be a hampering factor for productivity and portability. In this paper we propose PHAST, a modern C++, STL-like, single-source programming library and approach based on multi-dimensional dynamic containers and multi-layered functors that can be targeted on NVIDIA GPUs and multi-core CPUs. Its main purpose is to let programmers write code once for different architectures at a high level of abstraction, to reach high-performance while allowing fine parameter tuning and not shielding code from low-level target-specific optimizations. To assess the value of our proposal, we consider benchmarks from different application domains, and we evaluate their PHAST implementations against CUDA, OpenCL, Kokkos, and SYCL ones from both performance and productivity points of view. We show that PHAST can significantly reduce code complexity metrics while reaching very good performance. Biagio Peccerillo, Sandro Bartolini |
IEEE Trans. Parallel Distributed Syst. | 1 |