EDBT 2026 Demo / reviewers in the wild / expert
Giuliano Laccetti
dblp:92/842
· DBLP profile ↗
28ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0002-0057-2573ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 7 first-author · 5 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Massive Open Online Course (MOOC) on High-Performance Parallel Computing for Federica.eu Web-learning PlatformabstractMassive Open Online Courses (MOOCs) represent an accessible and user-friendly tool for disseminating innovative and cutting-edge topics to broad segments of civil society via online learning platforms, enabling users to learn at their own pace and on their own schedule. In this contribution, we describe the design and the implementation of a Massive Open Online Course on Parallel Computing and High-Performance Computing, developed for Federica Web Learning: the University Center for innovation, experimentation, and dissemination of multimedia teaching at the University of Naples Federico II. Giuliano Laccetti, Marco Lapegna, Ilaria Merciai |
PDP | 1 |
| 2023 | A novel approach for large-scale environmental data partitioning on cloud and on-premises storage for compute continuum applicationsabstractSummary Cloud‐based services have proved useful in several research fields, such as engineering, health science, and astrophysics, to mention a few examples. The computational environmental science community developed a strong need for cloud facilities to store, process, and manage data from observations and numerical models for simulations and forecasts. Weather forecast models and global sensor networks deal with multidimensional geo‐referenced data∖sets. However, environmental data consumer applications usually require a relatively small amount of multidimensional input data slice to analyze a specific area or time interval. Hence, reducing data dimension for information retrieval is mandatory. This paper presents a twofold solution: a technique to load and retrieve the sliced multidimensional data set on different cloud services such as Amazon Web Service (AWS), Google Cloud Platform, and Microsoft Azure. The experimental results performed on these cloud services highlight that the proposed method can significantly speed up the process of loading and retrieving the data slices compared to working with the entire data set in bulk or OPeNDAP server. Gennaro Mellone, Ciro Giuseppe De Vita, Dante D. Sánchez-Gallegos, Genaro Sanchez-Gallegos, Catherine Alessandra Torres Charles, Francisco Javier García Blas, Jesús Carretero 0001, José Luis González 0002, Giuliano Laccetti |
Concurr. Comput. Pract. Exp. | 9 |
| 2022 | Toward a high-performance clustering algorithm for securing edge computing environmentsabstractClustering algorithms are efficient tools for discov-ering correlations or affinities within large datasets and are the basis of several Machine Learning processes based on data generated by sensor networks. Recently, such algorithms have found an active application area closely correlated to the Edge Computing paradigm. The final aim is to transfer intelligence and decision-making ability near the edge of the networks to detect or prevent, as an example, attacks from insecure domains. In such a context, the present work introduces a new hybrid clustering algorithm for Edge Computing environments that can classify edge nodes taking into account their reliability. The algorithm is later evaluated from the points of view of the performance and energy consumption, comparing it with two high -end G PU - based computing systems. The achieved results confirm the possibility of designing intelligent sensors networks where decisions are taken at the data collection points. Giuliano Laccetti, Marco Lapegna, Raffaele Montella |
CCGRID | 1 |
| 2022 | Enabling the CUDA Unified Memory model in Edge, Cloud and HPC offloaded GPU kernelsabstractThe use of hardware accelerators, based on code and data offloading devoted to overcoming the CPU limitations in cores, is one of the main distinctive trends in high-end computing and related applications in the last decade. However, while code offloading is convenient for performance improvement, becoming a commonly used paradigm, memory access and management are a source of bottlenecks due to the need to interact with different address spaces. In this regard, NVidia introduced the CUDA Unified Memory model to avoid explicit memory copies between the machine hosting the accelerator device and the device itself and vice-versa. This paper shows a novel design and implementation of the support to the CUDA Unified Memory in open-source GPGPU virtualization services. The performance evaluation demonstrates that the overhead due to the virtualization and remoting is acceptable considering the possibility of sharing CUDA-enabled GPUs between various and heterogeneous machines hosted at the edge, in cloud infrastructures, or as accelerator nodes in an HPC scenario. A prototype implementation of the proposed solution is available as open-source. Raffaele Montella, Diana Di Luccio, Ciro Giuseppe De Vita, Gennaro Mellone, Marco Lapegna, Giuliano Laccetti, Sokol Kosta, Giulio Giunta |
CCGRID | 6 |
| 2022 | A hybrid clustering algorithm for high-performance edge computing devices [Short]abstractClustering algorithms are efficient tools for discovering correlations or affinities within large datasets and are the basis of several Artificial Intelligence processes based on data generated by sensor networks. Recently, such algorithms have found an active application area closely correlated to the Edge Computing paradigm. The final aim is to transfer intelligence and decision-making ability near the edge of the sensors networks, thus avoiding the stringent requests for low-latency and large-bandwidth networks typical of the Cloud Computing model. In such a context, the present work describes a new hybrid version of a clustering algorithm for the NVIDIA Jetson Nano board by integrating two different parallel strategies. The algorithm is later evaluated from the points of view of the performance and energy consumption, comparing it with two high-end GPU-based computing systems. The results confirm the possibility of creating intelligent sensor networks where decisions are taken at the data collection points. Giuliano Laccetti, Marco Lapegna, Diego Romano |
ISPDC | 1 |
| 2021 | Toward a multilevel scalable parallel Zielonka's algorithm for solving parity gamesabstractSummary In this work, we perform the feasibility analysis of a multi‐grained parallel version of the Zielonka Recursive (ZR) algorithm exploiting the coarse‐ and fine‐ grained concurrency. Coarse‐grained parallelism relies on a suitable splitting of the problem, that is, a graph decomposition based on its Strongly Connected Components (SCC) or a splitting of the formula generating the game, while fine‐grained parallelism is introduced inside the Attractor which is the most intensive computational kernel. This configuration is new and addressed for the first time in this article. Innovation goes from the introduction of properly defined metrics for the strong and weak scaling of the algorithm. These metrics conduct to an analysis of the values of these metrics for the fine grained algorithm, we can infer the expected performance of the multi‐grained parallel algorithm running in a distributed and hybrid computing environment. Results confirm that while a fine‐grained parallelism have a clear performance limitation, the performance gain we can expect to get by employing a multilevel parallelism is significant. Luisa D'Amore, Aniello Murano, Loredana Sorrentino, Rossella Arcucci, Giuliano Laccetti |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | Special Issue on High-end Heterogeneous Architectures, Methodologies, and Algorithms (HHAMA20)abstractTEST 02 - Elsevier's Scopus, the largest abstract and citation database of peer-reviewed literature. Search and access research from the science, technology, medicine, social sciences and arts and humanities fields. Sokol Kosta, Giuliano Laccetti, Marco Lapegna, Valeria Mele, Raffaele Montella |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Designing a GPU-parallel algorithm for raw SAR data compression: A focus on parallel performance estimation
Diego Romano, Marco Lapegna, Valeria Mele, Giuliano Laccetti |
Future Gener. Comput. Syst. | 4 |
| 2020 | Performance enhancement of a dynamic K-means algorithm through a parallel adaptive strategy on multicore CPUs
Giuliano Laccetti, Marco Lapegna, Valeria Mele, Diego Romano, Lukasz Szustak |
J. Parallel Distributed Comput. | 1 |
| 2019 | An adaptive algorithm for high-dimensional integrals on heterogeneous CPU-GPU systemsabstractSummary In this paper, we introduce an adaptive procedure for the numerical computation of a high‐dimensional integrals on HPC systems with heterogeneous nodes composed of multi‐core CPU and GPU devices. To this aim, we have integrated together two different approaches: a first one is in charge of a fair workload among the threads running on the multi‐core CPU, while a second one is in charge of an efficient execution of the computational kernels on the GPU. We tested the resulting algorithm on several test functions on a system where the nodes are provided with two Intel ten‐core CPU and one NVIDIA GPU device. Giuliano Laccetti, Marco Lapegna, Valeria Mele, Raffaele Montella |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | A Scalable Unified Model for Dynamic Data Structures in Message Passing (Clusters) and Shared Memory (multicore CPUs) Computing environmentsabstractConcurrent data structures are widely used in many software stack levels, ranging from high level parallel scientific applications to low level operating systems. The key issue of these objects is their concurrent use by several computing units (threads or process) so that the design of these structures is much more difficult compared to their sequential counterpart, because of their extremely dynamic nature requiring protocols to ensure data consistency, with a significant cost overhead. At this regard, several studies emphasize a tension between the needs of sequential correctness of the concurrent data structures and scalability of the algorithms, and in many cases it is evident the need to rethink the data structure design, using approaches based on randomization and/or redistribution techniques in order to fully exploit the computational power of the recent computing environments. The problem is grown in importance with the new generation High Performance Computing systems aimed to achieve extreme performance. It is easy to observe that such systems are based on heterogeneous architectures integrating several independent nodes in the form of clusters or MPP systems, where each node is composed by powerful computing elements (CPU core, GPUs or other acceleration devices) sharing resources in a single node. These systems therefore make massive use of communication libraries to exchange data among the nodes, as well as other tools for the management of the shared resources inside a single node. For such a reason, the development of algorithms and scientific software for dynamic data structures on these heterogeneous systems implies a suitable combination of several methodologies and tools to deal with the different kinds of parallelism corresponding to each specific device, so that to be aware of the underlying platform. The present work is aimed to introduce a scalable model to manage a special class of dynamic data structure known as heap based priority queue (or simply heap) on these heterogeneous architectures. A heap is generally used when the applications needs set of data not requiring a complete ordering, but only the access to some items tagged with high priority. In order to ensure a tradeoff between the correct access to high priority items by the several computing units with a low communication and synchronization overhead, a suitable reorganization of the heap is needed. More precisely we introduce a unified scalable model that can be used, with no modifications, to redeploy the items of a heap both in message passing environments (such as clusters and or MMP multicomputers with several nodes) as well as in shared memory environments (such as CPUs and multiprocessors with several cores) with an overhead independent of the number of computing units. Computational results related to the application of the proposed strategy on some numerical case studies are presented for different types of computing environments. Giuliano Laccetti, Marco Lapegna, Raffaele Montella |
CCGrid | 1 |
| 2018 | Models, algorithms, and tools for highly heterogeneous computing environmentsabstractModels, algorithms, and tools for highly heterogeneous computing environments Giuliano Laccetti, Marco Lapegna, Raffaele Montella, Sokol Kosta |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Marine bathymetry processing through GPGPU virtualization in high performance cloud computingabstractSummary Fast technology development has influenced the widespread use of low‐power devices in different scientific, environmental, and everyday life areas, giving birth to the Internet of Things. In this paper, we focus on the context of marine studies, addressing the problem of marine bathymetry data processing and analysis via pervasive and Internet‐connected sensors and low‐power distributed devices. Pervasive and Internet‐connected low‐power devices (as the components involved in the sensing and processing actions) made diverse and different “things” as a worldwide‐distributed system. Given the high complexity of the algorithms involved in these studies, which usually involve general‐purpose graphic processing unit (GPGPU) computation, it is impossible for the limited devices to perform the required calculations. To overcome these limitations, in this paper, we propose and implement a vertical application of GVirtuS, the open‐source GPGPU virtualization and remoting service, for achieving high performance geographical data interpolation in a high performance cloud computing scenario. We present an innovative implementation by comparing, in terms of performance and accuracy, the inverse distance weighting and kriging interpolation methods in their parallel implementations leveraging on CUDA‐enabled GPGPUs. We present a real‐world use case related to high‐resolution bathymetry interpolation in a crowdsource data context in the Bay of Pozzuoli, Italy. Raffaele Montella, Livia Marcellino, Ardelio Galletti, Diana Di Luccio, Sokol Kosta, Giuliano Laccetti, Giulio Giunta |
Concurr. Comput. Pract. Exp. | 6 |
| 2017 | Accelerating Linux and Android applications on low-power devices through remote GPGPU offloadingabstractSummary Low‐power devices are usually highly constrained in terms of CPU computing power, memory, and GPGPU resources for real‐time applications to run. In this paper, we describe RAPID, a complete framework suite for computation offloading to help low‐powered devices overcome these limitations. RAPID supports CPU and GPGPU computation offloading on Linux and Android devices. Moreover, the framework implements lightweight secure data transmission of the offloading operations. We present the architecture of the framework, showing the integration of the CPU and GPGPU offloading modules. We show by extensive experiments that the overhead introduced by the security layer is negligible. We present the first benchmark results showing that Java/Android GPGPU code offloading is possible. Finally, we show the adoption of the GPGPU offloading into BioSurveillance, a commercial real‐time face recognition application. The results show that, thanks to RAPID, BioSurveillance is being successfully adapted to run on low‐power devices. The proposed framework is highly modular and exposes a rich application programming interface to developers, making it highly versatile while hiding the complexity of the underlying networking layer. Raffaele Montella, Sokol Kosta, David Oro, Javier Vera, Carles Fernández, Carlo Palmieri, Diana Di Luccio, Giulio Giunta, Marco Lapegna, Giuliano Laccetti |
Concurr. Comput. Pract. Exp. | 10 |
| 2012 | Modelling the Behaviour of an Adaptive Scheduling ControllerabstractThe deployment, management and total cost of ownership of large computing environments always involve huge investments. These systems, once in production, have to meet the needs of users belonging to large and heterogeneous communities: only an efficient and effective use of these systems can repay the investment made. The heterogeneity of user communities implies that computational resources are used for different type of applications, traditional (sequential) or HPC (MPI and Open MP based), whose demands are often conflicting. In this document we report experiences in designing, implementing and validating an adaptive scheduling controller (ASC) that, by using an "adaptive" approach in scheduling policy, allows a balanced, effective and efficient use of computational resources. Giovanni Battista Barone, Vania Boccia, Davide Bottalico, Luisa Carracciuolo, Alessandra Doria, Giuliano Laccetti |
CISIS | 6 |
| 2010 | A multi-grained distributed implementation of the parallel Block Conjugate Gradient algorithmabstractAbstract The Block Conjugate Gradient algorithm (Block‐CG) was developed to solve sparse linear systems of equations that have multiple right‐hand sides. We have adapted it for use in heterogeneous, geographically distributed, parallel architectures. Once the main operations of the Block‐CG (Tasks) have been collected into smaller groups (subjobs), each subjob is matched by the middleware MJMS (MPI Jobs Management System) with a suitable resource selected among those which are available. Moreover, within each subjob, concurrency is introduced at two different levels and with two different granularities: the coarse‐grained parallelism to perform independent tasks and the fine‐grained parallelism within the execution of a task. We refer to this algorithm as to multi‐grained distributed implementation of the parallel Block‐CG. We compare the performance of a parallel implementation with the one of the distributed implementation running on a variety of Grid computing environments. The middleware MJMS—developed by some of the authors and built on top of Globus Toolkit and Condor‐G—was used for co‐allocation, synchronization, scheduling and resource selection. Copyright © 2010 John Wiley & Sons, Ltd. Almerico Murli, Luisa D'Amore, Giuliano Laccetti, Francesco Gregoretti, Gennaro Oliva |
Concurr. Comput. Pract. Exp. | 3 |
| 2009 | A fusion-based approach to digital movie restoration
Lucia Maddalena 0001, Alfredo Petrosino, Giuliano Laccetti |
Pattern Recognit. | 3 |
| 2008 | Five Dimension Environmental Data Resource Brokering on Computational Grids and Scientific CloudsabstractIn this paper we describe how the grid computing resource brokering approach, classically divided into matchmaking discovery and optimize selection, is applied to five dimensional environmental data distribution. This result is achieved thanks to the integration of a Resource Broker Service we developed from scratch. This service implements the key feature of autonomic mapping of Globus Toolkit Index Service resources on the Condor ClassAd resource description. Our Five Dimensional Distribution Data Service relays on this software infrastructure to advertise metadata and it provides the requested data leveraging on the web service resource framework. We provide an example based on a grid aware application component, showing the application of resource brokering to environmental problems. We focused our interests mainly about high resolution weather forecast data management and results quality assessment. Our approach would be both modular and distributed computing technology aware in order to provide the e-Science community of grid based cloud computing enabled tools fully benefiting of this technology. Giulio Giunta, Giuliano Laccetti, Raffaele Montella |
APSCC | 2 |
| 2008 | The MedIGrid PSE in an LCG/gLite environmentabstractIn this paper we are concerned with improvements and enhancements of a medical imaging grid-enabled infrastructure, named MedIGrid, oriented to the transparent use of resource-intensive applications for managing, processing and visualizing biomedical images. We describe an implementation of the MedIGrid PSE in an LCG/gLite environment. We’ll mainly focus on how to exploit the features of the new middleware environment to improve the efficiency and the services reliability of the PSE; further, some comments will be devoted to how to modify, extend and/or improve the underlying numerical components. Almerico Murli, Vania Boccia, Luisa Carracciuolo, Luisa D'Amore, Giuliano Laccetti, Marco Lapegna |
ISPA | 5 |
| 2008 | MGF: A grid-enabled MPI library
Francesco Gregoretti, Giuliano Laccetti, Almerico Murli, Gennaro Oliva, Umberto Scafuri |
Future Gener. Comput. Syst. | 2 |
| 2007 | A framework model for grid security
Giuliano Laccetti, Giovanni Schmid |
Future Gener. Comput. Syst. | 1 |
| 2004 | Integrating Scientific Software Libraries in Problem Solving Environments: A Case Study with ScaLAPACK
Luisa D'Amore, Mario Rosario Guarracino, Giuliano Laccetti, Almerico Murli |
ICCSA (2) | 3 |
| 2004 | Parallel/Distributed Film Line Scratch Restoration by Fusion Techniques
Giuliano Laccetti, Lucia Maddalena 0001, Alfredo Petrosino |
ICCSA (2) | 1 |
| 2000 | Yabat: An Administration and Monitoring Software for Beowulf Clusters
Mario Rosario Guarracino, Giuliano Laccetti, Gennaro Oliva, Umberto Scafuri |
CLUSTER | 2 |
| 2000 | Browsing Virtual Reality on a PC ClusterabstractVRML is the standard language to build a virtual reality system, i.e. a language to describe 3D objects and interactive worlds. To use such a system, among the other tools, a browser is needed in order to access the information in the VRML files and to visualize the virtual world from a human perspective. This paper describes the implementation of a VRML parallel browser based on a modified version of graphics libraries, such as OpenGL and GLUT. The obtained interactive parallel rendering engine has been developed for a message passing computing environment; a detailed analysis of its performance on a PC cluster is also described. Mario Rosario Guarracino, Giuliano Laccetti, Diego Romano |
CLUSTER | 2 |
| 1999 | PAMIHR. A Parallel FORTRAN Program for Multidimensional Quadrature on Distributed Memory Architectures
Giuliano Laccetti, Marco Lapegna |
Euro-Par | 1 |
| 1999 | An implementation of a Fourier series method for the numerical inversion of the Laplace transformabstractOur method is based on the numerical evaluation of the integral which occurs in the Riemann Inversion formula. The trapezoidal rule approximation to this integral reduces to a Fourier series. We analyze the corresponding discretization error and demonstrate how this expression can be used in the development of an automatic routine , one in which the user needs to specify only the required accuracy. Luisa D'Amore, Giuliano Laccetti, Almerico Murli |
ACM Trans. Math. Softw. | 2 |
| 1999 | Algorithm 796: a Fortran software package for the numerical inversion of the Laplace transform based on a Fourier series methodabstractA software package for the numerical inversion of a Laplace Transform function is described. Besides function values of F ( z ) for complex and real z , the user has only to provide the numerical value of the Laplace convergence abscissa σ 0 or, failing this, an upper bound to this quantity, and the accuracy he or she requires in the computed value of the inverse Transform. The method implemented is based on a Fourier series expansion of the inverse transform, and it is especially suitable when such inverse Laplace Transform is sectionally continuous. Luisa D'Amore, Giuliano Laccetti, Almerico Murli |
ACM Trans. Math. Softw. | 2 |