Raffaele Montella

dblp:15/6533 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-4767-2045ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Virtualized GPU Offloading for AI Inference: Enabling Deep Learning Frameworks on GPU-less Devices
Wenrui Yu, Theodoros Aslanidis, Darshan Kadiluvagilu Thimme Gowda, Dimitris Chatzopoulos, Raffaele Montella, Sokol Kosta
ICC6
2026 Streaming I/O for scientific workflow engine acceleration
abstract
Scientific workflows are increasingly characterized by complex task dependencies and large-scale data exchanges, which place significant pressure on the input/output (I/O) systems of traditional Workflow Engines (WFEs). These challenges are particularly evident in data-intensive and real-time processing contexts, where conventional disk-based I/O mechanisms often become performance bottlenecks. This paper presents an approach to enhancing the DAGonStar scientific workflow engine by integrating CAPIO, a middleware designed to support memory-based streaming I/O. The integration combines DAGonStar’s orchestration capabilities with CAPIO’s efficient data handling to better support workflows operating on continuous or large-scale datasets. We describe the architectural modifications introduced to enable this collaboration and provide an analysis of the resulting system. The proposed solution aims to improve the responsiveness and flexibility of scientific workflows by streamlining data transfers and simplifying task coordination. This work contributes to the evolution of workflow systems toward more efficient and scalable models for scientific computing. • Integration of memory-based streaming I/O into scientific workflow engines. • Automated generation of synchronization rules through workflow dependency analysis of DAGonStar. • Enhanced pipeline’s tasks execution efficiency via system call interception through the usage of CAPIO. • Benchmark evaluation showing up to 33% reduction in execution time with DAGonCAPIO. • Support for both local batch and SLURM-based distributed executions.
Simone Perrotta, Ciro Giuseppe De Vita, Gennaro Mellone, Marco Edoardo Santimaria, Massimo Torquati, Francisco Javier García Blas, Raffaele Montella
Future Gener. Comput. Syst.7
2026 A coupled Lagrangian-AI hierarchical and heterogeneous model for predicting bacteria contamination in farmed mussels
Ciro Giuseppe De Vita, Gennaro Mellone, Diana Di Luccio, Francisco Javier García Blas, Francesca Barchiesi, Raffaele Montella
Future Gener. Comput. Syst.6
2025 Preserving and improving the legacy of eScience: the GLOBO experience
abstract
Legacy scientific codes are essential to long-standing operational workflows but often cannot fully exploit modern heterogeneous high-performance computing (HPC) architectures. GLOBO, a global atmospheric circulation model developed at CNR-ISAC, is one such model that is still used for research and operational weather forecasting. In this work, we present an ongoing refactoring of GLOBO to improve its scalability and prepare it for GPU acceleration and cloud execution. The modernization focuses on three main areas: dynamic memory allocation, replacement of legacy point-to-point MPI communications with optimized collective operations, and OpenACC-based GPU offloading. We describe the redesign of the communication strategy, which replaces multiple MPI_Isend/MPI_Irecv exchanges with collective primitives (MPI_Bcast, MPI_Scatterv, MPI_Sendrecv, MPI_Reduce) to reduce overhead and improve parallel efficiency. Scalability experiments across multiple spatial resolutions show that a new collective communication strategy consistently matches or exceeds the original implementation, delivering up to 40% runtime reduction in high-resolution, multi-node configurations. These results lay the groundwork for refactoring GLOBO toward a fully dynamic, GPU-accelerated, cloud-ready execution, preserving its scientific heritage while ensuring sustainable performance on future HPC platforms.
Carmine Coppola, Guido Davoli, Federico Fabiano, Ciro Giuseppe De Vita, Diana Di Luccio, Pasquale Corvino, Antonella Pirozzi, Andrea Alessandri, Raffaele Montella
eScience9
2025 AI and HPC for intense rain event early warning leveraging real-time weather radar
abstract
Global changes are increasing the frequency and intensity of extreme weather events, posing challenges for forecasting localized phenomena with sub-grid resolution. The Hi-WeFAI project addresses this by combining high-performance computing, Federated Artificial Intelligence, and heterogeneous sensor networks to improve short-term precipitation forecasting and flood nowcasting. In this paper, we present preliminary results using a transformer-based radar prediction model coupled with a flood model to generate high-resolution early warning maps. The results from the Naples pilot site improved accuracy and detail, highlighting the potential of the hybrid AI and HPC approach to support quasi-real-time decision-making and disaster risk reduction.
Diana Di Luccio, Ciro Giuseppe De Vita, Gennaro Mellone, Dante D. Sánchez-Gallegos, Pasquale Corvino, Mario Di Sarno, Pasquale De Luca, Emanuel Di Nardo, Vincenzo Capozzi, Vincenzo Bucciero, Raffaele Montella
eScience11
2025 Directed Acyclic Graph on Cross-Application Programmable I/O: Adding streaming flavour to scientific workflows
abstract
This paper introduces DAGonCAPIO, a workflow framework that integrates the DAGonStar engine with the CAPIO middleware to enable I/O streaming in scientific workflows. DAGonStar utilizes a Directed Acyclic Graph (DAG) to orchestrate tasks. At the same time, CAPIO enables downstream tasks to process data as soon as partial outputs become available, without requiring modifications to application code. This integration reduces delays from traditional file-based communication. The system uses the workflow:// schema to define data dependencies and generate CAPIO coordination scripts. DAGonStar was modified to support early scratch directory naming and decoupled task execution. Experiments with a WRF-based weather forecasting workflow on a 256-core cluster demonstrate that DAGonCAPIO reduces time-to-first-result by up to 4171 seconds, achieving a nearly 10x speedup.
Simone Perrotta, Marco Edoardo Santimaria, Ciro Giuseppe De Vita, Massimo Torquati, Diana Di Luccio, Pasquale Corvino, Antonella Pirozzi, Raffaele Montella
eScience8
2025 G-Litter Marine Litter Dataset Augmentation with Diffusion Models and Large Language Models on GPU Acceleration
abstract
Marine litter detection is crucial for environmental monitoring, yet the imbalance in existing datasets limits model performance in identifying various types of waste accurately. This paper presents an efficient data augmentation pipeline that combines generative diffusion models (e.g., Stable Diffusion) and Large Language Models (LLMs) to expand the G-Litter dataset, a marine litter dataset designed for autonomous detection in heterogeneous environments. Leveraging scalable diffusion models for image generation and Alpaca LLMs for diverse prompt generation, our approach augments underrepresented classes by generating over 200 additional images per class, significantly improving the dataset’s balance. Training G-Litter augmented dataset using YOLOv8 for object detection demonstrated an increase in detection performance, improving recall by 7.82% and mAP50 by 3.87% (compared with baseline results). This study emphasizes the potential for combining generative AI with HPC resources to automate data augmentation on large-scale, unstructured datasets, particularly in edge computing contexts for real-time marine monitoring. The models were tested on real videos captured during simulated missions, demonstrating a superior ability to detect submerged objects in dynamic scenarios. These results highlight the potential of generative AI techniques to improve dataset quality and detection model performance, laying the foundation for further expansion in real-time marine monitoring.
Gennaro Mellone, Ciro Giuseppe De Vita, Emanuel Di Nardo, Giuseppe Coviello, Diana Di Luccio, Pietro Aucelli, Angelo Ciaramella, Raffaele Montella
PDP8
2024 Improving Real-Time Data Streams Performance on Autonomous Surface Vehicles using DataX
abstract
In the evolving Artificial Intelligence (AI) era, the need for real-time algorithm processing in marine edge en-vironments has become a crucial challenge. Data acquisition, analysis, and processing in complex marine situations require sophisticated and highly efficient platforms. This study optimizes real-time operations on a containerized distributed processing platform designed for Autonomous Surface Vehicles (ASV) to help safeguard the marine environment. The primary objective is to improve the efficiency and speed of data processing by adopting a microservice management system called DataX. DataX leverages containerization to break down operations into modular units, and resource coordination is based on Kubernetes. This combination of technologies enables more efficient resource management and real-time operations optimization, contributing significantly to the success of marine missions. The platform was developed to address the unique challenges of managing data and running advanced algorithms in a marine context, which often involves limited connectivity, high latencies, and energy restrictions. Finally, as a proof of concept to justify this platform's evolution, experiments were carried out using a cluster of single-board computers equipped with GPUs, running an AI-based marine litter detection application and demonstrating the tangible benefits of this solution and its suitability for the needs of maritime missions.
Gennaro Mellone, Ciro Giuseppe De Vita, Giuseppe Coviello, Pietro Aucelli, Angelo Ciaramella, Raffaele Montella
PDP6
2024 Special Issue on the pervasive nature of HPC (PN-HPC)
abstract
Summary This special issue on the Pervasive Nature of HPC (PN‐HPC) collects an extension of the most valuable works presented at the sixth Workshop on Models, Algorithms and Methodologies for Hybrid Parallelism in New HPC Systems (MAMHYP‐22), held in Gdansk (Poland) in September 2022, jointly with the 14th conference on Parallel Processing and Applied Mathematics (PPAM‐22). New original papers related to the workshop themes are also included. The final aim is to provide a glimpse of the current state of knowledge related to the development of efficient methodologies and algorithms for HPC systems with multiple forms of parallelism.
Marco Lapegna, Valeria Mele, Raffaele Montella, Lukasz Szustak
Concurr. Comput. Pract. Exp.3
2024 A high-performance, parallel, and hierarchically distributed model for coastal run-up events simulation and forecasting
abstract
Abstract The request for quickly available forecasts of intense weather and marine events impacting coastal areas is gradually increasing. High-performance computing (HPC) and artificial intelligence techniques are crucial in this application. Risk mitigation and coastal management must design scientific workflow appropriately and maintain them continuously updated and operational. Climate change accelerating increase trend of the past decades impacted on sea-level rise, together with broader factors such as geostatic effects and subsidence, reducing the effectiveness of coastal defenses. Due to this, the support tools, such as Early Warning Systems, have become increasingly more valuable because they can process data promptly and provide valuable indications for mitigation proposals. We developed the Shoreline Alert Model (SAM), an operational Python tool that produces simulation scenarios, ‘what-if’ assumptions, and coastal flooding forecasts to fill this gap in our study area. SAM aims to provide decision-makers, scientists, and engineers with new tools to help forecast significant weather-marine events and support related management or emergency responses. SAM aims to fill the gap between the wind-driven wave models, which produce simulations and forecasts of waves of significant height, period, and direction in deep or mid-water, and the run-up local models, which exstimulate marine ingression in the event of intense weather phenomena. It employs a parallelization scheme that allows users to run it on heterogeneous parallel architectures. It produced results approximately 24 times faster than the baseline when using shared memory with distributed memory, processing roughly 20,000 coastal cross-shore profiles along the coastline of the Campania region (Italy). Increasing the performance of this model and, at the same time, honoring the need for relatively modest HPC resources will enable the local manager and policymakers to enforce fast and effective responses to intense weather phenomena.
Diana Di Luccio, Ciro Giuseppe De Vita, Aniello Florio, Gennaro Mellone, Catherine Alessandra Torres Charles, Guido Benassai, Raffaele Montella
J. Supercomput.7
2023 A containerized distributed processing platform for autonomous surface vehicles: preliminary results for marine litter detection
abstract
Autonomous Surface Vehicles and their management represent one of the significant challenges in coastal and offshore surveying. Although the development of this kind of data acquisition device has skyrocketed in the last few years, line guides and technological solutions still need to come. On the other hand, this kind of robotic vessel's true potential has yet to be explored. This paper presents ArgonautAI, a containerized distributed processing platform for autonomous surface vehicles. The proposed ArgonautAI architecture leverage a cluster of single-board computers with diverse and different characteristics (computing power, CUDA GPUs, FPGAs, GPIOs, PWMs, specialized I/O) orchestrated using Kubernetes and a customized programming interface. Furthermore, the proposed solution introduces two different types of containers: 1) the platform containers hosting the software life support for the platform and 2) the mission containers defined to support the survey mission-specific scopes. The firsts manage the vehicle's instruments (e.g. position, attitude, environment, depth), the data storage, the vessel-to-shore communication, and so on; the latter host mission-specific software components. Finally, as proof of concept of the proposed platform, we present an AI-based marine litter detection application using a hierarchical computer vision approach on heterogenic onboard computing resources.
Gennaro Mellone, Ciro Giuseppe De Vita, Dante D. Sánchez-Gallegos, Diana Di Luccio, Gaia Mattei, Francesco Peluso, Pietro Aucelli, Angelo Ciaramella, Raffaele Montella
PDP9
2023 Message from the Organizing Committee Chairs: PDP 2023
abstract
On behalf of the Organizing Committee, we welcome you to the 31st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (PDP2023), organized by the Department of Science and Technology of the University of Naples “Parthenope”. The conference was hosted in Naples in the prestigious Villa Doria d'Angri from the 1st to the 3rd of March 2023.
Raffaele Montella, Angelo Ciaramella, Marco Lapegna, Marco Danelutto, Dora Blanco Heras
PDP1
2023 Message from the General Chairs: PDP 2023
abstract
Welcome to the 31st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing (PDP2023).
Raffaele Montella, Angelo Ciaramella, Marco Lapegna, Marco Danelutto, Dora Blanco Heras
PDP1
2023 A highly scalable high-performance Lagrangian transport and diffusion model for marine pollutants assessment
abstract
While using High-Performance Computing (HPC) for precise and accurate air quality forecasts is a common issue, similar services devoted to marine pollution in coastal areas remain challenging. This paper presents Water quality Community Model Plus Plus (WaComM++) leveraging a parallelization schema enabling the users to run it on heterogeneous parallel architectures. We evaluated the proposed model under several execution approaches using a real-world application for pollutants forecast in the Gulf of Napoli (Campania, Italy). As a result, WaComM++ has produced results 657K times faster than the sequential run (taking into account the Particles' Outer Cycle and not considering the particle domain distribution) when using distributed and shared memory with multi-GPUs dealing with about 25 million particles.
Raffaele Montella, Diana Di Luccio, Ciro Giuseppe De Vita, Gennaro Mellone, Marco Lapegna, Gloria Ortega, Livia Marcellino, Enrico Zambianchi, Giulio Giunta
PDP1
2023 Parallel and hierarchically-distributed Shoreline Alert Model (SAM)
abstract
In this paper, the Shoreline Alert Model (SAM) is presented as a component of a computation platform based on workflows dedicated to extreme weather/marine event simulation. The model aims to mitigate the effects of global change by providing decision-makers, scientists, and engineers with a novel, next-generation tool set for facing extreme weather events and implementing related management or emergency responses. SAM uses a parallelization schema, allowing users to run it on heterogeneous parallel architectures. As a result, SAM produces approximately 24 times faster results than the baseline when using shared memory with distributed memory and dealing with about 20,000 transects along the Campania coastline. The system is based on the algorithms of the open-source numerical models WRF (Weather Research and Forecasting) and WW3 (Wave-watch III) implemented with refraction and shoaling routines together with run-up equations to form the modeling chain used for coastal flooding assessment.
Ciro Giuseppe De Vita, Gennaro Mellone, Aniello Florio, Catherine Alessandra Torres Charles, Diana Di Luccio, Marco Lapegna, Guido Benassai, Giorgio Budillon, Raffaele Montella
PDP9
2023 PuzzleMesh: A Puzzle Model to Build Mesh of Agnostic Services for Edge-Fog-Cloud
abstract
This paper presents the design, development, and evaluation of PuzzleMesh, an agnostic service mesh composition model to process large volumes of data in edge-fog-cloud environments. This model is based on a puzzle metaphor where pieces, puzzles, and metapuzzles represent self-contained autonomous and reusable software artifacts encapsulated into containers and published as microservices. Apiecerepresents the integration of apps with I/O interfaces (loops/sockets), parallel processing, and management software. Apuzzlerepresents a processing structure (e.g., workflows) built coupling pieces through loops and sockets. Puzzles integrate structures with a microservice architecture, implicit continuous dataflows, and transparent data exchange management software. Ametapuzzlerepresents a recursive assemble of puzzles. A mesh represents a pool of pieces, puzzles, and metapuzzles available for designers to choose artifacts to build services. A prototype developed using PuzzleMesh model was evaluated through case studies about the automatic construction of processing services for the acquisition, pre-processing, manufacturing, preserving, and visualizing of satellite imagery. A qualitative comparison revealed that PuzzleMesh provides a flexible way to build reusable and portable services and to improve the usability of the services. The case study also revealed that PuzzleMesh yielded better performance results than other state-of-the-art tools.
Dante D. Sánchez-Gallegos, José Luis González 0002, Jesús Carretero 0001, Heidy Marisol Marín-Castro, Andrei Tchernykh, Raffaele Montella
IEEE Trans. Serv. Comput.6
2022 Toward a high-performance clustering algorithm for securing edge computing environments
abstract
Clustering algorithms are efficient tools for discov-ering correlations or affinities within large datasets and are the basis of several Machine Learning processes based on data generated by sensor networks. Recently, such algorithms have found an active application area closely correlated to the Edge Computing paradigm. The final aim is to transfer intelligence and decision-making ability near the edge of the networks to detect or prevent, as an example, attacks from insecure domains. In such a context, the present work introduces a new hybrid clustering algorithm for Edge Computing environments that can classify edge nodes taking into account their reliability. The algorithm is later evaluated from the points of view of the performance and energy consumption, comparing it with two high -end G PU - based computing systems. The achieved results confirm the possibility of designing intelligent sensors networks where decisions are taken at the data collection points.
Giuliano Laccetti, Marco Lapegna, Raffaele Montella
CCGRID3
2022 Enabling the CUDA Unified Memory model in Edge, Cloud and HPC offloaded GPU kernels
abstract
The use of hardware accelerators, based on code and data offloading devoted to overcoming the CPU limitations in cores, is one of the main distinctive trends in high-end computing and related applications in the last decade. However, while code offloading is convenient for performance improvement, becoming a commonly used paradigm, memory access and management are a source of bottlenecks due to the need to interact with different address spaces. In this regard, NVidia introduced the CUDA Unified Memory model to avoid explicit memory copies between the machine hosting the accelerator device and the device itself and vice-versa. This paper shows a novel design and implementation of the support to the CUDA Unified Memory in open-source GPGPU virtualization services. The performance evaluation demonstrates that the overhead due to the virtualization and remoting is acceptable considering the possibility of sharing CUDA-enabled GPUs between various and heterogeneous machines hosted at the edge, in cloud infrastructures, or as accelerator nodes in an HPC scenario. A prototype implementation of the proposed solution is available as open-source.
Raffaele Montella, Diana Di Luccio, Ciro Giuseppe De Vita, Gennaro Mellone, Marco Lapegna, Giuliano Laccetti, Sokol Kosta, Giulio Giunta
CCGRID1
2022 AIQUAM: Artificial Intelligence-based water QUAlity Model
abstract
Monitoring the impact of the pollutants on the sea is a crucial issue for coastal human activities, such as aquaculture. However, leveraging a continuous microbiological laboratory analysis is unfeasible for costs and practical reasons. Here we present a novel methodology finalized to predict water quality as categorized indexes leveraging an integrated approach between computational components and artificial intelligence techniques. As a paradigm demonstrator, we couple WaComM++ with AIQUAM. The use case presented is an application of AIQUAM in the Bay of Naples (Campania Region, Italy) for predicting bacteria contaminants in mussel farms. The results are encouraging as the model reached a correct prediction rate of 93%.
Ciro Giuseppe De Vita, Gennaro Mellone, Diana Di Luccio, Sokol Kosta, Angelo Ciaramella, Raffaele Montella
e-Science6
2021 The Italian research on HPC key technologies across EuroHPC
abstract
High-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects.
Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati
CF13
2021 Process mining-constrained scheduling in the hybrid cloud
abstract
Summary Hybrid cloud, typically a combination of public and private cloud deployment models, is a rising paradigm due to the benefits it offers: full control of data and applications in the private cloud and elastic computing resource availability in the public cloud. This combination however brings an extra layer of complexity that can potentially erode the benefits and present serious challenges if not managed well. Among the challenges, ensuring business constraint compliance across the combination of cloud deployment models is a growing concern. Our article brings a sensitive, data‐ and process‐aware framework to bear on task scheduling in hybrid clouds with compliance to business constraints. Our proposed approach utilizes data from a real hybrid cloud‐based hospital billing system that is governed by complex and dynamic data processing rules. Our system successfully employs a process mining controlled algorithm to schedule tasks in the hybrid cloud to comply with the given set of business constraints.
Kenneth Kwame Azumah, Lene Tolstrup, Raffaele Montella, Sokol Kosta
Concurr. Comput. Pract. Exp.3
2021 Special Issue on High-end Heterogeneous Architectures, Methodologies, and Algorithms (HHAMA20)
abstract
TEST 02 - Elsevier's Scopus, the largest abstract and citation database of peer-reviewed literature. Search and access research from the science, technology, medicine, social sciences and arts and humanities fields.
Sokol Kosta, Giuliano Laccetti, Marco Lapegna, Valeria Mele, Raffaele Montella
Concurr. Comput. Pract. Exp.5
2021 Vessel to shore data movement through the Internet of Floating Things: A microservice platform at the edge
abstract
Summary The rise of the Internet of Things has generated high expectations about the improvement in people's lifestyles. In the last decade, we saw several examples of instrumented cities where different types of data were gathered, processed, and made available to inspire the next generation of scientists and engineers. In this framework, sensors and actuators became leading actors of technologically pervasive urban environments. However, in coastal areas, marine data crowdsourcing is difficult to apply due to the challenging operational conditions, extremely unstable network connectivity, and security issues in data movement. To fill this gap, we present a novel version of our DYNAMO transfer protocol (DTP), a platform‐independent data mover framework where data collected on board of vessels are stored locally and then moved from the edge to the cloud when the operating conditions are favorable. We evaluate the performance of DTP in a controlled environment with a private cloud by measuring the time it takes for the clouds ide to process and store a fixed amount of data while varying the number of microservice instances. We show that the time decreases exponentially when the number of microservice instances goes from 1 to 16 and it remains constant above that number.
Diana Di Luccio, Sokol Kosta, Aniello Castiglione, Antonio Maratea, Raffaele Montella
Concurr. Comput. Pract. Exp.5
2021 Special issue on workflows in Support of Large-Scale Science
Anirban Mandal, Raffaele Montella
Future Gener. Comput. Syst.2
2021 An efficient pattern-based approach for workflow supporting large-scale science: The DagOnStar experience
Dante D. Sánchez-Gallegos, Diana Di Luccio, Sokol Kosta, José Luis González 0002, Raffaele Montella
Future Gener. Comput. Syst.5
2020 Using the FACE-IT portal and workflow engine for operational food quality prediction and assessment: An application to mussel farms monitoring in the Bay of Napoli, Italy
Raffaele Montella, Alison Brizius, Diana Di Luccio, Cheryl H. Porter, Joshua Elliott, Ravi K. Madduri, David Kelly, Angelo Riccio, Ian T. Foster
Future Gener. Comput. Syst.1
2020 A gearbox model for processing large volumes of data by using pipeline systems encapsulated into virtual containers
Miguel Santiago-Duran, José Luis González 0002, André Brinkmann, Hugo G. Reyes-Anastacio, Jesús Carretero 0001, Raffaele Montella, Gregorio Toscano Pulido
Future Gener. Comput. Syst.6
2019 DeepNautilus: A Deep Learning Based System for Nautical Engines' Live Vibration Processing
Rosario Carbone, Raffaele Montella, Fabio Narducci, Alfredo Petrosino
CAIP (2)2
2019 An adaptive algorithm for high-dimensional integrals on heterogeneous CPU-GPU systems
abstract
Summary In this paper, we introduce an adaptive procedure for the numerical computation of a high‐dimensional integrals on HPC systems with heterogeneous nodes composed of multi‐core CPU and GPU devices. To this aim, we have integrated together two different approaches: a first one is in charge of a fair workload among the threads running on the multi‐core CPU, while a second one is in charge of an efficient execution of the computational kernels on the GPU. We tested the resulting algorithm on several test functions on a system where the nodes are provided with two Intel ten‐core CPU and one NVIDIA GPU device.
Giuliano Laccetti, Marco Lapegna, Valeria Mele, Raffaele Montella
Concurr. Comput. Pract. Exp.4
2019 Workflow-based automatic processing for Internet of Floating Things crowdsourced data
Raffaele Montella, Diana Di Luccio, Livia Marcellino, Ardelio Galletti, Sokol Kosta, Giulio Giunta, Ian T. Foster
Future Gener. Comput. Syst.1
2018 A Scalable Unified Model for Dynamic Data Structures in Message Passing (Clusters) and Shared Memory (multicore CPUs) Computing environments
abstract
Concurrent data structures are widely used in many software stack levels, ranging from high level parallel scientific applications to low level operating systems. The key issue of these objects is their concurrent use by several computing units (threads or process) so that the design of these structures is much more difficult compared to their sequential counterpart, because of their extremely dynamic nature requiring protocols to ensure data consistency, with a significant cost overhead. At this regard, several studies emphasize a tension between the needs of sequential correctness of the concurrent data structures and scalability of the algorithms, and in many cases it is evident the need to rethink the data structure design, using approaches based on randomization and/or redistribution techniques in order to fully exploit the computational power of the recent computing environments. The problem is grown in importance with the new generation High Performance Computing systems aimed to achieve extreme performance. It is easy to observe that such systems are based on heterogeneous architectures integrating several independent nodes in the form of clusters or MPP systems, where each node is composed by powerful computing elements (CPU core, GPUs or other acceleration devices) sharing resources in a single node. These systems therefore make massive use of communication libraries to exchange data among the nodes, as well as other tools for the management of the shared resources inside a single node. For such a reason, the development of algorithms and scientific software for dynamic data structures on these heterogeneous systems implies a suitable combination of several methodologies and tools to deal with the different kinds of parallelism corresponding to each specific device, so that to be aware of the underlying platform. The present work is aimed to introduce a scalable model to manage a special class of dynamic data structure known as heap based priority queue (or simply heap) on these heterogeneous architectures. A heap is generally used when the applications needs set of data not requiring a complete ordering, but only the access to some items tagged with high priority. In order to ensure a tradeoff between the correct access to high priority items by the several computing units with a low communication and synchronization overhead, a suitable reorganization of the heap is needed. More precisely we introduce a unified scalable model that can be used, with no modifications, to redeploy the items of a heap both in message passing environments (such as clusters and or MMP multicomputers with several nodes) as well as in shared memory environments (such as CPUs and multiprocessors with several cores) with an overhead independent of the number of computing units. Computational results related to the application of the proposed strategy on some numerical case studies are presented for different types of computing environments.
Giuliano Laccetti, Marco Lapegna, Raffaele Montella
CCGrid3
2018 DYNAMO: Distributed Leisure Yacht-Carried Sensor-Network for Atmosphere and Marine Data Crowdsourcing Applications
abstract
Data crowdsourcing is a increasingly pervasive and lifestyle-changing technology, due to the flywheel effect that results from the interaction between the internet of things and cloud computing. In smart cities, for example, many initiatives harvest valuable data from citizen sensors. However, this paradigm has not seen significant use in coastal and marine monitoring and management due to the challenges of the marine environment. In this work, we describe how this situation can be overcome via the adoption of an open technology ecosystem that provides leisure vessels with a platform for on-board data acquisition, storage, and processing, leveraging off-the-shelf mobile technologies. We introduce DYNAMO, a infrastructure designed to collect marine environmental data from a distributed sensor network carried by leisure vessels. The resulting crowdsourced data can be used to improve operational weather and marine predictions via the use of data assimilation methods. We show our preliminary results about the DYNAMO Daemon, a SignalK server we embedded in the native level of the Android operating system enabling the data gathering and transfer from vessels to the cloud.
Raffaele Montella, Sokol Kosta, Ian T. Foster
IC2E1
2018 Models, algorithms, and tools for highly heterogeneous computing environments
abstract
Models, algorithms, and tools for highly heterogeneous computing environments
Giuliano Laccetti, Marco Lapegna, Raffaele Montella, Sokol Kosta
Concurr. Comput. Pract. Exp.3
2018 Marine bathymetry processing through GPGPU virtualization in high performance cloud computing
abstract
Summary Fast technology development has influenced the widespread use of low‐power devices in different scientific, environmental, and everyday life areas, giving birth to the Internet of Things. In this paper, we focus on the context of marine studies, addressing the problem of marine bathymetry data processing and analysis via pervasive and Internet‐connected sensors and low‐power distributed devices. Pervasive and Internet‐connected low‐power devices (as the components involved in the sensing and processing actions) made diverse and different “things” as a worldwide‐distributed system. Given the high complexity of the algorithms involved in these studies, which usually involve general‐purpose graphic processing unit (GPGPU) computation, it is impossible for the limited devices to perform the required calculations. To overcome these limitations, in this paper, we propose and implement a vertical application of GVirtuS, the open‐source GPGPU virtualization and remoting service, for achieving high performance geographical data interpolation in a high performance cloud computing scenario. We present an innovative implementation by comparing, in terms of performance and accuracy, the inverse distance weighting and kriging interpolation methods in their parallel implementations leveraging on CUDA‐enabled GPGPUs. We present a real‐world use case related to high‐resolution bathymetry interpolation in a crowdsource data context in the Bay of Pozzuoli, Italy.
Raffaele Montella, Livia Marcellino, Ardelio Galletti, Diana Di Luccio, Sokol Kosta, Giuliano Laccetti, Giulio Giunta
Concurr. Comput. Pract. Exp.1
2017 Accelerating Linux and Android applications on low-power devices through remote GPGPU offloading
abstract
Summary Low‐power devices are usually highly constrained in terms of CPU computing power, memory, and GPGPU resources for real‐time applications to run. In this paper, we describe RAPID, a complete framework suite for computation offloading to help low‐powered devices overcome these limitations. RAPID supports CPU and GPGPU computation offloading on Linux and Android devices. Moreover, the framework implements lightweight secure data transmission of the offloading operations. We present the architecture of the framework, showing the integration of the CPU and GPGPU offloading modules. We show by extensive experiments that the overhead introduced by the security layer is negligible. We present the first benchmark results showing that Java/Android GPGPU code offloading is possible. Finally, we show the adoption of the GPGPU offloading into BioSurveillance, a commercial real‐time face recognition application. The results show that, thanks to RAPID, BioSurveillance is being successfully adapted to run on low‐power devices. The proposed framework is highly modular and exposes a rich application programming interface to developers, making it highly versatile while hiding the complexity of the underlying networking layer.
Raffaele Montella, Sokol Kosta, David Oro, Javier Vera, Carles Fernández, Carlo Palmieri, Diana Di Luccio, Giulio Giunta, Marco Lapegna, Giuliano Laccetti
Concurr. Comput. Pract. Exp.1
2016 Enabling Android-Based Devices to High-End GPGPUs
Raffaele Montella, Carmine Ferraro, Sokol Kosta, Valentina Pelliccia, Giulio Giunta
ICA3PP1
2015 FACE-IT: A science gateway for food security research
abstract
Summary Progress in sustainability science is hindered by challenges in creating and managing complex data acquisition, processing, simulation, post‐processing, and intercomparison pipelines. To address these challenges, we developed the Framework to Advance Climate, Economic, and Impact Investigations with Information Technology (FACE‐IT) for crop and climate impact assessments. This integrated data processing and simulation framework enables data ingest from geospatial archives; data regridding, aggregation, and other processing prior to simulation; large‐scale climate impact simulations with agricultural and other models, leveraging high‐performance and cloud computing; and post‐processing to produce aggregated yields and ensemble variables needed for statistics, for model intercomparison, and to connect biophysical models to global and regional economic models. FACE‐IT leverages the capabilities of the Globus Galaxies platform to enable the capture of workflows and outputs in well‐defined, reusable, and comparable forms. We describe FACE‐IT and applications within the Agricultural Model Intercomparison and Improvement Project and the Center for Robust Decision‐making on Climate and Energy Policy. Copyright © 2015 John Wiley & Sons, Ltd.
Raffaele Montella, David Kelly, Wei Xiong 0003, Alison Brizius, Joshua Elliott, Ravi K. Madduri, Ketan Maheshwari, Cheryl H. Porter, Peter Vilter, Michael Wilde, Meng Zhang 0007, Ian T. Foster
Concurr. Comput. Pract. Exp.1
2012 Virtualizing General Purpose GPUs for High Performance Cloud Computing: An Application to a Fluid Simulator
abstract
In this work we present an hypervisor-independent GPU Virtualization Service named GVirtus. It instantiates virtual machines able to access to the GPU in a transparent way. GPUs allow to speed up calculations over CPUs. Therefore, virtualizing GPUs is a major trend and can be considered a revolutionary tool for HPC. To test the performances of GVirtus we used a fluid simulator. Morover to exploit the computational power of GPUs in cloud computing we virtualized three different plugins for GVirtus Framework : Cuda Runtime, Cuda Driver and OpenCL plugins. Our results show that the overhead introduced by virtualization is almost irrelevant.
Roberto Di Lauro, Flora Giannone, Luigi Ambrosio, Raffaele Montella
ISPA4
2012 SIaaS - Sensing Instrument as a Service Using Cloud Computing to Turn Physical Instrument into Ubiquitous Service
abstract
Sensing Instrument as a Service (SIaaS) is a new cloud service that allows users to use data acquisition instruments shared through cloud infrastructure. It offers a common interface to managing physical sensing instrument and it permits to take all advantages of using cloud computing technology in storing and processing acquired data. With SIaaS different research groups could share physical instrument, in a controlled way, event if they are geographically distributed and user can access to them using an internet connected device without the need to install any sort of program. With the interaction of different own framework we have created a new cloud service for transforming physical resources in a ubiquitous service.
Roberto Di Lauro, Francesca Lucarelli, Raffaele Montella
ISPA3
2010 A GPGPU Transparent Virtualization Component for High Performance Computing Clouds
Giulio Giunta, Raffaele Montella, Giuseppe Agrillo, Giuseppe Coviello
Euro-Par (1)2
2008 Five Dimension Environmental Data Resource Brokering on Computational Grids and Scientific Clouds
abstract
In this paper we describe how the grid computing resource brokering approach, classically divided into matchmaking discovery and optimize selection, is applied to five dimensional environmental data distribution. This result is achieved thanks to the integration of a Resource Broker Service we developed from scratch. This service implements the key feature of autonomic mapping of Globus Toolkit Index Service resources on the Condor ClassAd resource description. Our Five Dimensional Distribution Data Service relays on this software infrastructure to advertise metadata and it provides the requested data leveraging on the web service resource framework. We provide an example based on a grid aware application component, showing the application of resource brokering to environmental problems. We focused our interests mainly about high resolution weather forecast data management and results quality assessment. Our approach would be both modular and distributed computing technology aware in order to provide the e-Science community of grid based cloud computing enabled tools fully benefiting of this technology.
Giulio Giunta, Giuliano Laccetti, Raffaele Montella
APSCC3
2007 Development of a GT4-Based Resource Broker Service: An Application to On-demand Weather and Marine Forecasting
Raffaele Montella
GPC1
2006 A Grid Computing Based Virtual Laboratory for Environmental Simulations
I. Ascione, Giulio Giunta, Patrizio Mariani, Raffaele Montella, Angelo Riccio
Euro-Par4