EDBT 2026 Demo / reviewers in the wild / expert
Daniele Spiga
dblp:80/4154
· DBLP profile ↗
10ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-2991-6384ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | interTwin: Advancing Scientific Digital Twins through AI, Federated Computing and DataabstractData will be made available on request. Andrea Manzi, Raul Bardaji, Ivan Rodero, Germán Moltó, Sandro Fiore, Isabel Campos Plasencia, Donatello Elia, Francesco Sarandrea, A. Paul Millar, Daniele Spiga, Matteo Bunino, Gabriele Accarino, Lorenzo Asprea, Samuel Bernardo, Miguel Caballer, Charis Chatzikyriakou, Diego Ciangottini, Michele Claus, Andrea Cristofori, Davide Donno, Emanuele Donno, Iacopo Ferrario, Massimiliano Fronza, Alexander W. Jacob, Javad Komijani, Marina Krstic Marinkovic, Federica Legger, Ivan Palomo, Estíbaliz Parcero, Rakesh Sarma, Gaurav Sinha Ray, Sara Vallero, Juraj Zvolensky |
Future Gener. Comput. Syst. | 10 |
| 2025 | Developments on the "Machine Learning as a Service for High Energy Physics" Framework and Related Cloud Native SolutionabstractMachine Learning (ML) techniques have been successfully used in many areas of High Energy Physics (HEP) and will play a significant role in the success of upcoming High-Luminosity Large Hadron Collider (HL-LHC) program at CERN. An unprecedented amount of data at the exascale will be collected by LHC experiments in the next decade, and this effort will require novel approaches to train and use ML models. The work presented in this paper is focused on the developments of a ML as a Service (MLaaS) solution for HEP, aiming to provide a cloud service that allows HEP users to run ML pipelines via HTTPs calls. These pipelines are executed by using MLaaS4HEP framework, which allows reading data, processing data, and training ML models directly using ROOT files of arbitrary size from local or distributed data sources. In particular, new features implemented on the framework will be presented as well as updates on the architecture of an existing prototype of the MLaaS4HEP cloud service will be provided. This solution includes two OAuth2 proxy servers as authentication/authorization layer, a MLaaS4HEP server, an XRootD proxy server for enabling access to remote ROOT data, and the TensorFlow as a Service (TFaaS) service in charge of the inference phase. Luca Giommi, Daniele Spiga, Mattia Paladino, Valentin Kuznetsov, Daniele Bonacorsi |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | Smart Caching in a Data Lake for High Energy Physics AnalysisabstractAbstract The continuous growth of data production in almost all scientific areas raises new problems in data access and management, especially in a scenario where the end-users, as well as the resources that they can access, are worldwide distributed. This work is focused on the data caching management in a Data Lake infrastructure in the context of the High Energy Physics field. We are proposing an autonomous method, based on Reinforcement Learning techniques, to improve the user experience and to contain the maintenance costs of the infrastructure. Tommaso Tedeschi, Marco Baioletti, Diego Ciangottini, Valentina Poggioni, Daniele Spiga, Loriano Storchi, Mirco Tracolli |
J. Grid Comput. | 5 |
| 2022 | The BondMachine, a moldable computer architecture
Mirko Mariotti, Daniel Magalotti, Daniele Spiga, Loriano Storchi |
Parallel Comput. | 3 |
| 2020 | An Intelligent Cache Management for Data Analysis at CMS
Mirco Tracolli, Marco Baioletti, Diego Ciangottini, Valentina Poggioni, Daniele Spiga |
ICCSA (2) | 5 |
| 2020 | Effective Big Data Caching through Reinforcement LearningabstractIn the era of big data, data volumes continue to grow in several different domains, from business to scientific fields. Sensors, edge devices, scientific applications and detectors generate huge amounts of data that are distributed for their nature. In order to extract value from such data requires a typical pipeline made of two main steps: first, the processing and then the data access. One of the main features for data access is fast response time, whose order of magnitude can vary a lot depending on the specific type of processing as well as processing patterns. The optimization of the access layer becomes more and more important while dealing with a geographically distributed environment where data must be retrieved from remote servers of a data lake. From the infrastructural perspectives, caching systems are used to mitigate latency and to serve better popular data. Thus, the role of the cache becomes a key to have an effective and efficient data access. In this article, we propose a Reinforcement Learning approach, using the Q-Learning technique, to improve the performances of a cache system in terms of data management. The proposed method uses two agents with different objectives and actions to control the addition and the eviction of files in the cache. The aim of this system is to increase the throughput reducing, at the same time, the cache costs, such as the amount of data written, and network utilization. Moreover, we tested our method in a context of data analysis, with information taken from High Energy Physics (HEP) workflow. Mirco Tracolli, Marco Baioletti, Valentina Poggioni, Daniele Spiga |
ICMLA | 4 |
| 2018 | Distributed and On-demand Cache for CMS Experiment at LHCabstractIn the CMS [1] computing model the experiment owns dedicated resources around the world that, for the most part, are located in computing centers with a well defined Tier hierarchy. The geo-distributed storage is then controlled centrally by the CMS Computing Operations. In this architecture data are distributed and replicated across the centers following a preplacement model, mostly human controlled. Analysis jobs are then mostly executed on computing resources close to the data location. This of course allow to avoid CPU wasting due to I/O latency, although it does not allow to optimize the available job slots. Diego Ciangottini, Daniele Spiga, Tommaso Boccali, Giacinto Donvito, Daniele Cesini, Giuseppe Bagliesi, Enrico Mazzoni, Antonio Falabella |
eScience | 2 |
| 2018 | INDIGO-DataCloud: a Platform to Facilitate Seamless Access to E-InfrastructuresabstractThis paper describes the achievements of the H2020 project INDIGO-DataCloud. The project has provided e-infrastructures with tools, applications and cloud framework enhancements to manage the demanding requirements of scientific communities, either locally or through enhanced interfaces. The middleware developed allows to federate hybrid resources, to easily write, port and run scientific applications to the cloud. In particular, we have extended existing PaaS (Platform as a Service) solutions, allowing public and private e-infrastructures, including those provided by EGI, EUDAT, and Helix Nebula, to integrate their existing services and make them available through AAI services compliant with GEANT interfederation policies, thus guaranteeing transparency and trust in the provisioning of such services. Our middleware facilitates the execution of applications using containers on Cloud and Grid based infrastructures, as well as on HPC clusters. Our developments are freely downloadable as open source components, and are already being integrated into many scientific applications. Davide Salomoni, Isabel Campos Plasencia, Luciano Gaido, Jesús E. Marco de Lucas, P. Solagna, Jorge Gomes 0001, Ludek Matyska, P. Fuhrman, Marcus Hardt, Giacinto Donvito, Lukasz Dutka, Marcin Plóciennik, Roberto Barbera, Ignacio Blanquer, Andrea Ceccanti, Eva Cetinic, Mário David, Doina Cristina Duma, Álvaro López García, Germán Moltó, Pablo Orviz Fernández, Zdenek Sustr, Matthew Viljoen, Fernando Aguilar, Marica Antonacci, Lucio Angelo Antonelli, Stefano Bagnasco, A. Bonving, Riccardo Bruno, Alessandro Costa, Davor Davidovic, Benjamin Ertl, Marco Fargetta, Sandro Fiore, S. Gallozzi, Z. Kurkcuoglu, Lara Lloret Iglesias, J. Martins, Alessandra Nuzzo, Paola Nassisi, Cosimo Palazzo, João Murta Pina, Eva Sciacca, Daniele Spiga, Marco Antonio Tangaro, Michal Urbaniak, Sara Vallero, Bas Wegh, Valentina Zaccolo, Federico Zambelli, Tomasz Zok |
J. Grid Comput. | 46 |
| 2010 | Distributed Analysis in CMS
Alessandra Fanfani, M. Anzar Afaq, Jose Afonso Sanches, Julia Andreeva, Giuseppe Bagliesi, L. A. T. Bauerdick, Stefano Belforte, Patricia Bittencourt Sampaio, Kenneth Bloom, Barry Blumenfeld, Daniele Bonacorsi, Chris Brew, Marco Calloni, Daniele Cesini, Mattia Cinquilli, Giuseppe Codispoti, Jorgen D'Hondt, Danilo N. Dongiovanni, Giacinto Donvito, David Dykstra, Erik Edelmann, Ricky Egeland, Peter Elmer, Giulio Eulisse, Dave Evans, Federica Fanzago, Fabio Farina, Derek Feichtinger, Ian Fisk, Josep Flix, Claudio Grandi, Yuyi Guo, Kalle Happonen, José M. Hernández, Chih-Hao Huang, Kejing Kang, Edward Karavakis, Matthias Kasemann, Carlos Kavka, Akram Khan, Bockjoo Kim, Jukka Klem, Jesper Koivumäki, Thomas Kress, Peter Kreuzer, Tibor Kurca, Valentin Kuznetsov, Stefano Lacaprara, Kati Lassila-Perini, James Letts, Tomas Lindén, Lee Lueking, Joris Maes, Nicolò Magini, Gerhild Maier, Patricia McBride, Simon Metson, Vincenzo Miccio, Sanjay Padhi, Haifeng Pi, Hassen Riahi, Daniel Riley, Paul Rossman, Pablo Saiz, Andrea Sartirana, Andrea Sciabà, Vijay Sekhri, Daniele Spiga, Lassi A. Tuura, Eric Wayne Vaandering, Lukas Vanelderen, Petra Van Mulders, Aresh Vedaee, Ilaria Villella, Eric Wicklund, Tony Wildish, Christoph Wissing, Frank Würthwein |
J. Grid Comput. | 69 |
| 2007 | The CMS Remote Analysis Builder (CRAB)
Daniele Spiga, Stefano Lacaprara, W. Bacchi, Mattia Cinquilli, Giuseppe Codispoti, Marco Corvo, A. Dorigo, Alessandra Fanfani, Federica Fanzago, Fabio Farina, M. Merlo, Oliver Gutsche, Leonello Servoli, Carlos Kavka |
HiPC | 1 |