VLDB 2026 Research / reviewers in the wild / expert
Vincenzo Eduardo Padulano
dblp:295/9904
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-1209-3641ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ROOT's RNTuple and the Case for Custom Scientific Data FormatsabstractEach month, CERN collects and stores multiple petabytes of measurement data from the Large Hadron Collider (LHC) for further analysis. These data are stored in a custom columnar format provided by CERN’s open-source software framework ROOT. However, the increased availability of general-purpose columnar data formats raises the question whether developing and maintaining a custom data format for high-energy physics (HEP) is still worth it. To address this question, we compare ROOT’s RNTuple data format against Apache ORC and Apache Parquet. We find that RNTuple stores data 6-21% more efficiently, and processes physics data 3-6x faster. Florine Willemijn de Geus, Stijn Jongbloed, Vincenzo Eduardo Padulano, Ana Lucia Varbanescu |
eScience | 3 |
| 2025 | EVENTSETPROCESSOR: An Engine for Efficiently Combining High-Energy Physics DataabstractCERN’s Large Hadron Collider (LHC), the world’s largest high-energy physics (HEP) instrument, collects tens of petabytes of data per year. The LHC’s next phase is expected to produce up to ten times more data, which calls for novel, more efficient ways of storing and processing these data.HEP collider data are prepared and provided to physicists as read-only data sets, stored in a custom columnar data format. While traditionally all data needed for a particular analysis were captured in a single data set, the increasing scale of the LHC and the advent of modern analysis techniques now requires analysis workflows to use data from different data sets. However, the processing model established across the HEP community does not yet provide a straightforward way to achieve this and currently relies heavily on data duplication to produce the desired data sets. This leads to significant overhead in analysis workflows, both in runtime and storage.To reduce this overhead, we propose more efficient ways to combine HEP data sets. Specifically, we design union and join operations, as defined in relational algebra, to combine HEP data sets at runtime, eliminating therefore the need for data duplication. In this paper, we specify these operations for HEP data and introduce EVENTSETPROCESSOR – an engine that implements these operations for HEP data processing. Through a first prototype, we show that this engine integrates well in existing HEP workflows, and that it can perform up to twice as fast as the current approach. Florine Willemijn de Geus, Vincenzo Eduardo Padulano, Jakob Blomer, Hannes Mühleisen, Ana Lucia Varbanescu |
eScience | 2 |
| 2023 | Leveraging State-of-the-Art Engines for Large-Scale Data Analysis in High Energy PhysicsabstractAbstract The Large Hadron Collider (LHC) at CERN has generated a vast amount of information from physics events, reaching peaks of TB of data per day which are then sent to large storage facilities. Traditionally, data processing workflows in the High Energy Physics (HEP) field have leveraged grid computing resources. In this context, users have been responsible for manually parallelising the analysis, sending tasks to computing nodes and aggregating the partial results. Analysis environments in this field have had a common building block in the ROOT software framework. This is the de facto standard tool for storing, processing and visualising HEP data. ROOT offers a modern analysis tool called RDataFrame, which can parallelise computations from a single machine to a distributed cluster while hiding most of the scheduling and result aggregation complexity from users. This is currently done by leveraging Apache Spark as the distributed execution engine, but other alternatives are being explored by HEP research groups. Notably, Dask has rapidly gained popularity thanks to its ability to interface with batch queuing systems, widespread in HEP grid computing facilities. Furthermore, future upgrades of the LHC are expected to bring a dramatic increase in data volumes. This paper presents a novel implementation of the Dask backend for the distributed RDataFrame tool in order to address the aforementioned future trends. The scalability of the tool with both the new backend and the already available Spark backend is demonstrated for the first time on more than two thousand cores, testing a real HEP analysis. Vincenzo Eduardo Padulano, Ivan Donchev Kabadzhov, Enric Tejedor, Enrico Guiraud, Pedro Alonso 0002 |
J. Grid Comput. | 1 |
| 2023 | Leveraging an open source serverless framework for high energy physics computingabstractAbstract CERN (Centre Europeen pour la Recherce Nucleaire) is the largest research centre for high energy physics (HEP). It offers unique computational challenges as a result of the large amount of data generated by the large hadron collider. CERN has developed and supports a software called ROOT , which is the de facto standard for HEP data analysis. This framework offers a high-level and easy-to-use interface called RDataFrame , which allows managing and processing large data sets. In recent years, its functionality has been extended to take advantage of distributed computing capabilities. Thanks to its declarative programming model, the user-facing API can be decoupled from the actual execution backend . This decoupling allows physical analysis to scale automatically to thousands of computational cores over various types of distributed resources. In fact, the distributed RDataFrame module already supports the use of established general industry engines such as Apache Spark or Dask. Notwithstanding the foregoing, these current solutions will not be sufficient to meet future requirements in terms of the amount of data that the new projected accelerators will generate. It is of interest, for this reason, to investigate a different approach, the one offered by serverless computing. Based on a first prototype using AWS Lambda , this work presents the creation of a new backend for RDataFrame distributed over the OSCAR tool, an open source framework that supports serverless computing. The implementation introduces new ways, relative to the AWS Lambda -based prototype, to synchronize the work of functions. Vincenzo Eduardo Padulano, Pablo Oliver Cortés, Pedro Alonso 0002, Enric Tejedor, Sebastián Risco, Germán Moltó |
J. Supercomput. | 1 |
| 2022 | A Serverless Engine for High Energy Physics Distributed AnalysisabstractThe Large Hadron Collider (LHC) at CERN has generated in the last decade an unprecedented volume of data for the High-Energy Physics (HEP) field. Scientific collaborations interested in analysing such data very often require computing power beyond a single machine. This issue has been tackled traditionally by running analyses in distributed environments using stateful, managed batch computing systems. While this approach has been effective so far, current estimates for future computing needs of the field present large scaling challenges. Such a managed approach may not be the only viable way to tackle them and an interesting alternative could be provided by serverless architectures, to enable an even larger scaling potential. This work describes a novel approach to running real HEP scientific applications through a distributed serverless computing engine. The engine is built upon ROOT, a well-established HEP data analysis software, and distributes its computations to a large pool of concurrent executions on Amazon Web Services Lambda Serverless Platform. Thanks to the developed tool, physicists are able to access datasets stored at CERN (also those that are under restricted access policies) and process it on remote infrastructures outside of their typical environment. The analysis of the serverless functions is monitored at runtime to gather performance metrics, both for data- and computation-intensive workloads. Jacek Kusnierz, Vincenzo Eduardo Padulano, Maciej Malawski, Kamil Burkiewicz, Enric Tejedor, Pedro Alonso 0002, Michael Pitt, Valentina Avati |
CCGRID | 2 |