EDBT 2026 Demo / reviewers in the wild / expert
Luca Gagliardelli
dblp:172/1467
· DBLP profile ↗
14ranked-venue papers in the field
4as first author
8since 2021 · last 2026
0000-0001-5977-1078ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 10 (4 first)Knowledge Engineering, Semantic Web & Information Systems · 2Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Versatile Sketch-Based Attribute Filtering for Hybrid Vector SearchabstractThis work addresses the problem of hybrid search in vector databases, which store vectors together with some property attributes. Given a query that consists of a vector and some restrictions on its property attributes, we want to retrieve approximate nearest neighbor vectors for the query while ensuring compliance with predicate conditions, such as point or range filters on a specific vector property attribute. The challenge is compounded by the need to balance two competing requirements: on one hand, ensuring high accuracy in the vector search by leveraging a similarity-based index that is independent of specific attributes, allowing it to serve all queries; on the other hand, the impracticality of replicating such a structure for each attribute or predicate condition. To address these challenges, we propose an agnostic, attribute popularity-aware solution for predicate filtering in approximate nearest neighbor (ANN) search, leveraging the efficiency of graph-based indexing structures for vectors. Our method begins by clustering nodes within the underlying graph structure and constructing lightweight in-memory sketches for the predicates. During query processing, the search selectively applies a two-hop traversal strategy only when necessary, guided by the attribute popularity within the identified cluster. Experimental evaluation across five benchmark datasets demonstrates that our approach consistently outperforms state-of-the-art methods. Adeel Aslam, Luca Gagliardelli, El Kindi Rezig, George Konstantinidis 0001, Giovanni Simonini |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Evaluation of Dataframe Libraries for Data Preparation on a Single Machine
Angelo Mozzillo, Luca Zecchini, Luca Gagliardelli, Adeel Aslam, Sonia Bergamaschi, Giovanni Simonini |
EDBT | 3 |
| 2024 | Stream-aware indexing for distributed inequality join processing
Adeel Aslam, Giovanni Simonini, Luca Gagliardelli, Luca Zecchini, Sonia Bergamaschi |
Inf. Syst. | 3 |
| 2024 | GSM: A generalized approach to Supervised Meta-blocking for scalable entity resolutionabstractEntity Resolution (ER) constitutes a core data integration task that relies on Blocking in order to tame its quadratic time complexity. Schema-agnostic blocking achieves very high recall, requires no domain knowledge and applies to data of any structuredness and schema heterogeneity. This comes at the cost of many irrelevant candidate pairs (i.e., comparisons), which can be significantly reduced through Meta-blocking techniques, i.e., techniques that leverage the co-occurrence patterns of entities inside the blocks: first, a weighting scheme assigns a score to every pair of candidate entities in proportion to the likelihood that they are matching and then, a pruning algorithm discards the pairs with the lowest scores. Supervised Meta-blocking goes beyond this approach by combining multiple scores per comparison into a feature vector that is fed to a binary classifier. By using probabilistic classifiers, Generalized Supervised Meta-blocking associates every pair of candidates with a score that can be used: (i) by any pruning algorithm for retaining the set of candidate comparisons; and (ii) by state-of-the-art progressive ER methods to identify the most promising candidates as early as possible (when time is a critical component for the downstream applications that consume the data). For higher effectiveness, new weighting schemes are examined as features. Through an extensive experimental analysis, we identify the best pruning algorithms, their optimal sets of features as well as the minimum possible size of the training set. The resulting approaches achieve excellent performance across several established benchmark datasets. Luca Gagliardelli, George Papadakis 0001, Giovanni Simonini, Sonia Bergamaschi, Themis Palpanas |
Inf. Syst. | 1 |
| 2023 | A Big Data Platform for the Management of Local Energy Communities DataabstractIn this paper we present a big data platform designed to collect and analyze energy data of Local Energy Communities, with the goal to improve the conscious use of energy by the users. The platform, originally commissioned by ENEA, is designed to acquire and manage different kinds of data (e.g., energy consumption and production, weather data, etc.) coming from multiple sources in many formats. In this work, we present the designed architecture, and several dataflows which show the real-use cases that highlight the main strengths offered by our platform. Sonia Bergamaschi, Luca Gagliardelli |
IEEE Big Data | 2 |
| 2023 | HKS: Efficient Data Partitioning for Stateful Streaming
Adeel Aslam, Giovanni Simonini, Luca Gagliardelli, Angelo Mozzillo, Sonia Bergamaschi |
DaWaK | 3 |
| 2022 | Generalized Supervised Meta-blockingabstractEntity Resolution is a core data integration task that relies on Blocking to scale to large datasets. Schema-agnostic blocking achieves very high recall, requires no domain knowledge and applies to data of any structuredness and schema heterogeneity. This comes at the cost of many irrelevant candidate pairs (i.e., comparisons), which can be significantly reduced by Meta-blocking techniques that leverage the entity co-occurrence patterns inside blocks: first, pairs of candidate entities are weighted in proportion to their matching likelihood, and then, pruning discards the pairs with the lowest scores. Supervised Meta-blocking goes beyond this approach by combining multiple scores per comparison into a feature vector that is fed to a binary classifier. By using probabilistic classifiers, Generalized Supervised Meta-blocking associates every pair of candidates with a score that can be used by any pruning algorithm. For higher effectiveness, new weighting schemes are examined as features. Through extensive experiments, we identify the best pruning algorithms, their optimal sets of features, as well as the minimum possible size of the training set. Luca Gagliardelli, George Papadakis 0001, Giovanni Simonini, Sonia Bergamaschi, Themis Palpanas |
Proc. VLDB Endow. | 1 |
| 2021 | Reproducible experiments on Three-Dimensional Entity Resolution with JedAI
Georgios M. Mandilaras, George Papadakis 0001, Luca Gagliardelli, Giovanni Simonini, Emmanouil Thanos, George Giannakopoulos, Sonia Bergamaschi, Themis Palpanas, Manolis Koubarakis, Alicia Lara-Clares, Antonio Fariña |
Inf. Syst. | 3 |
| 2020 | RulER: Scaling Up Record-level Matching Rules
Luca Gagliardelli, Giovanni Simonini, Sonia Bergamaschi |
EDBT | 1 |
| 2020 | Three-dimensional Entity Resolution with JedAI
George Papadakis 0001, Georgios M. Mandilaras, Luca Gagliardelli, Giovanni Simonini, Emmanouil Thanos, George Giannakopoulos, Sonia Bergamaschi, Themis Palpanas, Manolis Koubarakis |
Inf. Syst. | 3 |
| 2019 | SparkER: Scaling Entity Resolution in SparkabstractWe present SparkER, an ER tool that can scale practitioners' favorite ER algorithms.SparkER has been devised to take full advantage of parallel and distributed computation as well (running on top of Apache Spark).The first SparkER version was focused on the blocking step and implements both schema-agnostic and Blast meta-blocking approaches (i.e. the state-of-the-art ones); a GUI for SparkER, to let non-expert users to use it in an unsupervised mode, was developed.The new version of SparkER to be shown in this demo, extends significantly the tool.Entity matching and Entity Clustering modules have been added.Moreover, in addition to the completely unsupervised mode of the first version, a supervised mode has been added.The user can be assisted in supervising the entire process and in injecting his knowledge in order to achieve the best result.During the demonstration, attendees will be shown how SparkER can significantly help in devising and debugging ER algorithms. Luca Gagliardelli, Giovanni Simonini, Domenico Beneventano, Sonia Bergamaschi |
EDBT | 1 |
| 2019 | Scaling entity resolution: A loosely schema-aware approach
Giovanni Simonini, Luca Gagliardelli, Sonia Bergamaschi, H. V. Jagadish |
Inf. Syst. | 2 |
| 2015 | Open Data for Improving Youth PoliciesabstractThe Open Data \textit{philosophy} is based on the idea that certain data should be made available to all citizens, in an open form, without any copyright restrictions, patents or other mechanisms of control. Various government have started to publish open data, first of all USA and UK in 2009, and in 2015, the Open Data Barometer project (www.opendatabarometer.org) states that on 77 diverse states across the world, over 55 percent have developed some form of Open Government Data initiative. We claim Public Administrations, that are the main producers and one of the consumers of Open Data, might effectively extract important information by integrating its own data with open data sources.This paper reports the activities carried on during a one-year research project on Open Data for Youth Policies. The project was mainly devoted to explore the youth situation in the municipalities and provinces of the Emilia Romagna region (Italy), in particular, to examine data on population, education and work.The project goals were: to identify interesting data sources both from the open data community and from the private repositories of local governments of Emilia Romagna region related to the Youth Policies; to integrate them and, to show up the result of the integration by means of a useful navigator tool; in the end, to publish new information on the web as Linked Open Data. This paper also reports the main issues encountered that may seriously affect the entire process of consumption, integration till the publication of open data. Domenico Beneventano, Sonia Bergamaschi, Luca Gagliardelli, Laura Po |
KEOD | 3 |
| 2015 | Driving Innovation in Youth Policies with Open Data
Domenico Beneventano, Sonia Bergamaschi, Luca Gagliardelli, Laura Po |
IC3K | 3 |