Daniele Dell'Aglio

dblp:57/8026 · DBLP profile ↗
← Back
31ranked-venue papers in the field
4as first author
12since 2021 · last 2025
0000-0003-4904-2511ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 18 (3 first)Information Retrieval & Web Search · 7 (1 first)Database Systems & Data Management · 4Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2025 Pasteur: Scaling Privacy-Aware Data Synthesis
Antheas Kapenekakis, Daniele Dell'Aglio, Martin Bøgsted, Minos N. Garofalakis, Katja Hose
ADBIS2
2025 The Yelp Collaborative Knowledge Graph
abstract
Yelp Open Dataset (YOD) is a widely used dataset for Recommender Systems (RS). Multiple Knowledge Graphs (KGs) have been built for YOD, but they have various issues: the conversion processes usually do not follow state-of-the-art methodologies, fail to properly link to other KGs, do not link to existing vocabularies, ignore important data, and are generally of small size. Instead, we present the Yelp Collaborative Knowledge Graph (YCKG), where we correctly integrating taxonomies, product categories, business locations, and the Yelp social network, through common practices within the semantic web community, overcoming all these issues. As a result, the YCKG includes 150k businesses and 16.9M reviews from 1.9M distinct real users, resulting in over 244 million triples, 144 distinct predicates, for about 72 million resources, with an average in-degree and out-degree of 3.3 and 12.2, respectively. Further, we release both the data and the code used to generate the KG for inspection and further extensions. This dataset can be used to develop and test both recommendation and data-mining algorithms able to exploit rich and semantically meaningful knowledge. We publicize the code for the CKG construction on: https://github.com/MadsCorfixen/The-Yelp-Collaborative-Knowledge-Graph.
Theis E. Jendal, Mads Corfixen, Magnus Olesen, Peter Dolog, Katja Hose, Daniele Dell'Aglio, Matteo Lissandrini
CIKM6
2025 Heterogeneous Graph Representation for Dataset Link Prediction on Dynamic and Sparse Scholarly Graphs
Ornella Irrera, Matteo Lissandrini, Daniele Dell'Aglio, Gianmaria Silvello
TPDL3
2025 PrivEval: a tool for interactive evaluation of privacy metrics in synthetic data generation
abstract
Synthetic data generation (SDG) is the process of generating a new synthetic dataset based on the statistical properties of a confidential existing dataset. Differential privacy is the property of a SDG mechanism that establishes how protected individuals whose sensitive data is part of the confidential dataset are, when sharing such data. To ensure a SDG is differentially private, noise is injected into the statistics learned from the dataset. Depending on the amount of noise injected, we witness a trade-off between privacy and utility. Privacy is then measured via a set of privacy metrics that usually establish a lower bound on a few aspects of the privacy-utility trade-off. Therefore, it is not possible to assess privacy based only on one metric. To close this gap, we demonstrate PrivEval, a tool to assist users in evaluating the privacy properties of a synthetic dataset. PrivEval implements several privacy metrics and validates them on both a single user and the overall dataset. Besides, PrivEval checks assumptions behind each metric. Hence, PrivEval is a first step to bridge the gap between privacy experts and the general public to make privacy estimation more transparent.
Frederik M. Trudslev, Matteo Lissandrini, Juan Manuel Rodriguez, Martin Bøgsted, Daniele Dell'Aglio
Proc. VLDB Endow.5
2024 Synthesizing Accurate Relational Data under Differential Privacy
abstract
Medical data is sensitive personal data which, according to GDPR and HIPAA, necessitates regulations concerning their use. Anonymizing this data prior to research would allow for broader access, due to a lower sensitivity. Privacy-aware data synthesis has been proposed as a solution. However, current algorithms face difficulties in synthesizing medical data while maintaining privacy and utility. This is due to the structure of medical data which consists of multiple interlinked tables with high dimensional columns containing sequential aspects of the patient trajectory. The resulting number of correlations is intractable to model naively and, if relational correlations are not accounted for, the resulting data has poor utility (e.g., leads to invalid patient trajectories). In this paper, we present MARE, a relational synthesis algorithm which focuses on a set of core correlations found in relational data while pruning others. The resulting lower computational complexity allows MARE to produce accurate relational data. We showcase that MARE can synthesize multiple medical datasets, which contain sequential aspects, while maintaining utility in form of inter-table and inter-row correlations and privacy guarantees.
Antheas Kapenekakis, Daniele Dell'Aglio, Charles Vesteghem, Laurids Poulsen, Martin Bøgsted, Minos N. Garofalakis, Katja Hose
IEEE Big Data2
2024 Reproducibility and Analysis of Scientific Dataset Recommendation Methods
abstract
Datasets play a central role in scholarly communications. However, scholarly graphs are often incomplete, particularly due to the lack of connections between publications and datasets. Therefore, the importance of dataset recommendation—identifying relevant datasets for a scientific paper, an author, or a textual query—is increasing. Although various methods have been proposed for this task, their reproducibility remains unexplored, making it difficult to compare them with new approaches. We reviewed current recommendation methods for scientific datasets, focusing on the most recent and competitive approaches, including an SVM-based model, a bi-encoder retriever, a method leveraging co-authors and citation network embeddings, and a heterogeneous variational graph autoencoder. These approaches underwent a comprehensive analysis under consistent experimental conditions. Our reproducibility efforts show that three methods can be reproduced, while the graph variational autoencoder is challenging due to unavailable code and test datasets. Hence, we re-implemented this method and performed a component-based analysis to examine its strengths and limitations. Furthermore, our study indicated that three out of four considered methods produce subpar results when applied to real-world data instead of specialized datasets with ad-hoc features.
Ornella Irrera, Matteo Lissandrini, Daniele Dell'Aglio, Gianmaria Silvello
RecSys3
2023 Towards the Web of Embeddings: Integrating multiple knowledge graph embedding spaces with FedCoder
abstract
The Semantic Web is distributed yet interoperable: Distributed since resources are created and published by a variety of producers, tailored to their specific needs and knowledge; Interoperable as entities are linked across resources, allowing to use resources from different providers in concord. Complementary to the explicit usage of Semantic Web resources, embedding methods made them applicable to machine learning tasks. Subsequently, embedding models for numerous tasks and structures have been developed, and embedding spaces for various resources have been published. The ecosystem of embedding spaces is distributed but not interoperable: Entity embeddings are not readily comparable across different spaces. To parallel the Web of Data with a Web of Embeddings, we must thus integrate available embedding spaces into a uniform space. Current integration approaches are limited to two spaces and presume that both of them were embedded with the same method — both assumptions are unlikely to hold in the context of a Web of Embeddings. In this paper, we present FedCoder— an approach that integrates multiple embedding spaces via a latent space. We assert that linked entities have a similar representation in the latent space so that entities become comparable across embedding spaces. FedCoder employs an autoencoder to learn this latent space from linked as well as non-linked entities. Our experiments show that FedCoder substantially outperforms state-of-the-art approaches when faced with different embedding models, that it scales better than previous methods in the number of embedding spaces, and that it improves with more graphs being integrated whilst performing comparably with current approaches that assumed joint learning of the embeddings and were, usually, limited to two sources. Our results demonstrate that FedCoder is well adapted to integrate the distributed, diverse, and large ecosystem of embeddings spaces into an interoperable Web of Embeddings.
Matthias Baumgartner, Daniele Dell'Aglio, Heiko Paulheim, Abraham Bernstein
J. Web Semant.2
2022 A framework for differentially-private knowledge graph embeddings
Xiaolin Han 0002, Daniele Dell'Aglio, Tobias Grubenmann, Reynold Cheng, Abraham Bernstein
J. Web Semant.2
2022 Visualising the effects of ontology changes and studying their understanding with ChImp
abstract
Due to the Semantic Web’s decentralised nature, ontology engineers rarely know all applications that leverage their ontology. Consequently, they are unaware of the full extent of possible consequences that changes might cause to the ontology. Our goal is to lessen the gap between ontology engineers and users by investigating ontology engineers’ understanding of ontology changes’ impact at editing time. Hence, this paper introduces the Protégé plugin ChImp which we use to reach our goal. We elicited requirements for ChImp through a questionnaire with ontology engineers. We then developed ChImp according to these requirements and it displays all changes of a given session and provides selected information on said changes and their effects. For each change, it computes a number of metrics on both the ontology and its materialisation. It displays those metrics on both the originally loaded ontology at the beginning of the editing session and the current state to help ontology engineers understand the impact of their changes. We investigated the informativeness of materialisation impact measures, the meaning of severe impact, and also the usefulness of ChImp in an online user study with 36 ontology engineers. We asked the participants to solve two ontology engineering tasks – with and without ChImp (assigned in random order) – and answer in-depth questions about the applied changes as well as the materialisation impact measures. We found that ChImp increased the participants’ understanding of change effects and that they felt better informed. Answers also suggest that the proposed measures were useful and informative. We also learned that the participants consider different outcomes of changes severe, but most would define severity based on the amount of changes to the materialisation compared to its size. The participants also acknowledged the importance of quantifying the impact of changes and that the study will affect their approach of editing ontologies.
Romana Pernisch, Daniele Dell'Aglio, Mirko Serbak, Rafael S. Gonçalves 0001, Abraham Bernstein
J. Web Semant.2
2021 Single Point Incremental Fourier Transform on 2D Data Streams
abstract
In radio astronomy, antennas monitor portions of the sky to collect radio signals. The antennas produce data streams that are of high volume and velocity (~2.5 GB/s) and the inverse Fourier transform is used to convert the collected signals into sky images that astrophysicists use to conduct their research. Applying the inverse Fourier transform in a streaming setting, however, is not ideal since its computational complexity is quadratic in the size of the image.In this article, we propose the Single Point Incremental Fourier Transform (SPIFT), a novel incremental algorithm to produce sequences of sky images. SPIFT computes the Fourier transform for a new signal in a linear number of complex multiplications by exploiting twiddle factors, multiplicative constant coefficients. We prove that twiddle factors are periodic and show how circular shifts can be exploited to reuse multiplication results. The cost of the additive operations can be curbed by exploiting the embarrassingly parallel nature of the additions, which modern big data streaming frameworks can leverage to compute slices of the image in parallel. Our experiments suggest that SPIFT can efficiently generate sequences of sky images: it computes the complex multiplications 4 to 12x faster than the Discrete Fourier Transform, and its parallelisation of the additive operations shows linear speedup.
Muhammad Saad 0006, Abraham Bernstein, Michael H. Böhlen, Daniele Dell'Aglio
ICDE4
2021 Toward Measuring the Resemblance of Embedding Models for Evolving Ontologies
abstract
Updates on ontologies affect the operations built on top of them. But not all changes are equal: some updates drastically change the result of operations; others lead to minor variations, if any. Hence, estimating the impact of a change ex-ante is highly important, as it might make ontology engineers aware of the consequences of their action during editing. However, in order to estimate the impact of changes, we need to understand how to measure them.
Romana Pernisch, Daniele Dell'Aglio, Abraham Bernstein
K-CAP2
2021 Beware of the hierarchy - An analysis of ontology evolution and the materialisation impact for biomedical ontologies
abstract
Ontologies are becoming a key component of numerous applications and research fields. But knowledge captured within ontologies is not static. Some ontology updates potentially have a wide ranging impact; others only affect very localised parts of the ontology and their applications. Investigating the impact of the evolution gives us insight into the editing behaviour but also signals ontology engineers and users how the ontology evolution is affecting other applications. However, such research is in its infancy. Hence, we need to investigate the evolution itself and its impact on the simplest of applications: the materialisation. In this work, we define impact measures that capture the effect of changes on the materialisation. In the future, the impact measures introduced in this work can be used to investigate how aware the ontology editors are about consequences of changes. By introducing five different measures, which focus either on the change in the materialisation with respect to the size or on the number of changes applied, we are able to quantify the consequences of ontology changes. To see these measures in action, we investigate the evolution and its impact on materialisation for nine open biomedical ontologies, most of which adhere to the EL++ description logic. Our results show that these ontologies evolve at varying paces but no statistically significant difference between the ontologies with respect to their evolution could be identified. We identify three types of ontologies based on the types of complex changes which are applied to them throughout their evolution. The impact on the materialisation is the same for the investigated ontologies, bringing us to the conclusion that the effect of changes on the materialisation can be generalised to other similar ontologies. Further, we found that the materialised concept inclusion axioms experience most of the impact induced by changes to the class inheritance of the ontology and other changes only marginally touch the materialisation.
Romana Pernisch, Daniele Dell'Aglio, Abraham Bernstein
J. Web Semant.2
2020 Differentially Private Stream Processing for the Semantic Web
abstract
Data often contains sensitive information, which poses a major obstacle to publishing it. Some suggest to obfuscate the data or only releasing some data statistics. These approaches have, however, been shown to provide insufficient safeguards against de-anonymisation. Recently, differential privacy (DP), an approach that injects noise into the query answers to provide statistical privacy guarantees, has emerged as a solution to release sensitive data. This study investigates how to continuously release privacy-preserving histograms (or distributions) from online streams of sensitive data by combining DP and semantic web technologies. We focus on distributions, as they are the basis for many analytic applications. Specifically, we propose SihlQL, a query language that processes RDF streams in a privacy-preserving fashion. SihlQL builds on top of SPARQL and the w-event DP framework. We show how some peculiarities of w-event privacy constrain the expressiveness of SihlQL queries. Addressing these constraints, we propose an extension of w-event privacy that provides answers to a larger class of queries while preserving their privacy. To evaluate SihlQL, we implemented a prototype engine that compiles queries to Apache Flink topologies and studied its privacy properties using real-world data from an IPTV provider and an online e-commerce web site.
Daniele Dell'Aglio, Abraham Bernstein
WWW1
2019 Collaborative Streaming: Trust Requirements for Price Sharing
abstract
Stream Processing (SP) is an important Big Data technology enabling continuous querying of data streams. The stream setting offers the opportunity to exploit synergies and, theoretically, share the access and processing costs between multiple different collaborators. But what should be the monetary contribution of each consumer when they do not trust each other and have varying valuations of the differing outcomes? In this article, we present Collaborative Stream Processing (CSP), a model where the costs, which are set exogenously by providers, are shared between multiple consumers, the collaborators. For this, we identify three important requirements for CSP to establish trust between the collaborators and propose a CSP algorithm, ENCSPA, adhering to these requirements. Based on the collaborators' outcome valuations and the costs of the raw data streams, ENCSPA computes the payment for each collaborator. At the same time, ENCSPA ensures that no collaborator has an incentive to manipulate the system by providing misinformation about her/his value, budget, or time limit. We show that ENCSPA can calculate payments in a reasonable amount of time for up to one thousand collaborators.
Tobias Grubenmann, Daniele Dell'Aglio, Abraham Bernstein
IEEE BigData2
2018 Efficient Temporal Reasoning on Streams of Events with DOTR
Alessandro Margara, Gianpaolo Cugola, Dario Collavini, Daniele Dell'Aglio
ESWC4
2018 Distributed Stream Consistency Checking
Shen Gao, Daniele Dell'Aglio, Jeff Z. Pan, Abraham Bernstein
ICWE2
2018 Aligning Knowledge Base and Document Embedding Models Using Regularized Multi-Task Learning
Matthias Baumgartner, Wen Zhang 0015, Bibek Paudel, Daniele Dell'Aglio, Huajun Chen, Abraham Bernstein
ISWC (1)4
2018 VoCaLS: Vocabulary and Catalog of Linked Streams
Riccardo Tommasini 0001, Yehia Abo Sedira, Daniele Dell'Aglio, Marco Balduini, Muhammad Intizar Ali, Danh Le Phuoc, Emanuele Della Valle, Jean-Paul Calbimonte
ISWC (2)3
2017 Break the Windows: Explicit State Management for Stream Processing Systems
Alessandro Margara, Daniele Dell'Aglio, Abraham Bernstein
EDBT2
2017 Computing Authoring Tests from Competency Questions: Experimental Validation
Matt Dennis, Kees van Deemter, Daniele Dell'Aglio, Jeff Z. Pan
ISWC (1)3
2016 A Query Model to Capture Event Pattern Matching in RDF Stream Processing Query Languages
Daniele Dell'Aglio, Minh Dao-Tran, Jean-Paul Calbimonte, Danh Le Phuoc, Emanuele Della Valle
EKAW1
2016 Heaven: A Framework for Systematic Comparative Research Approach for RSP Engines
Riccardo Tommasini 0001, Emanuele Della Valle, Marco Balduini, Daniele Dell'Aglio
ESWC4
2016 When a FILTER Makes the Difference in Continuously Answering SPARQL Queries on Streaming and Quasi-Static Linked Data
Shima Zahmatkesh, Emanuele Della Valle, Daniele Dell'Aglio
ICWE3
2016 TripleWave: Spreading RDF Streams on the Web
Andrea Mauri 0001, Jean-Paul Calbimonte, Daniele Dell'Aglio, Marco Balduini, Marco Brambilla 0001, Emanuele Della Valle, Karl Aberer
ISWC (2)3
2016 Planning Ahead: Stream-Driven Linked-Data Access Under Update-Budget Constraints
Shen Gao, Daniele Dell'Aglio, Soheila Dehghanzadeh, Abraham Bernstein, Emanuele Della Valle, Alessandra Mileo
ISWC (1)2
2015 Approximate Continuous Query Answering over Streams and Dynamic Linked Data Sets
Soheila Dehghanzadeh, Daniele Dell'Aglio, Shen Gao, Emanuele Della Valle, Alessandra Mileo, Abraham Bernstein
ICWE2
2014 RSP-QL Semantics: A Unifying Query Model to Explain Heterogeneity of RDF Stream Processing Systems
abstract
RDF and SPARQL are established standards for data interchange and querying on the Web. While they have been shown to be useful and applicable in many scenarios, they are not sufficiently adequate for dealing with streams of data and their intrinsic continuous nature. In the last years data and query languages have been proposed to extend both RDF and SPARQL for streams and continuous processing, under the name of RDF Stream Processing – RSP. These efforts resulted in several models and implementations that, at a first look, appear to propose alternative syntaxes but equivalent semantics. However, when asked to continuously answer the same queries on the same data streams, they provide different answers at disparate moments due to the heterogeneity of their operational semantics. These discrepancies render the process of understanding and comparing continuous query results complex and misleading. In this work, the authors propose RSP-QL, a comprehensive model that formally defines the semantics of an RSP system. RSP-QL makes explicit the hidden assumptions of currently available RSP systems, allows defining a formal notion of correctness for RSP query results and, thus, explains why available implementations provide different answers at disparate moments.
Daniele Dell'Aglio, Emanuele Della Valle, Jean-Paul Calbimonte, Óscar Corcho
Int. J. Semantic Web Inf. Syst.1
2013 Social Listening of City Scale Events Using the Streaming Linked Data Framework
Marco Balduini, Emanuele Della Valle, Daniele Dell'Aglio, Mikalai Tsytsarau, Themis Palpanas, Cristian Confalonieri
ISWC (2)3
2013 On Correctness in RDF Stream Processor Benchmarking
Daniele Dell'Aglio, Jean-Paul Calbimonte, Marco Balduini, Óscar Corcho, Emanuele Della Valle
ISWC (2)1
2012 Linking Smart Cities Datasets with Human Computation - The Case of UrbanMatch
Irene Celino, Simone Contessa, Marta Corubolo, Daniele Dell'Aglio, Emanuele Della Valle, Stefano Fumeo, Thorsten Krüger
ISWC (2)4
2012 BOTTARI: An augmented reality mobile application to deliver personalized and location-based recommendations by continuous analysis of social media streams
Marco Balduini, Irene Celino, Daniele Dell'Aglio, Emanuele Della Valle, Yi Huang 0002, Tony Kyung-il Lee, Seon-Ho Kim, Volker Tresp
J. Web Semant.3