Ekaterini Ioannou

dblp:68/3366 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
4since 2021 · last 2025
0000-0002-4922-6639ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 19 · 8 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 pyJedAI: A Library with Resolution-Related Structures and Procedures for Products
abstract
This work presents an open-source Python library, named pyJedAI, which provides functionalities supporting the creation of algorithms related to product entity resolution. Building over existing state-of-the-art resolution algorithms, the tool offers a plethora of important tasks required for processing product data collections. It can be easily used by researchers and practitioners for creating algorithms analyzing products, such as real-time ad bidding, sponsored search, or pricing determination. In essence, it allows users to easily import product data from the possible sources, compare products in order to detect either similar or identical products, generate a graph representation using the products and desired relationships, and either visualize or export the outcome in various forms. Our experimental evaluation on data from well-known online retailers illustrates high accuracy and low execution time for the supported tasks. To the best of our knowledge, this is the first Python package to focus on product entities and provide this range of product entity resolution functionalities. History: Accepted by Ted Ralphs, Area Editor for Software Tools. Funding: This was partially funded by the EU project STELAR (Horizon Europe) [Grant 101070122]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0410 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0410 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Ekaterini Ioannou, Konstantinos Nikoletos, George Papadakis 0001
INFORMS J. Comput.1
2024 Unveiling Dis-Integration
abstract
Entity Resolution (ER) has been extensively studied over the last decade, with a plethora of algorithmic solutions, techniques, and methodologies having been proposed [1]. The individual state-of-the-art ER algorithms are offered through open-source systems, such as Magellan [2] and JedAI [3], which typically implement end-to-end solutions through a sequence of workflow steps. Each workflow step requires its own special configuration and fine tuning, thus turning the creation of complete ER solutions into a non-trivial, time-consuming process that requires adapting, among others, to the characteristics of the data to be resolved (e.g., relational, semi-structured, etc.), to its intrinsic noise (e.g., misspellings, abbreviations, etc.) as well as to application constraints (e.g., execution time).
George Papadakis 0001, Ekaterini Ioannou, Yannis Velegrakis
ICDE2
2024 The Five Generations of Entity Resolution on Web Data
Konstantinos Nikoletos, Ekaterini Ioannou, George Papadakis 0001
ICWE2
2023 Self-configured Entity Resolution with pyJedAI
abstract
Entity Resolution has been an active research topic for the last three decades, with numerous algorithms proposed in the literature. However, putting them into practice is often a complex task that requires implementing, combining and configuring complementary individual algorithms into comprehensive end-to-end workflows. To facilitate this process, we are developing pyJedAI, a novel system that provides a unifying framework for any type of main works in the field (i.e., both unsupervised and learning-based ones). Our vision is to facilitate both novice and expert users to use and combine these algorithms through a series of principled approaches for automatically configuring and benchmarking end-to-end pipelines.
Vasilis Efthymiou, Ekaterini Ioannou, Manos Karvounis, Manolis Koubarakis, Jakub Maciejewski, Konstantinos Nikoletos, George Papadakis 0001, Dimitrios Skoutas 0001, Yannis Velegrakis, Alexandros Zeakis
IEEE Big Data2
2020 Entity Resolution: Past, Present and Yet-to-Come
abstract
Entity Resolution (ER) lies at the core of data integration, with a bulk of research focusing on its effectiveness and its time efficiency. Most past relevant works were crafted for addressing Veracity over structured (relational) data. They typically rely on schema, expert and external knowledge to maximize accuracy. Part of these methods have been recently extended to process large volumes of data through massive parallelization techniques, such as the MapReduce paradigm. With the present advent of Big Web Data, the scope moved towards Variety, aiming to handle semi-structured data collections, with noisy and highly heterogeneous information. Relevant works adopt a novel, loosely schema-aware functionality that emphasizes scalability and robustness to noise. Another line of present research focuses on Velocity, i.e., processing data collections of a continuously increasing volume. In this tutorial, we present the ER generations by discussing past, present, and yet-to-come mechanisms. For each generation, we outline the corresponding ER workflow along with the state-of-the-art methods per workflow step. Thus, we provide the participants with a deep understanding of the broad field of ER, highlighting the recent advances in crowd-sourcing and deep learning applications in this active research domain. We also equip them with practical skills in applying ER workflows through a hands-on session that involves our publicly available ER toolbox and data.
George Papadakis 0001, Ekaterini Ioannou, Themis Palpanas
EDBT2
2017 Holistic Query Evaluation over Information Extraction Pipelines
abstract
We introduce holistic in-database query processing over information extraction pipelines. This requires considering the joint conditional distribution over generic Conditional Random Fields that uses factor graphs to encode extraction tasks. Our approach introduces Canopy Factor Graphs , a novel probabilistic model for effectively capturing the joint conditional distribution given a canopy clustering of the data, and special query operators for retrieving resolution information. Since inference on such models is intractable, we introduce an approximate technique for query processing and optimizations that cut across the integrated tasks for reducing the required processing time. Effectiveness and scalability are verified through an extensive experimental evaluation using real and synthetic data.
Ekaterini Ioannou, Minos N. Garofalakis
Proc. VLDB Endow.1
2015 Query Analytics over Probabilistic Databases with Unmerged Duplicates
abstract
Recent entity resolution approaches exhibit benefits when addressing the problem through unmerged duplicates: instances describing real-world objects are not merged based on apriori thresholds or human intervention, instead relevant resolution information is employed for evaluating resolution decisions during query processing using “possible worlds” semantics. In this paper, we present the first known approach for efficiently handling complex analytical queries over probabilistic databases with unmerged duplicates. We propose the ENTITY-JOIN operator that allows expressing complex aggregation and iceberg/top-k queries over joins between tables with unmerged duplicates and other database tables. Our technical content includes a novel indexing structure for efficient access to the entity resolution information and novel techniques for the efficient evaluation of complex probabilistic queries that retrieve analytical and summarized information over a (potentially, huge) collection of possible resolution worlds. Our extensive experimental evaluation verifies the benefits of our approach.
Ekaterini Ioannou, Minos N. Garofalakis
IEEE Trans. Knowl. Data Eng.1
2013 A Blocking Framework for Entity Resolution in Highly Heterogeneous Information Spaces
abstract
In the context of entity resolution (ER) in highly heterogeneous, noisy, user-generated entity collections, practically all block building methods employ redundancy to achieve high effectiveness. This practice, however, results in a high number of pairwise comparisons, with a negative impact on efficiency. Existing block processing strategies aim at discarding unnecessary comparisons at no cost in effectiveness. In this paper, we systemize blocking methods for clean-clean ER (an inherently quadratic task) over highly heterogeneous information spaces (HHIS) through a novel framework that consists of two orthogonal layers: the effectiveness layer encompasses methods for building overlapping blocks with small likelihood of missed matches; the efficiency layer comprises a rich variety of techniques that significantly restrict the required number of pairwise comparisons, having a controllable impact on the number of detected duplicates. We map to our framework all relevant existing methods for creating and processing blocks in the context of HHIS, and additionally propose two novel techniques: attribute clustering blocking and comparison scheduling. We evaluate the performance of each layer and method on two large-scale, real-world data sets and validate the excellent balance between efficiency and effectiveness that they achieve.
George Papadakis 0001, Ekaterini Ioannou, Themis Palpanas, Claudia Niederée, Wolfgang Nejdl
IEEE Trans. Knowl. Data Eng.2
2012 Beyond 100 million entities: large-scale blocking-based resolution for heterogeneous data
abstract
A prerequisite for leveraging the vast amount of data available on the Web is Entity Resolution, i.e., the process of identifying and linking data that describe the same real-world objects. To make this inherently quadratic process applicable to large data sets, blocking is typically employed: entities (records) are grouped into clusters - the blocks - of matching candidates and only entities of the same block are compared. However, novel blocking techniques are required for dealing with the noisy, heterogeneous, semi-structured, user-generateddata in the Web, as traditional blocking techniques are inapplicable due to their reliance on schema information. The introduction of redundancy, improves the robustness of blocking methods but comes at the price of additional computational cost.
George Papadakis 0001, Ekaterini Ioannou, Claudia Niederée, Themis Palpanas, Wolfgang Nejdl
WSDM2
2011 Efficient discovery of frequent subgraph patterns in uncertain graph databases
abstract
Mining frequent subgraph patterns in graph databases is a challenging and important problem with applications in several domains. Recently, there is a growing interest in generalizing the problem to uncertain graphs, which can model the inherent uncertainty in the data of many applications. The main difficulty in solving this problem results from the large number of candidate subgraph patterns to be examined and the large number of subgraph isomorphism tests required to find the graphs that contain a given pattern. The latter becomes even more challenging, when dealing with uncertain graphs. In this paper, we propose a method that uses an index of the uncertain graph database to reduce the number of comparisons needed to find frequent subgraph patterns. The proposed algorithm relies on the apriori property for enumerating candidate subgraph patterns efficiently. Then, the index is used to reduce the number of comparisons required for computing the expected support of each candidate pattern. It also enables additional optimizations with respect to scheduling and early termination, that further increase the efficiency of the method. The evaluation of our approach on three real-world datasets as well as on synthetic uncertain graph databases demonstrates the significant cost savings with respect to the state-of-the-art approach.
Odysseas Papapetrou, Ekaterini Ioannou, Dimitrios Skoutas 0001
EDBT2
2011 LinkDB: a probabilistic linkage database system
abstract
Entity linkage deals with the problem of identifying whether two pieces of information represent the same real world object. The traditional methodology computes the similarity among the entities, and then merges those with similarity above some specific threshold. We demonstrate LinkDB, an original entity storage and querying system that deals with the entity linkage problem in a novel way. LinkDB is a probabilistic linkage database that uses existing linkage techniques to generate linkages among entities, but instead of performing the merges based on these linkages, it stores them alongside the data and performs only the required merges at run-time, by effectively taking into consideration the query specifications. We explain the technical challenges behind this kind of query answering, and we show how this new mechanism is able to provide answers that traditional entity linkage mechanisms cannot.
Ekaterini Ioannou, Wolfgang Nejdl, Claudia Niederée, Yannis Velegrakis
SIGMOD Conference1
2011 Efficient entity resolution for large heterogeneous information spaces
abstract
We have recently witnessed an enormous growth in the volume of structured and semi-structured data sets available on the Web. An important prerequisite for using and combining such data sets is the detection and merge of information that describes the same real-world entities, a task known as Entity Resolution. To make this quadratic task efficient, blocking techniques are typically employed. However, the high dynamics, loose schema binding, and heterogeneity of (semi-)structured data, impose new challenges to entity resolution. Existing blocking approaches become inapplicable because they rely on the homogeneity of the considered data and a-priory known schemata. In this paper, we introduce a novel approach for entity resolution, scaling it up for large, noisy, and heterogeneous information spaces. It combines an attribute-agnostic mechanism for building blocks with intelligent block processing techniques that boost blocks with high expected utility, propagate knowledge about identified matches, and preempt the resolution process when it gets too expensive. Our extensive evaluation on real-world, large, heterogeneous data sets verifies that the suggested approach is both effective and efficient.
George Papadakis 0001, Ekaterini Ioannou, Claudia Niederée, Peter Fankhauser
WSDM2
2010 From Web Data to Entities and Back
Zoltán Miklós 0001, Nicolas Bonvin, Paolo Bouquet, Michele Catasta, Daniele Cordioli, Peter Fankhauser, Julien Gaugaz, Ekaterini Ioannou, Hristo Koshutanski, Antonio Maña
CAiSE8
2010 Efficient Semantic-Aware Detection of Near Duplicate Resources
Ekaterini Ioannou, Odysseas Papapetrou, Dimitrios Skoutas 0001, Wolfgang Nejdl
ESWC (2)1
2010 Efficient Term Cloud Generation for Streaming Web Content
Odysseas Papapetrou, George Papadakis 0001, Ekaterini Ioannou, Dimitrios Skoutas 0001
ICWE3
2010 Enabling entity-based aggregators for web 2.0 data
abstract
Selecting and presenting content culled from multiple heterogeneous and physically distributed sources is a challenging task. The exponential growth of the web data in modern times has brought new requirements to such integration systems. Data is not any more produced by content providers alone, but also from regular users through the highly popular Web 2.0 social and semantic web applications. The plethora of the available web content increased its demand by regular users who could not any more wait the development of advanced integration tools. They wanted to be able to build in a short time their own specialized integration applications. Aggregators came to the risk of these users. They allowed them not only to combine distributed content, but also to process it in ways that generate new services available for further consumption.
Ekaterini Ioannou, Claudia Niederée, Yannis Velegrakis
WWW1
2010 On-the-Fly Entity-Aware Query Processing in the Presence of Linkage
abstract
Entity linkage is central to almost every data integration and data cleaning scenario. Traditional techniques use some computed similarity among data structure to perform merges and then answer queries on the merged data. We describe a novel framework for entity linkage with uncertainty. Instead of using the linkage information to merge structures a-priori, possible linkages are stored alongside the data with their belief value. A new probabilistic query answering technique is used to take the probabilistic linkage into consideration. The framework introduces a series of novelties: (i) it performs merges at run time based not only on existing linkages but also on the given query; (ii) it allows results that may contain structures not explicitly represented in the data, but generated as a result of a reasoning on the linkages; and (iii) enables an evaluation of the query conditions that spans across linked structures, offering a functionality not currently supported by any traditional probabilistic databases. We formally define the semantics, describe an efficient implementation and report on the findings of our experimental evaluation.
Ekaterini Ioannou, Wolfgang Nejdl, Claudia Niederée, Yannis Velegrakis
Proc. VLDB Endow.1
2010 Leveraging personal metadata for Desktop search: The Beagle++ system
Enrico Minack, Raluca Paiu, Stefania Costache 0001, Gianluca Demartini, Julien Gaugaz, Ekaterini Ioannou, Paul-Alexandru Chirita, Wolfgang Nejdl
J. Web Semant.6
2009 Entity Search with NECESSITY
Ekaterini Ioannou, Saket Sathe 0001, Nicolas Bonvin, Anshul Jain, Srikanth Bondalapati, Gleb Skobeltsyn, Claudia Niederée, Zoltán Miklós 0001
WebDB1
2008 Probabilistic Entity Linkage for Heterogeneous Information Spaces
Ekaterini Ioannou, Claudia Niederée, Wolfgang Nejdl
CAiSE1