George Papastefanatos

dblp:44/2379 · DBLP profile ↗
← Back
35ranked-venue papers in the field
4as first author
14since 2021 · last 2026
0000-0002-9273-9843ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 25 (2 first)Big Data, Cloud & Distributed Data Systems · 3Business Process & Enterprise Data · 3 (1 first)Information Retrieval & Web Search · 2Data Mining & Knowledge Discovery · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Interpretable Highlights for Experiment Tracking
Vassilis Stamatopoulos, Panagiotis Gidarakos, Stavros Maroulis, George Papastefanatos, Panos Vassiliadis
DOLAP4
2026 EmeraldMind: A Knowledge Graph-Augmented Framework for Greenwashing Detection
abstract
As AI and web agents become pervasive in decision-making, it is critical to design intelligent systems that not only support sustainability efforts but also guard against misinformation. Greenwashing, i.e., misleading corporate sustainability claims, poses a major challenge to environmental progress. To address this challenge, we introduce EmeraldMind, a fact-centric framework integrating a domain-specific knowledge graph with retrieval-augmented generation to automate greenwashing detection. EmeraldMind builds the EmeraldGraph from diverse corporate ESG (environmental, social, and governance) reports, surfacing verifiable evidence, often missing in generic knowledge bases, and supporting large language models in claim assessment. The framework delivers justification-centric classifications, presenting transparent, evidence-backed verdicts and abstaining responsibly when claims cannot be verified. Experiments on a new greenwashing claims dataset demonstrate that EmeraldMind achieves competitive accuracy, greater coverage, and superior explanation quality compared to generic LLMs, without the need for fine-tuning or retraining.
Georgios Kaoukis, Ioannis Aris Koufopoulos, Eleni Psaroudaki, Danae Pla Karidi, Evaggelia Pitoura, George Papastefanatos, Panayiotis Tsaparas
WWW6
2025 QueryER: A Framework for Fast Analysis-Aware Deduplication over Dirty Data
George Alexiou, George Papastefanatos, Vassilis Stamatopoulos, Georgia Koutrika, Nectarios Koziris
EDBT2
2025 GLOVES: Global Counterfactual-based Visual Explanations
Panagiotis Gidarakos, Nikolas Theologitis, Stavros Maroulis, Loukas Kavouras, Giorgos Giannopoulos, George Papastefanatos
EDBT6
2024 Visualization-aware Time Series Min-Max Caching with Error Bound Guarantees
abstract
This paper addresses the challenges in interactive visual exploration of large multi-variate time series data. Traditional data reduction techniques may improve latency but can distort visualizations. State-of-the-art methods aimed at 100% accurate visualization often fail to maintain interactive response times or require excessive preprocessing and additional storage. We propose an in-memory adaptive caching approach, MinMaxCache, that efficiently reuses previous query results to accelerate visualization performance within accuracy constraints. MinMaxCache fetches data at adaptively determined aggregation granularities to maintain interactive response times and generate approximate visualizations with accuracy guarantees. Our results show that it is up to 10 times faster than current solutions without significant accuracy compromise.
Stavros Maroulis, Vassilis Stamatopoulos, George Papastefanatos, Manolis Terrovitis
Proc. VLDB Endow.3
2023 Forecasting Resource Demand for Dynamic Datacenter Sizing in Telco Infrastructures
abstract
The deployment of emerging cloud computing technologies onto telecommunication (telco) infrastructures, coupled with highly stringent regulations relating to energy efficiency, power consumption, and CO2 emissions, put telco datacenters and their associated operations under a drastic transformation process. This paves the way for new optimization opportunities, such as dynamic datacenter sizing (DDS) with respect to power consumption constraints and volatile demands across multiple resource types. DDS boils down to determining the optimal subset of active servers in a datacenter, and it can be modeled as a planning problem where decisions regarding resource availability are based on projections of resource demands. This raises the need for accurate and efficient forecasting solutions, which shall be tailored to the problem at hand.To this end, our work focuses on the investigation, the development, and the evaluation of forecasting methods which predict the demand of 5G (and beyond) workloads across multiple resources. Concretely, using the daily pattern of load data pertaining to an operational data-plane function and the prevailing way of horizontal scaling in virtualized datacenters, we exemplify the evolution of resource demands of 5G applications residing at the network edge. This lets us simulate a dataset of demands across various resource types, upon which we study informative features, and then train and evaluate a wide suite of forecasting methods. Our experiments show that the intrinsic characteristics of the data render simple statistical algorithms capable of accurately capturing the underlying patterns, and even outperforming more complex algorithms. Last, we evaluate and discuss the trade-off between single- and multi-output forecasting models.
Dimitra Paranou, Angelos Pentelas, Dimitris Katsiros, Konstantinos Maidatsis, George Giannopoulos, Evangelos Angelou, Nikos Anastopoulos, George Papastefanatos
IEEE Big Data8
2023 Auditing for Spatial Fairness
Dimitris Sacharidis, Giorgos Giannopoulos, George Papastefanatos, Kostas Stefanidis
EDBT3
2023 Resource-aware adaptive indexing for in situ visual exploration and analytics
Stavros Maroulis, Nikos Bikakis, George Papastefanatos, Panos Vassiliadis, Yannis Vassiliou
VLDB J.3
2022 Machine Learning Platform for Extreme Scale Computing on Compressed IoT Data
abstract
With the lowering costs of sensors, high-volume and high-velocity data are increasingly being generated and analyzed, especially in IoT domains like energy and smart homes. Consequently, applications that require accurate short-term forecasts and predictions are also steadily increasing. In this paper, we provide an overview of a novel end-to-end platform that provides efficient ingestion, compression, transfer, query processing, and machine learning-based analytics for high-frequency and high-volume time series from IoT. The performance of the platform is evaluated using real-world dataset from RES installations. The results show the importance of high-frequency analytics and the surprisingly positive impact of error bounded lossy compression on machine learning in the form of AutoML. For example, when detecting yaw misalignments in wind turbines, an improvement of 9% in accuracy was observed for AutoML models on lossy compressed data compared to the current industry standard of 10-minute aggregated data. Thus, these small-scale experiments show the potential of the platform, and larger pilots are planned.
Seshu Tirupathi, Dhaval Salwala, Giulio Zizzo, Ambrish Rawat, Mark Purcell, Søren Kejser Jensen, Christian Thomsen 0001, Nguyen Ho, Carlos Muñiz Cuza, Jonas Brusokas, Torben Bach Pedersen, George Alexiou, Giorgos Giannopoulos, Panagiotis Gidarakos, Alexandros Kalimeris, Stavros Maroulis, George Papastefanatos, Ioannis Psarros, Vassilis Stamatopoulos, Manolis Terrovitis
IEEE Big Data17
2022 Relational schema optimization for RDF-based knowledge graphs
abstract
Characteristic sets (CS) organize RDF triples based on the set of properties associated with their subject nodes. This concept was recently used in indexing techniques, as it can capture the implicit schema of RDF data. While most CS-based approaches yield significant improvements in space and query performance, they fail to perform well when answering complex query workloads in the presence of schema heterogeneity, i.e., when the number of CSs becomes very large, resulting in a highly partitioned data organization. In this paper, we address this problem by introducing a novel technique, for merging CSs based on their hierarchical structure. Our method employs a lattice to capture the hierarchical relationships between CSs, identifies dense CSs and merges dense CSs with their ancestors. We have implemented our algorithm on top of a relational backbone, where each merged CS is stored in a relational table, and therefore, CS merging results in a smaller number of required tables to host the source triples of a dataset. Moreover, we perform an extensive experimental study to evaluate the performance and impact of merging to the storage and querying of RDF datasets, indicating significant improvements. We also conduct a sensitivity analysis to identify the stability and any possible weaknesses of our algorithm, and report on our results.
George Papastefanatos, Marios Meimaris, Panos Vassiliadis
Inf. Syst.1
2021 Evo-Path: Querying Data Evolution through Complex Changes
Theodora Galani, Yannis Stavrakas, George Papastefanatos, Yannis Vassiliou
DATA3
2021 Adaptive Indexing for In-situ Visual Exploration and Analytics
Stavros Maroulis, Nikos Bikakis, George Papastefanatos, Panos Vassiliadis, Yannis Vassiliou
DOLAP3
2021 RawVis: A System for Efficient In-situ Visual Analytics
abstract
In-situ processing has received a great deal of attention in recent years. In in-situ scenarios, big raw data files which do not fit in main memory, must be efficiently handled on-the-fly using commodity hardware, without the overhead of a preprocessing phase or the loading of data into a database system. This paper presents RawVis, an open source data visualization system for in-situ visual exploration and analytics over big raw data. RawVis implements novel indexing schemes and adaptive processing techniques allowing users to perform efficient visual and analytics operations directly over the data files. RawVis provides real-time interaction, reporting low response time, over large data files, using commodity hardware.
Stavros Maroulis, Nikos Bikakis, George Papastefanatos, Panos Vassiliadis, Yannis Vassiliou
SIGMOD Conference3
2021 In-situ visual exploration over big raw data
Nikos Bikakis, Stavros Maroulis, George Papastefanatos, Panos Vassiliadis
Inf. Syst.3
2020 Hierarchical Property Set Merging for SPARQL Query Optimization
Marios Meimaris, George Papastefanatos, Panos Vassiliadis
DOLAP2
2020 LinkZoo: A Collaborative Resource Management Tool Based on Linked Data
abstract
This article presents LinkZoo, a web-based, linked data enabled tool that supports collaborative management of information resources. LinkZoo addresses the modern needs of information-intensive collaboration environments to publish, manage, and share heterogeneous resources within user-driven contexts. Users create and manage diverse types of resources into common spaces such as files, web documents, people, datasets, and calendar events. They can interlink them, annotate them, and share them with other users, thus enabling collaborative editing, as well as enrich them with links to externally linked data resources. Resources are inherently modeled and published as resource description framework (RDF) and can be explicitly interlinked and dereferenced by external applications. LinkZoo supports creation of dynamic communities that enable web-based collaboration through resource sharing and annotating, exposing objects on the linked data Cloud under controlled vocabularies and permissions. The authors demonstrate the applicability of the tool on a popular collaboration use case scenario for sharing and organizing research resources.
George Alexiou, Marios Meimaris, George Papastefanatos, Ioannis Anagnostopoulos
Int. J. Semantic Web Inf. Syst.3
2018 RawVis: Visual Exploration over Raw Data
Nikos Bikakis, Stavros Maroulis, George Papastefanatos, Panos Vassiliadis
ADBIS3
2018 Computational methods and optimizations for containment and complementarity in web data cubes
Marios Meimaris, George Papastefanatos, Panos Vassiliadis, Ioannis Anagnostopoulos
Inf. Syst.2
2017 Extraction of Embedded Queries via Static Analysis of Host Code
Petros Manousis, Apostolos V. Zarras, Panos Vassiliadis, George Papastefanatos
CAiSE4
2017 Introducing Solon: A Semantic Platform for Managing Legal Sources
Marios Koniaris, George Papastefanatos, Marios Meimaris, George Alexiou
TPDL2
2017 Distance-Based Triple Reordering for SPARQL Query Optimization
abstract
SPARQL query optimization relies on the design and execution of query plans that involve reordering triple patterns, in the hopes of minimizing cardinality of intermediate results. In practice, this is not always effective, as many existing systems succeed in certain types of query patterns and fail in others. This kind of trade-off is often a derivative of the algorithms behind query planning. In this paper, we introduce a novel join reordering approach that translates a query into a multidimensional vector space and performs distance-based optimization by taking into account the relative differences between the triple patterns. Preliminary experiments on synthetic data show that our algorithm consistently outperforms established methodologies, providing better plans for many different types of query patterns.
Marios Meimaris, George Papastefanatos
ICDE2
2017 Extended Characteristic Sets: Graph Indexing for SPARQL Query Optimization
abstract
SPARQL query execution in state of the art RDF engines depends on, and is often limited by the underlying storage and indexing schemes. Typically, these systems exhaustively store permutations of the standard three-column triples table. However, even though RDF can give birth to datasets with loosely defined schemas, it is common for an emerging structure to appear in the data. In this paper, we introduce a novel indexing scheme for RDF data, that takes advantage of the inherent structure of triples. To this end, we define the Extended Characteristic Set (ECS), a schema abstraction that classifies triples based on the properties of their subjects and objects, and we discuss methods and algorithms for the identification and extraction of ECSs. We show how these can be used to assist query processing, and we implement axonDB, an RDF storage and querying engine based on ECS indexing. We perform an experimental evaluation on real world and synthetic datasets and observe that axonDB outperforms the competition by a few orders of magnitude.
Marios Meimaris, George Papastefanatos, Nikos Mamoulis, Ioannis Anagnostopoulos
ICDE2
2017 Parallel meta-blocking for scaling entity resolution over big heterogeneous data
Vasilis Efthymiou, George Papadakis 0001, George Papastefanatos, Kostas Stefanidis, Themis Palpanas
Inf. Syst.3
2016 Scaling Entity Resolution to Large, Heterogeneous Data with Enhanced Meta-blocking
abstract
Entity Resolution constitutes a quadratic task that typically scales to large entity collections through blocking. The resulting blocks can be restructured by Meta-blocking in order to significantly increase precision at a limited cost in recall. Yet, its processing can be time-consuming, while its precision remains poor for configurations with high recall. In this work, we propose new meta-blocking methods that improve precision by up to an order of magnitude at a negligible cost to recall. We also introduce two efficiency techniques that, when combined, reduce the overhead time of Metablocking by more than an order of magnitude. We evaluate our approaches through an extensive experimental study over 6 realworld, heterogeneous datasets. The outcomes indicate that our new algorithms outperform all meta-blocking techniques as well as the state-of-the-art methods for block processing in all respects.
George Papadakis 0001, George Papastefanatos, Themis Palpanas, Manolis Koubarakis
EDBT2
2016 Double Chain-Star: an RDF indexing scheme for fast processing of SPARQL joins
abstract
State of the art RDF stores often rely on exhaustive indexing and sequential (self-)joins for SPARQL query processing. However, query execution is dependent on, and often limited by the underlying storage and indexing schemes. Even though RDF can give birth to datasets with loosely defined schemas, it is common for an emerging structure to be present in the data. In this paper we introduce a novel indexing scheme, called Double Chain Star (DCS), that takes advantage of the inherent structure that is often found in RDF datasets by extending the notion of Characteristic Sets to cater for chain-star joins. DCS essentially reduces pairs of chain-star patterns that typically involve multiple self-joins, to mere index scans. We perform preliminary experiments and show promising results in comparison with Jena TDB and RDF-3X. © 2016, Copyright is with the authors.
Marios Meimaris, George Papastefanatos
EDBT2
2016 Efficient Computation of Containment and Complementarity in RDF Data Cubes
abstract
Multidimensional data are published in the web of data under common directives, such as the Resource Description Framework (RDF). The increasing volume and diversity of these data pose the challenge of finding relations between them in a most efficient and accurate way, by taking into advantage their overlapping schemes. In this paper we define two types of relationships between multidimensional RDF data, and we propose algorithms for efficient and scalable computation of these relationships. Specifically, we define the notions of containment and complementarity between points in multidimensional dataspaces, as different aspects of relatedness, and we propose a baseline method for computing them, as well as two alternative methods that target speed and scalability. We provide an experimental evaluation over real-world and synthetic datasets and we compare our approach to a SPARQL-based and a rule-based alternative, which prove to be inefficient for increasing input sizes. © 2016, Copyright is with the authors.
Marios Meimaris, George Papastefanatos, Panos Vassiliadis, Ioannis Anagnostopoulos
EDBT2
2016 graphVizdb: A scalable platform for interactive large graph visualization
abstract
We present a novel platform for the interactive visualization of very large graphs. The platform enables the user to interact with the visualized graph in a way that is very similar to the exploration of maps at multiple levels. Our approach involves an offline preprocessing phase that builds the layout of the graph by assigning coordinates to its nodes with respect to a Euclidean plane. The respective points are indexed with a spatial data structure, i.e., an R-tree, and stored in a database. Multiple abstraction layers of the graph based on various criteria are also created offline, and they are indexed similarly so that the user can explore the dataset at different levels of granularity, depending on her particular needs. Then, our system translates user operations into simple and very efficient spatial operations (i.e., window queries) in the backend. This technique allows for a fine-grained access to very large graphs with extremely low latency and memory requirements and without compromising the functionality of the tool. Our web-based prototype supports three main operations: (1) interactive navigation, (2) multi-level exploration, and (3) keyword search on the graph metadata.
Nikos Bikakis, John Liagouris, Maria Krommyda, George Papastefanatos, Timos K. Sellis
ICDE4
2015 Parallel meta-blocking: Realizing scalable entity resolution over large, heterogeneous data
abstract
Entity resolution constitutes a crucial task for many applications, but has an inherently quadratic complexity. Typically, it scales to large volumes of data through blocking: similar entities are clustered into blocks so that it suffices to perform comparisons only within each block. Meta-blocking further increases efficiency by cleaning the overlapping blocks from unnecessary comparisons. However, even Meta-blocking can be time-consuming: applying it to blocks with 7.4 million entities and 2.21011 comparisons takes almost 8 days on a modern high-end server. In this paper, we parallelize Meta-blocking based on MapReduce. We propose a simple strategy that explicitly creates the core concept of Meta-blocking, the blocking graph. We then describe an advanced strategy that creates the blocking graph implicitly, reducing the overhead of data exchange. We also introduce a load balancing algorithm that distributes the computationally intensive workload evenly among the available compute nodes. Our experimental analysis verifies the superiority of our advanced strategy and demonstrates an almost linear speedup for all meta-blocking techniques with respect to the number of available nodes.
Vasilis Efthymiou, George Papadakis 0001, George Papastefanatos, Kostas Stefanidis, Themis Palpanas
IEEE BigData3
2015 RDF Resource Search and Exploration with LinkZoo
abstract
The Linked Data paradigm is the most common practice for publishing, sharing and managing information in the Data Web. Linkzoo is an IT infrastructure for collaborative publishing, annotating and sharing of Data Web resources, and their publication as Linked Data. In this paper, we overview LinkZoo and its main components, and we focus on the search facilities provided to retrieve and explore RDF resources. Two search services are presented: (1) an interactive, two-step keyword search service, where live natural language query suggestions are given to the user based on the input keywords and the resource types they match within LinkZoo, and (2) a keyword search service for exploring remote SPARQL endpoints that automatically generates a set of candidate SPARQL queries, i.e., SPARQL queries that try to capture user's information needs as expressed by the keywords used. Finally, we demonstrate the search functionalities through a use case drawn from the life sciences domain.
Marios Meimaris, George Alexiou, Katerina Gkirtzou, George Papastefanatos, Theodore Dalamagas 0001
DATA4
2015 Schema-agnostic vs Schema-based Configurations for Blocking Methods on Homogeneous Data
abstract
Entity Resolution constitutes a core task for data integration that, due to its quadratic complexity, typically scales to large datasets through blocking methods. These can be configured in two ways. The schema-based configuration relies on schema information in order to select signatures of high distinctiveness and low noise, while the schema-agnostic one treats every token from all attribute values as a signature. The latter approach has significant potential, as it requires no fine-tuning by human experts and it applies to heterogeneous data. Yet, there is no systematic study on its relative performance with respect to the schema-based configuration. This work covers this gap by comparing analytically the two configurations in terms of effectiveness, time efficiency and scalability. We apply them to 9 established blocking methods and to 11 benchmarks of structured data. We provide valuable insights into the internal functionality of the blocking methods with the help of a novel taxonomy. Our studies reveal that the schema-agnostic configuration offers unsupervised and robust definition of blocking keys under versatile settings, trading a higher computational cost for a consistently higher recall than the schema-based one. It also enables the use of state-of-the-art blocking methods without schema knowledge.
George Papadakis 0001, George Alexiou, George Papastefanatos, Georgia Koutrika
Proc. VLDB Endow.3
2014 Supervised Meta-blocking
abstract
Entity Resolution matches mentions of the same entity. Being an expensive task for large data, its performance can be improved by blocking, i.e., grouping similar entities and comparing only entities in the same group. Blocking improves the run-time of Entity Resolution, but it still involves unnecessary comparisons that limit its performance. Meta-blocking is the process of restructuring a block collection in order to prune such comparisons. Existing unsupervised meta-blocking methods use simple pruning rules, which offer a rather coarse-grained filtering technique that can be conservative (i.e., keeping too many unnecessary comparisons) or aggressive (i.e., pruning good comparisons). In this work, we introduce supervised meta-blocking techniques that learn classification models for distinguishing promising comparisons. For this task, we propose a small set of generic features that combine a low extraction cost with high discriminatory power. We show that supervised meta-blocking can achieve high performance with small training sets that can be manually created. We analytically compare our supervised approaches with baseline and competitor methods over 10 large-scale datasets, both real and synthetic.
George Papadakis 0001, George Papastefanatos, Georgia Koutrika
Proc. VLDB Endow.2
2013 Automating the Adaptation of Evolving Data-Intensive Ecosystems
Petros Manousis, Panos Vassiliadis, George Papastefanatos
ER3
2010 HECATAEUS: Regulating schema evolution
abstract
HECATAEUS is an open-source software tool for enabling impact prediction, what-if analysis, and regulation of relational database schema evolution. We follow a graph theoretic approach and represent database schemas and database constructs, like queries and views, as graphs. Our tool enables the user to create hypothetical evolution events and examine their impact over the overall graph before these are actually enforced on it. It also allows definition of rules for regulating the impact of evolution via (a) default values for all the nodes of the graph and (b) simple annotations for nodes deviating from the default behavior. Finally, HECATAEUS includes a metric suite for evaluating the impact of evolution events and detecting crucial and vulnerable parts of the system.
George Papastefanatos, Panos Vassiliadis, Alkis Simitsis, Yannis Vassiliou
ICDE1
2008 Design Metrics for Data Warehouse Evolution
George Papastefanatos, Panos Vassiliadis, Alkis Simitsis, Yannis Vassiliou
ER1
2007 What-If Analysis for Data Warehouse Evolution
George Papastefanatos, Panos Vassiliadis, Alkis Simitsis, Yannis Vassiliou
DaWaK1