Manolis Koubarakis

dblp:40/4517 · DBLP profile ↗
← Back
83ranked-venue papers in the field
9as first author
21since 2021 · last 2026
0000-0002-1954-8338ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 38 (7 first)Knowledge Engineering, Semantic Web & Information Systems · 30 (1 first)Information Retrieval & Web Search · 12Big Data, Cloud & Distributed Data Systems · 2Business Process & Enterprise Data · 1 (1 first)
YearPublicationVenuePosition
2026 Privacy-preserving Record Linkage: Past, Present and Yet-to-Come
Lefteris Stetsikas, Dimitrios Karapiperis, George Papadakis 0001, Manolis Koubarakis
EDBT4
2026 ChatMatcher: End-to-end entity resolution with 7B LLM-based matching
Ioannis Arvanitis-Kasinikos, Lefteris Stetsikas, George Papadakis 0001, Manolis Koubarakis
Inf. Syst.4
2025 AvengER: Ensembling and Fine-Tuning LLMs for SELECT Prompts in Entity Resolution
Alexandros Zeakis, George Papadakis 0001, Dimitrios Skoutas 0001, Manolis Koubarakis
ESWC (1)4
2025 An in-depth analysis of pre-trained embeddings for entity resolution
Alexandros Zeakis, George Papadakis 0001, Dimitrios Skoutas 0001, Manolis Koubarakis
VLDB J.4
2024 Generating a Question Answering Dataset About Geographic Changes in a Knowledge Graph
Michalis Mitsios, Dharmen Punjani, Sara Abdollahi, Simon Gottschalk 0001, Eleni Tsalapati, Elena Demidova, Manolis Koubarakis
EKAW7
2024 Transformers in the Service of Description Logic-Based Contexts
Angelos Poulis, Eleni Tsalapati, Manolis Koubarakis
EKAW3
2024 Open benchmark for filtering techniques in entity resolution
Franziska Neuhof, Marco Fisichella, George Papadakis 0001, Konstantinos Nikoletos, Nikolaus Augsten, Wolfgang Nejdl, Manolis Koubarakis
VLDB J.7
2024 Three-dimensional Geospatial Interlinking with JedAI-spatial
abstract
Geospatial data constitutes a considerable part of Semantic Web data, but so far, its sources are inadequately interlinked in the Linked Open Data cloud. Geospatial Interlinking aims to cover this gap by associating geometries with topological relations like those of the Dimensionally Extended 9-Intersection Model. Due to its quadratic time complexity, various algorithms aim to carry out Geospatial Interlinking efficiently. We present JedAI-spatial, a novel, open-source system that organizes these algorithms according to three dimensions: (i) Space Tiling, which determines the approach that reduces the search space, (ii) Budget-awareness, which distinguishes interlinking algorithms into batch and progressive ones, and (iii) Execution mode, which discerns between serial algorithms, running on a single CPU-core, and parallel ones, running on top of Apache Spark. We analytically describe JedAI-spatial’s architecture and capabilities and perform thorough experiments to provide interesting insights about the relative performance of its algorithms.
Marios Papamichalopoulos, George Papadakis 0001, Georgios M. Mandilaras, Maria Despoina Siampou, Nikos Mamoulis, Manolis Koubarakis
J. Web Semant.6
2023 Self-configured Entity Resolution with pyJedAI
abstract
Entity Resolution has been an active research topic for the last three decades, with numerous algorithms proposed in the literature. However, putting them into practice is often a complex task that requires implementing, combining and configuring complementary individual algorithms into comprehensive end-to-end workflows. To facilitate this process, we are developing pyJedAI, a novel system that provides a unifying framework for any type of main works in the field (i.e., both unsupervised and learning-based ones). Our vision is to facilitate both novice and expert users to use and combine these algorithms through a series of principled approaches for automatically configuring and benchmarking end-to-end pipelines.
Vasilis Efthymiou, Ekaterini Ioannou, Manos Karvounis, Manolis Koubarakis, Jakub Maciejewski, Konstantinos Nikoletos, George Papadakis 0001, Dimitrios Skoutas 0001, Yannis Velegrakis, Alexandros Zeakis
IEEE Big Data4
2023 ExtremeEarth: Managing Water Availability for Crops Using Earth Observation and Machine Learning
Florian Appel, Heike Bach, Silke Migdall, Manolis Koubarakis, George Stamoulis 0001, Dimitris Bilidas, Despina-Athanasia Pantazi, Lorenzo Bruzzone, Claudia Paris, Giulio Weikmann
EDBT4
2023 Fire Risk Management using Data Cubes, Machine Learning and OBDA systems
abstract
We present a fire risk management system which takes input data from various sources (e.g., meteorological data, satellite indicators for vegetation, historical burned areas), produces a harmonized spatio-temporal data cube to compute fire risk and enables semantic querying to assist fire risk management. The distinguishing implementation features of the system is the use of data cubes, machine learning algorithms and, most importantly, geospatial ontology-based data access technologies. The system has been implemented in the European project DeepCube for the geographic area of Greece and can be used operationally to assist authorities to determine fire risk during the summer fire season.
Dimitris Bilidas, Anastasios Mantas, Filippos Yfantis, George Stamoulis 0001, Manolis Koubarakis, Spyros Kondylatos, Ioannis Prapas, Ioannis Papoutsis
SIGSPATIAL/GIS5
2023 Supervised Scheduling for Geospatial Interlinking
abstract
Geospatial Interlinking constitutes a crucial data integration task that associates pairs of geometries with topological relations. Its high computational cost, though, scales poorly to voluminous datasets. Progressive methods were recently proposed to reduce this cost by sacrificing recall to an affordable extent. They operate in a learning-free manner that relies on mere heuristics, which can be conservative (i.e., retaining too many unrelated pairs) or aggressive (i.e., discarding too many related pairs). In this work, we extend them with Supervised Scheduling, a quick and principled way of defining the processing order of the candidate geometry pairs that are likely to be topologically related, based on their classification probability. Our approach leverages generic features with low extraction cost but high discriminatory power. We integrate Supervised Scheduling into a progressive end-to-end algorithm that automatically labels the required training instances at a low computational cost. Thorough experiments verify the high performance and robustness of our features as well as the limited size of the training set that suffices for learning an accurate classification model. Our experiments also verify the superior performance of our approach in comparison to existing learning-free ones over five real, large datasets.
Maria Despoina Siampou, George Papadakis 0001, Nikos Mamoulis, Manolis Koubarakis
SIGSPATIAL/GIS4
2023 Benchmarking Geospatial Question Answering Engines Using the Dataset GeoQuestions1089
Sergios-Anestis Kefalidis, Dharmen Punjani, Eleni Tsalapati, Konstantinos Plas, Mariangela Pollali, Michail Mitsios, Myrto Tsokanaridou, Manolis Koubarakis, Pierre Maret
ISWC8
2023 Pre-trained Embeddings for Entity Resolution: An Experimental Analysis
abstract
Many recent works on Entity Resolution (ER) leverage Deep Learning techniques involving language models to improve effectiveness. This is applied to both main steps of ER, i.e., blocking and matching. Several pre-trained embeddings have been tested, with the most popular ones being fastText and variants of the BERT model. However, there is no detailed analysis of their pros and cons. To cover this gap, we perform a thorough experimental analysis of 12 popular language models over 17 established benchmark datasets. First, we assess their vectorization overhead for converting all input entities into dense embeddings vectors. Second, we investigate their blocking performance, performing a detailed scalability analysis, and comparing them with the state-of-the-art deep learning-based blocking method. Third, we conclude with their relative performance for both supervised and unsupervised matching. Our experimental results provide novel insights into the strengths and weaknesses of the main language models, facilitating researchers and practitioners to select the most suitable ones in practice.
Alexandros Zeakis, George Papadakis 0001, Dimitrios Skoutas 0001, Manolis Koubarakis
Proc. VLDB Endow.4
2022 Strabo 2: Distributed Management of Massive Geospatial RDF Datasets
Dimitris Bilidas, Theofilos Ioannidis, Nikos Mamoulis, Manolis Koubarakis
ISWC4
2022 TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching
abstract
Set similarity join is an important problem with many applications in data discovery, cleaning and integration. To increase robustness, fuzzy set similarity join calculates the similarity of two sets based on maximum weighted bipartite matching instead of set overlap. This allows pairs of elements, represented as sets or strings, to also match approximately rather than exactly, e.g., based on Jaccard similarity or edit distance. However, this significantly increases the verification cost, making even more important the need for efficient and effective filtering techniques to reduce the number of candidate pairs. The current state-of-the-art algorithm relies on similarity computations between pairs of elements to filter candidates. In this paper, we propose token-based instead of element-based filtering, showing that it is significantly more lightweight, while offering similar or even better pruning effectiveness. Moreover, we address the top- k variant of the problem, alleviating the need for a user-specified similarity threshold. We also propose early termination to reduce the cost of verification. Our experimental results on six real-world datasets show that our approach always outperforms the state of the art, being an order of magnitude faster on average.
Alexandros Zeakis, Dimitrios Skoutas 0001, Dimitris Sacharidis, Odysseas Papapetrou, Manolis Koubarakis
Proc. VLDB Endow.5
2021 Scalable Transformation of Big Geospatial Data into Linked Data
Georgios M. Mandilaras, Manolis Koubarakis
ISWC2
2021 Progressive, Holistic Geospatial Interlinking
abstract
Geospatial data constitute a considerable part of Semantic Web data, but at the moment, its sources are inadequately interlinked with topological relations in the Linked Open Data cloud. Geospatial Interlinking covers this gap with batch techniques that are restricted to individual topological relations, even though most operations are common for all main relations. In this work, we introduce a batch algorithm that simultaneously computes all topological relations and define the task of Progressive Geospatial Interlinking, which produces results in a pay-as-you-go manner when the available computational or temporal resources are limited. We propose two progressive algorithms and conduct a thorough experimental study over large, real datasets, demonstrating the superiority of our techniques over the current state-of-the-art.
George Papadakis 0001, Georgios M. Mandilaras, Nikos Mamoulis, Manolis Koubarakis
WWW4
2021 In-memory parallelization of join queries over large ontological hierarchies
Dimitris Bilidas, Manolis Koubarakis
Distributed Parallel Databases2
2021 Reproducible experiments on Three-Dimensional Entity Resolution with JedAI
Georgios M. Mandilaras, George Papadakis 0001, Luca Gagliardelli, Giovanni Simonini, Emmanouil Thanos, George Giannakopoulos, Sonia Bergamaschi, Themis Palpanas, Manolis Koubarakis, Alicia Lara-Clares, Antonio Fariña
Inf. Syst.9
2021 Handling redundant processing in OBDA query execution over relational sources
Dimitris Bilidas, Manolis Koubarakis
J. Web Semant.2
2020 JedAI3 : beyond batch, blocking-based Entity Resolution
abstract
JedAI is an open-source toolkit that allows for building and benchmarking thousands of schema-agnostic Entity Resolution (ER) pipelines through a non-learning, blocking-based end-to-end workflow. In this paper, we present its latest release, JedAI3 , which conveys two new end-to-end workflows: one for budgetagnostic ER that is based on similarity joins, and one for budgetaware (i.e., progressive) ER. This version also adds support for pre-trained word or character embeddings and connects JedAI to the Python data analysis ecosystem. Overall, these enhancements provide JedAI with features offered by no other ER tool, especially in the schema- and domain-agnostic context.
George Papadakis 0001, Leonidas Tsekouras, Emmanouil Thanos, Nikiforos Pittaras, Giovanni Simonini, Dimitrios Skoutas 0001, Paul Isaris, George Giannakopoulos, Themis Palpanas, Manolis Koubarakis
EDBT10
2020 Three-dimensional Entity Resolution with JedAI
George Papadakis 0001, Georgios M. Mandilaras, Luca Gagliardelli, Giovanni Simonini, Emmanouil Thanos, George Giannakopoulos, Sonia Bergamaschi, Themis Palpanas, Manolis Koubarakis
Inf. Syst.9
2019 The Copernicus App Lab project: Easy Access to Copernicus Data
Konstantina Bereta, Hervé Caumont, Ulrike Daniels, Erwin Goor, Manolis Koubarakis, Despina-Athanasia Pantazi, George Stamoulis 0001, Sam Ubels, Valentijn Venus, Firman Wahyudi
EDBT5
2019 Scalable Parallelization of RDF Joins on Multicore Architectures
Dimitris Bilidas, Manolis Koubarakis
EDBT2
2019 From Copernicus Big Data to Extreme Earth Analytics
abstract
Copernicus is the European programme for monitoring the Earth.It consists of a set of systems that collect data from satellites and in-situ sensors, process this data and provide users with reliable and up-to-date information on a range of environmental and security issues.The data and information processed and disseminated puts Copernicus at the forefront of the big data paradigm, giving rise to all relevant challenges, the so-called 5 Vs: volume, velocity, variety, veracity and value.In this short paper, we discuss the challenges of extracting information and knowledge from huge archives of Copernicus data.We propose to achieve this by scale-out distributed deep learning techniques that run on very big clusters offering virtual machines and GPUs.We also discuss the challenges of achieving scalability in the management of the extreme volumes of information and knowledge extracted from Copernicus data.The envisioned scientific and technical work will be carried out in the context of the H2020 project ExtremeEarth which starts in January 2019.
Manolis Koubarakis, Konstantina Bereta, Dimitris Bilidas, Konstantinos Giannousis, Theofilos Ioannidis, Despina-Athanasia Pantazi, George Stamoulis 0001, Jim Dowling, Seif Haridi, Vladimir Vlassov, Lorenzo Bruzzone, Claudia Paris, Torbjørn Eltoft, Thomas Krämer, Angelos Charalambidis, Vangelis Karkaletsis, Stasinos Konstantopoulos, Theofilos Kakantousis, Mihai Datcu, Corneliu Octavian Dumitru, Florian Appel, Heike Bach, Silke Migdall, Nicholas Hughes, David Arthurs, Andrew Fleming
EDBT1
2019 Comparative Analysis of Content-based Personalized Microblog Recommendations
Efi Karra Taniskidou, George Papadakis 0001, George Giannakopoulos, Manolis Koubarakis
EDBT4
2019 Extending the YAGO2 Knowledge Graph with Precise Geospatial Knowledge
abstract
We extend YAGO2 with geospatial information represented by geometries (e.g., lines, polygons, multipolygons, etc.) encoded by Open Geospatial Consortium standards. The new geospatial information comes from official sources such as the administrative divisions of countries but also from volunteered open data of OpenStreetMap. The resulting knowledge graph is currently the richest in terms of geospatial information publicly available, open source, knowledge graph.
Nikolaos Karalis, Georgios M. Mandilaras, Manolis Koubarakis
ISWC (2)3
2019 Ontop-spatial: Ontop of geospatial databases
Konstantina Bereta, Guohui Xiao 0001, Manolis Koubarakis
J. Web Semant.3
2018 Distributed Execution of Spatial SQL Queries
abstract
The volume of available spatial data that is generated and collected has significantly increased in the last few years. A number of applications based on Map-Reduce-like systems and cloud infrastructure have emerged. These applications offer a variety of features, however they differ in terms of spatial functions, partitioning and indexing. In this paper we present our own implementation that enables spatial support for distributed execution of spatial SQL queries as part of the system Exareme. Then, we evaluate some of the State-of-the-Art existing geospatial distributed systems, emphasizing on systems based on Apache Spark. We conduct detailed functional and performance benchmarks that include corner cases that stress the systems in comparison and reveal their advantages and weaknesses in both functionality and performance.
Konstantinos Giannousis, Konstantina Bereta, Nikolaos Karalis, Manolis Koubarakis
IEEE BigData4
2018 From Copernicus Big Data to Big Information and Big Knowledge: A Demo from the Copernicus App Lab Project
abstract
Copernicus is the European program for monitoring the Earth. It consists of a set of complex systems that collect data from satellites and in-situ sensors, process this data and provide users with reliable and up-to-date information on a range of environmental and security issues. The data collected by Copernicus is made available freely following an open access policy. Information extracted from Copernicus data is disseminated to users through the Copernicus services which address six thematic areas: land, marine, atmosphere, climate, emergency and security. We present a demo from the Horizon 2020 Copernicus App Lab project which takes big data from the Copernicus land service, makes it available on the Web as linked geospatial data and interlinks it with other useful public data to aid the development of applications by developers that might not be Earth Observation experts. Our demo targets a scenario where we want to study the "greenness" of Paris.
Konstantina Bereta, Hervé Caumont, Erwin Goor, Manolis Koubarakis, Despina-Athanasia Pantazi, George Stamoulis 0001, Sam Ubels, Valentijn Venus, Firman Wahyudi
CIKM4
2018 From Big Data to Big Information and Big Knowledge: the Case of Earth Observation Data
abstract
Some particularly important rich sources of open and free big geospatial data are the Earth observation (EO) programs of various countries such as the Landsat program of the US and the Copernicus programme of the European Union. EO data is a paradigmatic case of big data and the same is true for the big information and big knowledge extracted from it. EO data (satellite images and in-situ data), and the information and knowledge extracted from it, can be utilized in many applications with financial and environmental impact in areas such as emergency management, climate change, agriculture and security.
Konstantina Bereta, Manolis Koubarakis, Stefan Manegold, George Stamoulis 0001, Begüm Demir
CIKM2
2018 Modeling and Preserving Greek Government Decisions Using Semantic Web Technologies and Permissionless Blockchains
Themis Beris, Manolis Koubarakis
ESWC2
2018 The return of JedAI: End-to-End Entity Resolution for Structured and Semi-Structured Data
abstract
JedAI is an Entity Resolution toolkit that can be used in three ways: (i) as an open-source library that combines state-of-the-art methods into a plethora of end-to-end workflows, (ii) as a user-friendly desktop application with a wizardlike interface that provides complex, out-of-the-box solutions even to lay users, and (iii) as a workbench for comparing the performance of numerous workflows over both structured and semi-structured data. Here, we present its significant upgrade, JedAI 2.0, which enhances the original version in three important respects: (i) time efficiency , as the running time has been drastically reduced with the use of high performance data structures and multi-core processing, (ii) effectiveness , since we enriched its library with more established methods, a new layer that exploits loose schema binding as well as the automatic, data-driven configuration of individual methods or entire workflows, and (iii) usability , as the GUI now enables users to manually configure any method based on concrete guidelines, to store the matching results into any of the supported data formats and to visually explore both input and output data.
George Papadakis 0001, Leonidas Tsekouras, Emmanouil Thanos, George Giannakopoulos, Themis Palpanas, Manolis Koubarakis
Proc. VLDB Endow.6
2018 GeoTriples: Transforming geospatial data into RDF graphs using R2RML and RML mappings
Kostis Kyzirakos, Dimitrianos Savva, Ioannis Vlachopoulos, Alexandros Vasileiou 0003, Nikolaos Karalis, Manolis Koubarakis, Stefan Manegold
J. Web Semant.6
2017 Modeling and Querying Greek Legislation Using Semantic Web Technologies
Ilias Chalkidis, Charalampos Nikolaou, Panagiotis Soursos, Manolis Koubarakis
ESWC (1)4
2017 The BigDataEurope Platform - Supporting the Variety Dimension of Big Data
Sören Auer, Simon Scerri, Aad Versteden, Erika Pauwels, Angelos Charalambidis, Stasinos Konstantopoulos, Jens Lehmann 0001, Hajira Jabeen, Ivan Ermilov, Gezim Sejdiu, Andreas Ikonomopoulos, Spyros Andronopoulos, Mandy Vlachogiannis, Charalambos Pappas, Athanasios Davettas, Iraklis A. Klampanos, Efstathios Grigoropoulos, Vangelis Karkaletsis, Victor de Boer, Ronny Siebes, Mohamed Nadjib Mami, Sergio Albani, Michele Lazzarini, Paulo Nunes, Emanuele Angiuli, Nikiforos Pittaras, George Giannakopoulos, Giorgos Argyriou, George Stamoulis 0001, George Papadakis 0001, Manolis Koubarakis, Pythagoras Karampiperis, Axel-Cyrille Ngonga Ngomo, Maria-Esther Vidal
ICWE31
2017 Query Reorganization Algorithms for Efficient Boolean Information Filtering
abstract
In the information filtering paradigm, clients subscribe to a server with continuous queries that express their information needs and get notified every time appropriate information is published. To perform this task in an efficient way, servers employ indexing schemes that support fast matches of the incoming information with the query database. Such indexing schemes involve (i) main-memory trie-based data structures that cluster similar queries by capturing common elements between them and (ii) efficient filtering mechanisms that exploit this clustering to achieve high throughput and low filtering times. However, state-of-the-art indexing schemes are sensitive to the query insertion order and cannot adopt to an evolving query workload, degrading the filtering performance over time. In this paper, we present an adaptive trie-based algorithm that outperforms current methods by relying on query statistics to reorganise the query database. Contrary to previous approaches, we show that the nature of the constructed tries, rather than their compactness, is the determining factor for efficient filtering performance. Our algorithm does not depend on the order of insertion of queries in the database, manages to cluster queries even when clustering possibilities are limited, and achieves more than 96 percent filtering time improvement over its state-of-the-art competitors. Finally, we demonstrate that our solution is easily extensible to multi-core machines.
Lefteris Zervakis, Christos Tryfonopoulos, Spiros Skiadopoulos, Manolis Koubarakis
IEEE Trans. Knowl. Data Eng.4
2017 Ontology Based Data Access in Statoil
Evgeny Kharlamov, Dag Hovland, Martin G. Skjæveland, Dimitris Bilidas, Ernesto Jiménez-Ruiz, Guohui Xiao 0001, Ahmet Soylu, Davide Lanti, Martín Rezk, Dmitriy Zheleznyakov, Martin Giese, Hallstein Lie, Yannis E. Ioannidis, Yannis Kotidis, Manolis Koubarakis, Arild Waaler
J. Web Semant.15
2016 Scaling Entity Resolution to Large, Heterogeneous Data with Enhanced Meta-blocking
abstract
Entity Resolution constitutes a quadratic task that typically scales to large entity collections through blocking. The resulting blocks can be restructured by Meta-blocking in order to significantly increase precision at a limited cost in recall. Yet, its processing can be time-consuming, while its precision remains poor for configurations with high recall. In this work, we propose new meta-blocking methods that improve precision by up to an order of magnitude at a negligible cost to recall. We also introduce two efficiency techniques that, when combined, reduce the overhead time of Metablocking by more than an order of magnitude. We evaluate our approaches through an extensive experimental study over 6 realworld, heterogeneous datasets. The outcomes indicate that our new algorithms outperform all meta-blocking techniques as well as the state-of-the-art methods for block processing in all respects.
George Papadakis 0001, George Papastefanatos, Themis Palpanas, Manolis Koubarakis
EDBT4
2016 Ontology-Based Data Access for Maritime Security
Stefan Brüggemann, Konstantina Bereta, Guohui Xiao 0001, Manolis Koubarakis
ESWC4
2016 Full-Text Support for Publish/Subscribe Ontology Systems
Lefteris Zervakis, Christos Tryfonopoulos, Spiros Skiadopoulos, Manolis Koubarakis
ESWC4
2016 Ontop of Geospatial Databases
Konstantina Bereta, Manolis Koubarakis
ISWC (1)2
2015 Special issue of the Journal of Web Semantics on ontology-based data access
Diego Calvanese, Manolis Koubarakis, David Toman 0001
J. Web Semant.2
2015 Sextant: Visualizing time-evolving linked geospatial data
Charalampos Nikolaou, Kallirroi Dogani, Konstantina Bereta, George Garbis, Manos Karpathiotakis, Kostis Kyzirakos, Manolis Koubarakis
J. Web Semant.7
2014 Wildfire monitoring using satellite images, ontologies and linked geospatial data
Kostis Kyzirakos, Manos Karpathiotakis, George Garbis, Charalampos Nikolaou, Konstantina Bereta, Ioannis Papoutsis, Themos Herekakis, Dimitrios Michail 0001, Manolis Koubarakis, Charalambos Kontoes
J. Web Semant.9
2013 Real-time wildfire monitoring using scientific database and linked data technologies
abstract
We present a real-time wildfire monitoring service that exploits satellite images and linked geospatial data to detect hotspots and monitor the evolution of fire fronts. The service makes heavy use of scientific database technologies (array databases, SciQL, data vaults) and linked data technologies (ontologies, linked geospatial data, stSPARQL) and is implemented on top of MonetDB and Strabon. The service is now operational at the National Observatory of Athens and has been used during the previous summer by emergency managers monitoring wildfires in Greece.
Manolis Koubarakis, Charalambos Kontoes, Stefan Manegold
EDBT1
2013 Representation and Querying of Valid Time of Triples in Linked Geospatial Data
Konstantina Bereta, Panayiotis Smeros, Manolis Koubarakis
ESWC3
2013 Geographica: A Benchmark for Geospatial RDF Stores (Long Version)
George Garbis, Kostis Kyzirakos, Manolis Koubarakis
ISWC (2)3
2013 The Spatiotemporal RDF Store Strabon
Kostis Kyzirakos, Manos Karpathiotakis, Konstantina Bereta, George Garbis, Charalampos Nikolaou, Panayiotis Smeros, Stella Giannakopoulou, Kallirroi Dogani, Manolis Koubarakis
SSTD9
2013 Querying Incomplete Geospatial Information in RDF
Charalampos Nikolaou, Manolis Koubarakis
SSTD2
2012 Strabon: A Semantic Geospatial DBMS
Kostis Kyzirakos, Manos Karpathiotakis, Manolis Koubarakis
ISWC (1)3
2012 TELEIOS: A Database-Powered Virtual Earth Observatory
abstract
TELEIOS is a recent European project that addresses the need for scalable access to petabytes of Earth Observation data and the discovery and exploitation of knowledge that is hidden in them. TELEIOS builds on scientific database technologies (array databases, SciQL, data vaults) and Semantic Web technologies (stRDF and stSPARQL) implemented on top of a state of the art column store database system (MonetDB). We demonstrate a first prototype of the TELEIOS Virtual Earth Observatory (VEO) architecture, using a forest fire monitoring application as example.
Manolis Koubarakis, Kostis Kyzirakos, Manos Karpathiotakis, Charalampos Nikolaou, Stavros Vassos, George Garbis, Michael Sioutis, Konstantina Bereta, Dimitrios Michail 0001, Charalambos Kontoes, Ioannis Papoutsis, Themos Herekakis, Stefan Manegold, Martin L. Kersten, Milena Ivanova, Holger Pirk, Ying Zhang 0027, Mihai Datcu, Gottfried Schwarz, Corneliu Octavian Dumitru, Daniela Espinoza-Molina, Katrin Molch, Ugo Di Giammatteo, Manuela Sagona, Sergio Perelli, Thorsten Reitz, Eva Klien, Robert Gregor
Proc. VLDB Endow.1
2012 FoXtrot: Distributed structural and value XML filtering
abstract
Publish/subscribe systems have emerged in recent years as a promising paradigm for offering various popular notification services. In this context, many XML filtering systems have been proposed to efficiently identify XML data that matches user interests expressed as queries in an XML query language like XPath. However, in order to offer XML filtering functionality on an Internet-scale, we need to deploy such a service in a distributed environment, avoiding bottlenecks that can deteriorate performance. In this work, we design and implement FoXtrot, a system for filtering XML data that combines the strengths of automata for efficient filtering and distributed hash tables for building a fully distributed system. Apart from structural-matching, performed using automata, we also discuss different methods for evaluating value-based predicates. We perform an extensive experimental evaluation of our system, FoXtrot, on a local cluster and on the PlanetLab network and demonstrate that it can index millions of user queries, achieving a high indexing and filtering throughput. At the same time, FoXtrot exhibits very good load-balancing properties and improves its performance as we increase the size of the network.
Iris Miliaraki, Manolis Koubarakis
ACM Trans. Web2
2011 The Papyrus Digital Library: Discovering History in the News
Akrivi Katifori, Charalampos Nikolaou, Manolis Platakis, Yannis E. Ioannidis, A. Tympas, Manolis Koubarakis, Nikos Sarris, V. Tountopoulos, Efstratios Tzoannos, Siarhei Bykau, Nadzeya Kiyavitskaya, Chrisa Tsinaraki, Yannis Velegrakis
TPDL6
2011 A Semantically Enabled Service Architecture for Mashups over Streaming and Stored Data
Alasdair J. G. Gray, Raúl García-Castro, Kostis Kyzirakos, Manos Karpathiotakis, Jean-Paul Calbimonte, Kevin R. Page, Jason Sadler, Alex Frazer, Ixent Galpin, Alvaro A. A. Fernandes, Norman W. Paton, Óscar Corcho, Manolis Koubarakis, David De Roure, Kirk Martinez, Asunción Gómez-Pérez
ESWC (2)13
2010 Modeling and Querying Metadata in the Semantic Sensor Web: The Model stRDF and the Query Language stSPARQL
Manolis Koubarakis, Kostis Kyzirakos
ESWC (1)1
2010 SPARQL Query Optimization on Top of DHTs
Zoi Kaoudi, Kostis Kyzirakos, Manolis Koubarakis
ISWC (1)3
2010 Atlas: Storing, updating and querying RDF(S) data on top of DHTs
Zoi Kaoudi, Manolis Koubarakis, Kostis Kyzirakos, Iris Miliaraki, Matoula Magiridou, Antonios Papadakis-Pesaresi
J. Web Semant.2
2009 Information filtering and query indexing for an information retrieval model
abstract
In the information filtering paradigm, clients subscribe to a server with continuous queries or profiles that express their information needs. Clients can also publish documents to servers. Whenever a document is published, the continuous queries satisfying this document are found and notifications are sent to appropriate clients. This article deals with the filtering problem that needs to be solved efficiently by each server: Given a database of continuous queries db and a document d , find all queries q ∈ db that match d . We present data structures and indexing algorithms that enable us to solve the filtering problem efficiently for large databases of queries expressed in the model AWP . AWP is based on named attributes with values of type text, and its query language includes Boolean and word proximity operators.
Christos Tryfonopoulos, Manolis Koubarakis, Yannis Drougas
ACM Trans. Inf. Syst.2
2008 Continuous multi-way joins over distributed hash tables
abstract
This paper studies the problem of evaluating continuous multi-way joins on top of Distributed Hash Tables (DHTs). We present a novel algorithm, called recursive join (RJoin), that takes into account various parameters crucial in a distributed setting i.e., network traffic, query processing load distribution, storage load distribution etc. The key idea of RJoin is incremental evaluation: as relevant tuples arrive continuously, a given multi-way join is rewritten continuously into a join with fewer join operators, and is assigned continuously to different nodes of the network. In this way, RJoin distributes the responsibility of evaluating a continuous multi-way join to many network nodes by assigning parts of the evaluation of each binary join to a different node depending on the values of the join attributes. The actual nodes to be involved are decided by RJoin dynamically after taking into account the rate of incoming tuples with values equal to the values of the joined attributes. RJoin also supports sliding window joins which is a crucial feature, especially for long join paths, since it provides a mechanism to reduce the query processing state and thus keep the cost of handling incoming tuples stable. In addition, RJoin is able to handle message delays due to heavy network traffic. We present a detailed mathematical and experimental analysis of RJoin and study the performance tradeoffs that occur.
Stratos Idreos, Erietta Liarou, Manolis Koubarakis
EDBT3
2008 RDFS Reasoning and Query Answering on Top of DHTs
Zoi Kaoudi, Iris Miliaraki, Manolis Koubarakis
ISWC3
2008 Approximate Information Filtering in Peer-to-Peer Networks
Christian Zimmer 0001, Christos Tryfonopoulos, Klaus Berberich, Manolis Koubarakis, Gerhard Weikum
WISE4
2008 Xml data dissemination using automata on top of structured overlay networks
abstract
We present a novel approach for filtering XML documents using nondeterministic finite automata and distributed hash tables. Our approach differs architecturally from recent proposals that deal with distributed XML filtering; they assume an XML broker architecture, whereas our solution is built on top of distributed hash tables. The essence of our work is a distributed implementation of YFilter, a state-of-the-art automata-based XML filtering system on top of Chord. We experimentally evaluate our approach and demonstrate that our algorithms can scale to millions of XPath queries under various filtering scenarios, and also exhibit very good load balancing properties.
Iris Miliaraki, Zoi Kaoudi, Manolis Koubarakis
WWW3
2007 A Family of Directional Relation Models for Extended Objects
abstract
In this paper, we introduce a family of expressive models for qualitative spatial reasoning with directions. The proposed family is based on the cognitive plausible cone-based model. We formally define the directional relations that can be expressed in each model of the family. Then, we use our formal framework to study two interesting problems: computing the inverse of a directional relation and composing two directional relations. For the composition operator, in particular, we concentrate on two commonly used definitions, namely, consistency-based and existential composition. Our formal framework allows us to prove that our solutions are correct. The presented solutions are handled in a uniform manner and apply to all of the models of the family.
Spiros Skiadopoulos, Nikos Sarkas, Timos K. Sellis, Manolis Koubarakis
IEEE Trans. Knowl. Data Eng.4
2007 Correction to "A Family of Directional Relation Models for Extended Objects"
abstract
In the above titled paper (ibid., vol. 19, no. 8, pp. 1116-1130, Aug 07), some information appeared incorrectly. The corrections appear here.
Spiros Skiadopoulos, Nikos Sarkas, Timos K. Sellis, Manolis Koubarakis
IEEE Trans. Knowl. Data Eng.4
2006 Distributed Evaluation of Continuous Equi-join Queries over Large Structured Overlay Networks
abstract
We study the problem of continuous relational query processing in Internet-scale overlay networks realized by distributed hash tables. We concentrate on the case of continuous two-way equi-join queries. Joins are hard to evaluate in a distributed continuous query environment because data from more than one relations is needed, and this data is inserted in the network asynchronously. Each time a new tuple is inserted, the network nodes have to cooperate to check if this tuple can contribute to the satisfaction of a query when combined with previously inserted tuples. We propose a series of algorithms that initially index queries at network nodes using hashing. Then, they exploit the values of join attributes in incoming tuples to rewrite the given queries into simpler ones, and reindex them in the network where they might be satisfied by existing or future tuples. We present a detailed experimental evaluation in a simulated environment and we show that our algorithms are scalable, balance the storage and query processing load and keep the network traffic low.
Stratos Idreos, Christos Tryfonopoulos, Manolis Koubarakis
ICDE3
2006 Evaluating Conjunctive Triple Pattern Queries over Large Structured Overlay Networks
Erietta Liarou, Stratos Idreos, Manolis Koubarakis
ISWC3
2006 Logic and Computational Complexity for Boolean Information Retrieval
abstract
We study the complexity of query satisfiability and entailment for the Boolean information retrieval models WP and AWV using techniques from propositional logic and computational complexity. WP and AWV can be used to represent and query textual information under the Boolean model using the concept of attribute with values of type text, the concept of word, and word proximity constraints. Variations of WP and AWP are in use in most deployed digital libraries using the Boolean model, text extenders for relational database systems (e.g., Oracle 10g), search engines, and P2P systems for information retrieval and filtering
Manolis Koubarakis, Spiros Skiadopoulos, Christos Tryfonopoulos
IEEE Trans. Knowl. Data Eng.1
2005 RUL: A Declarative Update Language for RDF
Matoula Magiridou, S. Sahtouris, Vassilis Christophides, Manolis Koubarakis
ISWC4
2005 Publish/subscribe functionality in IR environments using structured overlay networks
abstract
We study the problem of offering publish/subscribe functionality on top of structured overlay networks using data models and languages from IR. We show how to achieve this by extending the distributed hash table Chord and present a detailed experimental evaluation of our proposals.
Christos Tryfonopoulos, Stratos Idreos, Manolis Koubarakis
SIGIR3
2005 Computing and Managing Cardinal Direction Relations
abstract
Qualitative spatial reasoning forms an important part of the commonsense reasoning required for building intelligent geographical information systems (GIS). Previous research has come up with models to capture cardinal direction relations for typical GIS data. In this paper, we target the problem of efficiently computing the cardinal direction relations between regions that are composed of sets of polygons and present two algorithms for this task. The first of the proposed algorithms is purely qualitative and computes, in linear time, the cardinal direction relations between the input regions. The second has a quantitative aspect and computes, also in linear time, the cardinal direction relations with percentages between the input regions. Our experimental evaluation indicates that the proposed algorithms outperform existing methodologies. The algorithms have been implemented and embedded in an actual system, CARDIRECT, that allows the user to 1) specify and annotate regions of interest in an image or a map, 2) compute cardinal direction relations between them, and 3) pose queries in order to retrieve combinations of interesting regions.
Spiros Skiadopoulos, Christos Giannoukos, Nikos Sarkas, Panos Vassiliadis, Timos K. Sellis, Manolis Koubarakis
IEEE Trans. Knowl. Data Eng.6
2004 P2P-DIET: One-Time and Continuous Queries in Super-Peer Networks
Stratos Idreos, Manolis Koubarakis, Christos Tryfonopoulos
EDBT2
2004 Computing and Handling Cardinal Direction Information
Spiros Skiadopoulos, Christos Giannoukos, Panos Vassiliadis, Timos K. Sellis, Manolis Koubarakis
EDBT5
2004 Filtering algorithms for information retrieval models with named attributes and proximity operators
abstract
In the selective dissemination of information (or publish/subscribe) paradigm, clients subscribe to a server with continuous queries (or profiles) that express their information needs. Clients can also publish documents to servers. Whenever a document is published, the continuous queries satisfying this document are found and notifications are sent to appropriate clients. This paper deals with the filtering problem that needs to be solved effciently by each server: Given a database of continuous queries db and a document d, find all queries q ∈ db that match d. We present data structures and indexing algorithms that enable us to solve the filtering problem efficiently for large databases of queries expressed in the model AWP which is based on named attributes with values of type text, and word proximity operators.
Christos Tryfonopoulos, Manolis Koubarakis, Yannis Drougas
SIGIR2
2004 P2P-DIET: An Extensible P2P Service that Unifies Ad-hoc and Continuous Querying in Super-Peer Networks
abstract
No abstract available.
Stratos Idreos, Manolis Koubarakis, Christos Tryfonopoulos
SIGMOD Conference2
2003 Towards High Performance Peer-to-Peer Content and Resource Sharing Systems
Peter Triantafillou, Chryssani Xiruhaki, Manolis Koubarakis, Nikos Ntarmos
CIDR3
2002 A formal framework for business process modelling and design
Manolis Koubarakis, Dimitris Plexousakis
Inf. Syst.1
2001 Composing Cardinal Direction Relations
Spiros Skiadopoulos, Manolis Koubarakis
SSTD2
2000 A Formal Model for Business Process Modeling and Design
Manolis Koubarakis, Dimitris Plexousakis
CAiSE1
1994 Database models for infinite and indefinite temporal information
Manolis Koubarakis
Inf. Syst.1
1993 Representation and Querying in Temporal Databases: the Power of Temporal Constraints
abstract
A temporal database model capable of representing absolute, relative, imprecise, and infinite temporal data is proposed. The model is based on temporal tables such as relation-like representations that can contain variables constrained by the formulas of a temporal theory. An algebraic query language for temporal tables is defined, and some problems related to query answering are discussed.>
Manolis Koubarakis
ICDE1
1990 Telos: Representing Knowledge About Information Systems
abstract
We describe Telos, a language intended to support the development of information systems. The design principles for the language are based on the premise that information system development is knowledge intensive and that the primary responsibility of any language intended for the task is to be able to formally represent the relevent knowledge. Accordingly, the proposed language is founded on concepts from knowledge representations. Indeed, the language is appropriate for representing knowledge about a variety of worlds related to a particular information system, such as the subject world (application domain), the usage world (user models, environments), the system world (software requirements, design), and the development world (teams, metodologies). We introduce the features of the language through examples, focusing on those provided for desribing metaconcepts that can then be used to describe knowledge relevant to a particular information system. Telos' fetures include an object-centered framework which supports aggregation, generalization, and classification; a novel treatment of attributes; an explicit representation of time; and facilities for specifying integrity constraints and deductive rules. We review actual applications of the language through further examples, and we sketch a formalization of the language.
John Mylopoulos, Alexander Borgida, Matthias Jarke, Manolis Koubarakis
ACM Trans. Inf. Syst.4