Panos Vassiliadis

dblp:13/1242 · DBLP profile ↗
← Back
95ranked-venue papers in the field
26as first author
23since 2021 · last 2026
0000-0003-0085-6776ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 70 (17 first)Business Process & Enterprise Data · 21 (9 first)Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 Multi-Query Optimization for the Novel Analyze Operator
Marios Iakovidis 0001, Panos Vassiliadis
DOLAP2
2026 Interpretable Highlights for Experiment Tracking
Vassilis Stamatopoulos, Panagiotis Gidarakos, Stavros Maroulis, George Papastefanatos, Panos Vassiliadis
DOLAP5
2026 Scalable Grid-based Computation of Kendall's Tau Correlation
Nikolaos Koutroumanis, Petros Karampas, Alexandros Karakasidis 0001, Nikos Mamoulis, Panos Vassiliadis
Proc. VLDB Endow.5
2025 Time-Related Patterns Of Schema Evolution
Panos Vassiliadis, Alexandros Karakasidis 0001
EDBT1
2025 Advances in databases and information systems - Selected papers from ADBIS 2023
Alberto Abelló, Ladjel Bellatreche, Oscar Romero 0001, Panos Vassiliadis, Robert Wrembel
Inf. Syst.4
2024 Cube query interestingness: Novelty, relevance, peculiarity and surprise
Dimos Gkitsakis, Spyridon Kaloudis, Eirini Mouselli, Verónika Peralta, Patrick Marcel, Panos Vassiliadis
Inf. Syst.6
2023 Assessment Methods for the Interestingness of Cube Queries
Dimos Gkitsakis, Spyridon Kaloudis, Eirini Mouselli, Verónika Peralta, Patrick Marcel, Panos Vassiliadis
DOLAP6
2023 The History, Present, and Future of ETL Technology (invited)
Alkis Simitsis, Spiros Skiadopoulos, Panos Vassiliadis
DOLAP3
2023 Cube Query Answering via the Results of Previous Cube Queries
Panos Vassiliadis
DOLAP1
2023 Joint Source and Schema Evolution: Insights from a Study of 195 FOSS Projects
Panos Vassiliadis, Fation Shehaj, George Kalampokis, Apostolos V. Zarras
EDBT1
2023 A Safari for Deviating GoF Pattern Definitions and Examples on the Web
Apostolos V. Zarras, Panos Vassiliadis
ER2
2023 Suggesting Assess Queries for Interactive Analysis of Multidimensional Data
abstract
Assessment is the process of comparing the actual to the expected behavior of a business phenomenon and judging the outcome of the comparison. The assess querying operator has been recently proposed to support assessment based on the results of a query on a data cube. This operator requires (i) the specification of an OLAP query to determine a target cube; (ii) the specification of a reference cube of comparison (benchmark), which represents the expected performance; (iii) the specification of how to perform the comparison, and (iv) a labeling function that classifies the result of this comparison. Despite the adoption of a SQL-like syntax that hides the complexity of the assessment process, writing a complete assess statement is not easy. In this paper we focus on making the user experience more comfortable by letting the system suggest suitable completions for partially-specified statements. To this end we propose two interaction modes: progressive refinement and auto-completion, both starting from an assess statement partially declared by the user. These two modes are evaluated both in terms of scalability and user experience, with the support of two experiments made with real users.
Matteo Francia, Matteo Golfarelli, Patrick Marcel, Stefano Rizzi, Panos Vassiliadis
IEEE Trans. Knowl. Data Eng.5
2023 Graph-Driven Federated Data Management
abstract
Modern data analysis applications, require the ability to provide on-demand integration of data sources while offering a flexible and user-friendly query interface. Traditional techniques for answering queries using views, focused on a rather static setting, fail to address such requirements. To overcome these issues, we propose a fully-fledged data integration approach based on graph-based constructs. The extensibility of graphs allows us to extend the traditional framework for data integration with view definitions. Furthermore, we also propose a query language based on subgraphs. We tackle query answering via a query rewriting algorithm based on well-known algorithms for answering queries using views. We experimentally show that the proposed method yields good performance and does not introduce a significant overhead.
Sergi Nadal, Alberto Abelló, Oscar Romero 0001, Stijn Vansummeren, Panos Vassiliadis
IEEE Trans. Knowl. Data Eng.5
2023 Resource-aware adaptive indexing for in situ visual exploration and analytics
Stavros Maroulis, Nikos Bikakis, George Papastefanatos, Panos Vassiliadis, Yannis Vassiliou
VLDB J.4
2022 Data Narrative Crafting via a Comprehensive and Well-Founded Process
Faten El Outa, Patrick Marcel, Verónika Peralta, Raphaël da Silva, Marie Chagnoux, Panos Vassiliadis
ADBIS6
2022 Graph-Driven Federated Data Management (Extended Abstract)
abstract
Modern data analysis applications require the ability to provide on-demand integration of data sources while offering a user-friendly query interface. Traditional methods for answering queries using views, focused on a rather static setting, fail to address such requirements. To overcome these issues, we propose a full fledged, GLAV-based data integration approach based on graph-based constructs. The extensibility of graphs allows us to extend the traditional framework for data integration with view definitions. Furthermore, we also propose a query language based on subgraphs. We tackle query answering via a query rewriting algorithm based on well-known algorithms for answering queries using views. We experimentally show that our method yields good performance with no significant overhead.
Sergi Nadal, Alberto Abelló, Oscar Romero 0001, Stijn Vansummeren, Panos Vassiliadis
ICDE5
2022 Relational schema optimization for RDF-based knowledge graphs
abstract
Characteristic sets (CS) organize RDF triples based on the set of properties associated with their subject nodes. This concept was recently used in indexing techniques, as it can capture the implicit schema of RDF data. While most CS-based approaches yield significant improvements in space and query performance, they fail to perform well when answering complex query workloads in the presence of schema heterogeneity, i.e., when the number of CSs becomes very large, resulting in a highly partitioned data organization. In this paper, we address this problem by introducing a novel technique, for merging CSs based on their hierarchical structure. Our method employs a lattice to capture the hierarchical relationships between CSs, identifies dense CSs and merges dense CSs with their ancestors. We have implemented our algorithm on top of a relational backbone, where each merged CS is stored in a relational table, and therefore, CS merging results in a smaller number of required tables to host the source triples of a dataset. Moreover, we perform an extensive experimental study to evaluate the performance and impact of merging to the storage and querying of RDF datasets, indicating significant improvements. We also conduct a sensitivity analysis to identify the stability and any possible weaknesses of our algorithm, and report on our results.
George Papastefanatos, Marios Meimaris, Panos Vassiliadis
Inf. Syst.3
2022 Taxa and super taxa of schema evolution and their relationship to activity, heartbeat and duration
Panos Vassiliadis, George Kalampokis
Inf. Syst.1
2021 Adaptive Indexing for In-situ Visual Exploration and Analytics
Stavros Maroulis, Nikos Bikakis, George Papastefanatos, Panos Vassiliadis, Yannis Vassiliou
DOLAP4
2021 Assess Queries for Interactive Analysis of Data Cubes
abstract
Assessment is the process of comparing the actual to the expected behavior of a business phenomenon and judging the outcome of the comparison. In this paper we propose assess, a novel querying operator that supports assessment based on the results of a query on a data cube. This operator requires (1) the specification of an OLAP query over a measure of a data cube, to define the target cube to be assessed; (2) the specification of a reference cube of comparison (benchmark), which represents the expected performance of the measure; (3) the specification of how to perform the comparison between the target cube and the benchmark, and (4) a labeling function that classifies the result of this comparison using a set of labels. After introducing an SQL-like syntax for our operator, we formally define its semantics in terms of a set of logical operators. To support the computation of assess we propose a basic plan as well as some optimization strategies, then we experimentally evaluate their performance using a prototype.
Matteo Francia, Matteo Golfarelli, Patrick Marcel, Stefano Rizzi, Panos Vassiliadis
EDBT5
2021 Profiles of Schema Evolution in Free Open Source Software Projects
abstract
In this paper, we present the findings of a large study of the evolution of the schema of 195 Free Open Source Software projects. We identify families of evolutionary behaviors, or taxa, in FOSS projects. A large percentage of the projects demonstrate very few, if any, actions of schema evolution. Two other taxa involve the evolution via focused actions, with either a single focused maintenance action, or a large percentage of evolution activity grouped in no more than a couple interventions. Schema evolution also involves moderate, and active evolution, with very different volumes of updates to the schema. To the best of our knowledge, this is the first study of this kind in the area of schema evolution, both in terms of presenting profiles of how schemata evolve, and, in terms of the dataset magnitude and the generalizability of the findings.
Panos Vassiliadis
ICDE1
2021 RawVis: A System for Efficient In-situ Visual Analytics
abstract
In-situ processing has received a great deal of attention in recent years. In in-situ scenarios, big raw data files which do not fit in main memory, must be efficiently handled on-the-fly using commodity hardware, without the overhead of a preprocessing phase or the loading of data into a database system. This paper presents RawVis, an open source data visualization system for in-situ visual exploration and analytics over big raw data. RawVis implements novel indexing schemes and adaptive processing techniques allowing users to perform efficient visual and analytics operations directly over the data files. RawVis provides real-time interaction, reporting low response time, over large data files, using commodity hardware.
Stavros Maroulis, Nikos Bikakis, George Papastefanatos, Panos Vassiliadis, Yannis Vassiliou
SIGMOD Conference4
2021 In-situ visual exploration over big raw data
Nikos Bikakis, Stavros Maroulis, George Papastefanatos, Panos Vassiliadis
Inf. Syst.4
2020 The Traveling Analyst Problem: Definition and Preliminary Study
Alexandre Chanson, Ben Crulis, Nicolas Labroche, Patrick Marcel, Verónika Peralta, Stefano Rizzi, Panos Vassiliadis
DOLAP7
2020 Hierarchical Property Set Merging for SPARQL Query Optimization
Marios Meimaris, George Papastefanatos, Panos Vassiliadis
DOLAP3
2020 A Study on the Effect of a Table's Involvement in Foreign Keys to its Schema Evolution
Konstantinos Dimolikas, Apostolos V. Zarras, Panos Vassiliadis
ER3
2020 Towards a Conceptual Model for Data Narratives
Faten El Outa, Matteo Francia, Patrick Marcel, Verónika Peralta, Panos Vassiliadis
ER5
2019 A Framework for Learning Cell Interestingness from Cube Explorations
Patrick Marcel, Verónika Peralta, Panos Vassiliadis
ADBIS3
2019 Towards a Benefit-based Optimizer for Interactive Data Analysis
Patrick Marcel, Nicolas Labroche, Panos Vassiliadis
DOLAP3
2019 An integration-oriented ontology to govern evolution in Big Data ecosystems
Sergi Nadal, Oscar Romero 0001, Alberto Abelló, Panos Vassiliadis, Stijn Vansummeren
Inf. Syst.4
2019 Beyond roll-up's and drill-down's: An intentional analytics model to reinvent OLAP
Panos Vassiliadis, Patrick Marcel, Stefano Rizzi
Inf. Syst.1
2018 RawVis: Visual Exploration over Raw Data
Nikos Bikakis, Stavros Maroulis, George Papastefanatos, Panos Vassiliadis
ADBIS4
2018 The Road to Highlights is Paved with Good Intentions: Envisioning a Paradigm Shift in OLAP Modeling
Panos Vassiliadis, Patrick Marcel
DOLAP1
2018 MDM: Governing Evolution in Big Data Ecosystems
abstract
On-demand integration of multiple data sources is a critical requirement in many Big Data settings. This has been coined as the data variety challenge, which refers to the complexity of dealing with an heterogeneous set of data sources to enable their integrated analysis. In Big Data settings, data sources are commonly represented by external REST APIs, which provide data in their original format and continously apply changes in their structure (i.e. schema). Thus, data analysts face the challenge to integrate such multiple sources, and then continuosly adapt their analytical processes to changes in the schema. To address this challenges, in this paper, we present the Metadata Management System, shortly MDM, a tool that supports data stewards and analysts to manage the integration and analysis of multiple heterogeneous sources under schema evolution. MDM adopts a vocabulary-based integration-oriented ontology to conceptualize the domain of interest and relies on local-as-view mappings to link it with the sources. MDM provides user-friendly mechanisms to manage the ontology and mappings. Finally, a query rewriting algorithm ensures that queries posed to the ontology are correctly resolved to the sources in the presence of multiple schema versions, a transparent process to data analysts. On-site, we will showcase using real-world examples how MDM facilitates the management of multiple evolving data sources and enables its integrated analysis.
Sergi Nadal, Alberto Abelló, Oscar Romero 0001, Stijn Vansummeren, Panos Vassiliadis
EDBT5
2018 Computational methods and optimizations for containment and complementarity in web data cubes
Marios Meimaris, George Papastefanatos, Panos Vassiliadis, Ioannis Anagnostopoulos
Inf. Syst.3
2017 Extraction of Embedded Queries via Static Analysis of Host Code
Petros Manousis, Apostolos V. Zarras, Panos Vassiliadis, George Papastefanatos
CAiSE3
2017 Survival in Schema Evolution: Putting the Lives of Survivor and Dead Tables in Counterpoint
Panos Vassiliadis, Apostolos V. Zarras
CAiSE1
2017 Schema Evolution and Foreign Keys: Birth, Eviction, Change and Absence
Panos Vassiliadis, Michail-Romanos Kolozoff, Maria Zerva, Apostolos V. Zarras
ER1
2017 Schema Evolution and Gravitation to Rigidity: A Tale of Calmness in the Lives of Structured Data
Panos Vassiliadis
MEDI1
2017 Gravitating to rigidity: Patterns of schema evolution - and its absence - in the lives of tables
Panos Vassiliadis, Apostolos V. Zarras, Ioannis Skoulis
Inf. Syst.1
2016 Keep Calm and Wait for the Spike! Insights on the Evolution of Amazon Services
Apostolos V. Zarras, Panos Vassiliadis, Ioannis Dinos
CAiSE2
2016 Schema Evolution for Relational Databases
Panos Vassiliadis
DATA1
2016 Efficient Computation of Containment and Complementarity in RDF Data Cubes
abstract
Multidimensional data are published in the web of data under common directives, such as the Resource Description Framework (RDF). The increasing volume and diversity of these data pose the challenge of finding relations between them in a most efficient and accurate way, by taking into advantage their overlapping schemes. In this paper we define two types of relationships between multidimensional RDF data, and we propose algorithms for efficient and scalable computation of these relationships. Specifically, we define the notions of containment and complementarity between points in multidimensional dataspaces, as different aspects of relatedness, and we propose a baseline method for computing them, as well as two alternative methods that target speed and scalability. We provide an experimental evaluation over real-world and synthetic datasets and we compare our approach to a SPARQL-based and a rule-based alternative, which prove to be inefficient for increasing input sizes. © 2016, Copyright is with the authors.
Marios Meimaris, George Papastefanatos, Panos Vassiliadis, Ioannis Anagnostopoulos
EDBT3
2015 How is Life for a Table in an Evolving Relational Schema? Birth, Death and Everything in Between
Panos Vassiliadis, Apostolos V. Zarras, Ioannis Skoulis
ER1
2015 CineCubes: Aiding data workers gain insights from OLAP queries
Dimitrios Gkesoulis, Panos Vassiliadis, Petros Manousis
Inf. Syst.2
2015 Growing up with stability: How open-source relational databases evolve
Ioannis Skoulis, Panos Vassiliadis, Apostolos V. Zarras
Inf. Syst.2
2014 Open-Source Databases: Within, Outside, or Beyond Lehman's Laws of Software Evolution?
Ioannis Skoulis, Panos Vassiliadis, Apostolos V. Zarras
CAiSE2
2014 Visual Maps for Data-Intensive Ecosystems
Efthymia Kontogiannopoulou, Petros Manousis, Panos Vassiliadis
ER3
2013 CineCubes: cubes as movie stars with little effort
abstract
In this paper we investigate how we can exploit the existence of a star schema in order to answer user OLAP queries with CineCube movies. Our method, implemented in an actual system, includes the following steps. The user submits a query over an underlying star schema. Taking this query as input, the system comes up with a set of queries complementing the information content of the original query, and executes them. Then, the system visualizes the query results and accompanies this presentation with a text commenting on the result highlights. Moreover, via a text-to-speech conversion the system automatically produces audio for the constructed text. Each combination of visualization, text and audio practically constitutes a cube movie, which is wrapped as a PowerPoint presentation and returned to the user.
Dimitrios Gkesoulis, Panos Vassiliadis
DOLAP2
2013 Automating the Adaptation of Evolving Data-Intensive Ecosystems
Petros Manousis, Panos Vassiliadis, George Papastefanatos
ER2
2013 Scheduling strategies for efficient ETL execution
Anastasios Karagiannis, Panos Vassiliadis, Alkis Simitsis
Inf. Syst.2
2012 Trading Privacy for Information Loss in the Blink of an Eye
Alexandra Pilalidou, Panos Vassiliadis
SSDBM2
2011 Efficient answering of set containment queries for skewed item distributions
abstract
In this paper we address the problem of efficiently evaluating containment (i.e., subset, equality, and superset) queries over set-valued data. We propose a novel indexing scheme, the Ordered Inverted File (OIF) which, differently from the state-of-the-art, indexes set-valued attributes in an ordered fashion. We introduce query processing algorithms that practically treat containment queries as range queries over the ordered postings lists of OIF and exploit this ordering to quickly prune unnecessary page accesses. OIF is simple to implement and our experiments on both real and synthetic data show that it greatly outperforms the current state-of-the-art methods for all three classes of containment queries.
Manolis Terrovitis, Panagiotis Bouros, Panos Vassiliadis, Timos K. Sellis, Nikos Mamoulis
EDBT3
2011 Similarity measures for multidimensional data
abstract
How similar are two data-cubes? In other words, the question under consideration is: given two sets of points in a multidimensional hierarchical space, what is the distance value between them? In this paper we explore various distance functions that can be used over multidimensional hierarchical spaces. We organize the discussed functions with respect to the properties of the dimension hierarchies, levels and values. In order to discover which distance functions are more suitable and meaningful to the users, we conducted two user study analysis. The first user study analysis concerns the most preferred distance function between two values of a dimension. The findings of this user study indicate that the functions that seem to fit better the user needs are characterized by the tendency to consider as closest to a point in a multidimensional space, points with the smallest shortest path with respect to the same dimension hierarchy. The second user study aimed in discovering which distance function between two data cubes, is mostly preferred by users. The two functions that drew the attention of users where (a) the summation of distances between every cell of a cube with the most similar cell of another cube and (b) the Hausdorff distance function. Overall, the former function was preferred by users than the latter; however the individual scores of the tests indicate that this advantage is rather narrow.
Eftychia Baikousi, Georgios Rogkakos, Panos Vassiliadis
ICDE3
2011 Managing contextual preferences
Kostas Stefanidis, Evaggelia Pitoura, Panos Vassiliadis
Inf. Syst.3
2010 HECATAEUS: Regulating schema evolution
abstract
HECATAEUS is an open-source software tool for enabling impact prediction, what-if analysis, and regulation of relational database schema evolution. We follow a graph theoretic approach and represent database schemas and database constructs, like queries and views, as graphs. Our tool enables the user to create hypothetical evolution events and examine their impact over the overall graph before these are actually enforced on it. It also allows definition of rules for regulating the impact of evolution via (a) default values for all the nodes of the graph and (b) simple annotations for nodes deviating from the default behavior. Finally, HECATAEUS includes a metric suite for evaluating the impact of evolution events and detecting crucial and vulnerable parts of the system.
George Papastefanatos, Panos Vassiliadis, Alkis Simitsis, Yannis Vassiliou
ICDE2
2010 Maintenance of top-k materialized views
Eftychia Baikousi, Panos Vassiliadis
Distributed Parallel Databases2
2010 Accelerating Web Service Workflow Execution via Intelligent Allocation of Services to Servers
abstract
The appropriate deployment of web service operations at the service provider site plays a critical role in the efficient provision of services to clients. In this paper, the authors assume that a service provider has several servers over which web service operations can be deployed. Given a workflow of web services and the topology of the servers, the most efficient mapping of operations to servers must then be discovered. Efficiency is measured in terms of two cost functions that concern the execution time of the workflow and the fairness of the load distribution among the servers. The authors study different topologies for the workflow structure and the server connectivity and propose a suite of greedy algorithms for each combination.
Konstantinos Stamkopoulos, Evaggelia Pitoura, Panos Vassiliadis, Apostolos V. Zarras
J. Database Manag.3
2009 View usability and safety for the answering of top-k queries via materialized views
abstract
In this paper, we investigate the problem of answering top-k queries via materialized views. We provide theoretical guarantees for the adequacy of a view to answer a top-k query, along with algorithmic techniques to compute the query via a view when this is possible. We explore the problem of answering a query via a combination of more than one view and show that it is impossible to improve our theoretical guarantees for the answering of a query via a combination of views. Finally, we experimentally assess our approach for its effectiveness and efficiency.
Eftychia Baikousi, Panos Vassiliadis
DOLAP2
2009 A taxonomy of ETL activities
abstract
Extract-Transform-Load (ETL) activities are software modules responsible for populating a data warehouse with operational data, which have undergone a series of transformations on their way to the warehouse. The whole process is very complex and of signifi-cant importance for the design and maintenance of the data ware-house. A plethora of commercial ETL tools are already available in the market. However, each one of them follows a different ap-proach for the modeling of ETL activities; i.e., of the building blocks of an ETL workflow. As a result, so far there is no standard or unified approach for describing such activities. In this paper, we are working towards the identification of generic properties that characterize ETL activities. In doing so, we follow a black-box approach and provide a taxonomy that characterizes ETL activities in terms of the relationship of their input to their output and provide a normal form that is based on interpreted semantics for the black box activities. Finally, we show how the proposed taxonomy can be used in the construction of larger modules, i.e., ETL archetype patterns, which can be used for the composition and optimization of ETL workflows.
Panos Vassiliadis, Alkis Simitsis, Eftychia Baikousi
DOLAP1
2008 Design Metrics for Data Warehouse Evolution
George Papastefanatos, Panos Vassiliadis, Alkis Simitsis, Yannis Vassiliou
ER2
2008 Meshing Streaming Updates with Persistent Data in an Active Data Warehouse
abstract
Active data warehousing has emerged as an alternative to conventional warehousing practices in order to meet the high demand of applications for up-to-date information. In a nutshell, an active warehouse is refreshed online and thus achieves a higher consistency between the stored information and the latest data updates. The need for online warehouse refreshment introduces several challenges in the implementation of data warehouse transformations, with respect to their execution time and their overhead to the warehouse processes. In this paper, we focus on a frequently encountered operation in this context, namely, the join of a fast stream 5" of source updates with a disk-based relation R, under the constraint of limited memory. This operation lies at the core of several common transformations such as surrogate key assignment, duplicate detection, or identification of newly inserted tuples. We propose a specialized join algorithm, termed mesh join (MESHJOIN), which compensates for the difference in the access cost of the two join inputs by 1) relying entirely on fast sequential scans of R and 2) sharing the I/O cost of accessing R across multiple tuples of 5". We detail the MESHJOIN algorithm and develop a systematic cost model that enables the tuning of MESHJOIN for two objectives: maximizing throughput under a specific memory budget or minimizing memory consumption for a specific throughput. We present an experimental study that validates the performance of MESHJOIN on synthetic and real-life data. Our results verify the scalability of MESHJOIN to fast streams and large relations and demonstrate its numerous advantages over existing join algorithms.
Neoklis Polyzotis, Spiros Skiadopoulos, Panos Vassiliadis, Alkis Simitsis, Nils-Erik Frantzell
IEEE Trans. Knowl. Data Eng.3
2007 What-If Analysis for Data Warehouse Evolution
George Papastefanatos, Panos Vassiliadis, Alkis Simitsis, Yannis Vassiliou
DaWaK2
2007 Deciding the physical implementation of ETL workflows
abstract
In this paper, we deal with the problem of determining the best possible physical implementation of an ETL workflow, given its logical-level description and an appropriate cost model as inputs. We formulate the problem as a state-space problem and provide a suitable solution for this task. We further extend this technique by intentionally introducing sorter activities in the workflow in order to search for alternative physical implementations with lower cost. We experimentally assess our method based on a principled organization of test suites.
Vasiliki Tziovara, Panos Vassiliadis, Alkis Simitsis
DOLAP2
2007 Supporting Streaming Updates in an Active Data Warehouse
abstract
Active data warehousing has emerged as an alternative to conventional warehousing practices in order to meet the high demand of applications for up-to-date information. In a nutshell, an active warehouse is refreshed on-line and thus achieves a higher consistency between the stored information and the latest data updates. The need for on-line warehouse refreshment introduces several challenges in the implementation of data warehouse transformations, with respect to their execution time and their overhead to the warehouse processes. In this paper, we focus on a frequently encountered operation in this context, namely, the join of a fast stream S of source updates with a disk-based relation R, under the constraint of limited memory. This operation lies at the core of several common transformations, such as, surrogate key assignment, duplicate detection or identification of newly inserted tuples. We propose a specialized join algorithm, termed mesh join (MeshJoin), that compensates for the difference in the access cost of the two join inputs by (a) relying entirely on fast sequential scans of R, and (b) sharing the I/O cost of accessing R across multiple tuples of S. We detail the Mesh Join algorithm and develop a systematic cost model that enables the tuning of Mesh Join for two objectives: maximizing throughput under a specific memory budget or minimizing memory consumption for a specific throughput. We present an experimental study that validates the performance of Mesh Join on synthetic and real-life data. Our results verify the scalability of Mesh-Join to fast streams and large relations, and demonstrate its numerous advantages over existing join algorithms.
Neoklis Polyzotis, Spiros Skiadopoulos, Panos Vassiliadis, Alkis Simitsis, Nils-Erik Frantzell
ICDE3
2007 Adding Context to Preferences
abstract
To handle the overwhelming amount of information currently available, personalization systems allow users to specify the information that interests them through preferences. Most often, users have different preferences depending on context. In this paper, we introduce a model for expressing such contextual preferences. Context is modeled as a set of multidimensional attributes. We formulate the context resolution problem as the problem of (a) identifying those preferences that qualify to encompass the context state of a query and (b) selecting the most appropriate among them. We also propose an algorithm for context resolution that uses a data structure, called the profile tree, that indexes preferences based on their associated context. Finally, we evaluate our approach from two perspectives: usability and performance.
Kostas Stefanidis, Evaggelia Pitoura, Panos Vassiliadis
ICDE3
2007 On Relaxing Contextual Preference Queries
abstract
Personalization systems exploit preferences for providing users with only relevant data from the huge volume of information that is currently available. We consider preferences that dependent on context, such as the location of the user. We model context as a set of attributes, each taking values from hierarchical domains. Often, the context of the query may be too specific to match any of the given preferences. In this paper, we consider possible expansions of the query context produced by relaxing one or more of its context attributes. A hierarchical attribute may be relaxed upwards by replacing its value by a more general one, downwards by replacing its value by a set of more specific values or sideways by replacing its value by sibling values in the hierarchy. We present an algorithm based on a prefix-based representation of context for identifying the preferences whose context matches the relaxed context of the query and some initial performance results.
Kostas Stefanidis, Evaggelia Pitoura, Panos Vassiliadis
MDM3
2007 Modeling and language support for the management of pattern-bases
Manolis Terrovitis, Panos Vassiliadis, Spiros Skiadopoulos, Elisa Bertino, Barbara Catania, Anna Maddalena, Stefano Rizzi
Data Knowl. Eng.2
2006 Modeling and Storing Context-Aware Preferences
Kostas Stefanidis, Evaggelia Pitoura, Panos Vassiliadis
ADBIS3
2006 A combination of trie-trees and inverted files for the indexing of set-valued attributes
abstract
Set-valued attributes frequently occur in contexts like market-basked analysis and stock market trends. Late research literature has mainly focused on set containment joins and data mining without considering simple queries on set valued attributes. In this paper we address superset, subset and equality queries and we propose a novel indexing scheme for answering them on set-valued attributes. The proposed index superimposes a trie-tree on top of an inverted file that indexes a relation with set-valued data. We show that we can efficiently answer the aforementioned queries by indexing only a subset of the most frequent of the items that occur in the indexed relation. Finally, we show through extensive experiments that our approach outperforms the state of the art mechanisms and scales gracefully as database size grows.
Manolis Terrovitis, Spyros Passas, Panos Vassiliadis, Timos K. Sellis
CIKM3
2005 Graph-Based Modeling of ETL Activities with Multi-level Transformations and Updates
Alkis Simitsis, Panos Vassiliadis, Manolis Terrovitis, Spiros Skiadopoulos
DaWaK2
2005 Blueprints and Measures for ETL Workflows
Panos Vassiliadis, Alkis Simitsis, Manolis Terrovitis, Spiros Skiadopoulos
ER1
2005 Optimizing ETL Processes in Data Warehouses
abstract
Extraction-transformation-loading (ETL) tools are pieces of software responsible for the extraction of data from several sources, their cleansing, customization and insertion into a data warehouse. Usually, these processes must be completed in a certain time window; thus, it is necessary to optimize their execution time. In this paper, we delve into the logical optimization of ETL processes, modeling it as a state-space search problem. We consider each ETL workflow as a state and fabricate the state space through a set of correct state transitions. Moreover, we provide algorithms towards the minimization of the execution cost of an ETL workflow.
Alkis Simitsis, Panos Vassiliadis, Timos K. Sellis
ICDE2
2005 A generic and customizable framework for the design of ETL scenarios
Panos Vassiliadis, Alkis Simitsis, Panos Georgantas, Manolis Terrovitis, Spiros Skiadopoulos
Inf. Syst.1
2005 State-Space Optimization of ETL Workflows
abstract
Extraction-transformation-loading (ETL) tools are pieces of software responsible for the extraction of data from several sources, their cleansing, customization, and insertion into a data warehouse. In this paper, we derive into the logical optimization of ETL processes, modeling it as a state-space search problem. We consider each ETL workflow as a state and fabricate the state space through a set of correct state transitions. Moreover, we provide an exhaustive and two heuristic algorithms toward the minimization of the execution cost of an ETL workflow. The heuristic algorithm with greedy characteristics significantly outperforms the other two algorithms for a large set of experimental cases.
Alkis Simitsis, Panos Vassiliadis, Timos K. Sellis
IEEE Trans. Knowl. Data Eng.2
2005 Computing and Managing Cardinal Direction Relations
abstract
Qualitative spatial reasoning forms an important part of the commonsense reasoning required for building intelligent geographical information systems (GIS). Previous research has come up with models to capture cardinal direction relations for typical GIS data. In this paper, we target the problem of efficiently computing the cardinal direction relations between regions that are composed of sets of polygons and present two algorithms for this task. The first of the proposed algorithms is purely qualitative and computes, in linear time, the cardinal direction relations between the input regions. The second has a quantitative aspect and computes, also in linear time, the cardinal direction relations with percentages between the input regions. Our experimental evaluation indicates that the proposed algorithms outperform existing methodologies. The algorithms have been implemented and embedded in an actual system, CARDIRECT, that allows the user to 1) specify and annotate regions of interest in an image or a map, 2) compute cardinal direction relations between them, and 3) pose queries in order to retrieve combinations of interesting regions.
Spiros Skiadopoulos, Christos Giannoukos, Nikos Sarkas, Panos Vassiliadis, Timos K. Sellis, Manolis Koubarakis
IEEE Trans. Knowl. Data Eng.4
2004 Computing and Handling Cardinal Direction Information
Spiros Skiadopoulos, Christos Giannoukos, Panos Vassiliadis, Timos K. Sellis, Manolis Koubarakis
EDBT3
2004 Data Mapping Diagrams for Data Warehouse Design with UML
Sergio Luján-Mora, Panos Vassiliadis, Juan Trujillo 0001
ER2
2004 Modeling and Language Support for the Management of Pattern-Bases
Manolis Terrovitis, Panos Vassiliadis, Spiros Skiadopoulos, Elisa Bertino, Barbara Catania, Anna Maddalena
SSDBM2
2003 A Framework for the Design of ETL Scenarios
Panos Vassiliadis, Alkis Simitsis, Panos Georgantas, Manolis Terrovitis
CAiSE1
2003 CPM: A Cube Presentation Model for OLAP
Andreas S. Maniatis, Panos Vassiliadis, Spiros Skiadopoulos, Yannis Vassiliou
DaWaK2
2003 Advanced visualization for OLAP
abstract
Data visualization is one of the big issues of database research. OLAP as a decision support technology is highly related to the developments of data visualization area. In this paper we demonstrate how the Cube Presentation Model (CPM), a novel presentational model for OLAP screens, can be naturally mapped on the Table Lens, which is an advanced visualization technique from the Human-Computer Interaction area, particularly tailored for cross-tab reports. We consider how the user interacts with an OLAP screen and based on the particularities of Table Lens, we propose an automated proactive users support. Finally, we discuss the necessity and the applicability of advanced visualization techniques in the presence of recent technological developments. Copyright 2003 ACM.
Andreas S. Maniatis, Panos Vassiliadis, Spiros Skiadopoulos, Yannis Vassiliou
DOLAP2
2003 Towards a Logical Model for Patterns
Stefano Rizzi, Elisa Bertino, Barbara Catania, Matteo Golfarelli, Maria Halkidi, Manolis Terrovitis, Panos Vassiliadis, Michalis Vazirgiannis, Euripides Vrachnos
ER7
2002 On the Logical Modeling of ETL Processes
Panos Vassiliadis, Alkis Simitsis, Spiros Skiadopoulos
CAiSE1
2002 Conceptual modeling for ETL processes
abstract
Extraction-Transformation-Loading (ETL) tools are pieces of software responsible for the extraction of data from several sources, their cleansing, customization and insertion into a data warehouse. In this paper, we focus on the problem of the definition of ETL activities and provide formal foundations for their conceptual representation. The proposed conceptual model is (a) customized for the tracing of inter-attribute relationships and the respective ETL activities in the early stages of a data warehouse project; (b) enriched with a 'palette' of a set of frequently used ETL activities, like the assignment of surrogate keys, the check for null values, etc; and (c) constructed in a customizable and extensible manner, so that the designer can enrich it with his own re-occurring patterns for ETL activities.
Panos Vassiliadis, Alkis Simitsis, Spiros Skiadopoulos
DOLAP1
2001 Data warehouse process management
Panos Vassiliadis, Christoph Quix, Yannis Vassiliou, Matthias Jarke
Inf. Syst.1
2001 ARKTOS: towards the modeling, design, control and execution of ETL processes
Panos Vassiliadis, Zografoula Vagena, Spiros Skiadopoulos, Nikos Karayannidis, Timos K. Sellis
Inf. Syst.1
2000 A Model for Data Warehouse Operational Processes
Panos Vassiliadis, Christoph Quix, Yannis Vassiliou, Matthias Jarke
CAiSE1
2000 Modelling and Optimisation Issues for Multidimensional Databases
Panos Vassiliadis, Spiros Skiadopoulos
CAiSE1
2000 Concept Based Design of Data Warehouses: The DWQ Demonstrators
abstract
The ESPRIT Project DWQ (Foundations of Data Warehouse Quality) aimed at improving the quality of DW design and operation through systematic enrichment of the semantic foundations of data warehousing. Logic-based knowledge representation and reasoning techniques were developed to control accuracy, consistency, and completeness via advanced conceptual modeling techniques for source integration, data reconciliation, and multi-dimensional aggregation. This is complemented by quantitative optimization techniques for view materialization, optimizing timeliness and responsiveness without losing the semantic advantages from the conceptual approach. At the operational level, query rewriting and materialization refreshment algorithms exploit the knowledge developed at design time. The demonstration shows the interplay of these tools under a shared metadata repository, based on an example extracted from an application at Telecom Italia.
Matthias Jarke, Christoph Quix, Diego Calvanese, Maurizio Lenzerini, Enrico Franconi, Spyros Ligoudistianos, Panos Vassiliadis, Yannis Vassiliou
SIGMOD Conference7
2000 Towards Quality-oriented Data Warehouse Usage and Evolution
Panos Vassiliadis, Mokrane Bouzeghoub, Christoph Quix
Inf. Syst.1
1999 Towards Quality-Oriented Data Warehouse Usage and Evolution
Panos Vassiliadis, Mokrane Bouzeghoub, Christoph Quix
CAiSE1
1999 Architecture and Quality in Data Warehouses: An Extended Repository Approach
Matthias Jarke, Manfred A. Jeusfeld, Christoph Quix, Panos Vassiliadis
Inf. Syst.4
1998 Architecture and Quality in Data Warehouses
Matthias Jarke, Manfred A. Jeusfeld, Christoph Quix, Panos Vassiliadis
CAiSE4
1998 Modeling Multidimensional Databases, Cubes and Cube Operations
abstract
Online analytical processing (OLAP) is a trend in database technology, which has attracted the interest of a lot of research work. OLAP is based on the multidimensional view of data, supported either by multidimensional databases (MOLAP) or relational engines (ROLAP). We propose a model for multidimensional databases. Dimensions, dimension hierarchies and cubes are formally introduced. We also introduce cube operations (changing of levels in the dimension hierarchy, function application, navigation etc.). The approach is based on the notion of the base cube, which is used for the calculation of the results of cube operations. We focus our approach on the support of a series of operations on cubes (i.e., the preservation of the results of previous operations and the applicability of aggregate functions in a series of operations). Furthermore, we provide a mapping of the multidimensional model to the relational model and to multidimensional arrays.
Panos Vassiliadis
SSDBM1