Myra Spiliopoulou

dblp:s/MyraSpiliopoulou · DBLP profile ↗
← Back
68ranked-venue papers in the field
13as first author
10since 2021 · last 2026
0000-0002-1828-5759ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 40 (9 first)Database Systems & Data Management · 11 (3 first)Information Retrieval & Web Search · 8Knowledge Engineering, Semantic Web & Information Systems · 5Other / Interdisciplinary · 4 (1 first)
YearPublicationVenuePosition
2026 Multimodal Sensing via Cost-Aware Progressive Inference for Stress Detection in Wearables
Sarun Varghese, Kumar Sai Jonnala, Myra Spiliopoulou
MDM3
2026 The Imitation Game: Evaluating Persona-Driven LLM Response Behavior in Web Surveys
Saijal Shahania, Myra Spiliopoulou, David Broneske
PAKDD (4)2
2025 Semi-supervised Learning with Pairwise Instance Comparisons for Medical Instance Classification
Anne Rother, Till Ittermann, Myra Spiliopoulou
IDA3
2023 A Similarity-Guided Framework for Error-Driven Discovery of Patient Neighbourhoods in EMA Data
Vishnu Unnikrishnan 0002, Miro Schleicher, Clara Puga, Rüdiger Pryss, Carsten Vogel, Winfried Schlee, Myra Spiliopoulou
IDA7
2023 WISHFUL - Website Extraction of Institutional Sources with Heterogeneous Factors and User-Driven Linkage
Saijal Shahania, Myra Spiliopoulou, David Broneske
iiWAS2
2022 Reducing Missingness in a Stream through Cost-Aware Active Feature Acquisition
abstract
Missing features can negatively impact the performance of machine learning solutions and past research has been focused on how to acquire the most predictive features in static scenarios. Active Feature Acquisition (AFA) for data streams extends the conventional static paradigm by taking into account that the importance of a feature may change as the stream drifts.In this study, we propose an AFA method that takes the cost of the features for each arriving instance into account and, at the same time, allows multiple features to be acquired at once. This reflects the fact that labels often depend on multiple features, so that acquiring many low-cost features may result in higher improvement than acquiring a single feature, as has been proposed in our earlier work [1]. We evaluated our approach on 7 real and 8 synthetic data sets. We investigated three different budget sizes and three different feature cost configurations for 7 different percentages of missing data on each of our data sets. We use the results of our experiments to elaborate on when the acquisition of multiple features is more or less beneficial than acquiring a single feature. Finally, we show the need for more sophisticated metrics to estimate a feature’s predictive quality.
Maik Büttner, Christian Beyer, Myra Spiliopoulou
DSAA3
2022 Expect the gap: A recommender approach to estimate the absenteeism of self-monitoring mHealth app users
abstract
Adherence and phases of non-adherence in the usage of mHealth apps for self-monitoring generate time series characterized by gaps of varying duration. These reduce the knowledge that can be gained from the data because important observations are not captured. In this paper, an approach is presented that allows experts to estimate the possible duration of a user’s absence in order to build on it and take appropriate action.For this purpose, the users’ time series (with gaps and of different lengths) are decomposed into sequences that have no gaps anymore. Each sequence is associated with the duration of the subsequent gap. Using unsupervised binning, these gap values are sorted into a predefined number of categories. Sequences with similar gap durations are thus assigned the same category labels, grouped using unsupervised time series clustering and are assigned with cluster labels. Thus triplets are derived consisting of user ID, cluster label and gap category label. These triplets can then be used for collaborative filtering with matrix factorization. It is now possible to estimate the gap duration even if the same sequence has not yet been observed for the app user.The results show that the quality of binning depends on the appropriate choice of the number of categories, on the technique used, and on the maximum length of the gaps. In this example, a 5-star rating, the Fischer-Jenks algorithm and a maximum gap length of 30 days. We demonstrate how clustering can be used to gradually adapt the time series to the conditions of collaborative filtering and how complex matrix factorization models respond better to the complex structures of the data and lead to better results. In this work, K-Medoids and Agglomerative Clustering obtained applicable results and for matrix factorization SVD++.We believe that our approach provides a valuable back-end tool for experts to better assess the adherence of users to a selfmonitoring app.
Miro Schleicher, Rüdiger Pryss, Johannes Schobel, Winfried Schlee, Myra Spiliopoulou
DSAA5
2021 Discovery of Patient Phenotypes through Multi-layer Network Analysis on the Example of Tinnitus
abstract
Electronic health records (EHR) often include multiple perspectives on a patient's current state of well-being (e.g. vital signs and subjective indicators measured by questionnaires). In this study, we use these perspectives to build phenotypes of chronic tinnitus patients and investigate how these phenotypes are associated with response to treatment. Therefore, we model patients as nodes in a network, where those perspectives are interpreted as layers of a multi-layer network. To identify phenotypes of patients in the network, we implement a community detection algorithm. Some of these communities can be considered as phenotypes if they represent subgroups of patients that are similar according to the investigated perspectives. Furthermore, we analyze the influence of the layers on the final community structure of patients. We then propose a method to add layers given their community structure similarity. Finally, we fit a model, per community, to predict the treatment outcome. In some communities, this prediction outperformed the baseline scenario where the predictor was fitted to all patients.
Clara Puga, Uli Niemann, Vishnu Unnikrishnan 0002, Miro Schleicher, Winfried Schlee, Myra Spiliopoulou
DSAA6
2021 A Framework for Authorial Clustering of Shorter Texts in Latent Semantic Spaces
Rafi Trad, Myra Spiliopoulou
IDA2
2021 Guest editorial: Special issue on mining for health
Myra Spiliopoulou, Panagiotis Papapetrou
Data Min. Knowl. Discov.1
2020 Multivariate Time Series as Images: Imputation Using Convolutional Denoising Autoencoder
abstract
Missing data is a common occurrence in the time series domain, for instance due to faulty sensors, server downtime or patients not attending their scheduled appointments. One of the best methods to impute these missing values is Multiple Imputations by Chained Equations (MICE) which has the drawback that it can only model linear relationships among the variables in a multivariate time series. The advancement of deep learning and its ability to model non-linear relationships among variables make it a promising candidate for time series imputation. This work proposes a modified Convolutional Denoising Autoencoder (CDA) based approach to impute multivariate time series data in combination with a preprocessing step that encodes time series data into 2D images using Gramian Angular Summation Field (GASF). We compare our approach against a standard feed-forward Multi Layer Perceptron (MLP) and MICE. All our experiments were performed on 5 UEA MTSC multivariate time series datasets, where 20 to 50% of the data was simulated to be missing completely at random. The CDA model outperforms all the other models in 4 out of 5 datasets and is tied for the best algorithm in the remaining case.
Abdullah Al Safi, Christian Beyer, Vishnu Unnikrishnan 0002, Myra Spiliopoulou
IDA4
2019 Mining and Model Understanding on Medical Data
abstract
What are the basic forms of healthcare data? How are Electronic Health Records and Cohorts structured? How can we identify the key variables in such data and how important are temporal abstractions? What are the main challenges in knowledge extraction from medical data sources? What are the key machine algorithms used for this purpose? What are the main questions that clinicians and medical experts pose to machine learning researchers?
Myra Spiliopoulou, Panagiotis Papapetrou
KDD1
2018 Learning under Feature Drifts in Textual Streams
abstract
Huge amounts of textual streams are generated nowadays, especially in social networks like Twitter and Facebook. As the discussion topics and user opinions on those topics change drastically with time, those streams undergo changes in data distribution, leading to changes in the concept to be learned, a phenomenon called concept drift. One particular type of drift, that has not yet attracted a lot of attention is feature drift, i.e., changes in the features that are relevant for the learning task at hand. In this work, we propose an approach for handling feature drifts in textual streams. Our approach integrates i) an ensemble-based mechanism to accurately predict the feature/word values for the next time-point by taking into account the different features might be subject to different temporal trends and ii) a sketch-based feature space maintenance mechanism that allows for a memory-bounded maintenance of the feature space over the stream. Experiments with textual streams from the sentiment analysis, email preference and spam detection demonstrate that our approach achieves significantly better or competitive performance compared to baselines.
Damianos P. Melidis, Myra Spiliopoulou, Eirini Ntoutsi
CIKM2
2018 Entity-Level Stream Classification: Exploiting Entity Similarity to Label the Future Observations Referring to an Entity
abstract
Stream classification algorithms traditionally treat arriving observations as independent. However, in many applications the arriving examples may depend on the "entity" that generated them, e.g. in product reviewing or in the interactions of users with an application server. In this study, we investigate the potential of this dependency by partitioning the original stream of observations into entity-centric substreams and by incorporating entity-specific information into the learning model. We propose a k Nearest Neighbour inspired stream classification approach (kNN), in which the label of an arriving observation is predicted by exploiting knowledge on the observations belonging to this entity and to entities similar to it. For the computation of entity similarity, we consider knowledge about the observations and knowledge about the entity, potentially transferred from another domain. To distinguish between cases where this kind of knowledge transfer is beneficial for stream classification and cases where the knowledge on the entities does not contribute to classifying the observations, we also propose a heuristic approach based on random sampling of substreams using k Random Entities (kRE). Our learning scenario is not fully supervised: after acquiring labels for the initial few observations of each entity, we assume that no additional labels arrive, and attempt to predict the labels of near-future and far-future observations from that initial seed. We report on our findings from three datasets.
Christian Beyer, Vishnu Unnikrishnan 0002, Pawel Matuszyk, Uli Niemann, Rüdiger Pryss, Winfried Schlee, Eirini Ntoutsi, Myra Spiliopoulou
DSAA8
2018 Predicting Worker Disagreement for More Effective Crowd Labeling
abstract
Crowdsourcing is a popular mechanism used for labeling tasks to produce large corpora for training. However, producing a reliable crowd labeled training corpus is challenging and resource consuming. Research on crowdsourcing has shown that label quality is much affected by worker engagement and expertise. In this study, we postulate that label quality can also be affected by inherent ambiguity of the documents to be labeled. Such ambiguities are not known in advance, of course, but, once encountered by the workers, they lead to disagreement in the labeling - a disagreement that cannot be resolved by employing more workers. To deal with this problem, we propose a crowd labeling framework: we train a disagreement predictor on a small seed of documents, and then use this predictor to decide which documents of the complete corpus should be labeled and which should be checked for document-inherent ambiguities before assigning (and potentially wasting) worker effort on them. We report on the findings of the experiments we conducted on crowdsourcing a Twitter corpus for sentiment classification.
Stefan Räbiger, Gizem Gezici, Yücel Saygin, Myra Spiliopoulou
DSAA4
2018 How do annotators label short texts? Toward understanding the temporal dynamics of tweet labeling
Stefan Räbiger, Myra Spiliopoulou, Yücel Saygin
Inf. Sci.2
2018 Forgetting techniques for stream-based matrix factorization in recommender systems
Pawel Matuszyk, João Vinagre, Myra Spiliopoulou, Alípio Mário Jorge, João Gama 0001
Knowl. Inf. Syst.3
2017 Introduction to the special issue dedicated to the Journal Track of ECML PKDD 2017
Kurt Driessens, Dragi Kocev, Marko Robnik-Sikonja, Myra Spiliopoulou
Data Min. Knowl. Discov.4
2016 Extracting opinionated (sub)features from a stream of product reviews using accumulated novelty and internal re-organization
Max Zimmermann, Eirini Ntoutsi, Myra Spiliopoulou
Inf. Sci.3
2015 Probabilistic Active Learning in Datastreams
Daniel Kottke, Georg Krempl, Myra Spiliopoulou
IDA3
2015 Medical Mining: KDD 2015 Tutorial
abstract
In year 2015, we experience a proliferation of scientific publications, conferences and funding programs on KDD for medicine and healthcare. However, medical scholars and practitioners work differently from KDD researchers: their research is mostly hypothesis-driven, not data-driven. KDD researchers need to understand how medical researchers and practitioners work, what questions they have and what methods they use, and how mining methods can fit into their research frame and their everyday business. Purpose of this tutorial is to contribute to this learning process. We address medicine and healthcare; there the expertise of KDD scholars is needed and familiarity with medical research basics is a prerequisite. We aim to provide basics for (1) mining in epidemiology and (2) mining in the hospital. We also address, to a lesser extent, the subject of (3) preparing and annotating Electronic Health Records for mining.
Myra Spiliopoulou, Pedro Pereira Rodrigues, Ernestina Menasalvas Ruiz
KDD1
2015 Ageing-Based Multinomial Naive Bayes Classifiers Over Opinionated Data Streams
Sebastian Wagner 0005, Max Zimmermann, Eirini Ntoutsi, Myra Spiliopoulou
ECML/PKDD (1)4
2014 Mining Longitudinal Epidemiological Data to Understand a Reversible Disorder
Tommy Hielscher, Myra Spiliopoulou, Henry Völzke, Jens-Peter Kühn
IDA2
2014 Interactive Medical Miner: Interactively Exploring Subpopulations in Epidemiological Datasets
Uli Niemann, Myra Spiliopoulou, Henry Völzke, Jens-Peter Kühn
ECML/PKDD (3)2
2013 Correcting the Usage of the Hoeffding Inequality in Stream Mining
Pawel Matuszyk, Georg Krempl, Myra Spiliopoulou
IDA3
2013 Framework for Storing and Processing Relational Entities in Stream Mining
Pawel Matuszyk, Myra Spiliopoulou
PAKDD (2)2
2013 MONIC and Followups on Modeling and Monitoring Cluster Transitions
Myra Spiliopoulou, Eirini Ntoutsi, Yannis Theodoridis, René Schult
ECML/PKDD (3)1
2012 Where Are We Going? Predicting the Evolution of Individuals
Zaigham Faraz Siddiqui, Márcia D. B. Oliveira, João Gama 0001, Myra Spiliopoulou
IDA4
2012 A Semi-supervised Incremental Clustering Algorithm for Streaming Data
Maria Halkidi, Myra Spiliopoulou, Aikaterini Pavlou
PAKDD (1)2
2012 Guest editorial: special issue on a decade of mining the Web
Myra Spiliopoulou, Bamshad Mobasher, Olfa Nasraoui, Osmar R. Zaïane
Data Min. Knowl. Discov.1
2011 Extracting cross references from life science databases for search result ranking
abstract
Scholars in life sciences have to process huge amounts of data in a disciplined and efficient way. These data are spread among thousands of databases which overlap in content but differ substantially with respect to interface, formats and data structure. Search engines have the potential of assisting in data retrieval from these structured sources but fall short of providing a relevance ranking of the results that reflects the needs of life science scholars. One such need is to acquire insights to cross-references among entities in the databases, whereby search hits with many cross-references are expected to be more informative than those with few cross-references. In this work, we investigate to what extend this expectation holds. We propose BioXREF, a method that extracts cross-references from multiple life science databases by combining targeted crawling, pointer chasing, sampling and information extraction. We study the retrieval quality of our method and the relationship between manually crafted relevance ranking and relevance ranking based on cross-references, and report on first, promising results.
Anja Bachmann, René Schult, Matthias Lange 0001, Myra Spiliopoulou
CIKM4
2011 Online Clustering of High-Dimensional Trajectories under Concept Drift
Georg Krempl, Zaigham Faraz Siddiqui, Myra Spiliopoulou
ECML/PKDD (2)3
2011 Query-Sets + + : A Scalable Approach for Modeling Web Sites
Barbara Poblete, Myra Spiliopoulou, Marcelo Mendoza
SPIRE2
2011 Summarization Meets Visualization on Online Social Networks
abstract
Getting an overview of a large online social net-work and deciding which communities to join is a challenging task for a new user. We propose a method that maps a large network into a smaller graph with two kinds of nodes: a node of the first kind is representative of a community, a node of the second kind is neighbor to a representative and rejects the semantics of that community. Our approach encompasses a learning and ranking algorithm that derives this smaller graph from the original one, and a visualization algorithm that returns a graph layout to the observer. We report on our results on inspecting the network of a folksonomy.
Hans-Henning Gabriel, Myra Spiliopoulou, Emmanouela Stachtiari, Athena Vakali
Web Intelligence2
2010 Visually Summarizing Semantic Evolution in Document Streams with Topic Table
André Gohr, Myra Spiliopoulou, Alexander Hinneburg
IC3K2
2010 Tree Induction over Perennial Objects
Zaigham Faraz Siddiqui, Myra Spiliopoulou
SSDBM2
2010 Density-based semi-supervised clustering
Myra Spiliopoulou, Ernestina Menasalvas Ruiz
Data Min. Knowl. Discov.2
2010 Guest editorial: Special Issue - 13th International Conference on Natural Language and Information Systems (NLDB 2008)
Epaminondas Kapetanios, Vijayan Sugumaran, Myra Spiliopoulou
Data Knowl. Eng.3
2010 Privacy-preserving query log mining for business confidentiality protection
abstract
We introduce the concern of confidentiality protection of business information for the publication of search engine query logs and derived data. We study business confidentiality, as the protection of nonpublic data from institutions, such as companies and people in the public eye. In particular, we relate this concern to the involuntary exposure of confidential Web site information, and we transfer this problem into the field of privacy-preserving data mining. We characterize the possible adversaries interested in disclosing Web site confidential data and the attack strategies that they could use. These attacks are based on different vulnerabilities found in query log for which we present several anonymization heuristics to prevent them. We perform an experimental evaluation to estimate the remaining utility of the log after the application of our anonymization techniques. Our experimental results show that a query log can be anonymized against these specific attacks while retaining a significant volume of useful data.
Barbara Poblete, Myra Spiliopoulou, Ricardo Baeza-Yates
ACM Trans. Web2
2009 Topic Evolution in a Stream of Documents
abstract
Document collections evolve over time, new topics emerge and old ones decline. At the same time, the terminology evolves as well. Much literature is devoted to topic evolution in finite document sequences assuming a fixed vocabulary. In this study, we propose “Topic Monitor” for the monitoring and understanding of topic and vocabulary evolution over an infinite document sequence, i.e. a stream. We use Probabilistic Latent Semantic Analysis (PLSA) for topic modeling and propose new folding-in techniques for topic adaptation under an evolving vocabulary. We extract a series of models, on which we detect index-based topic threads as human-interpretable descriptions of topic evolution.
André Gohr, Alexander Hinneburg, René Schult, Myra Spiliopoulou
SDM4
2009 Combining Multiple Interrelated Streams for Incremental Clustering
Zaigham Faraz Siddiqui, Myra Spiliopoulou
SSDBM2
2009 Spectral Clustering in Social-Tagging Systems
Alexandros Nanopoulos, Hans-Henning Gabriel, Myra Spiliopoulou
WISE3
2007 Domain Relevance on Term Weighting
Marko Brunzel, Myra Spiliopoulou
NLDB2
2007 DENGRAPH: A Density-based Community Detection Algorithm
abstract
Detecting densely connected subgroups in graphs such as communities in social networks is of interest in many research fields. Several methods have been developed to find communities but most of them have a high time complexity and are thus not applicable for large networks. Inspired by the clustering algorithm incremental DBSCAN we propose a density-based graph clustering algorithm DENGRAPH that is designed to deal with large dynamic datasets with noise and present first experimental results.
Tanja Falkowski, Anja Barth, Myra Spiliopoulou
Web Intelligence3
2006 Discovering Emerging Topics in Unlabelled Text Collections
René Schult, Myra Spiliopoulou
ADBIS2
2006 Discovering Semantic Sibling Associations from Web Documents with XTREEM-SP
Marko Brunzel, Myra Spiliopoulou
DaWaK2
2006 Discovering Semantic Sibling Groups from Web Documents with XTREEM-SG
Marko Brunzel, Myra Spiliopoulou
EKAW2
2006 MONIC: modeling and monitoring cluster transitions
abstract
There is much recent work on detecting and tracking change in clusters, often based on the study of the spatiotemporal properties of a cluster. For the many applications where cluster change is relevant, among them customer relationship management, fraud detection and marketing, it is also necessary to provide insights about the nature of cluster change: Is a cluster corresponding to a group of customers simply disappearing or are its members migrating to other clusters? Is a new emerging cluster reflecting a new target group of customers or does it rather consist of existing customers whose preferences shift? To answer such questions, we propose the framework MONIC for modeling and tracking of cluster transitions. Our cluster transition model encompasses changes that involve more than one cluster, thus allowing for insights on cluster change in the whole clustering. Our transition tracking mechanism is not based on the topological properties of clusters, which are only available for some types of clustering, but on the contents of the underlying data stream. We present our first results on monitoring cluster transitions over the ACM digital library.
Myra Spiliopoulou, Eirini Ntoutsi, Yannis Theodoridis, René Schult
KDD1
2006 Mining and Visualizing the Evolution of Subgroups in Social Networks
abstract
A social network consists of people who interact in some way such as members of online communities sharing information via the WWW. To learn more about how to facilitate community building e.g. in organizations, it is important to analyze the interaction behavior of their members over time. So far, many tools have been provided that allow for the analysis of static networks and some for the temporal analysis of networks - however only on the vertex and edge level. In this paper we propose two approaches to analyze the evolution of two different types of online communities on the level of subgroups. The first method consists of statistical analyses and visualizations that allow for an interactive analysis of subgroup evolutions in communities that exhibit a rather membership structure. The second method is designed for the detection of communities in an environment with highly fluctuating members. For both methods, we discuss results of experiments with real data from an online student community
Tanja Falkowski, Jörg Bartelheimer, Myra Spiliopoulou
Web Intelligence3
2005 RELFIN - Topic Discovery for Ontology Enhancement and Annotation
Markus Schaal, Roland M. Müller, Marko Brunzel, Myra Spiliopoulou
ESWC4
2004 Deriving Multiple Topics to Label Small Document Regions
Henner Graubitz, Myra Spiliopoulou
DaWaK2
2003 Efficient Monitoring of Patterns in Data Mining Environments
Steffan Baron, Myra Spiliopoulou, Oliver Günther 0001
ADBIS2
2002 Building and Exploiting Ad Hoc Concept Hierarchies for Web Log Analysis
Carsten Pohle, Myra Spiliopoulou
DaWaK2
2002 Structuring Domain-Specific Text Archives by Deriving a Probabilistic XML DTD
Karsten Winkler, Myra Spiliopoulou
PKDD2
2002 Web Mining
Ron Kohavi, Brij M. Masand, Myra Spiliopoulou, Jaideep Srivastava
Data Min. Knowl. Discov.3
2002 A Survey of Temporal Knowledge Discovery Paradigms and Methods
abstract
With the increase in the size of data sets, data mining has recently become an important research topic and is receiving substantial interest from both academia and industry. At the same time, interest in temporal databases has been increasing and a growing number of both prototype and implemented systems are using an enhanced temporal understanding to explain aspects of behavior associated with the implicit time-varying nature of the universe. This paper investigates the confluence of these two areas, surveys the work to date, and explores the issues involved and the outstanding problems in temporal data mining.
John F. Roddick, Myra Spiliopoulou
IEEE Trans. Knowl. Data Eng.2
2001 Monitoring Change in Mining Results
Steffan Baron, Myra Spiliopoulou
DaWaK2
2001 The DIAsDEM Framework for Converting Domain-Specific Texts into XML Documents with Data Mining Techniques
abstract
Modern organizations are accumulating huge volumes of textual documents. To turn archives into valuable knowledge sources, textual content must become explicit and able to be queried. Semantic tagging with markup languages such as XML satisfies both requirements. We thus introduce the DIAsDEM* framework for extracting semantics from structural text units (e.g., sentences), assigning XML tags to them and deriving a flat XML DTD for the archive. DIAsDEM focuses on archives characterized by a peculiar terminology and by an implicit structure such as court filings and company reports. In the knowledge discovery phase, text units are iteratively clustered by similarity of their content. Each iteration outputs clusters satisfying a set of quality criteria. Text units contained in these clusters are tagged with semiautomatically determined cluster labels and XML tags respectively. Additionally, extracted named entities (e.g., persons) serve as attributes of XML tags. We apply the framework in a case study on the German Commercial Register.
Henner Graubitz, Myra Spiliopoulou, Karsten Winkler
ICDM2
2001 Data Mining for Measuring and Improving the Success of Web Sites
Myra Spiliopoulou, Carsten Pohle
Data Min. Knowl. Discov.1
2000 Web mining for e-commerce (workshop session - title only)
abstract
No abstract available.
Ron Kohavi, Myra Spiliopoulou, Jaideep Srivastava
KDD2
2000 Analysis of Navigation Behaviour in Web Sites Integrating Multiple Information Systems
Bettina Berendt, Myra Spiliopoulou
VLDB J.2
1999 Managing Interesting Rules in Sequence Mining
Myra Spiliopoulou
PKDD1
1999 Data Mining for the Web
Myra Spiliopoulou
PKDD1
1998 WUM - A Tool for WWW Ulitization Analysis
Myra Spiliopoulou, Lukas C. Faulstich
WebDB1
1996 Parallel Optimization of Large Join Queries with Set Operators and Aggregates in a Parallel Environment Supporting Pipeline
abstract
Proposes a parallel optimizer for queries containing a large number of joins, as well as set operators and aggregate functions. The platform for the execution is a shared-disk multiprocessor machine supporting bushy parallelism and pipeline processing. Our model partitions the query into almost independent subtrees that can be optimized simultaneously, and it applies an enhanced variation of the iterative improvement technique on those subtrees which contain a large number of joins; this technique is parallelized, too. In order to estimate the cost of the states constructed during the optimization of join subtrees, cost formulae are developed that estimate the cost of relational algebra operators when executed across coalescing pipes.
Myra Spiliopoulou, Michael Hatzopoulos, Yannis Cotronis
IEEE Trans. Knowl. Data Eng.1
1995 Query Processing for Multimedia Applications on Optical Media
Myra Spiliopoulou, Yannis Cotronis, Michael Hatzopoulos
Inf. Process. Lett.1
1995 A Multimedia Title Development Environment (MTDE)
Aphrodite Tsalgatidou, Constantin Halatsis, Myra Spiliopoulou, Michael Hatzopoulos
Inf. Process. Manag.3
1992 Translation of SQL queries into a graph structure: query transformations and pre-optimization issues in a pipeline multiprocessor environment
Myra Spiliopoulou, Michael Hatzopoulos
Inf. Syst.1