Yannis Manolopoulos

dblp:m/YManolopoulos · DBLP profile ↗
← Back
218ranked-venue papers
26as first author
17since 2021 · last 2026
0000-0003-4026-4329ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 133 · 15 first-author · 5 since 2021Artificial intelligence and machine learning · 49 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 8 first-author · 7 since 2021Software engineering, systems software and programming languages · 13 · 1 first-authorTheory of computation · 8 · 2 first-authorSystems, architecture and hardware · 6Computer networks · 6Graphics, computer vision, multimedia, augmented reality and games · 6Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Comparison of CRISP - DM Versus OSEMN Methodologies Using Linear Regression and Statistical Analysis
abstract
ABSTRACT Artificial intelligence (AI) has transformed numerous sectors by offering innovative solutions to complex problems. However, effective implementation of AI projects requires a systematic and integrative approach to stay current with the latest developments in the field. Despite these improvements, realising successful AI and data science initiatives involves employing systematic approaches that can flexibly respond to changing technology and data environments. Two techniques are used to elucidate the life cycle of a high‐level data science project. The six‐phase methodology termed CRISP‐DM (Cross Industry Standard Process for Data Mining) precisely illustrates the data science life cycle. However, the entire workflow conducted by data scientists is performed using the OSEMN methodology (Obtain‐Scrub‐Explore‐Model‐iNterpret). Here, the CRISP‐DM and OSEMN frameworks have been evaluated and compared. An empirical study has been conducted experimenting with four study scenarios, each producing enlightening results with respect to the model fit and the prediction rate. According to four case studies, CRISP‐DM provides a more accurate and efficient method. All issues taken into account, this study contributes to better understanding the most effective approaches for choosing and applying data mining procedures, giving researchers and practitioners guidance on the most appropriate approach for their data analytic tasks. By comparing and contrasting these approaches, this study adds to the discussion of best practices within the data science community for researchers and practitioners when it comes to selecting the most suitable framework for data analysis.
Ketjona Shameti, Dhuratë Hyseni, Yannis Manolopoulos, Betim Cico
Expert Syst. J. Knowl. Eng.3
2025 Leveraging Ethical Narratives to Enhance LLM-AutoML Generated Machine Learning Models
abstract
ABSTRACT The growing popularity of generative AI and large language models (LLMs) has sparked innovation alongside debate, particularly around issues of plagiarism and intellectual property law. However, a less‐discussed concern is the quality of code generated by these models, which often contains errors and encourages poor programming practices. This paper proposes a novel solution by integrating LLMs with automated machine learning (AutoML). By leveraging AutoML's strengths in hyperparameter tuning and model selection, we present a framework for generating robust and reliable machine learning (ML) algorithms. Our approach incorporates natural language processing (NLP) and natural language understanding (NLU) techniques to interpret chatbot prompts, enabling more accurate and customisable ML model generation through AutoML. To ensure ethical AI practices, we have also introduced a filtering mechanism to address potential biases and enhance accountability. The proposed methodology not only demonstrates practical implementation but also achieves high predictive accuracy, offering a viable solution to current challenges in LLM‐based code generation. In summary, this paper introduces a new application of NLP and NLU to extract features from chatbot prompts, feeding them into an AutoML system to generate ML algorithms. This approach is framed within a rigorous ethical framework, addressing concerns of bias and accountability while enhancing the reliability of code generation.
Jordan Nelson, Michalis Pavlidis, Andrew Fish, Nikolaos Polatidis, Yannis Manolopoulos
Expert Syst. J. Knowl. Eng.5
2025 Special issue on the 15th International Conference on Management of Digital EcoSystems (MEDES 2023) and the 27th International Database Engineered Applications Symposium (IDEAS 2023)
Joe Tekli, Djamal Benslimane, Richard Chbeir, Yannis Manolopoulos, Ngoc Thanh Nguyen 0001
World Wide Web (WWW)4
2024 Thematic Editorial: The Ubiquitous Network
abstract
The historical paper on The Ubiquitous B-tree, authored by Douglas Comer in 1979, assists in paraphrasing and expressing the motto of The Ubiquitous Network. Initially, the keyword network was preceded by the keyword computer, but during the last two decades, the keyword network has acquired a newer, modern meaning. The terms Network Science, Complex Networks, Random Networks, Social Networks, Sensor Networks and Brain Network are the new protagonists in the wider spectrum of Computer Science research, attracting interest from several overlapping communities from Data Management and Data Mining to Artificial Intelligence and Deep Learning and to Algorithmic Graph Theory and Information Retrieval. Moreover, applications of Complex Networks span from Bioinformatics and Brain Science to Wireless Sensors and the Internet of Things, Transportation, Financial markets, Scientometrics and Social Sciences. Graph Theory is a field widely considered to have been founded in 1736 when Leonhard Euler solved the problem of the Bridges of Königsberg by representing the landscape with a graph. Essentially, Euler represented land pieces with nodes and the connecting bridges with edges. For centuries, Graph Theory was considered as an area of the discipline of Mathematics. With the emergence of the new science of Computer Science, and its quintessence, the algorithm, Algorithmic Graph Theory became a new subfield at the interface between Mathematics and Computer Science.
Yannis Manolopoulos
Comput. J.1
2024 Thematic editorial: sentiment analysis
abstract
With the advent of Web 2.0 it was possible for users to upload multimedia material, like text, image, audio or video, to numerous online social networks (OSN), such as Facebook (now Meta), Twitter (now X), Flickr, Instagram and so on. Thus, enormous data were then available for extraction and processing. For example, the total number of monetizable daily active Twitter users reaches 240 million, whereas the total number of tweets sent per day reaches 500 million (https://www.omnicoreagency.com/twitter-statistics). This data abundance is a paradise for text mining researchers. Multimodal Sentiment Analysis is a rich research area, which examines any kind of modality, extracts the sentiments behind, and classifies to positive, negative or neutral ones. For example, from visual and auditory instances, such as facial expressions or voice tones, a spectrum of sentiments can be conveyed. Also, by using natural language processing techniques on posted text, such as reviews, opinions, or comments related to services or products, such as films, books, restaurants, and so on, underlying sentiments can be extracted. Apparently, as every language is characterized by its own syntax, grammar and vocabulary, and thus there is room for localization adaptations in each case.
Yannis Manolopoulos
Comput. J.1
2024 Editorial on the Special Issue of the World Wide Web journal with selected papers from the 22nd International Conference on Web Information Systems Engineering (WISE)
Richard Chbeir, Zi Huang, Yannis Manolopoulos, Fabrizio Silvestri
World Wide Web (WWW)3
2024 FSSDroid: Feature subset selection for Android malware detection
abstract
Abstract Android malware has become an increasingly important threat to individuals, organizations, and society, posing significant risks to data security, privacy, and infrastructure. As malware evolves in sophistication and complexity, the detection and mitigation of these malicious software instances have become more challenging and time consuming since the required number of features to identify potential malware can be very high. To address this issue, we have developed an effective feature selection methodology for malware detection in Android. The critical concern in the field of malware detection is the complexity of algorithms and the use of features that are used to detect malware. The present paper delivers a methodology for pre-processing datasets to select the most optimal features that will allow detecting malware, while maintaining very high accuracy. The proposed methodology has been tested on two real world datasets and the results indicate that the number of features is significantly reduced from 489 to between 19 and 28 for the first dataset and from 9503 to between 9 and 27 for the second dataset, whilst the accuracy is maintained as if all features were used.
Nikolaos Polatidis, Stelios Kapetanakis, Marcello Trovati, Ioannis Korkontzelos, Yannis Manolopoulos
World Wide Web (WWW)5
2023 Recommendation Systems in Scholarly Publishing
Yannis Manolopoulos
IC3K1
2023 Recommendation Systems in Scholarly Publishing
Yannis Manolopoulos
WEBIST1
2023 Editorial Note to the special issue of the Information Systems journal on Web Engineering with selected papers from ICWE 2021 conference
Richard Chbeir, Flavius Frasincar, Yannis Manolopoulos
Inf. Syst.3
2022 Fast and Accurate Evaluation of Collaborative Filtering Recommendation Algorithms
Nikolaos Polatidis, Stelios Kapetanakis, Elias Pimenidis, Yannis Manolopoulos
ACIIDS (1)4
2022 An unsupervised distance-based model for weighted rank aggregation with list pruning
Leonidas Akritidis, Athanasios Fevgas, Panayiotis Bozanis, Yannis Manolopoulos
Expert Syst. Appl.4
2021 News Recommendations by Combining Intra-session with Inter-session and Content-Based Probabilistic Modelling
Panagiotis Symeonidis, Dmitry Chaltsev, Markus Zanker, Yannis Manolopoulos
ICCCI4
2021 Sink Group Betweenness Centrality
abstract
This article introduces the concept of Sink Group Node Betweenness centrality to identify those nodes in a network that can “monitor” the geodesic paths leading towards a set of subsets of nodes; it generalizes both the traditional node betweenness centrality and the sink betweenness centrality. We also provide extensions of the basic concept for node-weighted networks, and also describe the dual notion of Sink Group Edge Betweenness centrality. We exemplify the merits of these concepts and describe some areas where they can be applied.
Evangelia Fragkou 0001, Dimitrios Katsaros 0001, Yannis Manolopoulos
IDEAS3
2021 An overlapping clustering approach for precision, diversity and novelty-aware recommendations
ChemsEddine Berbague, Nour El Islem Karabadji, Hassina Seridi-Bouchelaghem, Panagiotis Symeonidis, Yannis Manolopoulos, Wajdi Dhifli
Expert Syst. Appl.5
2021 RELINE: point-of-interest recommendations using multiple network embeddings
Giannis Christoforidis, Pavlos Kefalas, Apostolos N. Papadopoulos, Yannis Manolopoulos
Knowl. Inf. Syst.4
2021 Indexing and progressive top-k similarity retrieval of trajectories
Nikolaos Pliakis, Eleftherios Tiakas, Yannis Manolopoulos
World Wide Web3
2020 Preface
abstract
[No abstract available]
Costin Badica, Mirjana Ivanovic, Yannis Manolopoulos, Riccardo Rosati 0001, Paolo Torroni
Fundam. Informaticae3
2020 Efficient distance join query processing in distributed spatial data management systems
Francisco García-García 0001, Antonio Corral, Luis Iribarne, Michael Vassilakopoulos, Yannis Manolopoulos
Inf. Sci.5
2020 A Data-Driven Unified Framework for Predicting Citation Dynamics
abstract
With the rising interest in predicting the scientific output, various efforts have been made to predict a scientist's h-index or the citation trajectory of a publication. In this work, we employ a dynamic categorization for scientists to ensure at each stage of their careers a comparison amongst their peers and combine this grouping with predictive models to estimate a scientist's future impact, as expressed by citation counts. Moreover, we investigate a wide range of factors identifying their importance in determining the future of science for different performance and academic levels with particular emphasis on features describing a scholar's position in multi-layered collaboration and citation networks. The robustness of the approach is examined on a longitudinal dataset centered around 700,302 data points representing Computer Scientists in various time periods with their complete networks of over 18 million collaboration links and 36 million citations. Our results indicate up to 30 percent improvement in prediction performance compared to baseline methods along with an average R2=0.96 for short term and R2=0.91 for long term predictions.
Antonia Gogoglou, Yannis Manolopoulos
IEEE Trans. Big Data2
2020 Indexing in flash storage devices: a survey on challenges, current approaches, and future trends
Athanasios Fevgas, Leonidas Akritidis, Panayiotis Bozanis, Yannis Manolopoulos
VLDB J.4
2019 Recommending Points of Interest in LBSNs Using Deep Learning Techniques
abstract
The representation of real-life problems by using k-partite graphs introduced a new era in Machine Learning. Moreover, the merge of virtual and physical layers through Location Based Social Networks (LBSN s) offers a different meaning into the constructed graphs. To this point, multiple models introduced in literature that aim to support users with personalized recommendations. These approaches represent the mathematical models that aim to understand users' behaviour by finding patterns on users' check-ins, reviews, ratings, friendships, etc. With this paper we describe and compare 20 of those state-of-the-art deep learning models to bring into the surface some of their strengths and shortcomings. First, we categorize them according to: data factors or features they use, data representation, methodologies used and recommendation types they support. Then, we highlight the existing limitations that tackles their performance. Finally, we introduce research trends and future directions.
Giannis Christoforidis, Pavlos Kefalas, Apostolos N. Papadopoulos, Yannis Manolopoulos
INISTA4
2019 Skyline-based dissimilarity of images
Nikolaos Georgiadis, Eleftherios Tiakas, Yannis Manolopoulos, Apostolos N. Papadopoulos
J. Intell. Inf. Syst.3
2018 The Range Skyline Query
abstract
The range skyline query retrieves the dynamic skyline for every individual query point in a range by generalizing the point-based dynamic skyline query. Its wide-ranging applications enable users to submit their preferences within an interval of 'ideally sought' values across every dimension, instead of being limited to submit their preference in relation to a single sought value. This paper considers the query as a hyper-rectangle iso-oriented towards the axes of the multi-dimensional space and proposes: (i) main-memory algorithmic strategies, which are simple to implement and (ii) secondary-memory pruning mechanisms for processing range skyline queries efficiently. The proposed approach is progressive and I/O optimal. A performance evaluation of the proposed technique demonstrates its robustness and practicability.
Theodoros Tzouramanis, Eleftherios Tiakas, Apostolos N. Papadopoulos, Yannis Manolopoulos
CIKM4
2018 Recommendation of Points-of-Interest Using Graph Embeddings
abstract
The rapid growth of Location-based Social Networks (LBSNs) has lead to the generation of massive datasets which are collected in an exponential rate. The collected information may be used to facilitate users' needs with recommendations related to their past preferences. Many recommendation models were introduced in the literature, which learn by the history of users and provide recommendations for Points-of-Interest. Unfortunately, most of them ignore the relation existing among the temporal properties, the spatial attributes and the periodicity of the check-ins. In this work, we present a novel methodology, named JLGE, that combines all aforementioned factors into one unified approach which facilitates POI recommendations. In particular, the model jointly learns the embeddings of six informational graphs i.e., two unipartite (user-user and POIPOI) and four bipartite (user-location, user-time, location-user, and location-time) into the same latent space and personalize the recommendations based on these embeddings. We have experimentally evaluated the accuracy of our model using two real-world datasets in terms of the top-n POIs recommendations. The performance evaluation results indicate a significant improvement in accuracy, in comparison to another state-of-theart graph-based approach.
Giannis Christoforidis, Pavlos Kefalas, Apostolos N. Papadopoulos, Yannis Manolopoulos
DSAA4
2018 The Science of Science and a Multilayer Network Approach to Scientists' Ranking
abstract
The deluge of data on scholarly output created unique opportunities for identifying the drivers of modern science, for studying career paths of scientists, and for measuring the research performance. These massive data and processing methodologies have given rise to an exciting new field, namely Science of Science (SoS) as the successor of what is called scientometrics or informetrics for many decades. Science of Science is the offspring of the fertile cooperation of many disciplines, such as network science, statistics, machine learning, mathematical analysis, sociology of science and so on. In this article, we provide a comprehensive coverage of recent advances in SoS related to network analysis, prediction and ranking, and investigate the issue of scientist ranking from a multilayer network perspective. Towards this goal, we contrast by experiments the well-known h-index and the recently proposed indicator C3-index to a generalization of PageRank for multilayer networks, namely BiPlex PageRank, which is based on solid tensor analysis. Both the obtained results and the brief survey of SoS will deepen our faith to SoS and stimulate further efforts in this transdisciplinary field.
Georgios Sideris, Dimitrios Katsaros 0001, Antonis Sidiropoulos 0001, Yannis Manolopoulos
IDEAS4
2018 Secure Reverse k-Nearest Neighbours Search over Encrypted Multi-dimensional Databases
abstract
The reverse k-nearest neighbours search is a fundamental primitive in multi-dimensional (i.e. multi-attribute) databases with applications in location-based services, online recommendations, statistical classification, pat-tern recognition, graph algorithms, computer games development, and so on. Despite the relevance and popularity of the query, no solution has yet been put forward that supports it in encrypted databases while protecting at the same time the privacy of both the data and the queries. With the outsourcing of massive datasets in the cloud, it has become urgent to find ways of ensuring the fast and secure processing of this query in untrustworthy cloud computing. This paper presents searchable encryption schemes which can efficiently and securely enable the processing of the reverse k-nearest neighbours query over encrypted multi-dimensional data, including index-based search schemes which can carry out fast query response that preserves data confidentiality and query privacy. The proposed schemes resist practical attacks operating on the basis of powerful background knowledge and their efficiency is confirmed by a theoretical analysis and extensive simulation experiments.
Theodoros Tzouramanis, Yannis Manolopoulos
IDEAS2
2018 Spatial Batch-Queries Processing Using xBR ^+ -trees in Solid-State Drives
George Roumelis, Michael Vassilakopoulos, Antonio Corral, Athanasios Fevgas, Yannis Manolopoulos
MEDI5
2018 Network Analysis of the Science of Science: A Case Study in SOFSEM Conference
Antonia Gogoglou, Theodora Tsikrika, Yannis Manolopoulos
SOFSEM3
2018 Efficient large-scale distance-based join queries in spatialhadoop
Francisco García-García 0001, Antonio Corral, Luis Iribarne, Michael Vassilakopoulos, Yannis Manolopoulos
GeoInformatica5
2018 A symbolic dynamics approach to Epileptic Chronnectomics: Employing strings to predict crisis onset
Nantia D. Iakovidou, Nikolaos A. Laskaris, Kostas Tsichlas, Yannis Manolopoulos, Manolis Christodoulakis, Eleftherios S. Papathanasiou, Savvas S. Papacostas, Georgios D. Mitsis
Theor. Comput. Sci.4
2018 Recommendations based on a heterogeneous spatio-temporal social network
Pavlos Kefalas, Panagiotis Symeonidis, Yannis Manolopoulos
World Wide Web3
2017 Predicting the Evolution of Scientific Output
Antonia Gogoglou, Yannis Manolopoulos
ICCCI (1)2
2017 Use-based Optimization of Spatial Access Methods
abstract
Spatial access methods have been extensively studied in the literature, during last decades. Access methods were designed for efficient processing of demanding queries and extensive comparisons between such methods have been presented. However, choosing the best values for the parameters that affect the performance of a spatial access method when such a method is expected to be utilized within a specific workload / context of operations has not been studied, so far. In this paper, we present the design and implementation of a framework to evaluate and optimize a spatial index. The (very popular) family of R-trees is chosen as the index of focus, though the same process can be applied for other (spatial, or not) indexes or combinations of them. We elaborate on the antagonizing aspects in the design of an R-tree, present the design and implementation of a benchmarking framework and develop a performance model for this index that incorporates benchmarking results. Next, we develop an optimization framework that uses this model to provide an optimized set up (node occupancies and node splitting method) for a specific use context (dataset type, tree size, number of queries and type of queries). We also present experimental results of an indicative use of the developed benchmarking framework and optimizer for a limited range of use contexts.
Nikolaos Athanasiou, Michael Vassilakopoulos, Antonio Corral, Yannis Manolopoulos
MEDES4
2017 Detecting Intrinsic Dissimilarities in Large Image Databases through Skylines
abstract
In this paper we try to detect dissimilar images in image databases without defining a similarity ranking function by capturing the intrinsic dissimilarities of image descriptor vectors. To this end, we apply the skyline operation using their multi-dimensional descriptor vectors. We implemented a number of skyline methods, combined with four hashing state-of-the-art algorithms for data partitioning to create efficient indexing to secondary memory. We compared and evaluated their results by using two real image datasets to measure the performance and the effectiveness of the implemented algorithms. Detailed results show the efficiency and effectiveness of our approach.
Nikolaos Georgiadis, Eleftherios Tiakas, Yannis Manolopoulos
MEDES3
2017 Bulk Insertions into xBR ^+ -trees
George Roumelis, Michael Vassilakopoulos, Antonio Corral, Yannis Manolopoulos
MEDI4
2017 The dbMark: A benchmarking system for watermarking methods for relational databases
abstract
Digital technology keeps falling prey to the ease with which the unauthorized reproduction and distribution of digital objects can be achieved. Research on copyright protection of digital data must counterbalance the failures of legal measures against digital piracy. Digital watermarking tops the list of technical countermeasures and recent research has focused on proposing watermarking methods for relational data as well as on the types of attacks which aim at removing or destroying the watermark. With the database field of research in urgent need of developing tools and platforms for the benchmarking of these watermarking methods, the present work has undertaken to construct a benchmarking system for the automatic evaluation and fair comparison of watermarking schemes for relational data. The new software tool is called dbMark and its construction has taken into consideration the characteristics of a large set of parameters across a wide range of database watermarking applications. This article presents and comments upon the results of several experiments that evaluate the performance of three different watermarking methods taking into consideration several parameters and types of attacks. There are indications that the dbMark can be used effectively as the primary platform for the experimental evaluation of watermarking methods for relational data. The source code and the executable file of the dbMark are provided free-of-charge on the Web.
Stavros Kyriakopoulos, Theodoros Tzouramanis, Yannis Manolopoulos
RCIS3
2017 A Hybrid Model for Linking Multiple Social Identities Across Heterogeneous Online Social Networks
Athanasios Kokkos, Theodoros Tzouramanis, Yannis Manolopoulos
SOFSEM3
2017 A time-aware spatio-textual recommender system
Pavlos Kefalas, Yannis Manolopoulos
Expert Syst. Appl.2
2017 Preference dynamics with multimodal user-item interactions in social media recommendation
Dimitrios Rafailidis, Pavlos Kefalas, Yannis Manolopoulos
Expert Syst. Appl.3
2017 Landmark selection for spectral clustering based on Weighted PageRank
Dimitrios Rafailidis, Eleni Constantinou, Yannis Manolopoulos
Future Gener. Comput. Syst.3
2017 Efficient query processing on large spatial databases: A performance study
George Roumelis, Michael Vassilakopoulos, Antonio Corral, Yannis Manolopoulos
J. Syst. Softw.4
2016 Enhancing SpatialHadoop with Closest Pair Queries
Francisco García-García 0001, Antonio Corral, Luis Iribarne, Michael Vassilakopoulos, Yannis Manolopoulos
ADBIS5
2016 A Scientist's Impact over Time: The Predictive Power of Clustering with Peers
abstract
The identification of latent patterns in big scholarly data that concern the performance of researchers is a significant task because it can potentially impact scientific careers since they are based in funding and promotion. This article investigates the temporal evolution of a scientist's impact. Instead of taking a detailed, microscopic view that examines the citation curves of every scientist's article, the article develops a scalable, macroscopic methodology that uses the articles' citation profiles to build a more abstract and high-level profile that characterizes a scientist. This profile is utilized to cluster scientists in a set of 'performance' clusters. To this end, established techniques such as Principal Component Analysis and Self-Organizing Map clustering are employed as well as a set of proposed heuristics. The effectiveness of the proposed methodology is examined by comparing the resulting rankings with the outcomes of the peer-review procedures that resulted in the E. F. Codd and the Turing awards. The good match between the outcomes of computerized and peer-review procedures provides solid evidence that the proposed techniques constitute a promising analysis method for big scholarly data.
Antonia Gogoglou, Antonis Sidiropoulos 0001, Dimitrios Katsaros 0001, Yannis Manolopoulos
IDEAS4
2016 Bulk-Loading xBR ^+ -trees
George Roumelis, Michael Vassilakopoulos, Antonio Corral, Yannis Manolopoulos
MEDI4
2016 New plane-sweep algorithms for distance-based join queries in spatial databases
George Roumelis, Antonio Corral, Michael Vassilakopoulos, Yannis Manolopoulos
GeoInformatica4
2016 Efficient and flexible algorithms for monitoring distance-based outliers over data streams
Maria Kontaki, Anastasios Gounaris, Apostolos N. Papadopoulos, Kostas Tsichlas, Yannis Manolopoulos
Inf. Syst.5
2016 A Graph-Based Taxonomy of Recommendation Algorithms and Systems in LBSNs
abstract
Recently, location-based social networks (LBSNs) gave the opportunity to users to share geo-tagged information along with photos, videos, and SMSs. Recommender systems can exploit this geographic information to provide much more accurate and reliable recommendations to users. In this paper, we present and compare 16 real life LBSNs, bringing into surface their advantages/ disadvantages, their special functionalities, and their impact in the mobile social Web. Moreover, we describe and compare extensively 43 state-of-the-art recommendation algorithms for LBSNs. We categorize these algorithms according to: personalization type, recommendation type, data factors/features, problem modeling methodology, and data representation. In addition to the above categorizations which cannot cover all algorithms in an integrated way, we also propose a hybrid k-partite graph taxonomy to categorize them based on the number of the involved k-partite graphs. Finally, we compare the recommendation algorithms with respect to their evaluation methodology (i.e., datasets and metrics) and we highlight new perspectives for future work in LBSNs.
Pavlos Kefalas, Panagiotis Symeonidis, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.3
2016 Processing Top-k Dominating Queries in Metric Spaces
abstract
Top - k dominating queries combine the natural idea of selecting the k best items with a comprehensive “goodness” criterion based on dominance. A point p 1 dominates p 2 if p 1 is as good as p 2 in all attributes and is strictly better in at least one. Existing works address the problem in settings where data objects are multidimensional points. However, there are domains where we only have access to the distance between two objects. In cases like these, attributes reflect distances from a set of input objects and are dynamically generated as the input objects change. Consequently, prior works from the literature cannot be applied, despite the fact that the dominance relation is still meaningful and valid. For this reason, in this work, we present the first study for processing top- k dominating queries over distance-based dynamic attribute vectors, defined over a metric space . We propose four progressive algorithms that utilize the properties of the underlying metric space to efficiently solve the problem and present an extensive, comparative evaluation on both synthetic and real-world datasets.
Eleftherios Tiakas, George Valkanas, Apostolos N. Papadopoulos, Yannis Manolopoulos, Dimitrios Gunopulos
ACM Trans. Database Syst.4
2015 Web Content Management Systems Archivability
Vangelis Banos, Yannis Manolopoulos
ADBIS2
2015 The xBR ^+ -tree: An Efficient Access Method for Points
George Roumelis, Michael Vassilakopoulos, Thanasis Loukopoulos, Antonio Corral, Yannis Manolopoulos
DEXA (1)5
2015 Editorial
Yannis Manolopoulos, Spyros Sioutas
Neurocomputing1
2015 Special Issue on Advanced Paradigms of Neural Networks' Learning Algorithms and Architectures in Engineering (APNNAAE)
Yannis Manolopoulos, Lazaros S. Iliadis
Inf. Sci.1
2015 D-P2P-Sim+: A novel distributed framework for P2P protocols performance testing
Spyros Sioutas, Evangelos Sakkopoulos, Alexandros Panaretos, Dimitrios Tsoumakos, Panagiotis Gerolymatos, Giannis Tzimas, Yannis Manolopoulos
J. Syst. Softw.7
2015 Extended feature combination model for recommendations in location-based mobile services
Masoud Sattari, Ismail Hakki Toroslu, Pinar Karagöz, Panagiotis Symeonidis, Yannis Manolopoulos
Knowl. Inf. Syst.5
2014 Metric-Based Top-k Dominating Queries
abstract
Top-k dominating queries combine the natural idea of se-lecting the k best items with a comprehensive \\goodness" criterion based on dominance. A point p1 dominates p2 if p1 is as good as p2 in all attributes and is strictly better in at least one. Existing works address the problem in settings where data objects are multidimensional points. However, there are domains where we only have access to the dis-tance between two objects. In cases like these, attributes re ect distances from a set of input objects and are dynam-ically generated as the input objects change. Consequently, prior works from the literature can not be applied, despite the fact that the dominance relation is still meaningful and valid. For this reason, in this work, we present the rst study for processing top-k dominating queries over distance-based dynamic asttribute vectors, dened over a metric space. We propose four progressive algorithms that utilize the proper-ties of the underlying metric space to eciently solve the problem, and present an extensive, comparative evaluation on both synthetic and real world data sets.
Eleftherios Tiakas, George Valkanas, Apostolos N. Papadopoulos, Yannis Manolopoulos
EDBT4
2014 A Bi-objective Cost Model for Database Queries in a Multi-cloud Environment
abstract
Cost models are broadly used in query processing to drive the query optimization process, accurately predict the query execution time, schedule database query tasks, apply admission control and derive resource requirements to name a few applications. The main role of cost models is to produce the time needed to run the query on a specific machine. In a multi-cloud environment, this is insufficient in two aspects: firstly, the machines employed are not defined a-priori, and secondly, time estimates need to be complemented with monetary cost information, because both the economic cost and the performance are of primary importance. This work addresses these two shortcomings and aims to serve as the first proposal for a bi-objective query cost model that is suitable for queries executed over resources provided by potentially multiple cloud providers. Moreover, our approach is applicable to more generic data flow graphs, the execution plans of which do not necessarily comprise relational operators.
Zisis Karampaglis, Anastasios Gounaris, Yannis Manolopoulos
MEDES3
2014 Scalable Spectral Clustering with Weighted PageRank
Dimitrios Rafailidis, Eleni Constantinou, Yannis Manolopoulos
MEDI3
2014 A New Plane-Sweep Algorithm for the K-Closest-Pairs Query
George Roumelis, Michael Vassilakopoulos, Antonio Corral, Yannis Manolopoulos
SOFSEM4
2014 Introduction to the special issue of the World Wide Web journal on "Social Media Preservation and Applications"
Alexandra I. Cristea, Dimitrios Katsaros 0001, Yannis Manolopoulos
World Wide Web3
2013 GeoSocialRec: Explaining Recommendations in Location-Based Social Networks
Panagiotis Symeonidis, Antonis Krinis, Yannis Manolopoulos
ADBIS3
2013 New perspectives for recommendations in location-based social networks: time, privacy and explainability
abstract
Online social networks have attracted users' attention in the last decade. Recommendation services constitute a critical functionality of such social platforms: users receive recommendations about resources (documents, pieces of music) and potential friends (people with the same interests). Recently, technological progressions in smart phones enabled the exploitation of geographical data information in social networks. Users can now receive recommendations about new Points of Interest (POIs), and new activities in POIs. Eventually, Location-based Social Networks (LBSNs) may become the 'Next Big Thing' of the Internet industry. This paper surveys the related work and current state-of-the-art algorithms in LBSNs. We also provide three new perspectives that concern recommendations in LBSNs: time-awareness, user's privacy issues, and explainability of recommendations. We present the latest work in LBSNs by comparing real systems and by categorizing them in multiple ways (platforms, personalization, etc.).
Pavlos Kefalas, Panagiotis Symeonidis, Yannis Manolopoulos
MEDES3
2013 Continuous outlier detection in data streams: an extensible framework and state-of-the-art algorithms
abstract
Anomaly detection is an important data mining task, aiming at the discovery of elements that show significant diversion from the expected behavior; such elements are termed as outliers. One of the most widely employed criteria for determining whether an element is an outlier is based on the number of neighboring elements within a fixed distance (R), against a fixed threshold (k). Such outliers are referred to as distance-based outliers and are the focus of this work. In this demo, we show both an extendible framework for outlier detection algorithms and specific outlier detection algorithms for the demanding case where outlier detection is continuously performed over a data stream. More specifically: i) first we demonstrate a novel flavor of an open-source publicly available tool for Massive Online Analysis (MOA) that is endowed with capabilities to encapsulate algorithms that continuously detect outliers and ii) second, we present four online outlier detection algorithms. Two of these algorithms have been designed by the authors of this demo, with a view to improving on key aspects related to outlier mining, such as running time, flexibility and space requirements.
Dimitrios Georgiadis, Maria Kontaki, Anastasios Gounaris, Apostolos N. Papadopoulos, Kostas Tsichlas, Yannis Manolopoulos
SIGMOD Conference6
2013 From biological to social networks: Link prediction based on multi-way spectral clustering
Panagiotis Symeonidis, Nantia D. Iakovidou, Nikolaos Mantas, Yannis Manolopoulos
Data Knowl. Eng.4
2013 ART: sub-logarithmic decentralized range query processing with probabilistic guarantees
Spyros Sioutas, Peter Triantafillou, George Papaloukopoulos, Evangelos Sakkopoulos, Kostas Tsichlas, Yannis Manolopoulos
Distributed Parallel Databases6
2012 Text Classification by Aggregation of SVD Eigenvectors
Panagiotis Symeonidis, Ivaylo Kehayov, Yannis Manolopoulos
ADBIS3
2012 Geo-activity recommendations by using improved feature combination
abstract
In this paper, we propose a new model to integrate additional data, which is obtained from geospatial resources other than original data set in order to improve Location/Activity recommendations. The data set that is used in this work is a GPS trajectory of some users, which is gathered over 2 years. In order to have more accurate predictions and recommendations, we present a model that injects additional information to the main data set and we aim to apply a mathematical method on the merged data. On the merged data set, singular value decomposition technique is applied to extract latent relations. Several tests have been conducted, and the results of our proposed method are compared with a similar work for the same data set.
Masoud Sattari, Murat Manguoglu, Ismail Hakki Toroslu, Panagiotis Symeonidis, Pinar Karagöz, Yannis Manolopoulos
UbiComp6
2012 A generalized taxonomy of explanations styles for traditional and social recommender systems
Alexis Papadimitriou, Panagiotis Symeonidis, Yannis Manolopoulos
Data Min. Knowl. Discov.3
2012 Edge betweenness centrality: A novel algorithm for QoS-based topology control over wireless sensor networks
Alfredo Cuzzocrea, Alexis Papadimitriou, Dimitrios Katsaros 0001, Yannis Manolopoulos
J. Netw. Comput. Appl.4
2012 Fast and accurate link prediction in social networking systems
Alexis Papadimitriou, Panagiotis Symeonidis, Yannis Manolopoulos
J. Syst. Softw.3
2012 Continuous Top-k Dominating Queries
abstract
Top-k dominating queries use an intuitive scoring function which ranks multidimensional points with respect to their dominance power, i.e., the number of points that a point dominates. The k points with the best (e.g., highest) scores are returned to the user. Both top-k and skyline queries have been studied in a streaming environment, where changes to the data set are very frequent. In such an environment, continuous query processing techniques are required toward efficient monitoring of query results, since periodic query re-execution is computationally intensive, and therefore, prohibitive. This work contains the first study of continuous top-k dominating queries over data streams. In comparison to continuous top-k and skyline queries, continuous top-k dominating queries pose additional challenges. Three exact algorithms (BFA, EVA, ADA) are studied, and among them ADA, which is enhanced with additional optimization techniques, shows the best overall performance. In some cases, we are willing to trade accuracy for speed. Toward this direction, two approximate algorithms are proposed (AHBA and AMSA). AHBA offers probabilistic guarantees regarding the accuracy of the result based on the Hoeffding bound, whereas AMSA performs a more aggressive computation resulting in more efficient processing. Evaluation results, based on real-life and synthetic data sets, show the efficiency and scalability of our techniques.
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.3
2011 NEFOS: Rapid Cache-Aware Range Query Processing with Probabilistic Guarantees
Spyros Sioutas, Kostas Tsichlas, Ioannis Karydis, Yannis Manolopoulos, Yannis Theodoridis
DEXA (1)4
2011 TAGs: scalable threshold-based algorithms for proximity computation in graphs
abstract
A fundamental and very useful operation in graphs is the computation of the proximity between nodes, i.e., the degree of dissimilarity (or similarity) between two nodes v and u. This is an important tool both in graph databases and graph mining applications, because it provides the base to support more complex tasks such as graph partitioning, clustering, classification, to name a few. All methods proposed in the literature assume that proximity is computed on a single graph by using a single distance measure. In addition, most of them focus on the proximity between node pairs. In this work, we present for the first time, scalable algorithms that: (i) they support proximity computation in multiple graph instances, (ii) they enable the utilization of several distance measures, (iii) they support proximity queries around a source node without limiting to node pairs and (iv) they support extensions for metric-based and skyline query processing. The main result of our work is the design of Threshold Algorithms for Graphs (denoted as TAGs), which are studied and evaluated experimentally by using real-life as well as synthetic graphs, based on both the G(n, p) Erdõs-Rényi model and power law degree distributions.
Apostolos Lyritsis, Apostolos N. Papadopoulos, Yannis Manolopoulos
EDBT3
2011 Continuous monitoring of distance-based outliers over data streams
abstract
Anomaly detection is considered an important data mining task, aiming at the discovery of elements (also known as outliers) that show significant diversion from the expected case. More specifically, given a set of objects the problem is to return the suspicious objects that deviate significantly from the typical behavior. As in the case of clustering, the application of different criteria lead to different definitions for an outlier. In this work, we focus on distance-based outliers: an object x is an outlier if there are less than k objects lying at distance at most R from x. The problem offers significant challenges when a stream-based environment is considered, where data arrive continuously and outliers must be detected on-the-fly. There are a few research works studying the problem of continuous outlier detection. However, none of these proposals meets the requirements of modern stream-based applications for the following reasons: (i) they demand a significant storage overhead, (ii) their efficiency is limited and (iii) they lack flexibility. In this work, we propose new algorithms for continuous outlier monitoring in data streams, based on sliding windows. Our techniques are able to reduce the required storage overhead, run faster than previously proposed techniques and offer significant flexibility. Experiments performed on real-life as well as synthetic data sets verify our theoretical study.
Maria Kontaki, Anastasios Gounaris, Apostolos N. Papadopoulos, Kostas Tsichlas, Yannis Manolopoulos
ICDE5
2011 Product recommendation and rating prediction based on multi-modal social networks
abstract
Online Social Rating Networks (SRNs) such as Epinions and Flixter, allow users to form several implicit social networks, through their daily interactions like co-commenting on the same products, or similarly co-rating products. The majority of earlier work in Rating Prediction and Recommendation of products (e.g. Collaborative Filtering) mainly takes into account ratings of users on products. However, in SRNs users can also built their explicit social network by adding each other as friends. In this paper, we propose Social-Union, a method which combines similarity matrices derived from heterogeneous (unipartite and bipartite) explicit or implicit SRNs. Moreover, we propose an effective weighting strategy of SRNs influence based on their structured density. We also generalize our model for combining multiple social networks. We perform an extensive experimental comparison of the proposed method against existing rating prediction and product recommendation algorithms, using synthetic and two real data sets (Epinions and Flixter). Our experimental results show that our Social-Union algorithm is more effective in predicting rating and recommending products in SRNs.
Panagiotis Symeonidis, Eleftherios Tiakas, Yannis Manolopoulos
RecSys3
2011 Decentralized execution of linear workflows over web services
Efthymia Tsamoura, Anastasios Gounaris, Yannis Manolopoulos
Future Gener. Comput. Syst.3
2011 Nonlinear dimensionality reduction for efficient and effective audio similarity searching
Dimitrios Rafailidis, Alexandros Nanopoulos, Yannis Manolopoulos
Multim. Tools Appl.3
2011 Progressive processing of subspace dominating queries
Eleftherios Tiakas, Apostolos N. Papadopoulos, Yannis Manolopoulos
VLDB J.3
2011 High performance, low complexity cooperative caching for wireless sensor networks
Nikos Dimokas, Dimitrios Katsaros 0001, Leandros Tassiulas, Yannis Manolopoulos
Wirel. Networks4
2010 Estimation of the Maximum Domination Value in Multi-dimensional Data Sets
Eleftherios Tiakas, Apostolos N. Papadopoulos, Yannis Manolopoulos
ADBIS3
2010 Brief announcement: ART--sub-logarithmic decentralized range query processing with probabilistic guarantees
abstract
We focus on range query processing on large-scale, typically distributed infrastructures. In this work we present the ART (Autonomous Range Tree) structure, which outperforms the most popular decentralized structures, including Chord (and some of its successors), BATON (and its successor) and Skip-Graphs. ART supports the join/leave and range query operations in O(log log N) and O(log2b logN +|A|) expected w.h.p number of hops respectively, where the base b is a double-exponentially power of two, N is the total number of peers and |A| the answer size.
Spyros Sioutas, George Papaloukopoulos, Evangelos Sakkopoulos, Kostas Tsichlas, Yannis Manolopoulos, Peter Triantafillou
PODC5
2010 Brief announcement: on the quest of optimal service ordering in decentralized queries
abstract
This paper deals with pipelined queries over services. The execution plan of such queries defines an order in which the services are called. We present the theoretical underpinnings of a newly proposed algorithm that produces the optimal linear ordering corresponding to a query being executed in a decentralized manner, i.e., when the services communicate directly with each other. The optimality is defined in terms of query response time, which is determined by the bottleneck service in the plan. The properties discussed in this work allow a branch-and-bound approach to be very efficient.
Efthymia Tsamoura, Anastasios Gounaris, Yannis Manolopoulos
PODC3
2010 Transitive node similarity for link prediction in social networks with positive and negative links
abstract
Online social networks (OSNs) like Facebook, and Myspace recommend new friends to registered users based on local features of the graph (i.e. based on the number of common friends that two users share). However, OSNs do not exploit the whole structure of the network. Instead, they consider only pathways of maximum length 2 between a user and his candidate friends. On the other hand, there are global approaches, which detect the overall path structure in a network, being computationally prohibitive for huge-size social networks. In this paper, we define a basic node similarity measure that captures effectively local graph features. We also exploit global graph features introducing transitive node similarity. Moreover, we derive variants of our method that apply in signed networks. We perform extensive experimental comparison of the proposed method against existing recommendation algorithms using synthetic and real data sets (Facebook, Hi5 and Epinions). Our experimental results show that our FriendTNS algorithm outperforms other approaches in terms of accuracy and it is also time efficient. We show that a significant accuracy improvement can be gained by using information about both positive and negative edges.
Panagiotis Symeonidis, Eleftherios Tiakas, Yannis Manolopoulos
RecSys3
2010 Continuous Processing of Preference Queries in Data Streams
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
SOFSEM3
2010 Cache consistency in Wireless Multimedia Sensor Networks
Nikos Dimokas, Dimitrios Katsaros 0001, Yannis Manolopoulos
Ad Hoc Networks3
2010 Energy-efficient distributed clustering in wireless sensor networks
Nikos Dimokas, Dimitrios Katsaros 0001, Yannis Manolopoulos
J. Parallel Distributed Comput.3
2010 MusicBox: Personalized Music Recommendation Based on Cubic Analysis of Social Tags
abstract
Social tagging is becoming increasingly popular in music information retrieval (MIR). It allows users to tag music items like songs, albums, or artists. Social tags are valuable to MIR, because they comprise a multifaced source of information about genre, style, mood, users' opinion, or instrumentation. In this paper, we examine the problem of personalized music recommendation based on social tags. We propose the modeling of social tagging data with three-order tensors, which capture cubic (three-way) correlations between users-tags-music items. The discovery of latent structure in this model is performed with the Higher Order Singular Value Decomposition (HOSVD), which helps to provide accurate and personalized recommendations, i.e., adapted to the particular users' preferences. To address the sparsity that incurs in social tagging data and further improve the quality of recommendation, we propose to enhance the model with a tag-propagation scheme that uses similarity values computed between the music items based on audio features. As a result, the proposed model effectively combines both information about social tags and audio features. The performance of the proposed method is examined experimentally with real data from Last.fm. Our results indicate the superiority of the proposed approach compared to existing methods that suppress the cubic relationships that are inherent in social tagging data. Additionally, our results suggest that the combination of social tagging data with audio features is preferable than the sole use of the former.
Alexandros Nanopoulos, Dimitrios Rafailidis, Panagiotis Symeonidis, Yannis Manolopoulos
IEEE Trans. Speech Audio Process.4
2010 A Unified Framework for Providing Recommendations in Social Tagging Systems Based on Ternary Semantic Analysis
abstract
Social tagging is the process by which many users add metadata in the form of keywords, to annotate and categorize items (songs, pictures, Web links, products, etc.). Social tagging systems (STSs) can provide three different types of recommendations: They can recommend 1) tags to users, based on what tags other users have used for the same items, 2) items to users, based on tags they have in common with other similar users, and 3) users with common social interest, based on common tags on similar items. However, users may have different interests for an item, and items may have multiple facets. In contrast to the current recommendation algorithms, our approach develops a unified framework to model the three types of entities that exist in a social tagging system: users, items, and tags. These data are modeled by a 3-order tensor, on which multiway latent semantic analysis and dimensionality reduction is performed using both the higher order singular value decomposition (HOSVD) method and the kernel-SVD smoothing technique. We perform experimental comparison of the proposed method against state-of-the-art recommendation algorithms with two real data sets (Last.fm and BibSonomy). Our results show significant improvements in terms of effectiveness measured through recall/precision.
Panagiotis Symeonidis, Alexandros Nanopoulos, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.3
2009 A novel distributed P2P simulator architecture: D-P2P-sim
abstract
In this paper we introduce a novel distributed simulation environment with GUI for P2P simulations (D-P2P-Sim). The key aim is to provide the appropriate integrated set of tools in a single software solution to evaluate the performance of various protocols. The basic architecture of the distributed P2P simulator is based on a multi-threading, asynchronous, message passing and distributed environment with graphical user interface to facilitate ease of use by both researchers and programmers.
Spyros Sioutas, George Papaloukopoulos, Evangelos Sakkopoulos, Kostas Tsichlas, Yannis Manolopoulos
CIKM5
2009 Personalized selection of web services for mobile environments: the m-scroutz solution
abstract
In this paper we discuss the integration of QoS-awa research and personalization algorithms in order to discover effectual Web Services for the case of mobile web users. We present a number of novel ranking algorithms specially designed for mobile web architectures. To validate and evaluate the proposed algorithms we developed a fully working prototype for mobile-PDA devices. The prototype enables PDA users to access electronic shops and to retrieve information about their products, while going shopping. Comparative experimental results have shown encouraging results that prove m-scroutz to be effective.
Evangelos Sakkopoulos, Poulia Adamopoulou, Athanasios K. Tsakalidis, Spyros Sioutas, Yannis Manolopoulos
MEDES5
2009 Building an efficient P2P overlay for energy-level queries in sensor networks
abstract
After the debunking of some myths about why P2P overlays are not feasible in sensornets, many such solutions have been proposed. None of the existing P2P overlays for sensornets provide "Energy-Level Application and Services". On this purpose and based on the efficient P2P method presented in [16], we design a novel P2P overlay for Energy Level discovery in a sensornet, the so-called ELDT (Energy Level Distributed Tree). Sensor nodes are mapped to peers based on their energy level. As the energy levels change, the sensor nodes would have to move from one peer to another and this oparation is the most crucial for the efficient scalability of the proposed system. Similarly, as the energy level of a sensor node becomes extremelly low, that node may want to forward it's task to another node with the desired energy level. The adaptation of the P2P index presented in [16] quarantees the best-known query performance of the above operation. We experimentally verify this performance via an appropriate simulator we have designed for this purpose.
Spyros Sioutas, George Papaloukopoulos, Michalis Nik Xenos, Yannis Manolopoulos
MEDES5
2009 An experimental performance comparison for indexing mobile objects on the plane
abstract
We present a time-efficient approach to index objects moving on the plane to efficiently answer range queries about their future positions. Each object is moving with non small velocity u, meaning that the velocity value distribution is skewed (Zipf) towards umin in some range [umin, umax], where umin is a positive lower threshold. Our algorithm enhances a previously described solution [18] by accommodating the ISB-tree access method as presented in [6]. Experimental evaluation shows the improved performance, scalability and efficiency of the new algorithm.
Spyros Sioutas, George Papaloukopoulos, Kostas Tsichlas, Yannis Manolopoulos
MEDES4
2009 MoviExplain: a recommender system with explanations
abstract
Providing justification to a recommendation gives credibility to a recommender system. Some recommender systems (Amazon.com etc.) try to explain their recommendations, in an effort to regain customer acceptance and trust. But their explanations are poor, because they are based solely on rating data, ignoring the content data. Our prototype system MoviExplain is a movie recommender system that provides both accurate and justifiable recommendations.
Panagiotis Symeonidis, Alexandros Nanopoulos, Yannis Manolopoulos
RecSys3
2009 High performance, low complexity cooperative caching for Wireless Sensor Networks
abstract
During the last decade, Wireless Sensor Networks (WSNs) have emerged and matured at such point that currently support several applications like environment control, intelligent buildings, target tracking in battlefields, and many more. The vast majority of these applications require an optimization to the communication among the sensors so as to serve data in short latency and with minimal energy consumption. Cooperative data caching has been proposed as an effective and efficient technique to achieve these goals concurrently. The essence of these protocols is the selection of the sensor nodes which will take special roles in running the caching and request forwarding decisions. This article introduces a new metric to aid in the selection of such nodes. Based on this metric, we propose a new cooperative caching protocol, which is compared against the state-of-the-art competing protocols. The simulation results attest the superiority of the proposed protocol; the proposed solution achieves on the average 20% improvement w.r.t. the competing method for the examined performance measures.
Nikos Dimokas, Dimitrios Katsaros 0001, Leandros Tassiulas, Yannis Manolopoulos
WOWMOM4
2009 Music search engines: Specifications and challenges
Alexandros Nanopoulos, Dimitrios Rafailidis, Maria M. Ruxanda, Yannis Manolopoulos
Inf. Process. Manag.4
2009 Node and edge selectivity estimation for range queries in spatial networks
Eleftherios Tiakas, Apostolos N. Papadopoulos, Alexandros Nanopoulos, Yannis Manolopoulos
Inf. Syst.4
2009 Searching for similar trajectories in spatial networks
Eleftherios Tiakas, Apostolos N. Papadopoulos, Alexandros Nanopoulos, Yannis Manolopoulos, Dragan Stojanovic, Slobodanka Djordjevic-Kajan
J. Syst. Softw.4
2009 CDNs Content Outsourcing via Generalized Communities
abstract
Content distribution networks (CDNs) balance costs and quality in services related to content delivery. Devising an efficient content outsourcing policy is crucial since, based on such policies, CDN providers can provide client-tailored content, improve performance, and result in significant economical gains. Earlier content outsourcing approaches may often prove ineffective since they drive prefetching decisions by assuming knowledge of content popularity statistics, which are not always available and are extremely volatile. This work addresses this issue, by proposing a novel self-adaptive technique under a CDN framework on which outsourced content is identified with no a-priori knowledge of (earlier) request statistics. This is employed by using a structure-based approach identifying coherent clusters of "correlated" Web server content objects, the so-called Web page communities. These communities are the core outsourcing unit and in this paper a detailed simulation experimentation has shown that the proposed technique is robust and effective in reducing user-perceived latency as compared with competing approaches, i.e., two communities-based approaches, Web caching, and non-CDN.
Dimitrios Katsaros 0001, George Pallis 0001, Konstantinos Stamos, Athena Vakali, Antonis Sidiropoulos 0001, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.6
2008 Predictive Join Processing between Regions and Moving Objects
Antonio Corral, Manuel Torres 0001, Michael Vassilakopoulos, Yannis Manolopoulos
ADBIS4
2008 Continuous Trend-Based Clustering in Data Streams
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
DaWaK3
2008 Ranking music data by relevance and importance
abstract
Due to the rapidly increasing availability of audio files on the Web, it is relevant to augment search engines with advanced audio search functionality. In this context, the ranking of the retrieved music is an important issue. This paper proposes a music ranking method capable of flexibly fusing the music based on its relevance and importance. The fusion is controlled by a single parameter, which can be intuitively tuned by the user. The notion of authoritative music among relevant music is introduced, and social media mined from the Web is used in an innovative manner to determine both the relevance and importance of music. The proposed method may support users with diverse needs when searching for music.
Maria M. Ruxanda, Alexandros Nanopoulos, Christian S. Jensen, Yannis Manolopoulos
ICME4
2008 SkyGraph: An Algorithm for Important Subgraph Discovery in Relational Graphs
Apostolos N. Papadopoulos, Apostolos Lyritsis, Yannis Manolopoulos
ECML/PKDD (1)3
2008 Tag recommendations based on tensor dimensionality reduction
abstract
Social tagging is the process by which many users add metadata in the form of keywords, to annotate and categorize information items (songs, pictures, web links, products etc.). Collaborative tagging systems recommend tags to users based on what tags other users have used for the same items, aiming to develop a common consensus about which tags best describe an item. However, they fail to provide appropriate tag recommendations, because: (i) users may have different interests for an information item and (ii) information items may have multiple facets. In contrast to the current tag recommendation algorithms, our approach develops a unified framework to model the three types of entities that exist in a social tagging system: users, items and tags. These data is represented by a 3-order tensor, on which latent semantic analysis and dimensionality reduction is performed using the Higher Order Singular Value Decomposition (HOSVD) technique. We perform experimental comparison of the proposed method against two state-of-the-art tag recommendations algorithms with two real data sets (Last.fm and BibSonomy). Our results show significant improvements in terms of effectiveness measured through recall/precision.
Panagiotis Symeonidis, Alexandros Nanopoulos, Yannis Manolopoulos
RecSys3
2008 SkyGraph: an algorithm for important subgraph discovery in relational graphs
Apostolos N. Papadopoulos, Apostolos Lyritsis, Yannis Manolopoulos
Data Min. Knowl. Discov.3
2008 Foreword for the DKE special issue with selected papers from the 8th International Conference on Enterprise Information Systems (ICEIS' 2006)
Joaquim Filipe, Yannis Manolopoulos
Data Knowl. Eng.2
2008 A new approach on indexing mobile objects on the plane
Spyros Sioutas, Konstantinos Tsakalidis, Kostas Tsichlas, Christos Makris 0001, Yannis Manolopoulos
Data Knowl. Eng.5
2008 Collaborative recommender systems: Combining effectiveness and efficiency
Panagiotis Symeonidis, Alexandros Nanopoulos, Apostolos N. Papadopoulos, Yannis Manolopoulos
Expert Syst. Appl.4
2008 Nearest-biclusters collaborative filtering based on constant and coherent values
Panagiotis Symeonidis, Alexandros Nanopoulos, Apostolos N. Papadopoulos, Yannis Manolopoulos
Inf. Retr.4
2008 Continuous subspace clustering in streaming time series
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
Inf. Syst.3
2008 Cooperative Caching in Wireless Multimedia Sensor Networks
Nikos Dimokas, Dimitrios Katsaros 0001, Yannis Manolopoulos
Mob. Networks Appl.3
2008 Music Retrieval Over Wireless Ad-Hoc Networks
abstract
Wireless networks introduce new opportunities for music delivery. The trend of using mobile devices on wireless networks can significantly extent the recent change of paradigm in the model of music distribution by allowing mobile clients to search for audio music in a network of wireless mobile hosts. This paper introduces the application of content-based music information retrieval (CBMIR) in wireless ad-hoc networks. We investigate, for the first time in the literature, the challenges posed by the wireless medium and recognize the factors that require optimization. We propose novel techniques that attain a significant reduction in both response time and network traffic, compared to naive approaches. Extensive experimental results illustrate the appropriateness, effectiveness, and efficiency of the proposed method to this bandwidth-starving and volatility due to mobility and environment.
Ioannis Karydis, Alexandros Nanopoulos, Apostolos N. Papadopoulos, Dimitrios Katsaros 0001, Yannis Manolopoulos
IEEE Trans. Speech Audio Process.5
2008 Providing Justifications in Recommender Systems
abstract
Recommender systems are gaining widespread acceptance in e-commerce applications to confront the ldquoinformation overloadrdquo problem. Providing justification to a recommendation gives credibility to a recommender system. Some recommender systems (Amazon.com, etc.) try to explain their recommendations, in an effort to regain customer acceptance and trust. However, their explanations are not sufficient, because they are based solely on rating or navigational data, ignoring the content data. Several systems have proposed the combination of content data with rating data to provide more accurate recommendations, but they cannot provide qualitative justifications. In this paper, we propose a novel approach that attains both accurate and justifiable recommendations. We construct a feature profile for the users to reveal their favorite features. Moreover, we group users into biclusters (i.e., groups of users which exhibit highly correlated ratings on groups of items) to exploit partial matching between the preferences of the target user and each group of users. We have evaluated the quality of our justifications with an objective metric in two real data sets (Reuters and MovieLens), showing the superiority of the proposed method over existing approaches.
Panagiotis Symeonidis, Alexandros Nanopoulos, Yannis Manolopoulos
IEEE Trans. Syst. Man Cybern. Part A3
2008 Prefetching in Content Distribution Networks via Web Communities Identification and Outsourcing
Antonis Sidiropoulos 0001, George Pallis 0001, Dimitrios Katsaros 0001, Konstantinos Stamos, Athena Vakali, Yannis Manolopoulos
World Wide Web6
2007 Adaptive k-Nearest-Neighbor Classification Using a Dynamic Number of Nearest Neighbors
Stefanos Ougiaroglou, Alexandros Nanopoulos, Apostolos N. Papadopoulos, Yannis Manolopoulos, Tatjana Welzer
ADBIS4
2007 Indexing Mobile Objects on the Plane Revisited
Spyros Sioutas, Konstantinos Tsakalidis, Kostas Tsichlas, Christos Makris 0001, Yannis Manolopoulos
ADBIS5
2007 Domination Mining and Querying
Apostolos N. Papadopoulos, Apostolos Lyritsis, Alexandros Nanopoulos, Yannis Manolopoulos
DaWaK4
2007 Adaptive similarity search in streaming time series with sliding windows
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
Data Knowl. Eng.3
2007 Data Mining techniques for the detection of fraudulent financial statements
Efstathios Kirkos, Charalambos Spathis, Yannis Manolopoulos
Expert Syst. Appl.3
2007 Mining association rules in very large clustered domains
Alexandros Nanopoulos, Apostolos N. Papadopoulos, Yannis Manolopoulos
Inf. Syst.3
2007 Finding maximum-length repeating patterns in music databases
Ioannis Karydis, Alexandros Nanopoulos, Yannis Manolopoulos
Multim. Tools Appl.3
2006 Generalized Indexing for Energy-Efficient Access to Partially Ordered Broadcast Data inWireless Networks
abstract
Energy conservation and access efficiency are two fundamental though competing goals in broadcast wireless networks. To tackle the energy penalty from sequential searching, the interleaving of index with data items has been proposed. Although, quite important contributions exist on providing broadcast indexes, they have one or more of the following problems. Firstly, all of them assume total ordering among broadcast data, and none considers the more general case of partial ordering. Secondly, they are balanced structure, which does not fit the "linear (one-dimensional) structure" of the wireless medium, in which imbalanced structures may offer significant advantages. Thirdly, they do not take into account the skewness in the access pattern, which prohibits larger performance gains to be reaped. Finally, they require all index items to be of equal size, which may not always give the optimal performance. To cope with all these problems, we introduce a new imbalanced tree-structured index. The new index is shown to be a generalization of two previously proposed high-performance indexes, and it introduces for the first time the problem of indexing partially ordered broadcast data. We present an experimental analysis of the proposed method, contrasting it with competing techniques. The analysis exhibits the efficiency of the proposed index in reducing the energy consumption without noticeably worsening the access latency
Dimitrios Katsaros 0001, Nikos Dimokas, Yannis Manolopoulos
IDEAS3
2006 Efficient Incremental Subspace Clustering in Data Streams
abstract
Performing data mining tasks in streaming data is considered a challenging research direction, due to the continuous data evolution. In this work, we focus on the problem of clustering streaming time series, based on the sliding window paradigm. More specifically, we use the concept of alpha-clusters in each time instance separately. A subspace alpha-cluster consists of a set of streams, whose value difference is less than a in a consecutive number of time instances (dimensions). The clusters can be continuously and incrementally updated as the streaming time series evolve. The proposed technique is based on a careful examination of pair-wise stream similarities for a subset of dimensions and then, it is generalized for more streams per cluster. Performance evaluation results show that the proposed pruning criteria are important for search space reduction, and that the cost of incremental cluster monitoring is computationally more efficient than reclustering
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
IDEAS3
2006 Collaborative Filtering Process in a Whole New Light
abstract
Collaborative filtering (CF) systems are gaining widespread acceptance in recommender systems and e-commerce applications. These systems combine information retrieval and data mining techniques to provide recommendations for products, based on suggestions of users with similar preferences. Nearest-neighbor CF process is influenced by several factors, which were not examined carefully in past work. In this paper, we bring to surface these factors in order to identify existing false beliefs. Moreover, by being able to view the "big picture" from the CF process, we propose new approaches that substantially improve the performance of CF algorithms. For instance, we obtain more than 40% percent increase in precision in comparison to widely-used CF algorithms. We perform an extensive experimental evaluation, with several real data sets, and produce results that invalidate some existing beliefs and illustrate the superiority of the proposed extensions
Panagiotis Symeonidis, Alexandros Nanopoulos, Apostolos N. Papadopoulos, Yannis Manolopoulos
IDEAS4
2006 Trajectory Similarity Search in Spatial Networks
abstract
In several applications, data objects are assumed to move on predefined spatial networks such as road segments, railways, and invisible air routes. Moving objects may exhibit similarity with respect to their traversed paths, and therefore two objects can be correlated based on their path similarity. In this paper, we study similarity search for moving object trajectories for spatial networks. The problem poses some important challenges, since it is quite different from the case where objects are allowed to move without motion restrictions. Experimental results performed on real-life spatial networks show that trajectory similarity can be supported in an effective and efficient manner by using metric-based access methods
Eleftherios Tiakas, Apostolos N. Papadopoulos, Alexandros Nanopoulos, Yannis Manolopoulos, Dragan Stojanovic, Slobodanka Djordjevic-Kajan
IDEAS4
2006 The Geodesic Broadcast Scheme for Wireless Ad Hoc Networks
abstract
Broadcasting is an effective means for disseminating information in wireless ad hoc networks. In this paper we propose a novel distributed broadcasting protocol in wireless ad hoc networks, which is based on an highly efficient metric for characterizing the importance of a node, with respect to its contribution in covering the local neighborhood. The protocol is reliable and achieves small communication complexity with linear in the number of nodes computation complexity. Experimental results for a large variety of network topologies show that the proposed algorithm is capable of generating small connected dominating sets, which guarantee a relatively small number of rebroadcasts
Dimitrios Katsaros 0001, Yannis Manolopoulos
WOWMOM2
2006 Processing Distance Join Queries with Constraints
abstract
Distance join queries are used in many modern applications, such as spatial databases, spatiotemporal databases and data mining. One of the most common distance join queries is the closest-pair query (CPQ). Given two datasets DA and DB the CPQ retrieves the pair (a, b), where a ∈ DA and b ∈ DB⁠, having the smallest distance between all pairs of objects. An extension to this problem is to generate the k closest pairs of objects (k-CPQ). In several cases spatial constraints are applied, and object pairs that are retrieved must also satisfy these constraints. Although the application of spatial constraints seems natural towards a more focused search, only recently they have been studied for the CPQ problem with the restriction that DA = DB⁠. In this work, we focus on constrained closest-pair queries, between two distinct datasets DA and DB⁠, where objects from DA must be enclosed by a spatial region R. Several algorithms are presented and evaluated using real-life and synthetic datasets. Among them, a heap-based method enhanced with batch capabilities outperforms the other approaches as it is demonstrated by an extensive performance evaluation.
Apostolos N. Papadopoulos, Alexandros Nanopoulos, Yannis Manolopoulos
Comput. J.3
2006 Cost models for distance joins queries using R-trees
Antonio Corral, Yannis Manolopoulos, Yannis Theodoridis, Michael Vassilakopoulos
Data Knowl. Eng.2
2006 Indexed-based density biased sampling for clustering applications
Alexandros Nanopoulos, Yannis Theodoridis, Yannis Manolopoulos
Data Knowl. Eng.3
2006 On past-time indexing of moving objects
Katerina Raptopoulou, Michael Vassilakopoulos, Yannis Manolopoulos
J. Syst. Softw.3
2006 Generalized comparison of graph-based ranking algorithms for publications and authors
Antonis Sidiropoulos 0001, Yannis Manolopoulos
J. Syst. Softw.2
2005 VA-Files vs. R*-Trees in Distance Join Queries
Antonio Corral, Alejandro D'Ermiliis, Yannis Manolopoulos, Michael Vassilakopoulos
ADBIS3
2005 Continuous Trend-Based Classification of Streaming Time Series
Maria Kontaki, Apostolos N. Papadopoulos, Yannis Manolopoulos
ADBIS3
2005 Audio Indexing for Efficient Music Information Retrieval
abstract
This paper presents an algorithm that efficiently retrieves audio data similar to an audio query. The proposed method utilises a feature extraction method for acoustical music sequences. The extracted features are grouped by Minimum Bounding Rectangles (MBRs) and indexed by means of a spatial access method. We also present a novel false alarm resolution method that utilises a reverse order schema while calculating the distance of the query and results, in order to avoid costly operations. Performance evaluation results show that the proposed technique achieves considerable performance improvement in comparison to an existing method.
Ioannis Karydis, Alexandros Nanopoulos, Apostolos N. Papadopoulos, Yannis Manolopoulos
MMM4
2005 A data mining approach for location prediction in mobile environments
Gökhan Yavas, Dimitrios Katsaros 0001, Özgür Ulusoy, Yannis Manolopoulos
Data Knowl. Eng.4
2005 Fast mining of frequent tree structures by hashing and indexing
Dimitrios Katsaros 0001, Alexandros Nanopoulos, Yannis Manolopoulos
Inf. Softw. Technol.3
2005 A new perspective to automatically rank scientific conferences using digital libraries
Antonis Sidiropoulos 0001, Yannis Manolopoulos
Inf. Process. Manag.2
2004 Towards Quadtree-Based Moving Objects Databases
Katerina Raptopoulou, Michael Vassilakopoulos, Yannis Manolopoulos
ADBIS3
2004 Algorithms for processing K-closest-pair queries in spatial databases
Antonio Corral, Yannis Manolopoulos, Yannis Theodoridis, Michael Vassilakopoulos
Data Knowl. Eng.2
2004 Broadcast program generation for Webcasting
Dimitrios Katsaros 0001, Yannis Manolopoulos
Data Knowl. Eng.2
2004 Benchmarking access methods for time-evolving regional data
Theodoros Tzouramanis, Michael Vassilakopoulos, Yannis Manolopoulos
Data Knowl. Eng.3
2004 Multi-Way Distance Join Queries in Spatial Databases
Antonio Corral, Yannis Manolopoulos, Yannis Theodoridis, Michael Vassilakopoulos
GeoInformatica2
2004 Memory-adaptive association rules mining
Alexandros Nanopoulos, Yannis Manolopoulos
Inf. Syst.2
2004 Special issue on ADBIS 2002: advances in databases and information systems
Pavol Návrat, Yannis Manolopoulos, Gottfried Vossen
Inf. Syst.2
2003 Distance Join Queries of Multiple Inputs in Spatial Databases
Antonio Corral, Yannis Manolopoulos, Yannis Theodoridis, Michael Vassilakopoulos
ADBIS2
2003 Compressing Large Signature Trees
Maria Kontaki, Yannis Manolopoulos, Alexandros Nanopoulos
ADBIS2
2003 Hierarchical Bitmap Index: An Efficient and Scalable Indexing Technique for Set-Valued Attributes
Mikolaj Morzy, Tadeusz Morzy, Alexandros Nanopoulos, Yannis Manolopoulos
ADBIS4
2003 Clustering Mobile Trajectories for Resource Allocation in Mobile Environments
Dimitrios Katsaros 0001, Alexandros Nanopoulos, Murat Karakaya, Gökhan Yavas, Özgür Ulusoy, Yannis Manolopoulos
IDA6
2003 LR-tree: a Logarithmic Decomposable Spatial Index Method
abstract
Since its introduction in 1984, R-tree has been proven to be one of the most practical and well-behaved data structures for accommodating dynamic massive sets of geometric objects and conducting a very diverse set of queries on such datasets in real-world applications. This success has led to a variety of versions, each one trying to tune the performance parameters of the original proposal. Among them, the most prominent one is R*-tree, which employs a number of carefully designed heuristics and is widely ccepted as achieving the best performance in most cases. However, in the presence of actively changing datasets, R*-tree still does not avoid performance tuning with forced reinsertion, i.e. a process that performs a kind of local rebuilding. The latter fact has motivated the investigation of the adaptation of a known dynamization technique, based on carefully triggered local rebuildings, for converting static or semi-dynamic, main memory data structures to dynamic ones onto R*-trees. In this paper, we present LR-trees, a new efficient scheme for dynamic manipulation of large datasets, which combines the search performance of the bulk-loaded R-trees with the updated performance of R*-trees. Experimental results provide evidence on the latter statement and illustrate the superiority of the proposed method.
Panayiotis Bozanis, Alexandros Nanopoulos, Yannis Manolopoulos
Comput. J.3
2003 Introduction to the Special Section on Spatiotemporal Databases
abstract
Yannis Manolopoulos; Introduction to the Special Section on Spatiotemporal Databases, The Computer Journal, Volume 46, Issue 6, 1 January 2003, Pages 662–663, h
Yannis Manolopoulos
Comput. J.1
2003 Performance Evaluation of Lazy Deletion Methods in R-trees
Alexandros Nanopoulos, Michael Vassilakopoulos, Yannis Manolopoulos
GeoInformatica3
2003 Fast Nearest-Neighbor Query Processing in Moving-Object Databases
Katerina Raptopoulou, Apostolos N. Papadopoulos, Yannis Manolopoulos
GeoInformatica3
2003 Efficient storage and querying of sequential patterns in database systems
Alexandros Nanopoulos, Maciej Zakrzewicz, Tadeusz Morzy, Yannis Manolopoulos
Inf. Softw. Technol.4
2003 Parallel bulk-loading of spatial data
Apostolos N. Papadopoulos, Yannis Manolopoulos
Parallel Comput.2
2003 A Data Mining Algorithm for Generalized Web Prefetching
abstract
Predictive Web prefetching refers to the mechanism of deducing the forthcoming page accesses of a client based on its past accesses. In this paper, we present a new context for the interpretation of Web prefetching algorithms as Markov predictors. We identify the factors that affect the performance of Web prefetching algorithms. We propose a new algorithm called WM,,, which is based on data mining and is proven to be a generalization of existing ones. It was designed to address their specific limitations and its characteristics include all the above factors. It compares favorably with previously proposed algorithms. Further, the algorithm efficiently addresses the increased number of candidates. We present a detailed performance evaluation of WM, with synthetic and real data. The experimental results show that WM/sub o/ can provide significant improvements over previously proposed Web prefetching algorithms.
Alexandros Nanopoulos, Dimitrios Katsaros 0001, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.3
2002 An efficient and effective algorithm for density biased sampling
abstract
In this paper we describe a new density-biased sampling algorithm. It exploits spatial indexes and the local density information they preserve, to provide improved quality of sampling result and fast access to elements of the dataset. It attains improved sampling quality, with respect to factors like skew, noise or dimensionality. Moreover, it has the advantage of efficiently handling dynamic updates, and it requires low execution times. The performance of the proposed method is examined experimentally. The comparative results illustrate its superiority over existing methods.
Alexandros Nanopoulos, Yannis Manolopoulos, Yannis Theodoridis
CIKM2
2002 Image indexing and retrieval using signature trees
Mario A. Nascimento, Eleni Tousidou, Vishal Chitkara, Yannis Manolopoulos
Data Knowl. Eng.4
2002 On the Generation of Time-Evolving Regional Data
Theodoros Tzouramanis, Michael Vassilakopoulos, Yannis Manolopoulos
GeoInformatica3
2002 Signature-based structures for objects with set-valued attributes
Eleni Tousidou, Panayiotis Bozanis, Yannis Manolopoulos
Inf. Syst.3
2002 Efficient similarity search for market basket data
Alexandros Nanopoulos, Yannis Manolopoulos
VLDB J.2
2001 The Impact of Buffering on Closest Pairs Queries Using R-Trees
Antonio Corral, Michael Vassilakopoulos, Yannis Manolopoulos
ADBIS3
2001 Time split linear quadtree for indexing image databases
abstract
The time split B-tree (TSBT) is modified for indexing a database of evolving binary images. This is accomplished by embedding ideas from linear region quadtrees that make the TSBT able to support spatio-temporal query processing. To improve query performance, additional pointers are added to the leaf-nodes of the TSBT. The resulting access method is called time split linear quadtree (TSLQ). Algorithms for processing five spatio-temporal queries have been adapted to the new structure. Such queries appear in multimedia systems, or geographical information systems (GIS), when searched by content. The TSLQ was implemented and results of extensive experiments on query time performance are presented, indicating that the proposed algorithmic approaches outbalance respective straightforward algorithms. The region data sets used in the experiments were real images of meteorological satellite views and synthetic raster images.
Theodoros Tzouramanis, Michael Vassilakopoulos, Yannis Manolopoulos
ICIP (2)3
2001 C2P: Clustering based on Closest Pairs
Alexandros Nanopoulos, Yannis Theodoridis, Yannis Manolopoulos
VLDB3
2001 A generalized comparison of linear representations of thematic layers
Yannis Manolopoulos, Enrico Nardelli, Guido Proietti, Eleni Tousidou
Data Knowl. Eng.1
2001 Mining patterns from graph traversals
Alexandros Nanopoulos, Yannis Manolopoulos
Data Knowl. Eng.2
2001 Distributed Processing of Similarity Queries
Apostolos N. Papadopoulos, Yannis Manolopoulos
Distributed Parallel Databases2
2000 Closest Pair Queries in Spatial Databases
abstract
This paper addresses the problem of finding the K closest pairs between two spatial data sets, where each set is stored in a structure belonging in the R-tree family. Five different algorithms (four recursive and one iterative) are presented for solving this problem. The case of 1 closest pair is treated as a special case. An extensive study, based on experiments performed with synthetic as well as with real point data sets, is presented. A wide range of values for the basic parameters affecting the performance of the algorithms, especially the effect of overlap between the two data sets, is explored. Moreover, an algorithmic as well as an experimental comparison with existing incremental algorithms addressing the same problem is presented. In most settings, the new algorithms proposed clearly outperform the existing ones.
Antonio Corral, Yannis Manolopoulos, Yannis Theodoridis, Michael Vassilakopoulos
SIGMOD Conference2
2000 Fringe Analysis of 2-3 Trees with Lazy Parent Split
abstract
B-trees with lazy parent split (lps) are B-tree variants, according to which parent splits are postponed until a future access of the latter node. This way, the number of splits during an insert is decreased and the number of locks is also decreased. Consequently, better concurrency is achieved. In this paper 2–3 trees with lps are studied. Fringe analysis is used to obtain bounds on some performance metrics of 2–3 trees with lps. The performance metrics of 2–3 tree with lps are compared with those of the classical 2–3 trees. The conclusions are that 2–3 trees with lps have slightly better performance, more keys in the fringe, larger storage utilization and a slightly shorter path length from the root to the leaves.
Antigoni Manousaka, Yannis Manolopoulos
Comput. J.2
2000 Improved Methods for Signature-Tree Construction
abstract
Signature-based tree structures which have been proposed in the past do not perform well for large databases. The problem arises from the fact that they are incapable of pruning searching, especially at the upper tree levels, and thus they have decreased selectivities. In this paper, we locate a number of reasons for this problem and propose several methods for node splitting and partial-tree restructuring, which lead to improved query-response times. We have implemented all methods and we present experimental results, which indicate that the proposed methods are superior in all cases to the standard one and up to 5–10 times better for medium and higher weights in inclusive (partial-match) queries. Additionally, we have developed new functions for the performance estimation of signature trees which, in contrast to a previous estimation function, are able to take into account the outcome of different split methods and to provide more accurate estimation.
Eleni Tousidou, Alexandros Nanopoulos, Yannis Manolopoulos
Comput. J.3
2000 Overlapping Linear Quadtrees and Spatio-Temporal Query Processing
abstract
In this paper, indexing in spatio-temporal databases by using the technique of overlapping is investigated. Overlapping has been previously applied in various access methods to combine consecutive structure instances into a single structure, without storing identical sub-structures. In this way, space is saved without sacrificing time performance. A new access method, overlapping linear quadtrees is introduced. This structure is able to store consecutive historical raster images, a database of evolving images. Moreover, it can be used to support query processing in such a database. Five such spatio-temporal queries along with the respective algorithms that take advantage of the properties of the new structure are introduced. The new access method was implemented and extensive experimental studies for space efficiency and query processing performance were conducted. A number of results of these experiments are presented. As far as space is concerned, these results indicate that, in the case of similar consecutive images, considerable storage is saved in comparison to independent linear quadtrees. In the case of query processing, the results indicate that the proposed algorithmic approaches outperform the respective straightforward algorithms, in most cases. The region data sets used in experiments were real images of meteorological satellite views and synthetic random images with specified aggregation.
Theodoros Tzouramanis, Michael Vassilakopoulos, Yannis Manolopoulos
Comput. J.3
2000 Performance Evaluation of Parallel S-Trees
abstract
The S-tree is a dynamic height-balanced tree similar in structure to B+trees. S-trees store fixed length bit-strings, which are called signatures. Signatures are used for indexing textbases, relational, object oriented and extensible databases as well as in data mining. In this article, methods of designing multi-disk B-trees are adapted to S-trees and new methods of parallelizing S-trees are developed. The resulting structures aim at achieving performance gain by accessing two or more disks simultaneously. In addition, two different searching techniques that exploit parallel disk accessing are devised. Performance results of experiments based on the new structures and searching techniques are also presented and discussed.
Eleni Tousidou, Michael Vassilakopoulos, Yannis Manolopoulos
J. Database Manag.3
2000 Data placement schemes in replicated mirrored disk systems
Athena Vakali, Yannis Manolopoulos
J. Syst. Softw.2
1999 Processing of Spatio-Temporal Queries in Image Databases
Theodoros Tzouramanis, Michael Vassilakopoulos, Yannis Manolopoulos
ADBIS3
1999 Overlapping B+-Trees: An Implementation of a Transaction Time Access Method
Theodoros Tzouramanis, Yannis Manolopoulos, Nikos A. Lorentzos
Data Knowl. Eng.2
1998 Multiple Range Query Optimization in Spatial Databases
Apostolos N. Papadopoulos, Yannis Manolopoulos
ADBIS2
1998 Replication in Mirrored Disk Systems
Athena Vakali, Yannis Manolopoulos
ADBIS2
1998 Indexing Time-Series Databases for Inverse Queries
Alexandros Nanopoulos, Yannis Manolopoulos
DEXA2
1998 Similarity Query Processing Using Disk Arrays
abstract
Similarity queries are fundamental operations that are used extensively in many modern applications, whereas disk arrays are powerful storage media of increasing importance. The basic trade-off in similarity query processing in such a system is that increased parallelism leads to higher resource consumptions and low throughput, whereas low parallelism leads to higher response times. Here, we propose a technique which is based on a careful investigation of the currently available data in order to exploit parallelism up to a point, retaining low response times during query processing. The underlying access method is a variation of the R*-tree, which is distributed among the components of a disk array, whereas the system is simulated using event-driven simulation. The performance results conducted, demonstrate that the proposed approach outperforms by factors a previous branch-and-bound algorithm and a greedy algorithm which maximizes parallelism as much as possible. Moreover, the comparison of the proposed algorithm to a hypothetical (non-existing) optimal one (with respect to the number of disk accesses) shows that the former is on average two times slower than the latter.
Apostolos N. Papadopoulos, Yannis Manolopoulos
SIGMOD Conference2
1998 Specifications for Efficient Indexing in Spatiotemporal Databases
abstract
A new issue that arises in modern applications involves the efficient manipulation of (static or moving) spatial objects, and the relationships among them. As a result, modern database systems should be able to efficiently support that type of data. Towards this goal, appropriate extensions of multidimensional access methods can be exploited in order to index and retrieve spatiotemporal objects, satisfying users' demands. This paper introduces the basic specifications such a spatiotemporal index structure should follow, evaluates existing proposals with respect to the above specifications, and illustrates issues of interest involving object representation, query processing, and index maintenance.
Yannis Theodoridis, Timos K. Sellis, Apostolos N. Papadopoulos, Yannis Manolopoulos
SSDBM4
1998 Comparison of Signature File Models with Superimposed Coding
Dimitrios Dervos, Yannis Manolopoulos, Panagiotis Linardis
Inf. Process. Lett.2
1997 S-Index: a Hybrid Structure for Text Retrieval
Dimitrios Dervos, Panagiotis Linardis, Yannis Manolopoulos
ADBIS3
1997 Performance of Nearest Neighbor Queries in R-Trees
Apostolos N. Papadopoulos, Yannis Manolopoulos
ICDT2
1997 On Sampling Regional Data
Michael Vassilakopoulos, Yannis Manolopoulos
Data Knowl. Eng.2
1997 Nearest Neighbor Queries in Shared-Nothing Environments
Apostolos N. Papadopoulos, Yannis Manolopoulos
GeoInformatica2
1997 Parallel data paths in two-headed disk systems
Athena Vakali, Yannis Manolopoulos
Inf. Softw. Technol.2
1997 An Exact Analysis on Expected Seeks in Shadowed Disks
Athena Vakali, Yannis Manolopoulos
Inf. Process. Lett.2
1997 MOF-Tree: A Spatial Access Method to Manipulate Multiple Overlapping Features
Yannis Manolopoulos, Enrico Nardelli, Apostolos N. Papadopoulos, Guido Proietti
Inf. Syst.1
1997 Analysis of the n-Dimensional Quadtree Decomposition for Arbitrary Hyperectangles
abstract
We give a closed-form expression for the average number of n-dimensional quadtree nodes ("pieces" or "blocks") required by an n-dimensional hyperrectangle aligned with the axes. Our formula includes as special cases the formulae of previous efforts for two-dimensional spaces. It also agrees with theoretical and empirical results that the number of blocks depends on the hypersurface of the hyperrectangle and not on its hypervolume. The practical use of the derived formula is that it allows the estimation of the space requirements of the n-dimensional quadtree decomposition. Quadtrees are used extensively in two-dimensional spaces (geographic information systems and spatial databases in general), as well in higher dimensionality spaces (as oct-trees for three-dimensional spaces, e.g., in graphics, robotics, and three-dimensional medical images). Our formula permits the estimation of the space requirements for data hyperrectangles when stored in an index structure like a (n-dimensional) quadtree, as well as the estimation of the search time for query hyperrectangles, for the so-called linear quadtrees. A theoretical contribution of the paper is the observation that the number of blocks is a piece-wise linear function of the sides of the hyperrectangle.
Christos Faloutsos, H. V. Jagadish, Yannis Manolopoulos
IEEE Trans. Knowl. Data Eng.3
1996 Global Page Replacement in Spatial Databases
Apostolos N. Papadopoulos, Yannis Manolopoulos
DEXA2
1996 Experimenting with Pattern-Matching Algorithms
Yannis Manolopoulos, Christos Faloutsos
Inf. Sci.1
1996 On the creation of quadtrees by using a branching process
Yannis Manolopoulos, Enrico Nardelli, Guido Proietti, Michael Vassilakopoulos
Image Vis. Comput.1
1995 On the Generation of Aggregated Random Spatial Regions
abstract
Traditionalrandom models for spatial two-dimensional data proposed in literature show their limits in generating in a satisfactory way instances of regions having a desired aggregation level.This is because none of them is really oriented to this aim.Rather, they are thought to model the behaviour of the constituting elements of the spatial data (so loosing sight of the context), or, alternatively, to model particular data structure for their representation, underestimating the fact that there is in general no semantic link between a region data and its representation.This means from one hand, the impossibility to produce meaningful the oretical results on time and space average performances of different data structures used to represent spatial regions, and, on the other hand, in an applicative context, the difficulty to generate instances of spatial regions having a statistical behaviour close to that of real data.To overcome this trouble, we introduce in our paper a new random model that provides the possibility to generate spatial regions having a desired aggregation.
Yannis Manolopoulos, Enrico Nardelli, Guido Proietti, Michael Vassilakopoulos
CIKM1
1995 Partial Match Retrieval in Two-Headed Disk Systems
Yannis Manolopoulos, Athena Vakali
DEXA1
1995 Functional Requirements for Historical and Interval Extensions to the Relational Model
Nikos A. Lorentzos, Yannis Manolopoulos
Data Knowl. Eng.2
1995 Dynamic Inverted Quadtree: A Structure for Pictorial Databases
Michael Vassilakopoulos, Yannis Manolopoulos
Inf. Syst.2
1995 A random model for analyzing region quadtrees
Michael Vassilakopoulos, Yannis Manolopoulos
Pattern Recognit. Lett.2
1994 Efficient Management of 2-d Interval Relations
Nikos A. Lorentzos, Yannis Manolopoulos
DEXA2
1994 Fast Subsequence Matching in Time-Series Databases
abstract
We present an efficient indexing method to locate 1-dimensional subsequences within a collection of sequences, such that the subsequences match a given (query) pattern within a specified tolerance. The idea is to map each data sequences into a small set of multidimensional rectangles in feature space. Then, these rectangles can be readily indexed using traditional spatial access methods, like the R*-tree [9]. In more detail, we use a sliding window over the data sequence and extract its features; the result is a trail in feature space. We propose an efficient and effective algorithm to divide such trails into sub-trails, which are subsequently represented by their Minimum Bounding Rectangles (MBRs). We also examine queries of varying lengths, and we show how to handle each case efficiently. We implemented our method and carried out experiments on synthetic and real data (stock price movements). We compared the method to sequential scanning, which is the only obvious competitor. The results were excellent: our method accelerated the search time from 3 times up to 100 times.
Christos Faloutsos, M. Ranganathan, Yannis Manolopoulos
SIGMOD Conference3
1994 Binary ranking for the signature file method
Dimitrios Dervos, Panagiotis Linardis, Yannis Manolopoulos
Inf. Softw. Technol.3
1994 Performance of Linear Hashing Schemes for Primary Key Retrieval
Yannis Manolopoulos, Nikos A. Lorentzos
Inf. Syst.1
1994 Analytical Comparison of Two Spatial Data Structures
Michael Vassilakopoulos, Yannis Manolopoulos
Inf. Syst.2
1994 Ranking the Validity of Block Candidacies in Signature Files
Dimitrios Dervos, Yannis Manolopoulos, Panagiotis Linardis
Inf. Sci.2
1994 B-Trees with Lazy Parent Split
Yannis Manolopoulos
Inf. Sci.1
1993 Analytical Results on the Quadtree Storage-Requirements
Michael Vassilakopoulos, Yannis Manolopoulos
CAIP2
1993 Overlapping quadtrees for the representation of similar images
Michael Vassilakopoulos, Yannis Manolopoulos, K. Economou
Image Vis. Comput.2
1992 Reverse Chaining for Answering Temporal Logical Queries (Short Note)
abstract
A possible structure for the past data of a partitioned temporal database is reverse field chaining. Under this technique field versions are chained by descending time. Exact analysis derives the expected number of block accesses when logical queries against a partitioned temporal database with reverse chaining are satisfied. Numerical results are given.
Yannis Manolopoulos
Comput. J.1
1992 File organizations with shared overflow blocks for variable length objects
Yannis Manolopoulos, Stavros Christodoulakis
Inf. Syst.1
1992 Probability distributions for seek time evaluation
Yannis Manolopoulos
Inf. Sci.1
1992 Algorithms for a hashed file with variable-length records
Yannis Manolopoulos, N. Fistas
Inf. Sci.1
1991 Seek Distances in Disks with Two Independent Heads Per Surface
Yannis Manolopoulos, Athena Vakali
Inf. Process. Lett.1
1990 The Optimum Execution Order of Queries in Linear Storage
John G. Kollias, Yannis Manolopoulos, Christos H. Papadimitriou
Inf. Process. Lett.2
1990 Efficient Expressions for Completely and Partly Unsuccessful Batched Search of Tree-Structured Files
abstract
Closed-form, nonrecurrent expressions for the cost of completely and partly unsuccessful batched searching are developed for complete j-ary tree files. These expressions are applied to both the replacement and nonreplacement models of the search queries. The expressions provide more efficient formulas than previously reported for calculating the cost of batched searching. The expressions can also be used to estimate the number of block accesses for hierarchical file structures.>
Sheau-Dong Lang, Yannis Manolopoulos
IEEE Trans. Software Eng.2
1989 Analysis of Overflow Handling for Variable Length Records
Stavros Christodoulakis, Yannis Manolopoulos, Per-Åke Larson
Inf. Syst.2
1989 Performance of a Two-Headed Disk System when Serving Database Queries Under the Scan Policy
abstract
Disk drives with movable two-headed arms are now commercially available. The two heads are separated by a fixed number of cylinders. A major problem for optimizing disk head movement, when answering database requests, is the specification of the optimum number of cylinders separating the two heads. An earlier analytical study assumed a FCFS model and concluded that the optimum separation distance should be equal to 0.44657 of the number of cylinders N of the disk. This paper considers that the SCAN scheduling policy is used in file access, and it applies combinatorial analysis to derive exact formulas for the expected head movement. Furthermore, it is proven that the optimum separation distance is N/2 - 1 (⌈ N /2 - 1⌉ and ⌊ N /2 - 1⌋) if N is even (odd). In addition, a comparison with a single-headed disk system operating under the same scheduling policy shows that if the two heads are optimally spaced, then the mean seek distance is less than one-half of the value obtained with one head. In fact that the SCAN policy is used for many database applications (for example,batching and secondary key retrieval) demonstrates the potential of two-headed disk systems for improving the performance of database systems.
Yannis Manolopoulos, John G. Kollias
ACM Trans. Database Syst.1
1989 Expressions for Completely and Partly Unsuccessful Batched Search of Sequential and Tree-Structured Files
abstract
A number of previous studies derived expressions for batched searching of sequential and tree-structured files on the assumption that all the keys in the batch exist in the file, i.e., all the searches are successful. Formulas for batched searching of sequential and tree-structured files are derived, but the assumption made is that either all or part of the keys in the batch do not exist in the file, i.e., the batched search is completely or partly unsuccessful.>
Yannis Manolopoulos, John G. Kollias
IEEE Trans. Software Eng.1
1987 Batched Interpolation Search
abstract
In a previous study an ordered array of N keys was considered and the problem of locating a batch of M requested keys was investigated by assuming both batched sequential and batched binary searching. This paper introduces the idea of batched interpolation search, and two variations of the method are presented. Comparisons with the two previously defined methods are also made.
Yannis Manolopoulos, John G. Kollias, F. Warren Burton
Comput. J.1
1987 A Model for an ISAM File with Multiple Overflow Chains
abstract
A model for an index sequential file employing multiple overflow chains per bucket is developed. This model is used to analyse the effects of insertions and deletions on the cost of successful and unsuccessful search in terms of block accesses. Numerical results are obtained illustrating the performance. The performance is also compared with that of an ISAM file using only one overflow chain per bucket.
Yannis Manolopoulos, Dimitris Kleftouris, Loukas Petrou
Comput. J.1
1986 Sequential vs. Binary Batched Searching
abstract
This study considers an ordered array of N keys and estimates the required number of key comparisons to locate M requested keys when a binary search is performed to find each key. The problem is analysed for the cases where the M requests (a) are performed individually on a First Come First Served basis, and (b) are treated as a batch. For the second case a break point is established which indicates whether it is preferable to apply binary or sequential search for a batch of M keys.
Yannis Manolopoulos, John G. Kollias, Michael Hatzopoulos
Comput. J.1
1986 Batched Search of Index Sequential Files
Yannis Manolopoulos
Inf. Process. Lett.1