EDBT 2026 Demo / reviewers in the wild / expert
Anastasios Arvanitis
dblp:08/7993
· DBLP profile ↗
10ranked-venue papers
6as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 6 first-authorArtificial intelligence and machine learning · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Query processing and optimization · 80% Data models and query languages · 20% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
preference query |
0.5 | 4 | 2014 | PrefDB: Supporting Preferences as First-Class Citizens in Relational Databases · IEEE Trans. Knowl. Data Eng. 2014 PrefDB: bringing preferences closer to the DBMS · SIGMOD Conference 2012 Towards Preference-aware Relational Databases · ICDE 2012 |
Query processing and optimization
query optimization |
0.4 | 3 | 2014 | PrefDB: Supporting Preferences as First-Class Citizens in Relational Databases · IEEE Trans. Knowl. Data Eng. 2014 PrefDB: bringing preferences closer to the DBMS · SIGMOD Conference 2012 Towards Preference-aware Relational Databases · ICDE 2012 |
Query processing and optimization › preference query
skyline query |
0.1 | 1 | 2010 | Probabilistic contextual skylines · ICDE 2010 |
Data models and query languages › relational model
extended relational model |
0.0 | 1 | 2012 | PrefDB: bringing preferences closer to the DBMS · SIGMOD Conference 2012 |
Methods — techniques the papers use, named apart from their topics
extended relational algebra · 0.3cost model · 0.2index-based algorithm · 0.1block nested loops · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Automated Performance Management for the Big Data Stack
Anastasios Arvanitis, Shivnath Babu, Eric Chu, Alkis Simitsis, Kevin Wilkinson |
CIDR | 1 |
| 2014 | Increasing recommendation accuracy and diversity via social networks hyperbolic embeddingabstractSeveral applications are built around sharing information by leveraging social network connections. For example, in social buying sites like Groupon, a deal is usually forwarded to interested recipients through their social graph. A primary goal is to improve user satisfaction by maximizing the relevance of the shared message to the target audience. In order to suggest more personalized products, one should consider offering not only accurate but also diverse recommendations, since diversification plays an important factor in increasing the users's satisfaction. In this work, we address this problem by proposing a social network hyperbolic embedding that exploits both social connections and user preferences aiming at increasing both the accuracy and the diversity of recommendations. Vasiliki Pouli, John S. Baras, Anastasios Arvanitis |
CCNC | 3 |
| 2014 | Efficient Concept-based Document RankingabstractRecently, there is increased interest in searching and computing the similarity between Electronic Medical Records (EMRs). A unique characteristic of EMRs is that they consist of ontological concepts derived from biomedical ontologies such as UMLS or SNOMED-CT. Medical researchers have found that it is effective to search and find similar EMRs using their concepts, and have proposed so-phisticated similarity measures. However, they have not addressed the performance and scalability challenges to support searching and computing similar EMRs using ontological concepts. In this paper, we formally define these important problems and show that they pose unique algorithmic challenges due to the nature of the search and similarity semantics and the multi-level relationships between the concepts. In particular, the similarity between two EMRs is a function of the minimum semantic distance from each concept of one document to a concept of the other and vice versa. We present an efficient algorithm to compute the similarity between two EMRs. Then, we propose an early-termination algorithm to search for the top-k most relevant EMRs to a set of concepts, and to find the top-k most similar EMRs to a given EMR. We experi-mentally evaluate the performance and scalability of our methods on a large real EMR data set. 1. Anastasios Arvanitis, Matthew T. Wiley, Vagelis Hristidis |
EDBT | 1 |
| 2014 | Multi-Query Diversification in Microblogging PostsabstractEffectively exploring data generated by microblogging services is challenging due to its high volume and production rate. To ad-dress this issue, we propose a solution that helps users effectively consume information from a microblogging stream, by filtering out redundant data. We formalize our approach as a novel optimization problem termed Multi-Query Diversification Problem (MQDP). In MQDP, the input consists of a list of microblogging posts and a set of user queries (e.g. news topics), where each query matches a subset of posts. The objective is to compute the smallest subset of posts that cover all other posts with respect to a “diversity di-mension ” that may represent time or, say, sentiment. Roughly, the solution (cover) has the property that each covered post has nearby posts in the cover that are collectively related to all queries relevant to this covered post. This is distinct from previous single-query diversity problems, as we may have two nearby posts that are related to intersecting but not nested sets of queries, in which case none covers the other. Another key difference is that we do not define diversity in terms of post similarity, since posts are too short for this approach to be meaningful; instead, we focus on finding representative posts for ordered diversity dimensions like time and sentiment, which are critical in microblogging. For example, for time as the diversity dimension, the selected posts will show how certain news events unfolded over time. We prove that MQDP is NP-hard and we propose an exact dy-namic programming algorithm that is feasible for small problem instances. We also propose two approximate algorithms with prov-able approximation bounds, and show how they can be adapted for a streaming setting. Through comprehensive experiments on real data, we show that our algorithms efficiently and effectively gener-ate diverse and representative posts. 1. Shiwen Cheng, Anastasios Arvanitis, Marek Chrobak, Vagelis Hristidis |
EDBT | 2 |
| 2014 | PrefDB: Supporting Preferences as First-Class Citizens in Relational DatabasesabstractIn this paper, we argue that preference-aware query processing needs to be pushed closer to the DBMS. We introduce a preference-aware relational data model that extends database tuples with preferences and an extended algebra that captures the essence of processing queries with preferences. Based on a set of algebraic properties and a cost model that we propose, we provide several query optimization strategies for extended query plans. Further, we describe a query execution algorithm that blends preference evaluation with query execution, while making effective use of the native query engine. We have implemented our framework and methods in a prototype system, PrefDB. PrefDB allows transparent and efficient evaluation of preferential queries on top of a relational DBMS. Our extensive experimental evaluation on two real-world datasets demonstrates the feasibility and advantages of our framework. Anastasios Arvanitis, Georgia Koutrika |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | How fresh do you want your search results?abstractResearchers have recognized the importance of utilizing temporal features for improving the performance of information retrieval systems. Specifically, the timeliness of a web document can be a significant factor for determining whether it is relevant for a search query. Previous works have proposed time-aware retrieval models with particular focus on news queries, where recent web documents related with a real-world event are generally preferable. These queries typically exhibit bursts in the volume of published documents or submitted queries. However, no work has studied the role of time in queries such as "credit card overdraft fees" that have no major spikes in either document or query volumes over time, yet they still favor more recently published documents. In this work, we focus on this class of queries that we refer to as "timely queries". We show that the change in the terms distribution of results of timely queries over time is strongly correlated with the users' perception of time sensitivity. Based on this observation, we propose a method to estimate the query timeliness requirements and we propose principled ways to incorporate document freshness into the ranking model. Our study shows that our method yields a more accurate estimation of timeliness compared to volume-based approaches. We experimentally compare our ranking strategy with other time-sensitive and non time-sensitive ranking algorithms and we show that it improves the results' retrieval quality for timely queries. Shiwen Cheng, Anastasios Arvanitis, Vagelis Hristidis |
CIKM | 2 |
| 2012 | Efficient influence-based processing of market research queriesabstractThe rapid growth of social web has contributed vast amounts of user preference data. Analyzing this data and its relationships with products could have several practical applications, such as personalized advertising, market segmentation, product feature promotion etc. In this work we develop novel algorithms for efficiently processing two important classes of queries involving user preferences, i.e. potential customers identification and product positioning. With regards to the first problem, we formulate product attractiveness based on the notion of reverse skyline queries. We then present a new algorithm, termed as RSA, that significantly reduces the I/O cost, as well as the computation cost, when compared to the state-of-the-art reverse skyline algorithm, while at the same time being able to quickly report the first results. Several real-world applications require processing of a large number of queries, in order to identify the product characteristics that maximize the number of potential customers. Motivated by this problem, we also develop a batched extension of our RSA algorithm that significantly improves upon processing multiple queries individually, by grouping contiguous candidates, exploiting I/O commonalities and enabling shared processing. Our experimental study using both real and synthetic data sets demonstrates the superiority of our proposed algorithms for the studied classes of queries. Anastasios Arvanitis, Antonios Deligiannakis, Yannis Vassiliou |
CIKM | 1 |
| 2012 | Towards Preference-aware Relational DatabasesabstractIn implementing preference-aware query processing, a straightforward option is to build a plug-in on top of the database engine. However, treating the DBMS as a black box affects both the expressivity and performance of queries with preferences. In this paper, we argue that preference-aware query processing needs to be pushed closer to the DBMS. We present a preference-aware relational data model that extends database tuples with preferences and an extended algebra that captures the essence of processing queries with preferences. A key novelty of our preference model itself is that it defines a preference in three dimensions showing the tuples affected, their preference scores and the credibility of the preference. Our query processing strategies push preference evaluation inside the query plan and leverage its algebraic properties for finer-grained query optimization. We experimentally evaluate the proposed strategies. Finally, we compare our framework to a pure plug-in implementation and we show its feasibility and advantages. Anastasios Arvanitis, Georgia Koutrika |
ICDE | 1 |
| 2012 | PrefDB: bringing preferences closer to the DBMSabstractIn this demonstration we present a preference-aware relational query answering system, termed PrefDB. The key novelty of PrefDB is the use of an extended relational data model and algebra that allow expressing different flavors of preferential queries. Furthermore, unlike existing approaches that either treat the DBMS as a black box or require modifications of the database core, PrefDB's hybrid implementation enables operator-level query optimizations without being obtrusive to the database engine. We showcase the flexibility and efficiency of PrefDB using PrefDBAdmin, a graphical tool that we have built aiming at assisting application designers in the task of building, testing and tuning queries with preferences. Anastasios Arvanitis, Georgia Koutrika |
SIGMOD Conference | 1 |
| 2010 | Probabilistic contextual skylinesabstractThe skyline query returns the most interesting tuples according to a set of explicitly defined preferences among attribute values. This work relaxes this requirement, and allows users to pose meaningful skyline queries without stating their choices. To compensate for missing knowledge, we first determine a set of uncertain preferences based on user profiles, i.e., information collected for previous contexts. Then, we define a probabilistic contextual skyline query (p-CSQ) that returns the tuples which are interesting with high probability. We emphasize that, unlike past work, uncertainty lies within the query and not the data, i.e., it is in the relationships among tuples rather than in their attribute values. Furthermore, due to the nature of this uncertainty, popular skyline methods, which rely on a particular tuple visit order, do not apply for p-CSQs. Therefore, we present novel non-indexed and index-based algorithms for answering p-CSQs. Our experimental evaluation concludes that the proposed techniques are significantly more efficient compared to a standard block nested loops approach. Dimitris Sacharidis, Anastasios Arvanitis, Timos K. Sellis |
ICDE | 2 |