EDBT 2026 Demo / reviewers in the wild / expert
Georgia Koutrika
dblp:k/GeorgiaKoutrika
· DBLP profile ↗
in reviewer pool
← Back
93ranked-venue papers in the field
39as first author
36since 2021 · last 2026
0000-0002-7377-0116ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 82 (35 first)Information Retrieval & Web Search · 5 (4 first)Data Mining & Knowledge Discovery · 4Knowledge Engineering, Semantic Web & Information Systems · 1Business Process & Enterprise Data · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Drives Learned Optimizer Performance? A Systematic Evaluation
Kostas Mparmparousis, Christos Tsapelas, Georgia Koutrika |
EDBT | 3 |
| 2026 | Query-Driven Data Exploration with Heterogeneous Treatment Effects
Antonis Mandamadiotis, Sihem Amer-Yahia, Georgia Koutrika |
ICDE | 3 |
| 2026 | Editorial: Special Issue for Selected Papers of VLDB 2023
Georgia Koutrika, Jun Yang 0001 |
VLDB J. | 1 |
| 2025 | QueryER: A Framework for Fast Analysis-Aware Deduplication over Dirty Data
George Alexiou, George Papastefanatos, Vassilis Stamatopoulos, Georgia Koutrika, Nectarios Koziris |
EDBT | 4 |
| 2025 | Towards Reliable Conversational Data Analytics
Sihem Amer-Yahia, Jasmina Bogojeska, Roberta Facchinetti, Valeria Franceschi, Aristides Gionis, Katja Hose, Georgia Koutrika, Roger D. Kouyos, Matteo Lissandrini, Silviu Maniu, Katsiaryna Mirylenka, Davide Mottin, Themis Palpanas, Mattia Rigotti, Yannis Velegrakis |
EDBT | 7 |
| 2025 | Analysis of Text-to-SQL Benchmarks: Limitations, Challenges and Opportunities
Anna Mitsopoulou, Georgia Koutrika |
EDBT | 2 |
| 2024 | AI and Human in Data Analytics: Who leads this Dance?
Georgia Koutrika |
DOLAP | 1 |
| 2024 | QPSeeker: An Efficient Neural Planner combining both data and queries through Variational Inference
Christos Tsapelas, Georgia Koutrika |
EDBT | 2 |
| 2024 | FreqyWM: Frequency Watermarking for the New Data EconomyabstractWe present a novel technique for modulating the appearance frequency of a few tokens within a dataset for encoding an invisible watermark that can be used to protect ownership rights upon data. We develop optimal as well as fast heuristic algorithms for creating and verifying such watermarks. We also demonstrate the robustness of our technique against various attacks and derive analytical bounds for the false positive probability of erroneously “detecting” a watermark on a dataset that does not carry it. Our technique is applicable to both single dimensional and multidimensional datasets, is independent of token type, allows for a fine control of the introduced distortion, and can be used in a variety of use cases that involve buying and selling data in contemporary data marketplaces. Devris Isler, Elisa Cabana, Álvaro García-Recuero, Georgia Koutrika, Nikolaos Laoutaris |
ICDE | 4 |
| 2024 | Guided SQL-Based Data Exploration with User FeedbackabstractThe exploration of large, real-world databases poses major challenges to users due to their volume and complexity. SQL is the preferred language for data exploration. However, the process of iteratively refining SQL queries is tedious and time consuming. We formulate the automation of personalized SQL-based data exploration as the problem of suggesting the most relevant query and accounting for user feedback at each step. We develop an end-to-end solution and a system to assist users in exploring different components of a complex database. We instantiate our solution using Multi-Armed Bandits, a category of algorithms that are suitable for interactive online learning by balancing exploration with exploitation. We design a lightweight algorithm to personalize stepwise SQL recommendations that efficiently discovers the current user preferences in coordination with that user's feedback and what other users prefer. We run extensive experiments that demonstrate the utility of our approach for large-scale data exploration. Antonis Mandamadiotis, Georgia Koutrika, Sihem Amer-Yahia |
ICDE | 2 |
| 2024 | Natural Language Data Interfaces: A Data Access Odyssey (Invited Talk)abstractBack in 1970’s, E. F. Codd worked on a prototype of a natural language question and answer application that would sit on top of a relational database system. Soon, natural language interfaces for databases (NLIDBs) became the holy grail for the database community. Different approaches have been proposed from the database, machine learning and NLP communities. Interest in the topic has had its peaks and valleys. After a long and adventurous journey of almost 50 years, there is a rekindled interest in NLIDBs in recent years, fueled by the need for democratizing data access and by the recent advances in deep learning and natural language processing in particular. There is a surge of works on natural language interfaces for databases using neural translation, and suddenly it becomes hard to keep up with advancements in the field. Are we close to finding the holy grail of data access? What are the lurking challenges that we need to surpass and what research opportunities arise? Finally, what is the role of the database community? Georgia Koutrika |
ICDT | 1 |
| 2023 | Data Democratisation with Deep Learning: The Anatomy of a Natural Language Data InterfaceabstractIn the age of the Digital Revolution, almost all human activities, from industrial and business operations to medical and academic research, are reliant on the constant integration and utilisation of ever-increasing volumes of data. However, the explosive volume and complexity of data makes data querying and exploration challenging even for experts, and makes the need to democratise the access to data, even for non-technical users, all the more evident. It is time to lift all technical barriers, by empowering users to access relational databases through conversation. We consider 3 main research areas that a natural language data interface is based on: Text-to-SQL, SQL-to-Text, and Data-to-Text. The purpose of this tutorial is a deep dive into these areas, covering state-of-the-art techniques and models, and explaining how the progress in the deep learning field has led to impressive advancements. We will present benchmarks that sparked research and competition, and discuss open problems and research opportunities with one of the most important challenges being the integration of these 3 research areas into one conversational system. George Katsogiannis-Meimarakis, Mike Xydas, Georgia Koutrika |
WSDM | 3 |
| 2023 | Natural Language Interfaces for Databases with Deep LearningabstractIn the age of the Digital Revolution, almost all human activities, from industrial and business operations to medical and academic research, are reliant on the constant integration and utilisation of ever-increasing volumes of data. However, the explosive volume and complexity of data makes data querying and exploration challenging even for experts, and makes the need to democratise the access to data, even for non-technical users, all the more evident. It is time to lift all technical barriers, by empowering users to access relational databases through conversation. We consider 3 main research areas that a natural language data interface is based on: Text-to-SQL, SQL-to-Text, and Data-to-Text. The purpose of this tutorial is a deep dive into these areas, covering state-of-the-art techniques and models, and explaining how the progress in the deep learning field has led to impressive advancements. We will present benchmarks that sparked research and competition, and discuss open problems and research opportunities with one of the most important challenges being the integration of these 3 research areas into one conversational system. George Katsogiannis-Meimarakis, Mike Xydas, Georgia Koutrika |
Proc. VLDB Endow. | 3 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001, Manos Athanassoulis, Kostas Stefanidis, Ju Fan, Abdul Quamar, Yuanyan Tian, Alekh Jindal, Carsten Binnig, Jennie Rogers, Senjuti Basu Roy, Steven Euijong Whang, Matthias Boehm 0001, Aaron J. Elmore, Vasilis Efthymiou, Xiao Hu 0005, Xiaofang Zhou 0001, Alan D. Fekete |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsabstractNatural Language to SQL systems (NL-to-SQL) have recently shown improved accuracy (exceeding 80%) for natural language to SQL query translation due to the emergence of transformer-based language models, and the popularity of the Spider benchmark. However, Spider mainly contains simple databases with few tables, columns, and entries, which do not reflect a realistic setting. Moreover, complex real-world databases with domain-specific content have little to no training data available in the form of NL/SQL-pairs leading to poor performance of existing NL-to-SQL systems. In this paper, we introduce ScienceBenchmark , a new complex NL-to-SQL benchmark for three real-world, highly domain-specific databases. For this new benchmark, SQL experts and domain experts created high-quality NL/SQL-pairs for each domain. To garner more data, we extended the small amount of human-generated data with synthetic data generated using GPT-3. We show that our benchmark is highly challenging, as the top performing systems on Spider achieve a very low performance on our benchmark. Thus, the challenge is many-fold: creating NL-to-SQL systems for highly complex domains with a small amount of hand-made training data augmented with synthetic data. To our knowledge, ScienceBenchmark is the first NL-to-SQL benchmark designed with complex real-world scientific databases, containing challenging training and test data carefully validated by domain experts. Yi Zhang 0142, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten, Georgia Koutrika, Kurt Stockinger |
Proc. VLDB Endow. | 5 |
| 2023 | A survey on deep learning approaches for text-to-SQLabstractAbstract To bridge the gap between users and data, numerous text-to-SQL systems have been developed that allow users to pose natural language questions over relational databases. Recently, novel text-to-SQL systems are adopting deep learning methods with very promising results. At the same time, several challenges remain open making this area an active and flourishing field of research and development. To make real progress in building text-to-SQL systems, we need to de-mystify what has been done, understand how and when each approach can be used, and, finally, identify the research challenges ahead of us. The purpose of this survey is to present a detailed taxonomy of neural text-to-SQL systems that will enable a deeper study of all the parts of such a system. This taxonomy will allow us to make a better comparison between different approaches, as well as highlight specific challenges in each step of the process, thus enabling researchers to better strategise their quest towards the “holy grail” of database accessibility. George Katsogiannis-Meimarakis, Georgia Koutrika |
VLDB J. | 2 |
| 2022 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2022 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2022 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2022 | Front Matter
Georgia Koutrika, Jun Yang 0001 |
Proc. VLDB Endow. | 1 |
| 2022 | Fairness in rankings and recommendations: an overviewabstractAbstract We increasingly depend on a variety of data-driven algorithmic systems to assist us in many aspects of life. Search engines and recommender systems among others are used as sources of information and to help us in making all sort of decisions from selecting restaurants and books, to choosing friends and careers. This has given rise to important concerns regarding the fairness of such systems. In this work, we aim at presenting a toolkit of definitions, models and methods used for ensuring fairness in rankings and recommendations. Our objectives are threefold: (a) to provide a solid framework on a novel, quickly evolving and impactful domain, (b) to present related methods and put them into perspective and (c) to highlight open challenges and research paths for future work. Evaggelia Pitoura, Kostas Stefanidis, Georgia Koutrika |
VLDB J. | 3 |
| 2021 | Deep Learning Approaches for Text-to-SQL Systems
George Katsogiannis-Meimarakis, Georgia Koutrika |
EDBT | 2 |
| 2021 | Fairness in Rankings and Recommenders: Models, Methods and Research DirectionsabstractWe increasingly depend on a variety of data-driven algorithmic systems to assist us in many aspects of life. Search engines and recommendation systems amongst others are used as sources of information and to help us in making all sort of decisions from selecting restaurants and books, to choosing friends and careers. This has given rise to important concerns regarding the fairness of such systems. This tutorial aims at presenting a toolkit of definitions, models and methods used for ensuring fairness in rankings and recommendations. Our objectives are three-fold: (a) to provide a solid framework on a novel, quickly evolving, and impactful domain, (b) to present related methods and put them into perspective, and (c) to highlight challenges and research paths for researchers and practitioners that work in data management and applications. Evaggelia Pitoura, Kostas Stefanidis, Georgia Koutrika |
ICDE | 3 |
| 2021 | Fairness-aware Methods in Rankings and RecommendersabstractWe increasingly depend on a variety of data-driven algorithmic systems to assist us in many aspects of life. Search engines and recommender systems amongst others are used as sources of information and to help us in making all sort of decisions from selecting restaurants and books, to choosing friends and careers. This has given rise to important concerns regarding the fairness of such systems. In this tutorial, we aim at presenting a toolkit of methods used for ensuring fairness in rankings and recommendations. Our objectives are two-fold: (a) to present related methods of this novel, quickly evolving and impactful domain, and put them into perspective, and (b) to highlight open challenges and research paths for future work. Evaggelia Pitoura, Kostas Stefanidis, Georgia Koutrika |
MDM | 3 |
| 2021 | An In-Depth Benchmarking of Text-to-SQL SystemsabstractText-to-SQL systems allow users to explore relational databases by posing free-form queries, alleviating the need for using structured languages, such as SQL. Although numerous systems have been developed so far, existing system evaluations lack in rigour. In this work, we build a text-to-SQL benchmark that covers different classes of queries, and we evaluate the effectiveness of several systems in the field. To evaluate system efficiency, we measure execution time and resource consumption for the different query classes. Our comprehensive evaluation aims at filling in a big gap in understanding the capabilities and boundaries of existing systems and it reveals several open challenges. Orest Gkini, Theofilos Belmpas, Georgia Koutrika, Yannis E. Ioannidis |
SIGMOD Conference | 3 |
| 2021 | PyExplore: Query Recommendations for Data Exploration without Query LogsabstractHelping users explore data becomes increasingly more important as databases get larger and more complex. In this demo, we present PyExplore, a data exploration tool aimed at helping end users formulate queries over new datasets. PyExplore takes as input an initial query from the user along with some parameters and provides interesting queries by leveraging data correlations and diversity. Apostolos Glenis, Georgia Koutrika |
SIGMOD Conference | 2 |
| 2021 | A Deep Dive into Deep Learning Approaches for Text-to-SQL SystemsabstractData is a prevalent part of every business and scientific domain,but its explosive volume and increasing complexity make data querying challenging even for experts. For this reason, numerous text-to-SQL systems have been developed that enable querying relational databases using natural language. The recent advances on deep neural networks along with the creation of two large datasets specifically made for training text-to-SQL systems, have paved the path for a novel and very promising research area. The purpose of this tutorial is a deep dive into this area, covering state-of-the-art techniques for natural language representation in neural networks,benchmarks that sparked research and competition, recent text-to-SQL systems using deep learning techniques, as well as open problems and research opportunities. George Katsogiannis-Meimarakis, Georgia Koutrika |
SIGMOD Conference | 2 |
| 2021 | DatAgent: The Imminent Age of Intelligent Data AssistantsabstractIn this demonstration, we present DatAgent , an intelligent data assistant system that allows users to ask queries in natural language, and can respond in natural language as well. Moreover, the system actively guides the user using different types of recommendations and hints, and learns from user actions. We will demonstrate different exploration scenarios that show how the system and the user engage in a human-like interaction inspired by the interaction paradigm of chatbots and virtual assistants. Antonis Mandamadiotis, Georgia Koutrika, Stavroula Eleftherakis, Apostolos Glenis, Dimitrios Skoutas 0001, Yannis Stavrakas |
Proc. VLDB Endow. | 2 |
| 2020 | Fairness in Rankings and RecommendersabstractPeer reviewed Evaggelia Pitoura, Georgia Koutrika, Kostas Stefanidis |
EDBT | 2 |
| 2020 | Recommendations as Graph ExplorationsabstractWe argue that most recommendation approaches can be abstracted as a graph exploration problem. In particular, we describe a graph-theoretic framework with two primary parts: (a) a recommendation graph, modeling all the elements of an (application) domain from a recommendation perspective, including the subjects and objects of recommendations as well as the relationships between them; (b) a set of path operations, inferring new edges, i.e., implicit or unknown relationships, by traversing and combining paths on the graph. The resulting path algebra model provides an abstraction and a common foundation that is beneficial to three aspects of recommendations: (a) expressive power - expression and subsequent use of several significantly different, existing but also novel recommendation approaches is reduced to parameterizing a unique model; (b) usability - by capturing part of the recommendation mechanisms in the underlying path algebra semantics, specification of recommendation approaches becomes easier and less tedious; (c) processing speed - implementing recommender systems on top of graph engines opens up the door for several optimizations that speed up execution. We demonstrate the above benefits by expressing several categories of recommendation approaches in the path algebra model and benchmarking some of them in a recommender system implemented on top of Neo4J, a widely used graph system. Marialena Kyriakidi, Georgia Koutrika, Yannis E. Ioannidis |
RecSys | 2 |
| 2020 | Analysis of Database Search Systems with THORabstractNumerous search systems have been implemented that allow users to pose unstructured queries over databases without the need to use a query language, such as SQL. Unfortunately, the landscape of efforts is fragmented with no clear sight of which system is best, and what open challenges we should pursue in our research. To help towards this direction, we present THOR that makes 4 important contributions: a query benchmark, a framework for comparing different systems, several search system implementations, and a highly interactive tool for comparing different search systems. Theofilos Belmpas, Orest Gkini, Georgia Koutrika |
SIGMOD Conference | 3 |
| 2019 | The Evolution of Search Over Structured Data: From SQL to Natural Language
Georgia Koutrika |
DATA | 1 |
| 2019 | Figuring out the User in a Few Steps: Bayesian Multifidelity Active Search with CokrigingabstractCan a system discover what a user wants without the user explicitly issuing a query? A recommender system proposes items of potential interest based on past user history. On the other hand, active search incites, and learns from, user feedback, in order to recommend items that meet a user's current tacit interests, hence promises to offer up-to-date recommendations going beyond those of a recommender system. Yet extant active search methods require an overwhelming amount of user input, relying solely on such input for each item they pick. In this paper, we propose MF-ASC, a novel active search mechanism that performs well with minimal user input. MF-ASC combines cheap, low-fidelity evaluations in the style of a recommender system with the user's high-fidelity input, using Gaussian process regression with multiple target variables (cokriging). To our knowledge, this is the first application of cokriging to active search. Our empirical study with synthetic and real-world data shows that MF-ASC outperforms the state of the art in terms of result relevance within a budget of interactions. Nikita Klyuchnikov, Davide Mottin, Georgia Koutrika, Emmanuel Müller, Panagiotis Karras |
KDD | 3 |
| 2018 | Recent Advances in Recommender Systems: Matrices, Bandits, and Blenders
Georgia Koutrika |
EDBT | 1 |
| 2018 | Modeling and Exploiting Goal and Action Associations for Recommendations
Dimitra Papadimitriou, Yannis Velegrakis, Georgia Koutrika |
EDBT | 3 |
| 2018 | Finding Related Forum Posts through Content Similarity over Intention-Based Segmentation (Extended Abstract)abstractWe study the problem of finding related forum posts to a post at hand. We developed a multi-segment matching technique that considers posts as a set of segments each one written with a different goal in its author mind and computes the relatedness between two posts based on the similarity of their respective segments that are intended for the same goal. The questions are how our method identifies such segments, how it figures out for what each segment is intended and how it exploits this information to rank the posts. We experimentally illustrate the effectiveness and efficiency of our segmentation method and overall approach of finding related forum posts. Dimitra Papadimitriou, Georgia Koutrika, Yannis Velegrakis, John Mylopoulos |
ICDE | 2 |
| 2018 | Modern Recommender Systems: from Computing Matrices to Thinking with NeuronsabstractStarting with the Netflix Prize, which fueled much recent progress in the field of collaborative filtering, recent years have witnessed rapid development of new recommendation algorithms and increasingly more complex systems, which greatly differ from their early content-based and collaborative filtering systems. Modern recommender systems leverage several novel algorithmic approaches: from matrix factorization methods and multi-armed bandits to deep neural networks. In this tutorial, we will cover recent algorithmic advances in recommender systems, highlight their capabilities, and their impact. We will give many examples of industrial-scale recommender systems that define the future of the recommender systems area. We will discuss related evaluation issues, and outline future research directions. The ultimate goal of the tutorial is to encourage the application of novel recommendation approaches to solve problems that go beyond user consumption and to further promote research in the intersection of recommender systems and databases. Georgia Koutrika |
SIGMOD Conference | 1 |
| 2017 | GnosisMiner: Reading Order Recommendations over Document Collections
Georgia Koutrika, Alkis Simitsis, Yannis E. Ioannidis |
EDBT | 1 |
| 2017 | Finding Related Forum Posts through Content Similarity over Intention-Based SegmentationabstractWe study the problem of finding related forum posts to a post at hand. In contrast to traditional approaches for finding related documents that perform content comparisons across the content of the posts as a whole, we consider each post as a set of segments, each written with a different goal in mind. We advocate that the relatedness between two posts should be based on the similarity of their respective segments that are intended for the same goal, i.e., are conveying the same intention. This means that it is possible for the same terms to weigh differently in the relatedness score depending on the intention of the segment in which they are found. We have developed a segmentation method that by monitoring a number of text features can identify the parts of a post where significant jumps occur indicating a point where a segmentation should take place. The generated segments of all the posts are clustered to form intention clusters and then similarities across the posts are calculated through similarities across segments with the same intention. We experimentally illustrate the effectiveness and efficiency of our segmentation method and our overall approach of finding related forum posts. Dimitra Papadimitriou, Georgia Koutrika, Yannis Velegrakis, John Mylopoulos |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | A Study of Web Print: What People Print in the Digital EraabstractThis article analyzes a proprietary log of printed web pages and aims at answering questions regarding the content people print (what), the reasons they print (why), as well as attributes of their print profile (who). We present a classification of pages printed based on their print intent and we describe our methodology for processing the print dataset used in this study. In our analysis, we study the web sites, topics, and print intent of the pages printed along the following aspects: popularity, trends, activity, user diversity, and consistency. We present several findings that reveal interesting insights into printing. We analyze our findings and discuss their impact and directions for future work. Georgia Koutrika, Qian Lin 0001 |
ACM Trans. Web | 1 |
| 2016 | The Goal Behind the Action: Toward Goal-Aware Systems and ApplicationsabstractHuman activity is almost always intentional, be it in a physical context or as part of an interaction with a computer system. By understanding why user-generated events are happening and what purposes they serve, a system can offer a significantly improved and more engaging experience. However, goals cannot be easily captured. Analyzing user actions such as clicks and purchases can reveal patterns and behaviors, but understanding the goals behind these actions is a different and challenging issue. Our work presents a unified, multidisciplinary viewpoint for goal management that covers many different cases where goals can be used and techniques with which they can be exploited. Our purpose is to provide a common reference point to the concepts and challenging tasks that need to be formally defined when someone wants to approach a data analysis problem from a goal-oriented point of view. This work also serves as a springboard to discuss several open challenges and opportunities for goal-oriented approaches in data management, analysis, and sharing systems and applications. Dimitra Papadimitriou, Georgia Koutrika, John Mylopoulos, Yannis Velegrakis |
ACM Trans. Database Syst. | 2 |
| 2015 | Generating reading orders over document collectionsabstractGiven a document collection, existing systems allow users to browse the collection or perform searches that return lists of documents ranked based on their relevance to the user query. While these approaches work fine when a user is trying to locate specific documents, they are insufficient when users need to access the pertinent documents in some logical order, for example for learning or editorial purposes. We present a system that automatically organizes a collection of documents in a tree from general to more specific documents, and allows a user to choose a reading sequence over the documents. This a novel way to content consumption that departs from the typical ranked lists of documents based on their relevance to a user query and from static navigational interfaces. We present a set of algorithms that solve the problem and we evaluate their performance as well as the reading trees generated. Georgia Koutrika, Steven J. Simske |
ICDE | 1 |
| 2015 | LearningAssistant: A novel learning resource recommendation systemabstractReading online content for educational, learning, training or recreational purposes has become a very popular activity. While reading, people may have difficulty understanding a passage or wish to learn more about the topics covered by it, hence they may naturally seek additional or supplementary resources for the particular passage. These resources should be close to the passage both in terms of the subject matter and the reading level. However, using a search engine to find such resources interrupts the reading flow. It is also an inefficient, trial-and-error process because existing web search and recommendation systems do not support large queries, they do not understand semantic topics, and they do not take into account the reading level of the original document a person is reading. In this demo, we present LearningAssistant, a novel system that enables online reading material to be smoothly enriched with additional resources that can supplement or explain any passage from the original material for a reader on demand. The system facilitates the learning process by recommending learning resources (documents, videos, etc) for selected text passages of any length. The recommended resources are ranked based on two criteria (a) how they match the different topics covered within the selected passage, and (b) the reading level of the original text where the selected passage comes from. User feedback from students who use our system in two real pilots, one with a high school and one with a university, for their courses suggest that our system is promising and effective. Georgia Koutrika, Shanchan Wu |
ICDE | 2 |
| 2015 | Goals in Social Media, information retrieval and intelligent agentsabstractThis tutorial provides a comprehensive and cohesive overview of goal modeling and recognition approaches by the Information Retrieval, the Artificial Intelligence and the Social Media communities. We will examine how these fields restrict the domain of study and how they capture notions easily perceived by humans' intuition but difficult to be formally defined and handled algorithmically. It is the purpose of this tutorial to provide a solid framework for placing existing work into perspective and highlight critical open challenges that will act as a springboard for researchers and practitioners in database systems, social data, and the Web, as well as developers of web-based, database-driven, and social applications, to work towards more user-centric systems and applications. Dimitra Papadimitriou, Yannis Velegrakis, Georgia Koutrika, John Mylopoulos |
ICDE | 3 |
| 2015 | Schema-agnostic vs Schema-based Configurations for Blocking Methods on Homogeneous DataabstractEntity Resolution constitutes a core task for data integration that, due to its quadratic complexity, typically scales to large datasets through blocking methods. These can be configured in two ways. The schema-based configuration relies on schema information in order to select signatures of high distinctiveness and low noise, while the schema-agnostic one treats every token from all attribute values as a signature. The latter approach has significant potential, as it requires no fine-tuning by human experts and it applies to heterogeneous data. Yet, there is no systematic study on its relative performance with respect to the schema-based configuration. This work covers this gap by comparing analytically the two configurations in terms of effectiveness, time efficiency and scalability. We apply them to 9 established blocking methods and to 11 benchmarks of structured data. We provide valuable insights into the internal functionality of the blocking methods with the help of a novel taxonomy. Our studies reveal that the schema-agnostic configuration offers unsupervised and robust definition of blocking keys under versatile settings, trading a higher computational cost for a consistently higher recall than the schema-based one. It also enables the use of state-of-the-art blocking methods without schema knowledge. George Papadakis 0001, George Alexiou, George Papastefanatos, Georgia Koutrika |
Proc. VLDB Endow. | 4 |
| 2014 | Learn2Learn: A Visual Educational System for Study Planning
Jishang Wei, Georgia Koutrika, Shanchan Wu |
EDBT | 2 |
| 2014 | Supervised Meta-blockingabstractEntity Resolution matches mentions of the same entity. Being an expensive task for large data, its performance can be improved by blocking, i.e., grouping similar entities and comparing only entities in the same group. Blocking improves the run-time of Entity Resolution, but it still involves unnecessary comparisons that limit its performance. Meta-blocking is the process of restructuring a block collection in order to prune such comparisons. Existing unsupervised meta-blocking methods use simple pruning rules, which offer a rather coarse-grained filtering technique that can be conservative (i.e., keeping too many unnecessary comparisons) or aggressive (i.e., pruning good comparisons). In this work, we introduce supervised meta-blocking techniques that learn classification models for distinguishing promising comparisons. For this task, we propose a small set of generic features that combine a low extraction cost with high discriminatory power. We show that supervised meta-blocking can achieve high performance with small training sets that can be manually created. We analytically compare our supervised approaches with baseline and competitor methods over 10 large-scale datasets, both real and synthetic. George Papadakis 0001, George Papastefanatos, Georgia Koutrika |
Proc. VLDB Endow. | 3 |
| 2014 | PrefDB: Supporting Preferences as First-Class Citizens in Relational DatabasesabstractIn this paper, we argue that preference-aware query processing needs to be pushed closer to the DBMS. We introduce a preference-aware relational data model that extends database tuples with preferences and an extended algebra that captures the essence of processing queries with preferences. Based on a set of algebraic properties and a cost model that we propose, we provide several query optimization strategies for extended query plans. Further, we describe a query execution algorithm that blends preference evaluation with query execution, while making effective use of the native query engine. We have implemented our framework and methods in a prototype system, PrefDB. PrefDB allows transparent and efficient evaluation of preferential queries on top of a relational DBMS. Our extensive experimental evaluation on two real-world datasets demonstrates the feasibility and advantages of our framework. Anastasios Arvanitis, Georgia Koutrika |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Meta-Blocking: Taking Entity Resolutionto the Next LevelabstractEntity Resolution is an inherently quadratic task that typically scales to large data collections through blocking. In the context of highly heterogeneous information spaces, blocking methods rely on redundancy in order to ensure high effectiveness at the cost of lower efficiency (i.e., more comparisons). This effect is partially ameliorated by coarse-grained block processing techniques that discard entire blocks either a-priori or during the resolution process. In this paper, we introduce meta-blocking as a generic procedure that intervenes between the creation and the processing of blocks, transforming an initial set of blocks into a new one with substantially fewer comparisons and equally high effectiveness. In essence, meta-blocking aims at extracting the most similar pairs of entities by leveraging the information that is encapsulated in the block-to-entity relationships. To this end, it first builds an abstract graph representation of the original set of blocks, with the nodes corresponding to entity profiles and the edges connecting the co-occurring ones. During the creation of this structure all redundant comparisons are discarded, while the superfluous ones can be removed by pruning of the edges with the lowest weight. We analytically examine both procedures, proposing a multitude of edge weighting schemes, graph pruning algorithms as well as pruning criteria. Our approaches are schema-agnostic, thus accommodating any type of blocks. We evaluate their performance through a thorough experimental study over three large-scale, real-world data sets, with the outcomes verifying significant efficiency enhancements at a negligible cost in effectiveness. George Papadakis 0001, Georgia Koutrika, Themis Palpanas, Wolfgang Nejdl |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Mirror mirror on the wall, which query's fairest of them all?
Georgia Koutrika, Alkis Simitsis |
CIDR | 1 |
| 2013 | HIL: a high-level scripting language for entity integrationabstractWe introduce HIL, a high-level scripting language for entity resolution and integration. HIL aims at providing the core logic for complex data processing flows that aggregate facts from large collections of structured or unstructured data into clean, unified entities. Such flows typically include many stages of processing that start from the outcome of information extraction and continue with entity resolution, mapping and fusion. A HIL program captures the overall integration flow through a combination of SQL-like rules that link, map, fuse and aggregate entities. A salient feature of HIL is the use of logical indexes in its data model to facilitate the modular construction and aggregation of complex entities. Another feature is the presence of a flexible, open type system that allows HIL to handle input data that is irregular, sparse or partially known. Mauricio A. Hernández, Georgia Koutrika, Rajasekar Krishnamurthy, Lucian Popa 0001, Ryan Wisnesky |
EDBT | 2 |
| 2013 | The farm: where pig scripts are bred and raisedabstractEven though scripting languages like Pig allow for simpler coding, performing analytics over Big Data using Map-Reduce engines remains challenging. To further assist developers, and support novice users, we offer "The Farm", a catalog of scriptable services supporting creation, discovery, composition, and optimized execution. Each Pig script added to The Farm becomes an executable service, with inputs and outputs defined by relation schemas. Those services are discoverable using natural language search, and composable using a drag-and-drop interface. To support efficient execution, composed services are automatically merged to a single executable script, which can then be run by a growing selection of platform-specific optimizers and interpreters. Craig Sayers, Alkis Simitsis, Georgia Koutrika, Alejandro Guerrero Gonzalez, David Tamez Cantu, Meichun Hsu |
SIGMOD Conference | 3 |
| 2013 | User Analytics with UbeOne: Insights into Web PrintingabstractAs web and mobile applications become more sensitive to the user context, there is a shift from purely off-line processing of user actions (log analysis) to real-time user analytics that can generate information about the user context to be instantly leveraged by the application. Ubeone is a system that enables both real-time and aggregate analytics from user data. The system is designed as a set of lightweight, composeable mechanisms that can progressively and collectively analyze a user action, such as pinning, saving or printing a web page. We will demonstrate the system capabilities on analyzing a live feed of URLs printed through a proprietary, web browser plug-in. This is in fact the first analysis of web printing activity. We will also give a taste of how the system can enable instant recommendations based on the user context. Georgia Koutrika, Qian Lin 0001, Jerry Liu |
Proc. VLDB Endow. | 1 |
| 2012 | Towards Preference-aware Relational DatabasesabstractIn implementing preference-aware query processing, a straightforward option is to build a plug-in on top of the database engine. However, treating the DBMS as a black box affects both the expressivity and performance of queries with preferences. In this paper, we argue that preference-aware query processing needs to be pushed closer to the DBMS. We present a preference-aware relational data model that extends database tuples with preferences and an extended algebra that captures the essence of processing queries with preferences. A key novelty of our preference model itself is that it defines a preference in three dimensions showing the tuples affected, their preference scores and the credibility of the preference. Our query processing strategies push preference evaluation inside the query plan and leverage its algebraic properties for finer-grained query optimization. We experimentally evaluate the proposed strategies. Finally, we compare our framework to a pure plug-in implementation and we show its feasibility and advantages. Anastasios Arvanitis, Georgia Koutrika |
ICDE | 2 |
| 2012 | Surfacing time-critical insights from social mediaabstractWe propose to demonstrate an end-to-end framework for leveraging time-sensitive and critical social media information for businesses. More specifically, we focus on identifying, structuring, integrating, and exposing timely insights that are essential to marketing services and monitoring reputation over social media. Our system includes components for information extraction from text, entity resolution and integration, analytics, and a user interface. Alexe Dumitru-Bogdan, Mauricio A. Hernández, Kirsten Hildrum, Rajasekar Krishnamurthy, Georgia Koutrika, Meena Nagarajan, Haggai Roitman, Michal Shmueli-Scheuer, Ioana Stanoi, Chitra Venkatramani, Rohit Wagle |
SIGMOD Conference | 5 |
| 2012 | PrefDB: bringing preferences closer to the DBMSabstractIn this demonstration we present a preference-aware relational query answering system, termed PrefDB. The key novelty of PrefDB is the use of an extended relational data model and algebra that allow expressing different flavors of preferential queries. Furthermore, unlike existing approaches that either treat the DBMS as a black box or require modifications of the database core, PrefDB's hybrid implementation enables operator-level query optimizations without being obtrusive to the database engine. We showcase the flexibility and efficiency of PrefDB using PrefDBAdmin, a graphical tool that we have built aiming at assisting application designers in the task of building, testing and tuning queries with preferences. Anastasios Arvanitis, Georgia Koutrika |
SIGMOD Conference | 2 |
| 2012 | Logos: a system for translating queries into narrativesabstractThis paper presents Logos, a system that provides natural language translations for relational queries expressed in SQL. Our translation mechanism is based on a graph-based approach to the query translation problem. We represent various forms of structured queries as directed graphs and we annotate the graph edges with template labels using an extensible template mechanism. Logos uses different graph traversal strategies for efficiently exploring these graphs and composing textual query descriptions. The audience may interactively explore Logos using various database schemata and issuing either sample or ad hoc queries. Andreas Kokkalis, Panagiotis Vagenas, Alexandros Zervakis, Alkis Simitsis, Georgia Koutrika, Yannis E. Ioannidis |
SIGMOD Conference | 5 |
| 2011 | On the selection of tags for tag cloudsabstractWe examine the creation of a tag cloud for exploring and understanding a set of objects (e.g., web pages, documents). In the first part of our work, we present a formal system model for reasoning about tag clouds. We then present metrics that capture the structural properties of a tag cloud, and we briefly present a set of tag selection algorithms that are used in current sites (e.g., del.icio.us, Flickr, Technorati) or that have been described in recent work. In order to evaluate the results of these algorithms, we devise a novel synthetic user model. This user model is specifically tailored for tag cloud evaluation and assumes an "ideal" user. We evaluate the algorithms under this user model, as well as the model itself, using two datasets: CourseRank (a Stanford social tool containing information about courses) and del.icio.us (a social bookmarking site). The results yield insights as to when and why certain selection schemes work best. Petros Venetis, Georgia Koutrika, Hector Garcia-Molina |
WSDM | 2 |
| 2011 | A survey on representation, composition and application of preferences in database systemsabstractPreferences have been traditionally studied in philosophy, psychology, and economics and applied to decision making problems. Recently, they have attracted the attention of researchers in other fields, such as databases where they capture soft criteria for queries. Databases bring a whole fresh perspective to the study of preferences, both computational and representational. From a representational perspective, the central question is how we can effectively represent preferences and incorporate them in database querying. From a computational perspective, we can look at how we can efficiently process preferences in the context of database queries. Several approaches have been proposed but a systematic study of these works is missing. The purpose of this survey is to provide a framework for placing existing works in perspective and highlight critical open challenges to serve as a springboard for researchers in database systems. We organize our study around three axes: preference representation, preference composition, and preference query processing. Kostas Stefanidis, Georgia Koutrika, Evaggelia Pitoura |
ACM Trans. Database Syst. | 2 |
| 2010 | Representation, composition and application of preferences in databasesabstractThis tutorial provides an overview of the key research results in the area of user preferences from a database perspective. The objective is to survey in a systematic and holistic way a number of approaches for preference representation and composition, querying with preferences and preference learning. Open research problems are also presented. Georgia Koutrika, Evaggelia Pitoura, Kostas Stefanidis |
ICDE | 1 |
| 2010 | Explaining structured queries in natural languageabstractMany applications offer a form-based environment for nai¿ve users for accessing databases without being familiar with the database schema or a structured query language. User interactions are translated to structured queries and executed. However, as a user is unlikely to know the underlying semantic connections among the fields presented in a form, it is often useful to provide her with a textual explanation of the query. In this paper, we take a graph-based approach to the query translation problem. We represent various forms of structured queries as directed graphs and we annotate the graph edges with template labels using an extensible template mechanism. We present different graph traversal strategies for efficiently exploring these graphs and composing textual query descriptions. Finally, we present experimental results for the efficiency and effectiveness of the proposed methods. Georgia Koutrika, Alkis Simitsis, Yannis E. Ioannidis |
ICDE | 1 |
| 2010 | Recsplorer: recommendation algorithms based on precedence miningabstractWe study recommendations in applications where there are temporal patterns in the way items are consumed or watched. For example, a student who has taken the Advanced Algorithms course is more likely to be interested in Convex Optimization, but a student who has taken Convex Optimization need not be interested in Advanced Algorithms in the future. Similarly, a person who has purchased the Godfather I DVD on Amazon is more likely to purchase Godfather II sometime in the future (though it is not strictly necessary to watch/purchase Godfather I beforehand). We propose a precedence mining model that estimates the probability of future consumption based on past behavior. We then propose Recsplorer: a suite of recommendation algorithms that exploit the precedence information. We evaluate our algorithms, as well as traditional recommendation ones, using a real course planning system. We use existing transcripts to evaluate how well the algorithms perform. In addition, we augment our experiments with a user study on the live system where users rate their recommendations. Aditya G. Parameswaran, Georgia Koutrika, Benjamin Bercovitz, Hector Garcia-Molina |
SIGMOD Conference | 2 |
| 2010 | Guest editorial: Special issue on collective intelligence
Epaminondas Kapetanios, Georgia Koutrika |
Inf. Sci. | 2 |
| 2010 | Personalizing queries based on networks of composite preferencesabstractPeople's preferences are expressed at varying levels of granularity and detail as a result of partial or imperfect knowledge. One may have some preference for a general class of entities, for example, liking comedies, and another one for a fine-grained, specific class, such as disliking recent thrillers with Al Pacino. In this article, we are interested in capturing such complex, multi-granular preferences for personalizing database queries and in studying their impact on query results. We organize the collection of one's preferences in a preference network (a directed acyclic graph), where each node refers to a subclass of the entities that its parent refers to, and whenever they both apply, more specific preferences override more generic ones. We study query personalization based on networks of preferences and provide efficient algorithms for identifying relevant preferences, modifying queries accordingly, and processing personalized queries. Finally, we present results of both synthetic and real-user experiments, which: (a) demonstrate the efficiency of our algorithms, (b) provide insight as to the appropriateness of the proposed preference model, and (c) show the benefits of query personalization based on composite preferences compared to simpler preference representations. Georgia Koutrika, Yannis E. Ioannidis |
ACM Trans. Database Syst. | 1 |
| 2009 | Social Systems: Can We Do More Than Just Poke Friends?
Georgia Koutrika, Benjamin Bercovitz, Robert Ikeda, Filip Kaliszan, Henry Liou, Zahra Mohammadi Zadeh, Hector Garcia-Molina |
CIDR | 1 |
| 2009 | Data clouds: summarizing keyword search results over structured dataabstractKeyword searches are attractive because they facilitate users searching structured databases. On the other hand, tag clouds are popular for navigation and visualization purposes over unstructured data because they can highlight the most significant concepts and hidden relationships in the underlying content dynamically. In this paper, we propose coupling the flexibility of keyword searches over structured data with the summarization and navigation capabilities of tag clouds to help users access a database. We propose using clouds over structured data (data clouds) to summarize the results of keyword searches over structured data and to guide users to refine their searches. The cloud presents the most significant words associated with the search results. Our keyword search model allows searching for entities than can span multiple tables in the database rather than just tuples, as existing keyword searches over databases do. We present several methods to compute the scores both for the entities and for the terms in the search results. We describe algorithms for keyword searches with data clouds and we present our system, CourseCloud, that offers a unified search and browse interface to a course database. We present experimental results showing (a) the appropriateness of the methods used for scoring terms, (b) the performance of the proposed algorithms, and (c) the effectiveness of CourseCloud compared to typical search and browse interfaces to a course database. Georgia Koutrika, Zahra Mohammadi Zadeh, Hector Garcia-Molina |
EDBT | 1 |
| 2009 | CourseCloud: summarizing and refining keyword searches over structured dataabstractIn this demo, we show data clouds that summarize the results of keyword searches over structured data. Data clouds provide insight into the database contents, hints for query modification and refinement and can lead to serendipitous discoveries of diverse results. In this demo paper: Georgia Koutrika, Zahra Mohammadi Zadeh, Hector Garcia-Molina |
EDBT | 1 |
| 2009 | Flexible Recommendations for Course PlanningabstractMost recommendation methods are "hard-wired" into the system and support only fixed recommendations. The purpose of this demo is to show the expressivity of flexible recommendation workflows, how flexible recommendations can be processed over relational data, and to show flexible recommendations in action through a real system used for course planning. Georgia Koutrika, Benjamin Bercovitz, Robert Ikeda, Filip Kaliszan, Henry Liou, Hector Garcia-Molina |
ICDE | 1 |
| 2009 | CourseRank: A Closed-Community Social System through the Magnifying Glass
Georgia Koutrika, Benjamin Bercovitz, Filip Kaliszan, Henry Liou, Hector Garcia-Molina |
ICWSM | 1 |
| 2009 | CourseRank: a social system for course planningabstractSpecial-purpose social sites can offer valuable services to well-defined, closed, communities, e.g., in a university or in a corporation. The purpose of this demo is to show the challenges, special features and potential of a focused social system in action through CourseRank, a course evaluation and planning social system. Benjamin Bercovitz, Filip Kaliszan, Georgia Koutrika, Henry Liou, Zahra Mohammadi Zadeh, Hector Garcia-Molina |
SIGMOD Conference | 3 |
| 2009 | FlexRecs: expressing and combining flexible recommendationsabstractRecommendation systems have become very popular but most recommendation methods are `hard-wired' into the system making experimentation with and implementation of new recommendation paradigms cumbersome. In this paper, we propose FlexRecs, a framework that decouples the definition of a recommendation process from its execution and supports flexible recommendations over structured data. In FlexRecs, a recommendation approach can be defined declaratively as a high-level parameterized workflow comprising traditional relational operators and new operators that generate or combine recommendations. We describe a prototype flexible recommendation engine that realizes the proposed framework and we present example workflows and experimental results that show its potential for capturing multiple, existing or novel, recommendations easily and having a flexible recommendation system that combines extensibility with reasonable performance. Georgia Koutrika, Benjamin Bercovitz, Hector Garcia-Molina |
SIGMOD Conference | 1 |
| 2009 | Entity resolution with iterative blockingabstractEntity Resolution (ER) is the problem of identifying which records in a database refer to the same real-world entity. An exhaustive ER process involves computing the similarities between pairs of records, which can be very expensive for large datasets. Various blocking techniques can be used to enhance the performance of ER by dividing the records into blocks in multiple ways and only comparing records within the same block. However, most blocking techniques process blocks separately and do not exploit the results of other blocks. In this paper, we propose an iterative blocking framework where the ER results of blocks are reflected to subsequently processed blocks. Blocks are now iteratively processed until no block contains any more matching records. Compared to simple blocking, iterative blocking may achieve higher accuracy because reflecting the ER results of blocks to other blocks may generate additional record matches. Iterative blocking may also be more efficient because processing a block now saves the processing time for other blocks. We implement a scalable iterative blocking system and demonstrate that iterative blocking can be more accurate and efficient than blocking for large datasets. Steven Euijong Whang, David Menestrina, Georgia Koutrika, Martin Theobald, Hector Garcia-Molina |
SIGMOD Conference | 3 |
| 2008 | Synthesizing structured text from logical database subsetsabstractIn the classical database world, information access has been based on a paradigm that involves structured, schema-aware, queries and tabular answers. In the current environment, however, where information prevails in most activities of society, serving people, applications, and devices in dramatically increasing numbers, this paradigm has proved to be very limited. On the query side, much work has been done on moving towards keyword queries over structured data. In our previous work, we have touched the other side as well, and have proposed a paradigm that generates entire databases in response to keyword queries. In this paper, we continue in the same direction and propose synthesizing textual answers in response to queries of any kind over structured data. In particular, we study the transformation of a dynamically-generated logical database subset into a narrative through a customizable, extensible, and templatebased process. In doing so, we exploit the structured nature of database schemas and describe three generic translation modules for different formations in the schema, called unary, split, and join modules. We have implemented the proposed translation procedure into our own database front end and have performed several experiments evaluating the textual answers generated as several features and parameters of the system are varied. We have also conducted a set of experiments measuring the effectiveness of such answers on users. The overall results are very encouraging and indicate the promise that our approach has for several applications. Alkis Simitsis, Georgia Koutrika, Yannis Alexandrakis, Yannis E. Ioannidis |
EDBT | 2 |
| 2008 | Flexible recommendations over rich dataabstractCourseRank is a course planning tool aimed at helping students at Stanford. Recommendations comprise an integral part of it. However, implementing existing recommendation methods leads to fixed recommendations that cannot adapt to each particular student's changing requirements and do not help exploit the full extent of the available learning opportunities at the university. In this paper, we describe the concept of a flexible recommendation workflow, i.e., a high-level description of a parameterized process for computing recommendations. The input parameters of a flexible recommendation process comprise the "knobs" that control the final output and hence generate flexible recommendations. We describe how flexible recommendations can be expressed over a relational database and we present our prototype system that allows defining and executing different, fully-parameterized, recommendation workflows over relational data. Finally, we describe a user interface in CourseRank that allows students customize recommendations. Georgia Koutrika, Robert Ikeda, Benjamin Bercovitz, Hector Garcia-Molina |
RecSys | 1 |
| 2008 | OLAP Cubes for Social Searches: Standing on the Shoulders of Giants?
Konstantinos Morfonios, Georgia Koutrika |
WebDB | 2 |
| 2008 | Can social bookmarking improve web search?abstractSocial bookmarking is a recent phenomenon which has the potential to give us a great deal of data about pages on the web. One major question is whether that data can be used to augment systems like web search. To answer this question, over the past year we have gathered what we believe to be the largest dataset from a social bookmarking site yet analyzed by academic researchers. Our dataset represents about forty million bookmarks from the social bookmarking site del.icio.us. We contribute a characterization of posts to del.icio. us: how many bookmarks exist (about 115 million), how fast is it growing, and how active are the URLs being posted about (quite active). We also contribute a characterization of tags used by bookmarkers. We found that certain tags tend to gravitate towards certain domains, and vice versa. We also found that tags occur in over 50 percent of the pages that they annotate, and in only 20 percent of cases do they not occur in the page text, backlink page text, or forward link page text of the pages they annotate. We conclude that social bookmarking can provide search data not currently provided by other sources, though it may currently lack the size and distribution of tags necessary to make a significant impact Paul Heymann, Georgia Koutrika, Hector Garcia-Molina |
WSDM | 2 |
| 2008 | Combating spam in tagging systems: An evaluationabstractTagging systems allow users to interactively annotate a pool of shared resources using descriptive strings called tags . Tags are used to guide users to interesting resources and help them build communities that share their expertise and resources. As tagging systems are gaining in popularity, they become more susceptible to tag spam : misleading tags that are generated in order to increase the visibility of some resources or simply to confuse users. Our goal is to understand this problem better. In particular, we are interested in answers to questions such as: How many malicious users can a tagging system tolerate before results significantly degrade? What types of tagging systems are more vulnerable to malicious attacks? What would be the effort and the impact of employing a trusted moderator to find bad postings? Can a system automatically protect itself from spam, for instance, by exploiting user tag patterns? In a quest for answers to these questions, we introduce a framework for modeling tagging systems and user tagging behavior. We also describe a method for ranking documents matching a tag based on taggers' reliability. Using our framework, we study the behavior of existing approaches under malicious attacks and the impact of a moderator and our ranking method. Georgia Koutrika, Frans Adjie Effendi, Zoltán Gyöngyi, Paul Heymann, Hector Garcia-Molina |
ACM Trans. Web | 1 |
| 2008 | Précis: from unstructured keywords as queries to structured databases as answers
Alkis Simitsis, Georgia Koutrika, Yannis E. Ioannidis |
VLDB J. | 2 |
| 2007 | Generalized Précis Queries for Logical Database Subset CreationabstractAs a large fraction of available information resides in databases, the need for facilitating access for the large majority of users becomes increasingly more important. Precis queries are free-form queries that generate entire multi-relation databases, which are logical subsets of existing ones. A logical subset contains not only items directly related to the given query selections but also items implicitly related to them in various ways with the purpose of providing to the user much greater insight into the original data. This paper is concerned with the definition and generation of logical database subsets based on precis queries under a generalized perspective that removes several restrictions of previous work and handles queries containing multiple terms combined using the operators AND, OR, and NOT. Alkis Simitsis, Georgia Koutrika, Yannis E. Ioannidis |
ICDE | 2 |
| 2006 | Comprehensible Answers to Précis Queries
Alkis Simitsis, Georgia Koutrika |
CAiSE | 2 |
| 2006 | Précis: The Essence of a Query AnswerabstractWide spread use of database systems in modern society has brought the need to provide inexperienced users with the ability to easily search a database with no specific knowledge of a query language. Several recent research efforts have focused on supporting keyword-based searches over relational databases. This paper presents an alternative proposal and introduces the idea of précis queries. These are free-form queries whose answer (a précis) is a synthesis of results, containing not only information directly related to the query selections but also information implicitly related to them in various ways. Our approach to précis queries includes two additional novelties: (a) queries do not generate individual relations but entire multi-relation databases; and (b) query results are personalized to user-specific and/or domain requirements. We develop a framework and system architecture for supporting such queries in the context of a relational database system and describe algorithms that implement the required functionality. Finally, we present a set of experimental results that evaluate the proposed algorithms and show the potential of this work. Georgia Koutrika, Alkis Simitsis, Yannis E. Ioannidis |
ICDE | 1 |
| 2005 | Personalized Queries under a Generalized Preference ModelabstractQuery personalization is the process of dynamically enhancing a query with related user preferences stored in a user profile with the aim of providing personalized answers. The underlying idea is that different users may find different things relevant to a search due to different preferences. Essential ingredients of query personalization are: (a) a model for representing and storing preferences in user profiles, and (b) algorithms for the generation of personalized answers using stored preferences. Modeling the plethora of preference types is a challenge. In this paper, we present a preference model that combines expressivity and concision. In addition, we provide efficient algorithms for the selection of preferences related to a query, and an algorithm for the progressive generation of personalized results, which are ranked based on user interest. Several classes of ranking functions are provided for this purpose. We present results of experiments both synthetic and with real users (a) demonstrating the efficiency of our algorithms, (b) showing the benefits of query personalization, and (c) providing insight as to the appropriateness of the proposed ranking functions. Georgia Koutrika, Yannis E. Ioannidis |
ICDE | 1 |
| 2005 | Constrained Optimalities in Query PersonalizationabstractPersonalization is a powerful mechanism that helps users to cope with the abundance of information on the Web. Database query personalization achieves this by dynamically constructing queries that return results of high interest to the user. This, however, may conflict with other constraints on the query execution time and/or result size that may be imposed by the search context, such as the device used, the network connection, etc. For example, if the user is accessing information using a mobile phone, then it is desirable to construct a personalized query that executes quickly and returns a handful of answers. Constrained Query Personalization (CQP) is an integrated approach to database query answering that dynamically takes into account the queries issued, the user's interest in the results, response time, and result size in order to build personalized queries. In this paper, we introduce CQP as a family of constrained optimization problems, where each time one of the parameters of concern is optimized while the others remain within the bounds of range constraints. Taking into account some key (exact or approximate) properties of these parameters, we map CQP to a state search problem and provide several algorithms for the discovery of optimal solutions. Experimental results demonstrate the effectiveness of the proposed techniques and the appropriateness of the overall approach. Georgia Koutrika, Yannis E. Ioannidis |
SIGMOD Conference | 1 |
| 2005 | Personalized Systems: Models and Methods from an IR and DB Perspective
Yannis E. Ioannidis, Georgia Koutrika |
VLDB | 2 |
| 2004 | Personalization of Queries in Database SystemsabstractAs information becomes available in increasing amounts to a wide spectrum of users, the need for a shift towards a more user-centered information access paradigm arises. We develop a personalization framework for database systems based on user profiles and identify the basic architectural modules required to support it. We define a preference model that assigns to each atomic query condition a personal degree of interest and provide a mechanism to compute the degree of interest in any complex query condition based on the degrees of interest in the constituent atomic ones. Preferences are stored in profiles. At query time, personalization proceeds in two steps: (a) preference selection and (b) preference integration into the original user query. We formulate the main personalization step, i.e. preference selection, as a graph computation problem and provide an efficient algorithm for it. We also discuss results of experimentation with a prototype query personalization system. Georgia Koutrika, Yannis E. Ioannidis |
ICDE | 1 |