Yannis Katsis

dblp:75/4070 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0002-1733-6227ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
12 papers
Data integration and cleaning · 45% Query processing and optimization · 18% Knowledge graphs · 16%
Human-computer interaction and pervasive computing
4 papers
Human-AI interaction · 54% Usability and user experience research · 22% User interface design and tools · 18%
Artificial intelligence
2 papers
Efficient and distributed learning · 38% Question answering and dialogue systems · 38% Language models and text generation · 12%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 22 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
User interface design and tools › personalization
Personalized AI
0.712023
Exploring the Use of Personalized AI for Identifying Misinformation on Social Media · CHI 2023
Usability and user experience research › evaluation methodology
evaluation framework
0.612022
A Simulation-Based Evaluation Framework for Interactive AI Systems and Its Application · AAAI 2022
Human-AI interaction
simulation-based evaluation
0.612022
InteractEva: A Simulation-Based Evaluation Framework for Interactive AI Systems · AAAI 2022
Machine learning › Efficient and distributed learning
automated machine learning
0.512021
AutoText: An End-to-End AutoAI Framework for Text · AAAI 2021
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.512021
KAAPA: Knowledge Aware Answers from PDF Analysis · AAAI 2021
Knowledge graphs
knowledge graph construction
0.512021
KAAPA: Knowledge Aware Answers from PDF Analysis · AAAI 2021
Data integration and cleaning › data extraction
table extraction
0.412020
Table Extraction and Understanding for Scientific and Enterprise Applications · Proc. VLDB Endow. 2020
Data integration and cleaning
table understanding
0.412020
Table Extraction and Understanding for Scientific and Enterprise Applications · Proc. VLDB Endow. 2020
Query processing and optimization › query rewriting
query answering using views
0.322014
Complete yet practical search for minimal query reformulations under constraints · SIGMOD Conference 2014
Exporting and interactively querying Web service-accessed sources: The CLIDE System · ACM Trans. Database Syst. 2007
Query processing and optimization › view maintenance
incremental view maintenance
0.212015
Utilizing IDs to Accelerate Incremental View Maintenance · SIGMOD Conference 2015
Collaborative and social computing
social media
0.212023
Exploring the Use of Personalized AI for Identifying Misinformation on Social Media · CHI 2023
Information retrieval
query reformulation
0.212014
Complete yet practical search for minimal query reformulations under constraints · SIGMOD Conference 2014
Query processing and optimization › semantic query processing
semantic query optimization
0.212014
Complete yet practical search for minimal query reformulations under constraints · SIGMOD Conference 2014
Information retrieval
fact-checking
0.212013
Fact checking and analyzing the web · SIGMOD Conference 2013
Data integration and cleaning › schema mapping
object-relational mapping
0.212013
Query containment in entity SQL · SIGMOD Conference 2013
Database theory
query containment
0.212013
Query containment in entity SQL · SIGMOD Conference 2013
Natural language and speech › Information extraction and text analysis › document analysis › document information extraction
table extraction
0.112021
KAAPA: Knowledge Aware Answers from PDF Analysis · AAAI 2021
Data integration and cleaning
collaborative data management
0.112010
Inconsistency resolution in online databases · ICDE 2010
Data integration and cleaning › data fusion
conflict resolution
0.112010
Inconsistency resolution in online databases · ICDE 2010
Data integration and cleaning
schema mapping
0.122008
Interactive source registration in community-oriented information integration · Proc. VLDB Endow. 2008
RIDE: a tool for interactive source registration in community-oriented information integration · Proc. VLDB Endow. 2008
Data integration and cleaning › data quality
inconsistency detection
0.012008
Interactive source registration in community-oriented information integration · Proc. VLDB Endow. 2008
Database theory › query answering
certain answers
0.012005
Determining source contribution in integration systems · PODS 2005

Methods — techniques the papers use, named apart from their topics

user simulation · 1.1user study · 0.7personalized machine learning · 0.7simulation-based evaluation · 0.6neural architecture search · 0.5hyperparameter optimization · 0.5provenance · 0.2constructive containment checking · 0.2query-by-example interface · 0.1answering queries using views · 0.1entropy maximization · 0.1source-to-target constraints · 0.1certain answer semantics · 0.1
YearPublicationVenuePosition
2025 Diagnosing and Prioritizing Issues in Automated Order-Taking Systems: A Machine-Assisted Error Discovery Approach
Maeda F. Hanafi, Frederick Reiss 0001, Yannis Katsis, Mohammad Hassan Falakmasir, Pauline Wang, Changchang Liu
CHI3
2025 mtRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
abstract
Abstract Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is asked to generate a response to a question in the context of a preceding conversation, is an important and often overlooked task with several additional challenges. We present mtRAG, an end-to-end human-generated multi-turn RAG benchmark that reflects several real-world properties across diverse dimensions for evaluating the full RAG pipeline. mtRAG contains 110 conversations averaging 7.7 turns each across four domains for a total of 842 tasks. We also explore automation paths via synthetic data and LLM-as-a-Judge evaluation. Our human and automatic evaluations show that even state-of-the-art LLM RAG systems struggle on mtRAG. We demonstrate the need for strong retrieval and generation systems that can handle later turns, unanswerable questions, non-standalone questions, and multiple domains. mtRAG is available at https://github.com/ibm/mt-rag-benchmark.
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis, R. Chulaka Gunasekara, Young-Suk Lee 0001, Lucian Popa 0001, Vraj Shah, Huaiyu Zhu 0001, Danish Contractor, Marina Danilevsky
Trans. Assoc. Comput. Linguistics1
2023 Exploring the Use of Personalized AI for Identifying Misinformation on Social Media
abstract
This work aims to explore how human assessments and AI predictions can be combined to identify misinformation on social media. To do so, we design a personalized AI which iteratively takes as training data a single user’s assessment of content and predicts how the same user would assess other content. We conduct a user study in which participants interact with a personalized AI that learns their assessments of a feed of tweets, shows its predictions of whether a user would find other tweets (in)accurate, and evolves according to the user feedback. We study how users perceive such an AI, and whether the AI predictions influence users’ judgment. We find that this influence does exist and it grows larger over time, but it is reduced when users provide reasoning for their assessment. We draw from our empirical observations to identify design implications and directions for future work.
Farnaz Jahanbakhsh, Yannis Katsis, Dakuo Wang, Lucian Popa 0001, Michael J. Muller
CHI2
2022 A Simulation-Based Evaluation Framework for Interactive AI Systems and Its Application
abstract
Interactive AI (IAI) systems are increasingly popular as the human-centered AI design paradigm is gaining strong traction. However, evaluating IAI systems, a key step in building such systems, is particularly challenging, as their output highly depends on the performed user actions. Developers often have to rely on limited and mostly qualitative data from ad-hoc user testing to assess and improve their systems. In this paper, we present InteractEva; a systematic evaluation framework for IAI systems. We also describe how we have applied InteractEva to evaluate a commercial IAI system, leading to both quality improvements and better data-driven design decisions.
Maeda F. Hanafi, Yannis Katsis, Martín Santillán Cooper, Yunyao Li 0001
AAAI2
2022 InteractEva: A Simulation-Based Evaluation Framework for Interactive AI Systems
abstract
Evaluating interactive AI (IAI) systems is a challenging task, as their output highly depends on the performed user actions. As a result, developers often depend on limited and mostly qualitative data derived from user testing to improve their systems. In this paper, we present InteractEva; a systematic evaluation framework for IAI systems. InteractEva employs (a) a user simulation backend to test the system against different use cases and user interactions at scale with (b) an interactive frontend allowing developers to perform important quantitative evaluation tasks, including acquiring a performance overview, performing error analysis, and conducting what-if studies. The framework has supported the evaluation and improvement of an industrial IAI text extraction system, results of which will be presented during our demonstration.
Yannis Katsis, Maeda F. Hanafi, Martín Santillán Cooper, Yunyao Li 0001
AAAI1
2022 SPOT: Knowledge-Enhanced Language Representations for Information Extraction
abstract
Knowledge-enhanced pre-trained models for language representation have been shown to be more effective in knowledge base construction tasks (i.e.,~relation extraction) than language models such as BERT. These knowledge-enhanced language models incorporate knowledge into pre-training to generate representations of entities or relationships. However, existing methods typically represent each entity with a separate embedding. As a result, these methods struggle to represent out-of-vocabulary entities and a large amount of parameters, on top of their underlying token models (i.e., the transformer), must be used and the number of entities that can be handled is limited in practice due to memory constraints. Moreover, existing models still struggle to represent entities and relationships simultaneously. To address these problems, we propose a new pre-trained model that learns representations of both entities and relationships from token spans and span pairs in the text respectively. By encoding spans efficiently with span modules, our model can represent both entities and their relationships but requires fewer parameters than existing models. We pre-trained our model with the knowledge graph extracted from Wikipedia and test it on a broad range of supervised and unsupervised information extraction tasks. Results show that our model learns better representations for both entities and relationships than baselines, while in supervised settings, fine-tuning our model outperforms RoBERTa consistently and achieves competitive results on information extraction tasks.
Jiacheng Li 0003, Yannis Katsis, Tyler Baldwin, Ho-Cheol Kim, Andrew Bartko, Julian J. McAuley, Chun-Nan Hsu
CIKM2
2022 Theoretical Rule-based Knowledge Graph Reasoning by Connectivity Dependency Discovery
abstract
Discovering precise and interpretable rules from knowledge graphs is regarded as an essential challenge, which can improve the performances of many downstream tasks and even provide new ways to approach some Natural Language Processing research topics. In this paper, we present a fundamen-tal theory for rule-based knowledge graph reasoning, based on which the connectivity dependencies in the graph are captured via multiple rule types. It is the first time for some of these rule types in a knowledge graph to be considered. Based on these rule types, our theory can provide precise interpretations to unknown triples. Then, we implement our theory by what we call the RuleDict model. Results show that our RuleDict model not only provides precise rules to interpret new triples, but also achieves state-of-the-art performances on one benchmark knowledge graph completion task, and is competitive on other tasks.
Canlin Zhang, Chun-Nan Hsu, Yannis Katsis, Ho-Cheol Kim, Yoshiki Vazquez-Baeza
IJCNN3
2021 AutoText: An End-to-End AutoAI Framework for Text
abstract
Building models for natural language processing (NLP) tasks remains a daunting task for many, requiring significant technical expertise, efforts, and resources. In this demonstration, we present AutoText, an end-to-end AutoAI framework for text, to lower the barrier of entry in building NLP models. AutoText combines state-of-the-art AutoAI optimization techniques and learning algorithms for NLP tasks into a single extensible framework. Through its simple, yet powerful UI, non-AI experts (e.g., domain experts) can quickly generate performant NLP models with support to both control (e.g., via specifying constraints) and understand learned models.
Arunima Chaudhary, Alayt Issak, Kiran Kate, Yannis Katsis, Abel N. Valente, Dakuo Wang, Alexandre V. Evfimievski, Sairam Gurajada, Ban Kawas, Cristiano Malossi, Lucian Popa 0001, Tejaswini Pedapati, Horst Samulowitz, Martin Wistuba, Yunyao Li 0001
AAAI4
2021 KAAPA: Knowledge Aware Answers from PDF Analysis
abstract
We present KaaPa (Knowledge Aware Answers from Pdf Analysis), an integrated solution for machine reading comprehension over both text and tables extracted from PDFs. KaaPa enables interactive question refinement using facets generated from an automatically induced Knowledge Graph. In addition it provides a concise summary of the supporting evidence for the provided answers by aggregating information across multiple sources. KaaPa can be applied consistently to any collection of documents in English with zero domain adaptation effort. We showcase the use of KaaPa for QA on scientific literature using the COVID-19 Open Research Dataset.
Nicolas R. Fauceglia, Mustafa Canim, Alfio Massimiliano Gliozzo, Jennifer J. Liang, Nancy Xin Ru Wang, Douglas Burdick, Nandana Mihindukulasooriya, Vittorio Castelli, Guy Feigenblat, David Konopnicki, Yannis Katsis, Radu Florian, Yunyao Li 0001, Salim Roukos, Avirup Sil
AAAI11
2020 Table Extraction and Understanding for Scientific and Enterprise Applications
abstract
Valuable high-precision data are often published in the form of tables in both scientific and business documents. While humans can easily identify, interpret and contextualize tables, developing general-purpose automated techniques for extraction of information from tables is difficult due to the wide variety of table formats employed across corpora. To extract useful data from tables, data cells must be correctly extracted and linked to all relevant headers, units of measure and in-text references. Table extraction involves identifying the border and cell structure for each document table, while table understanding provides context by linking cells with semantic information inside and outside the table, such as row and column headers, footnotes, titles, and references in surrounding text. The objective of this tutorial is to provide a detailed synopsis of existing approaches for table extraction and understanding, highlight open research problems, and provide an overview of potential applications.
Douglas Burdick, Marina Danilevsky, Alexandre V. Evfimievski, Yannis Katsis, Nancy Xin Ru Wang
Proc. VLDB Endow.4
2015 Combining Databases and Signal Processing in Plato
Yannis Katsis, Yoav Freund, Yannis Papakonstantinou
CIDR1
2015 Utilizing IDs to Accelerate Incremental View Maintenance
abstract
Prior Incremental View Maintenance (IVM) algorithms specify the view tuples that need to be modified by computing diff sets, which we call tuple-based diffs since a diff set contains one diff tuple for each to-be-modified view tuple. idIVM assumes the base tables have keys and performs IVM by computing ID-based diff sets that compactly identify the to-be-modified tuples through their IDs.
Yannis Katsis, Kian Win Ong, Yannis Papakonstantinou, Kevin Keliang Zhao
SIGMOD Conference1
2014 Complete yet practical search for minimal query reformulations under constraints
abstract
We revisit the Chase&Backchase (C&B) algorithm for query reformulation under constraints, which provides a uniform solution to such particular-case problems as view-based rewriting under constraints, semantic query optimization, and physical access path selection in query optimization. For an important class of queries and constraints, C&B has been shown to be complete, i.e. guaranteed to find all (join-)minimal reformulations under constraints. C&B is based on constructing a canonical rewriting candidate called a universal plan, then inspecting its exponentially many sub-queries in search for minimal reformulations, essentially removing redundant joins in all possible ways. This inspection involves chasing the subquery. Because of the resulting exponentially many chases, the conventional wisdom has held that completeness is a concept of mainly theoretical interest. We show that completeness can be preserved at practically relevant cost by introducing Prov-C&B, a novel reformulation algorithm that instruments the chase to maintain provenance information connecting the joins added during the chase to the universal plan subqueries responsible for adding these joins. This allows it to directly "read off" the minimal reformulations from the result of a single chase of the universal plan, saving exponentially many chases of its subqueries. We exhibit natural scenarios yielding speedups of over two orders of magnitude between the execution of the best view-based rewriting found by a commercial query optimizer and that of the best rewriting found by Prov-C&B (which the optimizer misses because of limited reasoning about constraints).
Ioana Ileana, Bogdan Cautis, Alin Deutsch, Yannis Katsis
SIGMOD Conference4
2013 DELPHI: Data E-platform for personalized population health
abstract
Recent studies recognize that health is influenced broadly by a multitude of factors of different types, including medical, genetic, environmental, social and behavioral factors. Developing successful health interventions therefore requires taking into account all these factors as well as the interactions between them. However, intervention designers have traditionally had access only to a very limited subset of health data (typically medical record data). Other health data, such as environmental or physical activity data, although already collected and stored, have been very difficult to access, since they are maintained by different providers and isolated in their own proprietary silos. This prevents physicians and intervention designers from acquiring a true overview of all factors influencing a condition and acting towards its prevention or cure. To solve this problem, we propose DELPHI: a platform allowing the integration of disparate health data into a single Whole Health Information Model (WHIM), providing a 360-degree view of an individual's health. DELPHI supports the integration of data and thus enables the design of applications and services that utilize the WHIM to offer the next generation of health services. In this paper, we describe DELPHI's architecture, outline the technical challenges encountered and describe an asthma management use case that will be enabled by DELPHI.
Yannis Katsis, Chaitanya K. Baru, Ted Chan, Sanjoy Dasgupta, Claudiu Farcas, William G. Griswold, Jeannie Huang, Lucila Ohno-Machado, Yannis Papakonstantinou, Fred Raab, Kevin Patrick 0001
Healthcom1
2013 Fact checking and analyzing the web
abstract
Fact checking and data journalism are currently strong trends. The sheer amount of data at hand makes it difficult even for trained professionals to spot biased, outdated or simply incorrect information. We propose to demonstrate FactMinder, a fact checking and analysis assistance application. SIGMOD attendees will be able to analyze documents using FactMinder and experience how background knowledge and open data repositories help build insightful overviews of current topics.
François Goasdoué, Konstantinos Karanasos, Yannis Katsis, Julien Leblay, Ioana Manolescu, Stamatis Zampetakis
SIGMOD Conference3
2013 Query containment in entity SQL
abstract
We describe a software architecture we have developed for a constructive containment checker of Entity SQL queries defined over extended ER schemas expressed in Microsoft's Entity Data Model. Our application of interest is compilation of object-to-relational mappings for Microsoft's ADO.NET Entity Framework, which has been shipping since 2007. The supported language includes several features which have been individually addressed in the past but, to the best of our knowledge, they have not been addressed all at once before. Moreover, when embarking on an implementation, we found no guidance in the literature on how to modularize the software or apply published algorithms to a commercially-supported language. This paper reports on our experience in addressing these real-world challenges.
Guillem Rull, Philip A. Bernstein, Ivo Garcia dos Santos, Yannis Katsis, Sergey Melnik 0001, Ernest Teniente
SIGMOD Conference4
2013 On the equivalence of distributed systems with queries and communication
Serge Abiteboul, Balder ten Cate, Yannis Katsis
J. Comput. Syst. Sci.3
2013 Growing triples on trees: an XML-RDF hybrid model for annotated documents
François Goasdoué, Konstantinos Karanasos, Yannis Katsis, Julien Leblay, Ioana Manolescu, Stamatis Zampetakis
VLDB J.3
2011 On the equivalence of distributed systems with queries and communication
abstract
Distributed data management systems consist of peers that store, exchange and process data in order to collaboratively achieve a common goal, such as evaluate some query. We study the equivalence of such systems. We model a distributed system by a collection of Active XML documents, i.e., trees augmented with function calls for performing tasks such as sending, receiving and querying data. As our model is quite general, the equivalence problem turns out to be undecidable. However, we exhibit several restrictions of the model, for which equivalence can be effectively decided. We also study the computational complexity of the equivalence problem, and present an axiomatization of equivalence, in the form of a set of equivalence-preserving rewrite rules allowing us to optimize a system by rewriting it into an equivalent, but possibly more efficient system.
Serge Abiteboul, Balder ten Cate, Yannis Katsis
ICDT3
2010 Inconsistency resolution in online databases
abstract
Shared online databases allow community members to collaboratively maintain knowledge. Collaborative editing though inevitably leads to inconsistencies as different members enter erroneous data or conflicting opinions. Ideally community members should be able to see and resolve these inconsistencies in a collaborative fashion. However most current online databases do not support inconsistency resolution. Instead they try to by-pass the problem by either ignoring inconsistencies and treating data as if they were not conflicting or by requiring inconsistencies to be resolved outside the system. To address this limitation, we propose Ricolla; an online database system that, by treating inconsistencies as first-class citizens, supports a natural workflow for the management of conflicting data. The system captures inconsistencies (so that community members can easily inspect them) and remains fully functional in their presence, thus enabling inconsistency resolution in an ¿as-you-go¿ fashion. Moreover it supports several schemes for the resolution of inconsistencies, allowing among others users to collaboratively resolve certain conflicts while disagreeing on others.
Yannis Katsis, Alin Deutsch, Yannis Papakonstantinou, Vasilis Vassalos
ICDE1
2008 Interactive source registration in community-oriented information integration
abstract
Modern Internet communities need to integrate and query structured information. Employing current information integration infrastructure, data integration is still a very costly effort, since source registration is performed by a central authority which becomes a bottleneck. We propose the community-based integration paradigm which pushes the source registration task to the independent community members. This creates new challenges caused by each community member's lack of a global overview on how her data interacts with the application queries of the community and the data from other sources. How can the source owner maximize the visibility of her data to existing applications, while minimizing the clean-up and reformatting cost associated with publishing? Does her data contradict (or could it contradict in the future) the data of other sources? We introduce RIDE, a visual registration tool that extends schema mapping interfaces like that of MS Biz Talk Server and IBM's Clio with a suggestion component that guides the source owner in the autonomous registration, assisting her in answering these questions. RIDE's implementation features efficient procedures for deciding various levels of self-reliance of a GLAV-style source registration for contributing answers to an application query and checking potential and definite inconsistency across sources.
Yannis Katsis, Alin Deutsch, Yannis Papakonstantinou
Proc. VLDB Endow.1
2008 RIDE: a tool for interactive source registration in community-oriented information integration
abstract
Modern Internet communities need to integrate and query structured information. Employing current information integration infrastructure, data integration is still a very costly effort, since source registration is performed by a central authority which becomes a bottleneck. We propose the community-based integration paradigm which pushes the source registration task to the independent community members. This creates new challenges caused by each member's lack of a global overview on how her data interacts with the application queries of the community and the data from other sources. How can the source owner maximize the visibility of her data to existing applications, while minimizing the clean-up and reformatting cost associated with publishing? Does her data contradict (or could it contradict in the future) the data of other sources?
Yannis Katsis, Alin Deutsch, Yannis Papakonstantinou, Kevin Keliang Zhao
Proc. VLDB Endow.1
2007 Exporting and interactively querying Web service-accessed sources: The CLIDE System
abstract
The CLIDE System assists the owners of sources that participate in Web service-based data publishing systems to publish a restricted set of parameterized queries over the schema of their sources and package them as WSDL services. The sources may be relational databases, which naturally have a schema, or ad hoc information/application systems whereas the owner publishes a virtual schema. CLIDE allows information clients to pose queries over the published schema and utilizes prior work on answering queries using views to answer queries that can be processed by combining and processing the results of one or more Web service calls. These queries are called feasible . Contrary to prior work, where infeasible queries are rejected without an explanatory feedback, leading the user into a frustrating trial-and-error cycle, CLIDE features a query formulation interface, which extends the QBE-like query builder of Microsoft's SQL Server with a color scheme that guides the user toward formulating feasible queries. CLIDE guarantees that the suggested query edit actions are complete (i.e., each feasible query can be built by following only suggestions), rapidly convergent (the suggestions are tuned to lead to the closest feasible completions of the query), and suitably summarized (at each interaction step, only a minimal number of actions needed to preserve completeness are suggested). We present the algorithms, implementation, and performance evaluation showing that CLIDE is a viable on-line tool.
Michalis Petropoulos, Alin Deutsch, Yannis Papakonstantinou, Yannis Katsis
ACM Trans. Database Syst.4
2005 Determining source contribution in integration systems
abstract
Owners of sources registered in an information integration system, which provides answers to a (potentially evolving) set of client queries, need to know their contribution to the query results. We study the problem of deciding, given a client query Q and a source registration R, whether R is (i) "self-sufficient" (can contribute to the result of Q even if it is the only source in the system) or (ii) "now complementary" (can contribute, but only in cooperation with other specific existing sources), or (iii)"later complementary" (can contribute if in the future appropriate new sources join the system). We consider open-world integration systems in which registrations are expressed using source-to-target constraints, and queries are answered under "certain answer" semantics.
Alin Deutsch, Yannis Katsis, Yannis Papakonstantinou
PODS2