Christoph Lofi

dblp:52/980 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
6since 2021 · last 2023
0000-0001-5641-5510ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 23 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2023 On the Popularity of Classical Music Composers on Community-Driven Platforms
Ioannis Petros Samiotis, Andrea Mauri 0001, Christoph Lofi, Alessandro Bozzon
ICWE3
2022 How can Explainability Methods be Used to Support Bug Identification in Computer Vision Models?
abstract
Deep learning models for image classification suffer from dangerous issues often discovered after deployment. The process of identifying bugs that cause these issues remains limited and understudied. Especially, explainability methods are often presented as obvious tools for bug identification. Yet, the current practice lacks an understanding of what kind of explanations can best support the different steps of the bug identification process, and how practitioners could interact with those explanations. Through a formative study and an iterative co-creation process, we build an interactive design probe providing various potentially relevant explainability functionalities, integrated into interfaces that allow for flexible workflows. Using the probe, we perform 18 user-studies with a diverse set of machine learning practitioners. Two-thirds of the practitioners engage in successful bug identification. They use multiple types of explanations, e.g. visual and textual ones, through non-standardized sequences of interactions including queries and exploration. Our results highlight the need for interactive, guiding, interfaces with diverse explanations, shedding light on future research directions.
Agathe Balayn, Natasa Rikalo, Christoph Lofi, Jie Yang 0028, Alessandro Bozzon
CHI3
2021 Exploring the Music Perception Skills of Crowd Workers
abstract
Music content annotation campaigns are common on paid crowdsourcing platforms. Crowd workers are expected to annotate complicated music artefacts, which can demand certain skills and expertise. Traditional methods of participant selection are not designed to capture these kind of domain-specific skills and expertise, and often domain-specific questions fall under the general demographics category. Despite the popularity of such tasks, there is a general lack of deeper understanding of the distribution of musical properties - especially auditory perception skills - among workers. To address this knowledge gap, we conducted a user study (N=100) on Prolific. We asked workers to indicate their musical sophistication through a questionnaire and assessed their music perception skills through an audio-based skill test. The goal of this work is to better understand the extent to which crowd workers possess higher perceptions skills, beyond their own musical education level and self reported abilities. Our study shows that untrained crowd workers can possess high perception skills on the music elements of melody, tuning, accent and tempo; skills that can be useful in a plethora of annotation tasks in the music domain.
Ioannis Petros Samiotis, Sihang Qiu, Christoph Lofi, Jie Yang 0028, Ujwal Gadiraju, Alessandro Bozzon
HCOMP3
2021 Valentine: Evaluating Matching Techniques for Dataset Discovery
abstract
Data scientists today search large data lakes to discover and integrate datasets. In order to bring together disparate data sources, dataset discovery methods rely on some form of schema matching: the process of establishing correspondences between datasets. Traditionally, schema matching has been used to find matching pairs of columns between a source and a target schema. However, the use of schema matching in dataset discovery methods differs from its original use. Nowadays schema matching serves as a building block for indicating and ranking inter-dataset relationships. Surprisingly, although a discovery method's success relies highly on the quality of the underlying matching algorithms, the latest discovery methods employ existing schema matching algorithms in an ad-hoc fashion due to the lack of openly-available datasets with ground truth, reference method implementations, and evaluation metrics.In this paper, we aim to rectify the problem of evaluating the effectiveness and efficiency of schema matching methods for the specific needs of dataset discovery. To this end, we propose Valentine, an extensible open-source experiment suite to execute and organize large-scale automated matching experiments on tabular data. Valentine includes implementations of seminal schema matching methods that we either implemented from scratch (due to absence of open source code) or imported from open repositories. The contributions of Valentine are: i) the definition of four schema matching scenarios as encountered in dataset discovery methods, ii) a principled dataset fabrication process tailored to the scope of dataset discovery methods and iii) the most comprehensive evaluation of schema matching techniques to date, offering insight on the strengths and weaknesses of existing techniques, that can serve as a guide for employing schema matching in future dataset discovery methods.
Christos Koutras, Georgios Siachamis 0001, Andra Ionescu, Kyriakos Psarakis, Jerry Brons, Marios Fragkoulis, Christoph Lofi, Angela Bonifati, Asterios Katsifodimos
ICDE7
2021 What do You Mean? Interpreting Image Classification with Crowdsourced Concept Extraction and Analysis
abstract
Global interpretability is a vital requirement for image classification applications. Existing interpretability methods mainly explain a model behavior by identifying salient image patches, which require manual efforts from users to make sense of, and also do not typically support model validation with questions that investigate multiple visual concepts. In this paper, we introduce a scalable human-in-the-loop approach for global interpretability. Salient image areas identified by local interpretability methods are annotated with semantic concepts, which are then aggregated into a tabular representation of images to facilitate automatic statistical analysis of model behavior. We show that this approach answers interpretability needs for both model validation and exploration, and provides semantically more diverse, informative, and relevant explanations while still allowing for scalable and cost-efficient execution.
Agathe Balayn, Panagiotis Soilis, Christoph Lofi, Jie Yang 0028, Alessandro Bozzon
WWW3
2021 Managing bias and unfairness in data for decision support: a survey of machine learning and data engineering approaches to identify and mitigate bias and unfairness within data management and analytics systems
abstract
Abstract The increasing use of data-driven decision support systems in industry and governments is accompanied by the discovery of a plethora of bias and unfairness issues in the outputs of these systems. Multiple computer science communities, and especially machine learning, have started to tackle this problem, often developing algorithmic solutions to mitigate biases to obtain fairer outputs. However, one of the core underlying causes for unfairness is bias in training data which is not fully covered by such approaches. Especially, bias in data is not yet a central topic in data engineering and management research. We survey research on bias and unfairness in several computer science domains, distinguishing between data management publications and other domains. This covers the creation of fairness metrics, fairness identification, and mitigation methods, software engineering approaches and biases in crowdsourcing activities. We identify relevant research gaps and show which data management activities could be repurposed to handle biases and which ones might reinforce such biases. In the second part, we argue for a novel data-centered approach overcoming the limitations of current algorithmic-centered methods. This approach focuses on eliciting and enforcing fairness requirements and constraints on data that systems are trained, validated, and used on. We argue for the need to extend database management systems to handle such constraints and mitigation methods. We discuss the associated future research directions regarding algorithms, formalization, modelling, users, and systems.
Agathe Balayn, Christoph Lofi, Geert-Jan Houben
VLDB J.2
2020 LOREM: Language-consistent Open Relation Extraction from Unstructured Text
abstract
We introduce a Language-consistent multi-lingual Open Relation Extraction Model (LOREM) for finding relation tuples of any type between entities in unstructured texts. LOREM does not rely on language-specific knowledge or external NLP tools such as translators or PoS-taggers, and exploits information and structures that are consistent over different languages. This allows our model to be easily extended with only limited training efforts to new languages, but also provides a boost to performance for a given single language. An extensive evaluation performed on 5 languages shows that LOREM outperforms state-of-the-art mono-lingual and cross-lingual open relation extractors. Moreover, experiments on languages with no or only little training data indicate that LOREM generalizes to other languages than the languages that it is trained on.
Tom Harting, Sepideh Mesbah, Christoph Lofi
WWW3
2019 Training Data Augmentation for Detecting Adverse Drug Reactions in User-Generated Content
abstract
Sepideh Mesbah, Jie Yang, Robert-Jan Sips, Manuel Valle Torre, Christoph Lofi, Alessandro Bozzon, Geert-Jan Houben. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sepideh Mesbah, Jie Yang 0028, Robert-Jan Sips, Manuel Valle Torre, Christoph Lofi, Alessandro Bozzon, Geert-Jan Houben
EMNLP/IJCNLP (1)5
2019 Coner: A Collaborative Approach for Long-Tail Named Entity Recognition in Scientific Publications
Daniel Vliegenthart, Sepideh Mesbah, Christoph Lofi, Akiko Aizawa, Alessandro Bozzon
TPDL3
2019 Evaluating Neural Text Simplification in the Medical Domain
abstract
Health literacy, i.e. the ability to read and understand medical text, is a relevant component of public health. Unfortunately, many medical texts are hard to grasp by the general population as they are targeted at highly-skilled professionals and use complex language and domain-specific terms. Here, automatic text simplification making text commonly understandable would be very beneficial. However, research and development into medical text simplification is hindered by the lack of openly available training and test corpora which contain complex medical sentences and their aligned simplified versions. In this paper, we introduce such a dataset to aid medical text simplification research. The dataset is created by filtering aligned health sentences using expert knowledge from an existing aligned corpus and a novel simple, language independent monolingual text alignment method. Furthermore, we use the dataset to train a state-of-the-art neural machine translation model, and compare it to a model trained on a general simplification dataset using an automatic evaluation, and an extensive human-expert evaluation.
Laurens Van den Bercken, Robert-Jan Sips, Christoph Lofi
WWW3
2018 Concept Focus: Semantic Meta-Data for Describing MOOC Content
Sepideh Mesbah, Guanliang Chen, Manuel Valle Torre, Alessandro Bozzon, Christoph Lofi, Geert-Jan Houben
EC-TEL5
2018 Can I Have a Mooc2Go, Please? On the Viability of Mobile vs. Stationary Learning
Yue Zhao 0001, Tarmo Robal, Christoph Lofi, Claudia Hauff
EC-TEL3
2018 Webcam-based Attention Tracking in Online Learning: A Feasibility Study
abstract
A main weakness of the open online learning movement is retention: a small minority of learners (on average 5-10%, in extreme cases <1%) that start a so-called Massive Open Online Course (MOOC) complete it successfully. There are many reasons why learners are unsuccessful, among the most important ones is the lack of self-regulation: learners are often not able to self-regulate their learning behavior. Designing tools that provide learners with a greater awareness of their learning is vital to the future success of MOOC environments. Detecting learners' loss of focus during learning is particularly important, as this can allow us to intervene and return the learners' attention to the learning materials. One technological affordance to detect such loss of focus are webcams---ubiquitous pieces of hardware available in almost all laptops today. In recent years, researchers have begun to exploit eye tracking and gaze data generated from webcams as part of complex machine learning solutions to detect inattention or loss of focus. Those approaches however tend to have a high detection lag, can be inaccurate, and are complex to design and maintain. In contrast, in this paper, we explore the possibility of a simple alternative---the presence or absence of a face---to detect a loss of focus in the online learning setting. To this end, we evaluate the performance of three consumer and professional eye/face-tracking frameworks using a benchmark suite we designed specifically for this purpose: it contains a set of common xMOOC user activities and behaviours. The results of our study show that even this basic approach poses a significant challenge to current hardware and software-based tracking solutions.
Tarmo Robal, Yue Zhao 0001, Christoph Lofi, Claudia Hauff
IUI3
2018 TSE-NER: An Iterative Approach for Long-Tail Entity Extraction in Scientific Publications
Sepideh Mesbah, Christoph Lofi, Manuel Valle Torre, Alessandro Bozzon, Geert-Jan Houben
ISWC (1)2
2017 Scalable Mind-Wandering Detection for MOOCs: A Webcam-Based Approach
Yue Zhao 0001, Christoph Lofi, Claudia Hauff
EC-TEL2
2017 Facet Embeddings for Explorative Analytics in Digital Libraries
Sepideh Mesbah, Kyriakos Fragkeskos, Christoph Lofi, Alessandro Bozzon, Geert-Jan Houben
TPDL3
2017 Semantic Annotation of Data Processing Pipelines in Scientific Publications
Sepideh Mesbah, Kyriakos Fragkeskos, Christoph Lofi, Alessandro Bozzon, Geert-Jan Houben
ESWC (1)3
2017 Sequences of Diverse Song Recommendations: An Exploratory Study in a Commercial System
abstract
This paper presents an exploratory study of the perceptions users have of diversity and ordering in playlist recommendations. There is a match between the diversification approach used in the system, and importance that users placed on the item properties. Surprisingly, participants had no expectations of the songs being in a particular order in a playlist. We discuss possible explanations for this finding, refining the research agenda to consider which ordering choices are perceptible to users, and influence user satisfaction.
Nava Tintarev, Christoph Lofi, Cynthia C. S. Liem
UMAP2
2016 Benchmarking Semantic Capabilities of Analogy Querying Algorithms
Christoph Lofi, Athiq Ahamed, Pratima Kulkarni, Ravi Thakkar
DASFAA (1)1
2015 "I would like to watch something like 'The Terminator'..." Cooperative Query Personalization Based on Perceptual Similarity
abstract
In this paper, we showcase a privacy-preserving query personalization system for experience items like movies, music, games, or books. Personalizing queries for such items is notoriously difficult as meaningful query attributes are either missing in the database or would require extensive domain knowledge not available to most users. For this reason, state-of-the-art content provision platforms as e.g., Netflix or Amazon usually rely on recommender systems to support their users, and are often working in parallel with traditional SQL-style queries. Unfortunately, recommender systems have several shortcomings as for example high barriers for new users joining the system, which first have to setup a preference profile in a lengthy process, the inability to pose meaningful queries beyond recommendations matching the personal profile, and severe privacy concerns due to storing personal rating data for all users long-term. In order to provide an alternative, we present in this demonstration paper a powerful and intuitive query-by-example (QBE) interaction system. Bayesian Navigation is used to personalize a user’s query on the fly. The central challenge when using QBE is the selection of features to represent the items in the database. Here, we rely on a high-dimensional feature space which was mined from rating data of a large number of users, allowing us to measure perceived similarity between items to steer the query process. This also addresses many issues of recommender systems as our query capabilities can be used by any user anonymously in a drive-by fashion. In our proposed demo, users can try our never before presented system hands-on, and can use it to discover interesting movies tailored to their preferences with a pleasantly simple and enjoyable user experience.
Christoph Lofi, Christian Nieke
EDBT1
2015 Towards Narrative Information Systems
Philipp Wille, Christoph Lofi, Wolf-Tilo Balke
WAIM2
2015 Crowdsourcing Twitter annotations to identify first-hand experiences of prescription drug use
abstract
Self-reported patient data has been shown to be a valuable knowledge source for post-market pharmacovigilance. In this paper we propose using the popular micro-blogging service Twitter to gather evidence about adverse drug reactions (ADRs) after firstly having identified micro-blog messages (also know as "tweets") that report first-hand experience. In order to achieve this goal we explore machine learning with data crowdsourced from laymen annotators. With the help of lay annotators recruited from CrowdFlower we manually annotated 1548 tweets containing keywords related to two kinds of drugs: SSRIs (eg. Paroxetine), and cognitive enhancers (eg. Ritalin). Our results show that inter-annotator agreement (Fleiss' kappa) for crowdsourcing ranks in moderate agreement with a pair of experienced annotators (Spearman's Rho=0.471). We utilized the gold standard annotations from CrowdFlower for automatically training a range of supervised machine learning models to recognize first-hand experience. F-Score values are reported for 6 of these techniques with the Bayesian Generalized Linear Model being the best (F-Score=0.64 and Informedness=0.43) when combined with a selected set of features obtained by using information gain criteria.
Nestor Alvaro, Mike Conway, Son Doan, Christoph Lofi, John P. Overington, Nigel Collier
J. Biomed. Informatics4
2014 Discriminating Rhetorical Analogies in Social Media
abstract
Analogies are considered to be one of the core concepts of human cognition and communication, and are very efficient at encoding complex information in a natural fashion.However, computational approaches towards largescale analysis of the semantics of analogies are hampered by the lack of suitable corpora with real-life example of analogies.In this paper we therefore propose a workflow for discriminating and extracting natural-language analogy statements from the Web, focusing on analogies between locations mined from travel reports, blogs, and the Social Web.For realizing this goal, we employ feature-rich supervised learning models which we extensively evaluate.We also showcase a crowd-supported workflow for building a suitable Gold dataset used for this purpose.The resulting system is able to successfully learn to identify analogies to a high degree of accuracy (F-Score 0.9) by using a high-dimensional subsequence feature space.
Christoph Lofi, Christian Nieke, Nigel Collier
EACL1
2014 Exploiting Perceptual Similarity: Privacy-Preserving Cooperative Query Personalization
Christoph Lofi, Christian Nieke
WISE (1)1
2013 Skyline queries in crowd-enabled databases
abstract
Skyline queries are a well-established technique for database query personalization and are widely acclaimed for their intuitive query formulation mechanisms. However, when operating on incomplete datasets, skylines queries are severely hampered and often have to resort to highly error-prone heuristics. Unfortunately, incomplete datasets are a frequent phenomenon, especially when datasets are generated automatically using various information extraction or information integration approaches. Here, the recent trend of crowd-enabled databases promises a powerful solution: during query execution, some database operators can be dynamically outsourced to human workers in exchange for monetary compensation, therefore enabling the elicitation of missing values during runtime. Unfortunately, this powerful feature heavily impacts query response times and (monetary) execution costs. In this paper, we present an innovative hybrid approach combining dynamic crowd-sourcing with heuristic techniques in order to overcome current limitations. We will show that by assessing the individual risk a tuple poses with respect to the overall result quality, crowd-sourcing efforts for eliciting missing values can be narrowly focused on only those tuples that may degenerate the expected quality most strongly. This leads to an algorithm for computing skyline sets on incomplete data with maximum result quality, while optimizing crowd-sourcing costs.
Christoph Lofi, Kinda El Maarry, Wolf-Tilo Balke
EDBT1
2013 Skyline Queries over Incomplete Data - Error Models for Focused Crowd-Sourcing
Christoph Lofi, Kinda El Maarry, Wolf-Tilo Balke
ER1
2012 Malleability-Aware Skyline Computation on Linked Open Data
Christoph Lofi, Ulrich Güntzer, Wolf-Tilo Balke
DASFAA (2)1
2012 iParticipate: Automatic Tweet Generation from Local Government Data
Christoph Lofi, Ralf Krestel
DASFAA (2)1
2012 Pushing the Boundaries of Crowd-enabled Databases with Query-driven Schema Expansion
abstract
By incorporating human workers into the query execution process crowd-enabled databases facilitate intelligent, social capabilities like completing missing data at query time or performing cognitive operators. But despite all their flexibility, crowd-enabled databases still maintain rigid schemas. In this paper, we extend crowd-enabled databases by flexible query-driven schema expansion, allowing the addition of new attributes to the database at query time. However, the number of crowd-sourced mini-tasks to fill in missing values may often be prohibitively large and the resulting data quality is doubtful. Instead of simple crowd-sourcing to obtain all values individually, we leverage the usergenerated data found in the Social Web: By exploiting user ratings we build perceptual spaces , i.e., highly-compressed representations of opinions, impressions, and perceptions of large numbers of users. Using few training samples obtained by expert crowd sourcing, we then can extract all missing data automatically from the perceptual space with high quality and at low costs. Extensive experiments show that our approach can boost both performance and quality of crowd-enabled databases, while also providing the flexibility to expand schemas in a query-driven fashion.
Joachim Selke, Christoph Lofi, Wolf-Tilo Balke
Proc. VLDB Endow.2
2010 Highly Scalable Multiprocessing Algorithms for Preference-Based Database Retrieval
Joachim Selke, Christoph Lofi, Wolf-Tilo Balke
DASFAA (2)2
2010 Efficient computation of trade-off skylines
abstract
When selecting alternatives from large amounts of data, trade-offs play a vital role in everyday decision making. In databases this is primarily reflected by the top-k retrieval paradigm. But recently it has been convincingly argued that it is almost impossible for users to provide meaningful scoring functions for top-k retrieval, subsequently leading to the adoption of the skyline paradigm. Here users just specify the relevant attributes in a query and all suboptimal alternatives are filtered following the Pareto semantics. Up to now the intuitive concept of compensation, however, cannot be used in skyline queries, which also contributes to the often unmanageably large result set sizes. In this paper we discuss an innovative and efficient method for computing skylines allowing the use of qualitative trade-offs. Such trade-offs compare examples from the database on a focused subset of attributes. Thus, users can provide information on how much they are willing to sacrifice to gain an improvement in some other attribute(s). Our contribution is the design of the first skyline algorithm allowing for qualitative compensation across attributes. Moreover, we also provide an novel trade-off representation structure to speed up retrieval. Indeed our experiments show efficient performance allowing for focused skyline sets in practical applications. Moreover, we show that the necessary amount of object comparisons can be sped up by an order of magnitude using our indexing techniques.
Christoph Lofi, Ulrich Güntzer, Wolf-Tilo Balke
EDBT1
2009 Efficient Skyline Refinement using Trade-Offs
abstract
Skyline queries have received a lot of attention due to their intuitive query formulation. Following the concept of Pareto optimality all dasiabestpsila database items satisfying different aspects of the query are returned to the user. However, this often results in huge result set sizes. In everyday's life users face the same problem. But here, when confronted with a too large variety of choices users tend to focus only on some aspects of the attribute space at a time and try to figure out acceptable compromises between these attributes. Such trade-offs are not reflected by the Pareto paradigm. Incorporating them into user preferences and adjusting skyline results accordingly thus needs special algorithms beyond traditional skylining. In this paper we propose a novel algorithm for efficiently incorporating such typical trade-off information into preference orders. Our experiments on both real world and synthetic data sets show the impact of our techniques: impractical skyline sizes efficiently become manageable with a minimum amount of user interaction. Additionally, we also design a method to elicit especially interesting trade-offs promising a high reduction of skyline sizes. At any point, the user can choose whether to provide individual trade-offs, or accept those suggested by the system. The benefit of incorporating trade-offs into the strict Pareto semantics is clear: result sets become manageable, while additionally getting more focused on the users' information needs.
Christoph Lofi, Wolf-Tilo Balke, Ulrich Güntzer
RCIS1
2008 Efficiently performing consistency checks for multi-dimensional preference trade-offs
abstract
Skyline queries have recently received a lot of attention due to their intuitive query capabilities. Following the concept of Pareto optimality all dasiabestpsila database objects are returned to the user. However, this often results in unmanageable large result set sizes hampering the success of this innovative paradigm. As an effective remedy for this problem, trade-offs provide a natural concept for dealing with incomparable choices. Such trade-offs, however, are not reflected by the Pareto paradigm. Thus, incorporating them into the userspsila preference orders and adjusting skyline results accordingly needs special algorithms beyond traditional skylining. For the actual integration of trade-offs into skylines, the problem of ensuring the consistency of arbitrary trade-off sets poses a demanding challenge. Consistency is a crucial aspect when dealing with multi-dimensional trade-offs spanning over several attributes. If the consistency should be violated, cyclic preferences may occur in the result set. But such cyclic preferences cannot be resolved by information systems in a sensible way. Often, this problem is circumvented by restricting the trade-offspsila expressiveness, e.g. by altogether ignoring some classes of possibly inconsistent trade-offs. In this paper, we will present a new algorithm capable of efficiently verifying the consistency of any arbitrary set of trade-offs. After motivating its basic concepts and introducing the algorithm itself, we will also show that it exhibits superior average-case performance. The benefits of our approach promise to pave the way towards personalized and cooperative information systems.
Christoph Lofi, Wolf-Tilo Balke, Ulrich Güntzer
RCIS1
2007 Eliciting Matters - Controlling Skyline Sizes by Incremental Integration of User Preferences
Wolf-Tilo Balke, Ulrich Güntzer, Christoph Lofi
DASFAA3
2007 User Interaction Support for Incremental Refinement of Preference-Based Queries
Wolf-Tilo Balke, Ulrich Güntzer, Christoph Lofi
RCIS3
2007 A Model for Competence Gap Analysis
Juri Luca De Coi, Eelco Herder, Arne Wolf Koesling, Christoph Lofi, Daniel Olmedilla, Odysseas Papapetrou, Wolf Siberski
WEBIST (3)4