EDBT 2026 Demo / reviewers in the wild / expert
Michaela Geierhos
dblp:51/4677
· DBLP profile ↗
20ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0002-8180-5606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Curation of Benchmark Templates for Measuring Gender Bias in Named Entity Recognition ModelsabstractNamed Entity Recognition (NER) constitutes a popular machine learning technique that empowers several natural language processing applications. As with other machine learning applications, NER models have been shown to be susceptible to gender bias. The latter is often assessed using benchmark datasets, which in turn are curated specifically for a given Natural Language Processing (NLP) task. In this work, we investigate the robustness of benchmark templates to detect gender bias and propose a novel method to improve the curation of such datasets. The method, based on masked token prediction, aims to filter out benchmark templates with a higher probability of detecting gender bias in NER models. We tested the method for English and German, using the corresponding fine-tuned BERT base model (cased) as the NER model. The gender gaps detected with templates classified as appropriate by the method were statistically larger than those detected with inappropriate templates. The results were similar for both languages and support the use of the proposed method in the curation of templates designed to detect gender bias. Ana Cimitan, Ana Alves-Pinto, Michaela Geierhos |
LREC/COLING | 3 |
| 2024 | Combining Frequency-Based Smoothing and Salient Masking for Performant and Imperceptible Adversarial Samples
Amon Soares de Souza, Andreas Meißner, Michaela Geierhos |
ICPR (22) | 3 |
| 2023 | Towards comparable ratings: Exploring bias in German physician reviewsabstractIn this study, we evaluate the impact of gender-biased data from German-language physician reviews on the fairness of fine-tuned language models. For two different downstream tasks, we use data reported to be gender biased and aggregate it with annotations. First, we propose a new approach to aspect-based sentiment analysis that allows identifying, extracting, and classifying implicit and explicit aspect phrases and their polarity within a single model. The second task we present is grade prediction, where we predict the overall grade of a review on the basis of the review text. For both tasks, we train numerous transformer models and evaluate their performance. The aggregation of sensitive attributes, such as a physician’s gender and migration background, with individual text reviews allows us to measure the performance of the models with respect to these sensitive groups. These group-wise performance measures act as extrinsic bias measures for our downstream tasks. In addition, we translate several gender-specific templates of the intrinsic bias metrics into the German language and evaluate our fine-tuned models. Based on this set of tasks, fine-tuned models, and intrinsic and extrinsic bias measures, we perform correlation analyses between intrinsic and extrinsic bias measures. In terms of sensitive groups and effect sizes, our bias measure results show different directions. Furthermore, correlations between measures of intrinsic and extrinsic bias can be observed in different directions. This leads us to conclude that gender-biased data does not inherently lead to biased models. Other variables, such as template dependency for intrinsic measures and label distribution in the data, must be taken into account as they strongly influence the metric results. Therefore, we suggest that metrics and templates should be chosen according to the given task and the biases to be assessed. Joschka Kersting, Falk Maoro, Michaela Geierhos |
Data Knowl. Eng. | 3 |
| 2023 | The problem of varying annotations to identify abusive language in social media contentabstractAbstract With the increase of user-generated content on social media, the detection of abusive language has become crucial and is therefore reflected in several shared tasks that have been performed in recent years. The development of automatic detection systems is desirable, and the classification of abusive social media content can be solved with the help of machine learning. The basis for successful development of machine learning models is the availability of consistently labeled training data. But a diversity of terms and definitions of abusive language is a crucial barrier. In this work, we analyze a total of nine datasets—five English and four German datasets—designed for detecting abusive online content. We provide a detailed description of the datasets, that is, for which tasks the dataset was created, how the data were collected, and its annotation guidelines. Our analysis shows that there is no standard definition of abusive language, which often leads to inconsistent annotations. As a consequence, it is difficult to draw cross-domain conclusions, share datasets, or use models for other abusive social media language tasks. Furthermore, our manual inspection of a random sample of each dataset revealed controversial examples. We highlight challenges in data annotation by discussing those examples, and present common problems in the annotation process, such as contradictory annotations and missing context information. Finally, to complement our theoretical work, we conduct generalization experiments on three German datasets. Nina Seemann, Yeong Su Lee, Julian Höllig, Michaela Geierhos |
Nat. Lang. Eng. | 4 |
| 2022 | Keep It Simple: Local Search-based Latent Space Editing
Andreas Meißner, Andreas Fröhlich, Michaela Geierhos |
IJCCI | 3 |
| 2021 | Well-Being in Plastic Surgery: Deep Learning Reveals Patients' Evaluations
Joschka Kersting, Michaela Geierhos |
DATA | 2 |
| 2021 | Influence of Training Data on the Invertability of Neural Networks for Handwritten Digit RecognitionabstractModel inversion attacks aim to extract details of training data from a trained model, potentially revealing sensitive information about a person’s identity. To abide with protection of personal privacy requirements, it is important to understand the mechanisms that increase the privacy of training data. In this work, we systematically investigated the impact of the training data on a model’s susceptibility to model inversion attacks for models trained at the task of hand-written digit recognition with the openly available MNIST dataset. Using an optimization-based inversion approach, we studied the impacts of the quantity and diversity of training data, and the number and selection of classes on the susceptibility of models to inversion. Our model inversion attack strategy was less successful for models with a larger number of training data and greater training data diversity. Moreover, atypical training records provided additional protection against model inversion. We discovered that not every class was equally susceptible to model inversion attacks and that the inversion results of one class were changed when models were trained with a different selection of classes. However, we did not detect a clear relationship between the number of classes and a model’s susceptibility to inversion. Our study shows that the inversion susceptibility of a model depends on the training data-not only the data used to train the class that is inverted, but also the data used to train the other classes. Antonia Adler, Michaela Geierhos, Eleanor Hobley |
ICMLA | 2 |
| 2021 | Semantic Text Segment Classification of Structured Technical Content
Julian Höllig, Philipp Dufter, Michaela Geierhos, Wolfgang Ziegler, Hinrich Schütze |
NLDB | 3 |
| 2021 | Human Language Comprehension in Aspect Phrase Extraction with Importance Weighting
Joschka Kersting, Michaela Geierhos |
NLDB | 2 |
| 2020 | Aspect Phrase Extraction in Sentiment Analysis with Deep LearningabstractThis paper deals with aspect phrase extraction and classification in sentiment analysis. We summarize current approaches and datasets from the domain of aspect-based sentiment analysis. This domain detects sentiments expressed for individual aspects in unstructured text data. So far, mainly commercial user reviews for products or services such as restaurants were investigated. We here present our dataset consisting of German physician reviews, a sensitive and linguistically complex field. Furthermore, we describe the annotation process of a dataset for supervised learning with neural networks. Moreover, we introduce our model for extracting and classifying aspect phrases in one step, which obtains an F1-score of 80%. By applying it to a more complex domain, our approach and results outperform previous approaches. Joschka Kersting, Michaela Geierhos |
ICAART (1) | 2 |
| 2020 | Detection of Privacy Disclosure in the Medical Domain: A SurveyabstractWhen it comes to increased digitization in the health care domain, privacy is a relevant topic nowadays. This relates to patient data, electronic health records or physician reviews published online, for instance. There exist different approaches to the protection of individuals’ privacy, which focus on the anonymization and masking of personal information subsequent to their mining. In the medical domain in particular, measures to protect the privacy of patients are of high importance due to the amount of sensitive data that is involved (e.g. age, gender, illnesses, medication). While privacy breaches in structured data can be detected more easily, disclosure in written texts is more difficult to find automatically due to the unstructured nature of natural language. Therefore, we take a detailed look at existing research on areas related to privacy protection. Likewise, we review approaches to the automatic detection of privacy disclosure in different types of medical data. We provide a survey of several studies concerned with privacy breaches in the medical domain with a focus on Physician Review Websites (PRWs). Finally, we briefly develop implications and directions for further research. Bianca Buff, Joschka Kersting, Michaela Geierhos |
ICPRAM | 3 |
| 2020 | What Reviews in Local Online Labour Markets Reveal about the Performance of Multi-service ProvidersabstractThis paper deals with online customer reviews of local multi-service providers. While many studies investigate product reviews and online labour markets with service providers delivering intangible products “over the wire”, we focus on websites where providers offer multiple distinct services that can be booked, paid and reviewed online but are performed locally offline. This type of service providers has so far been neglected in the literature. This paper analyses reviews and applies sentiment analysis. It aims to gain new insights into local multi-service providers’ performance. There is a broad literature range presented with regard to the topics addressed. The results show, among other things, that providers with good ratings continue to perform well over time. We find that many positive reviews seem to encourage sales. On average, quantitative star ratings and qualitative ratings in the form of review texts match. Further results are also achieved in this study. Joschka Kersting, Michaela Geierhos |
ICPRAM | 2 |
| 2019 | In Reviews We Trust: But Should We? Experiences with Physician Review WebsitesabstractThe ability to openly evaluate products, locations and services is an achievement of the Web 2.0. It has never been easier to inform oneself about the quality of products or services and possible alternatives. Forming one’s own opinion based on the impressions of other people can lead to better experiences. However, this presupposes trust in one’s fellows as well as in the quality of the review platforms. In previous work on physician reviews and the corresponding websites, it was observed that there occurs faulty behavior by some reviewers and there were noteworthy differences in the technical implementation of the portals and in the efforts of site operators to maintain high quality reviews. These experiences raise new questions regarding what trust means on review platforms, how trust arises and how easily it can be destroyed. Joschka Kersting, Frederik S. Bäumer, Michaela Geierhos |
IoTBDS | 3 |
| 2018 | How to Deal with Inaccurate Service Descriptions in On-The-Fly Computing: Open Challenges
Frederik S. Bäumer, Michaela Geierhos |
NLDB | 2 |
| 2017 | Internet of Things Architecture for Handling Stream Air Pollution DataabstractIn this paper, we present an IoT architecture which handles stream sensor data of air pollution. Particle pollution is known as a serious threat to human health. Along with developments in the use of wireless sensors and the IoT, we propose an architecture that flexibly measures and processes stream data collected in real-time by movable and low-cost IoT sensors. Thus, it enables a wide-spread network of wireless sensors that can follow changes in human behavior. Apart from stating reasons for the need of such a development and its requirements, we provide a conceptual design as well as a technological design of such an architecture. The technological design consists of Kaa and Apache Storm which can collect air pollution information in real-time and solve various problems to process data such as missing data and synchronization. This enables us to add a simulation in which we provide issues that might come up when having our architecture in use. Together with these issues, we state r easons for choosing specific modules among candidates. Our architecture combines wireless sensors with the Kaa IoT framework, an Apache Kafka pipeline and an Apache Storm Data Stream Management System among others. We even provide open-government data sets that are freely available. Joschka Kersting, Michaela Geierhos, Hanmin Jung, Taehong Kim |
IoTBDS | 2 |
| 2016 | On- and Off-Topic Classification and Semantic Annotation of User-Generated Software RequirementsabstractUsers prefer natural language software requirements because of their usability and accessibility.When they describe their wishes for software development, they often provide off-topic information.We therefore present REaCT 1 , an automated approach for identifying and semantically annotating the on-topic parts of requirement descriptions.It is designed to support requirement engineers in the elicitation process on detecting and analyzing requirements in user-generated content.Since no lexical resources with domain-specific information about requirements are available, we created a corpus of requirements written in controlled language by instructed users and uncontrolled language by uninstructed users.We annotated these requirements regarding predicate-argument structures, conditions, priorities, motivations and semantic roles and used this information to train classifiers for information extraction purposes.REaCT achieves an accuracy of 92% for the on-and off-topic classification task and an F 1measure of 72% for the semantic annotation. Markus Dollmann, Michaela Geierhos |
EMNLP | 2 |
| 2016 | How to Complete Customer Requirements - Using Concept Expansion for Requirement Refinement
Michaela Geierhos, Frederik S. Bäumer |
NLDB | 1 |
| 2015 | What Did You Mean? - Facing the Challenges of User-generated Software Requirements
Michaela Geierhos, Sabine Schulze, Frederik S. Bäumer |
ICAART (1) | 1 |
| 2015 | Filtering Reviews by Random Individual Error
Michaela Geierhos, Frederik S. Bäumer, Sabine Schulze, Valentina Stuß |
IEA/AIE | 1 |
| 2009 | Business Specific Online Information Extraction from German Websites
Yeong Su Lee, Michaela Geierhos |
CICLing | 2 |