Roman Klinger

dblp:21/4183 · DBLP profile ↗
← Back
52ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0002-2014-6619ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 3 first-author · 22 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Categorical Emotions or Appraisals - Which Emotion Model Explains Argument Convincingness Better?
abstract
The convincingness of an argument does not only depend on its structure (logos), the person who makes the argument (ethos), but also on the emotion that it causes in the recipient (pathos). While the overall intensity and categorical values of emotions in arguments have received considerable attention in the research community, we argue that the emotion an argument evokes in a recipient is subjective. It depends on the recipient's goals, standards, prior knowledge, and stance. Appraisal theories lend themselves as a link between the subjective cognitive assessment of events and emotions. They have been used in event-centric emotion analysis, but their suitability for assessing argument convincingness remains unexplored. In this paper, we evaluate whether appraisal theories are suitable for emotion analysis in arguments by considering subjective cognitive evaluations of the importance and impact of an argument on its receiver. Based on the annotations in the recently published ContArgA corpus, we perform zero-shot prompting experiments to evaluate the importance of gold-annotated and predicted emotions and appraisals for the assessment of the subjective convincingness labels. We find that, while categorical emotion information does improve convincingness prediction, the improvement is more pronounced with appraisals. This work presents the first systematic comparison between emotion models for convincingness prediction, demonstrating the advantage of appraisals, providing insights for theoretical and practical applications in computational argumentation.
Lynn Greschner, Meike Bauer, Sabine Weber, Roman Klinger
LREC4
2026 Trust Me, I Can Convince You: The Contextualized Argument Appraisal Framework and the ContArgA Corpus
Lynn Greschner, Sabine Weber, Roman Klinger
LREC3
2026 PARL: Prompt-based Agents for Reinforcement Learning
Yarik Menchaca Resendiz, Roman Klinger
LREC2
2026 Entity-Level Sentiment Analysis with Sentence Relevance Detection
Egil Rønningstad, Roman Klinger, Lilja Øvrelid, Erik Velldal
LREC2
2026 Disambiguation of Emotion Annotations by Contextualizing Events in Plausible Narratives
Johannes Schäfer, Roman Klinger
LREC2
2026 Less Is More? The Role of Demographic Author Information in Emotion Classification of Ambiguous Text
Sabine Weber, Lynn Greschner, Roman Klinger
LREC3
2025 Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts
abstract
Accurate modeling of subjective phenomena such as emotion expression requires data annotated with authors’ intentions. Commonly such data is collected by asking study participants to donate and label genuine content produced in the real world, or create content fitting particu- lar labels during the study. Asking participants to create content is often simpler to implement and presents fewer risks to participant privacy than data donation. However, it is unclear if and how study-created content may differ from genuine content, and how differences may impact models. We collect study-created and genuine multimodal social media posts labeled for emotion and compare them on several dimen- sions, including model performance. We find that compared to genuine posts, study-created posts are longer, rely more on their text and less on their images for emotion expression, and focus more on emotion-prototypical events. The samples of participants willing to donate versus create posts are demographically different. Study-created data is valuable to train models that generalize well to genuine data, but realistic effectiveness estimates require genuine data.
Christopher Bagdon, Aidan Combs, Carina Silberer, Roman Klinger
ACL (1)4
2025 Which Demographics do LLMs Default to During Annotation?
abstract
Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li, Nadine Probol, Lynn Greschner, Sean Papay, Yarik Menchaca Resendiz, Aswathy Velutharambath, Amelie Wuehrl, Sabine Weber, Roman Klinger. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Johannes Schäfer, Aidan Combs, Christopher Bagdon, Nadine Probol, Lynn Greschner, Sean Papay, Yarik Menchaca Resendiz, Aswathy Velutharambath, Amelie Wührl, Sabine Weber, Roman Klinger
ACL (1)12
2025 MOPO: Multi-Objective Prompt Optimization for Affective Text Generation
abstract
How emotions are expressed depends on the context and domain. On X (formerly Twitter), for instance, an author might simply use the hashtag #anger, while in a news headline, emotions are typically written in a more polite, indirect manner. To enable conditional text generation models to create emotionally connotated texts that fit a domain, users need to have access to a parameter that allows them to choose the appropriate way to express an emotion. To achieve this, we introduce MOPO, a Multi-Objective Prompt Optimization methodology. MOPO optimizes prompts according to multiple objectives (which correspond here to the output probabilities assigned by emotion classifiers trained for different domains). In contrast to single objective optimization, MOPO outputs a set of prompts, each with a different weighting of the multiple objectives. Users can then choose the most appropriate prompt for their context. We evaluate MOPO using three objectives, determined by various domain-specific emotion classifiers. MOPO improves performance by up to 15 pp across all objectives with a minimal loss (1–2 pp) for any single objective compared to single-objective optimization. These minor performance losses are offset by a broader generalization across multiple objectives – which is not possible with single-objective optimization. Additionally, MOPO reduces computational requirements by simultaneously optimizing for multiple objectives, eliminating separate optimization procedures for each objective.
Yarik Menchaca Resendiz, Roman Klinger
COLING2
2024 Can Factual Statements Be Deceptive? The DeFaBel Corpus of Belief-based Deception
abstract
If a person firmly believes in a non-factual statement, such as “The Earth is flat”, and argues in its favor, there is no inherent intention to deceive. As the argumentation stems from genuine belief, it may be unlikely to exhibit the linguistic properties associated with deception or lying. This interplay of factuality, personal belief, and intent to deceive remains an understudied area. Disentangling the influence of these variables in argumentation is crucial to gain a better understanding of the linguistic properties attributed to each of them. To study the relation between deception and factuality, based on belief, we present the DeFaBel corpus, a crowd-sourced resource of belief-based deception. To create this corpus, we devise a study in which participants are instructed to write arguments supporting statements like “eating watermelon seeds can cause indigestion”, regardless of its factual accuracy or their personal beliefs about the statement. In addition to the generation task, we ask them to disclose their belief about the statement. The collected instances are labelled as deceptive if the arguments are in contradiction to the participants’ personal beliefs. Each instance in the corpus is thus annotated (or implicitly labelled) with personal beliefs of the author, factuality of the statement, and the intended deceptiveness. The DeFaBel corpus contains 1031 texts in German, out of which 643 are deceptive and 388 are non-deceptive. It is the first publicly available corpus for studying deception in German. In our analysis, we find that people are more confident in the persuasiveness of their arguments when the statement is aligned with their belief, but surprisingly less confident when they are generating arguments in favor of facts. The DeFaBel corpus can be obtained from https://www.ims.uni-stuttgart.de/data/defabel .
Aswathy Velutharambath, Roman Klinger, Amelie Wührl
LREC/COLING2
2024 EmoProgress: Cumulated Emotion Progression Analysis in Dreams and Customer Service Dialogues
abstract
Emotion analysis often involves the categorization of isolated textual units, but these are parts of longer discourses, like dialogues or stories. This leads to two different established emotion classification setups: (1) Classification of a longer text into one or multiple emotion categories. (2) Classification of the parts of a longer text (sentences or utterances), either (2a) with or (2b) without consideration of the context. None of these settings, does, however, enable to answer the question which emotion is presumably experienced at a specific moment in time. For instance, a customer’s request of “My computer broke.” would be annotated with anger. This emotion persists in a potential follow-up reply “It is out of warranty.” which would also correspond to the global emotion label. An alternative reply “We will send you a new one.” might, in contrast, lead to relief. Modeling these label relations requires classification of textual parts under consideration of the past, but without access to the future. Consequently, we propose a novel annotation setup for emotion categorization corpora, in which the annotations reflect the emotion up to the annotated sentence. We ensure this by uncovering the textual parts step-by-step to the annotator, asking for a label in each step. This perspective is important to understand the final, global emotion, while having access to the individual sentence’s emotion contributions to this final emotion. In modeling experiments, we use these data to check if the context is indeed required to automatically predict such cumulative emotion progressions.
Eileen Wemmer, Sofie Labat, Roman Klinger
LREC/COLING3
2024 What Makes Medical Claims (Un)Verifiable? Analyzing Entity and Relation Properties for Fact Verification
abstract
Amelie Wuehrl, Yarik Menchaca Resendiz, Lara Grimminger, Roman Klinger. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Amelie Wührl, Yarik Menchaca Resendiz, Lara Grimminger, Roman Klinger
EACL (1)4
2024 "You are an expert annotator": Automatic Best-Worst-Scaling Annotations for Emotion Intensity Modeling
abstract
Christopher Bagdon, Prathamesh Karmalkar, Harsha Gurulingappa, Roman Klinger. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Christopher Bagdon, Prathamesh Karmalkar, Harsha Gurulingappa, Roman Klinger
NAACL-HLT4
2023 Affective Natural Language Generation of Event Descriptions through Fine-grained Appraisal Conditions
abstract
Models for affective text generation have shown a remarkable progress, but they commonly rely only on basic emotion theories or valance/arousal values as conditions.This is appropriate when the goal is to create explicit emotion statements ("The kid is happy.").Emotions are, however, commonly communicated implicitly.For instance, the emotional interpretation of an event ("Their dog died.")does often not require an explicit emotion statement.In psychology, appraisal theories explain the link between a cognitive evaluation of an event and the potentially developed emotion.They put the assessment of the situation on the spot, for instance regarding the own control or the responsibility for what happens.We hypothesize and subsequently show that including appraisal variables as conditions in a generation framework comes with two advantages.(1) The generation model is informed in greater detail about what makes a specific emotion and what properties it has.This leads to text generation that better fulfills the condition.(2) The variables of appraisal allow a user to perform a more fine-grained control of the generated text, by stating properties of a situation instead of only providing the emotion category.Our Bart and T5-based experiments with 7 emotions (Anger, Disgust, Fear, Guilt, Joy, Sadness, Shame), and 7 appraisals (Attention, Responsibility, Control, Circumstance, Pleasantness, Effort, Certainty) show that (1) adding appraisals during training improves the accurateness of the generated texts by 10 pp in F 1 .Further, (2) the texts with appraisal variables are longer and contain more details.This exemplifies the greater control for users.
Yarik Menchaca Resendiz, Roman Klinger
INLG2
2023 Dimensional Modeling of Emotions in Text with Appraisal Theories: Corpus Creation, Annotation Reliability, and Prediction
abstract
Abstract The most prominent tasks in emotion analysis are to assign emotions to texts and to understand how emotions manifest in language. An important observation for natural language processing is that emotions can be communicated implicitly by referring to events alone, appealing to an empathetic, intersubjective understanding of events, even without explicitly mentioning an emotion name. In psychology, the class of emotion theories known as appraisal theories aims at explaining the link between events and emotions. Appraisals can be formalized as variables that measure a cognitive evaluation by people living through an event that they consider relevant. They include the assessment if an event is novel, if the person considers themselves to be responsible, if it is in line with their own goals, and so forth. Such appraisals explain which emotions are developed based on an event, for example, that a novel situation can induce surprise or one with uncertain consequences could evoke fear. We analyze the suitability of appraisal theories for emotion analysis in text with the goal of understanding if appraisal concepts can reliably be reconstructed by annotators, if they can be predicted by text classifiers, and if appraisal concepts help to identify emotion categories. To achieve that, we compile a corpus by asking people to textually describe events that triggered particular emotions and to disclose their appraisals. Then, we ask readers to reconstruct emotions and appraisals from the text. This set-up allows us to measure if emotions and appraisals can be recovered purely from text and provides a human baseline to judge a model’s performance measures. Our comparison of text classification methods to human annotators shows that both can reliably detect emotions and appraisals with similar performance. Therefore, appraisals constitute an alternative computational emotion analysis paradigm and further improve the categorization of emotions in text with joint models.
Enrica Troiano, Laura Oberländer, Roman Klinger
Comput. Linguistics3
2023 From theories on styles to their transfer in text: Bridging the gap with a hierarchical survey
abstract
Abstract Humans are naturally endowed with the ability to write in a particular style. They can, for instance, rephrase a formal letter in an informal way, convey a literal message with the use of figures of speech or edit a novel by mimicking the style of some well-known authors. Automating this form of creativity constitutes the goal of style transfer. As a natural language generation task, style transfer aims at rewriting existing texts, and specifically, it creates paraphrases that exhibit some desired stylistic attributes. From a practical perspective, it envisions beneficial applications, like chatbots that modulate their communicative style to appear empathetic, or systems that automatically simplify technical articles for a non-expert audience. Several style-aware paraphrasing methods have attempted to tackle style transfer. A handful of surveys give a methodological overview of the field, but they do not support researchers to focus on specific styles. With this paper, we aim at providing a comprehensive discussion of the styles that have received attention in the transfer task. We organize them in a hierarchy, highlighting the challenges for the definition of each of them and pointing out gaps in the current research landscape. The hierarchy comprises two main groups. One encompasses styles that people modulate arbitrarily, along the lines of registers and genres. The other group corresponds to unintentionally expressed styles, due to an author’s personal characteristics. Hence, our review shows how these groups relate to one another and where specific styles, including some that have not yet been explored, belong in the hierarchy. Moreover, we summarize the methods employed for different stylistic families, hinting researchers towards those that would be the most fitting for future research.
Enrica Troiano, Aswathy Velutharambath, Roman Klinger
Nat. Lang. Eng.3
2022 Natural Language Inference Prompts for Zero-shot Emotion Classification in Text across Corpora
abstract
Within textual emotion classification, the set of relevant labels depends on the domain and application scenario and might not be known at the time of model development. This conflicts with the classical paradigm of supervised learning in which the labels need to be predefined. A solution to obtain a model with a flexible set of labels is to use the paradigm of zero-shot learning as a natural language inference task, which in addition adds the advantage of not needing any labeled training data. This raises the question how to prompt a natural language inference model for zero-shot learning emotion classification. Options for prompt formulations include the emotion name anger alone or the statement “This text expresses anger”. With this paper, we analyze how sensitive a natural language inference-based zero-shot-learning classifier is to such changes to the prompt under consideration of the corpus: How carefully does the prompt need to be selected? We perform experiments on an established set of emotion datasets presenting different language registers according to different sources (tweets, events, blogs) with three natural language inference models and show that indeed the choice of a particular prompt formulation needs to fit to the corpus. We show that this challenge can be tackled with combinations of multiple prompts. Such ensemble is more robust across corpora than individual prompts and shows nearly the same performance as the individual best prompt for a particular corpus.
Flor Miriam Plaza del Arco, María Teresa Martín Valdivia, Roman Klinger
COLING3
2022 Constraining Linear-chain CRFs to Regular Languages
Sean Papay, Roman Klinger, Sebastian Padó
ICLR2
2022 CoVERT: A Corpus of Fact-checked Biomedical COVID-19 Tweets
abstract
During the first two years of the COVID-19 pandemic, large volumes of biomedical information concerning this new disease have been published on social media. Some of this information can pose a real danger, particularly when false information is shared, for instance recommendations how to treat diseases without professional medical advice. Therefore, automatic fact-checking resources and systems developed specifically for medical domain are crucial. While existing fact-checking resources cover COVID-19 related information in news or quantify the amount of misinformation in tweets, there is no dataset providing fact-checked COVID-19 related Twitter posts with detailed annotations for biomedical entities, relations and relevant evidence. We contribute CoVERT, a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19 related (mis)information. The corpus consists of 300 tweets, each annotated with named entities and relations. We employ a novel crowdsourcing methodology to annotate all tweets with fact-checking labels and supporting evidence, which crowdworkers search for online. This methodology results in substantial inter-annotator agreement. Furthermore, we use the retrieved evidence extracts as part of a fact-checking pipeline, finding that the real-world evidence is more useful than the knowledge directly available in pretrained language models.
Isabelle Mohr, Amelie Wührl, Roman Klinger
LREC3
2022 x-enVENT: A Corpus of Event Descriptions with Experiencer-specific Emotion and Appraisal Annotations
abstract
Emotion classification is often formulated as the task to categorize texts into a predefined set of emotion classes. So far, this task has been the recognition of the emotion of writers and readers, as well as that of entities mentioned in the text. We argue that a classification setup for emotion analysis should be performed in an integrated manner, including the different semantic roles that participate in an emotion episode. Based on appraisal theories in psychology, which treat emotions as reactions to events, we compile an English corpus of written event descriptions. The descriptions depict emotion-eliciting circumstances, and they contain mentions of people who responded emotionally. We annotate all experiencers, including the original author, with the emotions they likely felt. In addition, we link them to the event they found salient (which can be different for different experiencers in a text) by annotating event properties, or appraisals (e.g., the perceived event undesirability, the uncertainty of its outcome). Our analysis reveals patterns in the co-occurrence of people’s emotions in interaction. Hence, this richly-annotated resource provides useful data to study emotions and event evaluations from the perspective of different roles, and it enables the development of experiencer-specific emotion and appraisal classification systems.
Enrica Troiano, Laura Oberländer, Maximilian Wegge, Roman Klinger
LREC4
2022 Recovering Patient Journeys: A Corpus of Biomedical Entities and Relations on Twitter (BEAR)
abstract
Text mining and information extraction for the medical domain has focused on scientific text generated by researchers. However, their access to individual patient experiences or patient-doctor interactions is limited. On social media, doctors, patients and their relatives also discuss medical information. Individual information provided by laypeople complements the knowledge available in scientific text. It reflects the patient’s journey making the value of this type of data twofold: It offers direct access to people’s perspectives, and it might cover information that is not available elsewhere, including self-treatment or self-diagnose. Named entity recognition and relation extraction are methods to structure information that is available in unstructured text. However, existing medical social media corpora focused on a comparably small set of entities and relations. In contrast, we provide rich annotation layers to model patients’ experiences in detail. The corpus consists of medical tweets annotated with a fine-grained set of medical entities and relations between them, namely 14 entity (incl. environmental factors, diagnostics, biochemical processes, patients’ quality-of-life descriptions, pathogens, medical conditions, and treatments) and 20 relation classes (incl. prevents, influences, interactions, causes). The dataset consists of 2,100 tweets with approx. 6,000 entities and 2,200 relations.
Amelie Wührl, Roman Klinger
LREC2
2022 Embarrassingly Simple Performance Prediction for Abductive Natural Language Inference
abstract
The task of abductive natural language inference (αNLI), to decide which hypothesis is the more likely explanation for a set of observations, is a particularly difficult type of NLI.Instead of just determining a causal relationship, it requires common sense to also evaluate how reasonable an explanation is.All recent competitive systems build on top of contextualized representations and make use of transformer architectures for learning an NLI model.When somebody is faced with a particular NLI task, they need to select the best model that is available.This is a time-consuming and resource-intense endeavour.To solve this practical problem, we propose a simple method for predicting the performance without actually fine-tuning the model.We do this by testing how well the pre-trained models perform on the αNLI task when just comparing sentence embeddings with cosine similarity to what the performance that is achieved when training a classifier on top of these embeddings.We show that the accuracy of the cosine similarity approach correlates strongly with the accuracy of the classification approach with a Pearson correlation coefficient of 0.65.Since the similarity computation is orders of magnitude faster to compute on a given dataset (less than a minute vs. hours), our method can lead to significant time savings in the process of model selection.
Emils Kadikis, Vaibhav Srivastav, Roman Klinger
NAACL-HLT3
2020 Appraisal Theories for Emotion Classification in Text
abstract
Automatic emotion categorization has been predominantly formulated as text classification in which textual units are assigned to an emotion from a predefined inventory, for instance following the fundamental emotion classes proposed by Paul Ekman (fear, joy, anger, disgust, sadness, surprise) or Robert Plutchik (adding trust, anticipation).This approach ignores existing psychological theories to some degree, which provide explanations regarding the perception of events.For instance, the description that somebody discovers a snake is associated with fear, based on the appraisal as being an unpleasant and non-controllable situation.This emotion reconstruction is even possible without having access to explicit reports of a subjective feeling (for instance expressing this with the words "I am afraid.").Automatic classification approaches therefore need to learn properties of events as latent variables (for instance that the uncertainty and the mental or physical effort associated with the encounter of a snake leads to fear).With this paper, we propose to make such interpretations of events explicit, following theories of cognitive appraisal of events, and show their potential for emotion classification when being encoded in classification models.Our results show that high quality appraisal dimension assignments in event descriptions lead to an improvement in the classification of discrete emotion categories.We make our corpus of appraisal-annotated emotion-associated event descriptions publicly available.
Jan Hofmann, Enrica Troiano, Kai Sassenberg, Roman Klinger
COLING4
2020 Lost in Back-Translation: Emotion Preservation in Neural Machine Translation
abstract
Machine translation provides powerful methods to convert text between languages, and is therefore a technology enabling a multilingual world.An important part of communication, however, takes place at the non-propositional level (e.g., politeness, formality, emotions), and it is far from clear whether current MT methods properly translate this information.This paper investigates the specific hypothesis that the non-propositional level of emotions is at least partially lost in MT.We carry out a number of experiments in a back-translation setup and establish that (1) emotions are indeed partially lost during translation; (2) this tendency can be reversed almost completely with a simple re-ranking approach informed by an emotion classifier, taking advantage of diversity in the n-best list; (3) the re-ranking approach can also be applied to change emotions, obtaining a model for emotion style transfer.An in-depth qualitative analysis reveals that there are recurring linguistic changes through which emotions are toned down or amplified, such as change of modality.
Enrica Troiano, Roman Klinger, Sebastian Padó
COLING2
2020 Dissecting Span Identification Tasks with Performance Prediction
abstract
Span identification (in short, span ID) tasks such as chunking, NER, or code-switching detection, ask models to identify and classify relevant spans in a text.Despite being a staple of NLP, and sharing a common structure, there is little insight on how these tasks' properties influence their difficulty, and thus little guidance on what model families work well on span ID tasks, and why.We analyze span ID tasks via performance prediction, estimating how well neural architectures do on different tasks.Our contributions are: (a) we identify key properties of span ID tasks that can inform performance prediction; (b) we carry out a large-scale experiment on English data, building a model to predict performance for unseen span ID tasks that can support architecture choices; (c), we investigate the parameters of the meta model, yielding new insights on how model and task properties interact to affect span ID performance.We find, e.g., that span frequency is especially important for LSTMs, and that CRFs help when spans are infrequent and boundaries non-distinctive.
Sean Papay, Roman Klinger, Sebastian Padó
EMNLP (1)2
2020 GoodNewsEveryone: A Corpus of News Headlines Annotated with Emotions, Semantic Roles, and Reader Perception
abstract
Most research on emotion analysis from text focuses on the task of emotion classification or emotion intensity regression. Fewer works address emotions as a phenomenon to be tackled with structured learning, which can be explained by the lack of relevant datasets. We fill this gap by releasing a dataset of 5000 English news headlines annotated via crowdsourcing with their associated emotions, the corresponding emotion experiencers and textual cues, related emotion causes and targets, as well as the reader’s perception of the emotion of the headline. This annotation task is comparably challenging, given the large number of classes and roles to be identified. We therefore propose a multiphase annotation procedure in which we first find relevant instances with emotional content and then annotate the more fine-grained aspects. Finally, we develop a baseline for the task of automatic prediction of semantic role structures and discuss the results. The corpus we release enables further research on emotion classification, emotion intensity prediction, emotion cause detection, and supports further qualitative studies.
Laura Ana Maria Bostan, Evgeny Kim, Roman Klinger
LREC3
2020 PO-EMO: Conceptualization, Annotation, and Modeling of Aesthetic Emotions in German and English Poetry
abstract
Most approaches to emotion analysis of social media, literature, news, and other domains focus exclusively on basic emotion categories as defined by Ekman or Plutchik. However, art (such as literature) enables engagement in a broader range of more complex and subtle emotions. These have been shown to also include mixed emotional responses. We consider emotions in poetry as they are elicited in the reader, rather than what is expressed in the text or intended by the author. Thus, we conceptualize a set of aesthetic emotions that are predictive of aesthetic appreciation in the reader, and allow the annotation of multiple labels per line to capture mixed emotions within their context. We evaluate this novel setting in an annotation experiment both with carefully trained experts and via crowdsourcing. Our annotation with experts leads to an acceptable agreement of k = .70, resulting in a consistent dataset for future large scale analysis. Finally, we conduct first emotion classification experiments based on BERT, showing that identifying aesthetic emotions is challenging in our data, with up to .52 F1-micro on the German subset. Data and resources are available at https://github.com/tnhaider/poetry-emotion.
Thomas N. Haider, Steffen Eger, Evgeny Kim, Roman Klinger, Winfried Menninghaus
LREC4
2020 Automatic Section Recognition in Obituaries
abstract
Obituaries contain information about people’s values across times and cultures, which makes them a useful resource for exploring cultural history. They are typically structured similarly, with sections corresponding to Personal Information, Biographical Sketch, Characteristics, Family, Gratitude, Tribute, Funeral Information and Other aspects of the person. To make this information available for further studies, we propose a statistical model which recognizes these sections. To achieve that, we collect a corpus of 20058 English obituaries from TheDaily Item, Remembering.CA and The London Free Press. The evaluation of our annotation guidelines with three annotators on 1008 obituaries shows a substantial agreement of Fleiss κ = 0.87. Formulated as an automatic segmentation task, a convolutional neural network outperforms bag-of-words and embedding-based BiLSTMs and BiLSTM-CRFs with a micro F1 = 0.81.
Valentino Sabbatino, Laura Ana Maria Bostan, Roman Klinger
LREC3
2020 Learning soft domain constraints in a factor graph model for template-based information extraction
Hendrik ter Horst, Matthias Hartung, Philipp Cimiano, Nicole Brazda, Hans Werner Müller, Roman Klinger
Data Knowl. Eng.6
2019 Crowdsourcing and Validating Event-focused Emotion Corpora for German and English
abstract
Sentiment analysis has a range of corpora available across multiple languages.For emotion analysis, the situation is more limited, which hinders potential research on crosslingual modeling and the development of predictive models for other languages.In this paper, we fill this gap for German by constructing deISEAR, a corpus designed in analogy to the well-established English ISEAR emotion dataset.Motivated by Scherer's appraisal theory, we implement a crowdsourcing experiment which consists of two steps.In step 1, participants create descriptions of emotional events for a given emotion.In step 2, five annotators assess the emotion expressed by the texts.We show that transferring an emotion classification model from the original English ISEAR to the German crowdsourced deISEAR via machine translation does not, on average, cause a performance drop.
Enrica Troiano, Sebastian Padó, Roman Klinger
ACL (1)3
2019 Embedding Projection for Targeted Cross-lingual Sentiment: Model Comparisons and a Real-World Study
abstract
Sentiment analysis benefits from large, hand-annotated resources in order to train and test machine learning models, which are often data hungry. While some languages, e.g., English, have a vast arrayof these resources, most under-resourced languages do not, especially for fine-grained sentiment tasks, such as aspect-level or targeted sentiment analysis. To improve this situation, we propose a cross-lingual approach to sentiment analysis that is applicable to under-resourced languages and takes into account target-level information. This model incorporates sentiment information into bilingual distributional representations, byjointly optimizing them for semantics and sentiment, showing state-of-the-art performance at sentence-level when combined with machine translation. The adaptation to targeted sentiment analysis on multiple domains shows that our model outperforms other projection-based bilingual embedding methods on binary targetedsentiment tasks. Our analysis on ten languages demonstrates that the amount of unlabeled monolingual data has surprisingly little effect on the sentiment results. As expected, the choice of a annotated source language for projection to a target leads to better results for source-target language pairs which are similar. Therefore, our results suggest that more efforts should be spent on the creation of resources for less similar languages tothose which are resource-rich already. Finally, a domain mismatch leads to a decreased performance. This suggests resources in any language should ideally cover varieties of domains.
Jeremy Barnes 0001, Roman Klinger
J. Artif. Intell. Res.2
2018 Bilingual Sentiment Embeddings: Joint Projection of Sentiment Across Languages
abstract
Sentiment analysis in low-resource languages suffers from a lack of annotated corpora to estimate high-performing models.Machine translation and bilingual word embeddings provide some relief through cross-lingual sentiment approaches.However, they either require large amounts of parallel data or do not sufficiently capture sentiment information.We introduce Bilingual Sentiment Embeddings (BLSE), which jointly represent sentiment information in a source and target language.This model only requires a small bilingual lexicon, a source-language corpus annotated for sentiment, and monolingual word embeddings for each language.We perform experiments on three language combinations (Spanish, Catalan, Basque) for sentencelevel cross-lingual sentiment classification and find that our model significantly outperforms state-of-the-art methods on four out of six experimental setups, as well as capturing complementary information to machine translation.Our analysis of the resulting embedding space provides evidence that it represents sentiment information in the resource-poor target language without any annotated data in that language.
Jeremy Barnes 0001, Roman Klinger, Sabine Schulte im Walde
ACL (1)2
2018 Projecting Embeddings for Domain Adaption: Joint Modeling of Sentiment Analysis in Diverse Domains
abstract
Domain adaptation for sentiment analysis is challenging due to the fact that supervised classifiers are very sensitive to changes in domain. The two most prominent approaches to this problem are structural correspondence learning and autoencoders. However, they either require long training times or suffer greatly on highly divergent domains. Inspired by recent advances in cross-lingual sentiment analysis, we provide a novel perspective and cast the domain adaptation problem as an embedding projection task. Our model takes as input two mono-domain embedding spaces and learns to project them to a bi-domain space, which is jointly optimized to (1) project across domains and to (2) predict sentiment. We perform domain adaptation experiments on 20 source-target domain pairs for sentiment classification and report novel state-of-the-art results on 11 domain pairs, including the Amazon domain adaptation datasets and SemEval 2013 and 2016 datasets. Our analysis shows that our model performs comparably to state-of-the-art approaches on domains that are similar, while performing significantly better on highly divergent domains. Our code is available at https://github.com/jbarnesspain/domain_blse
Jeremy Barnes 0001, Roman Klinger, Sabine Schulte im Walde
COLING2
2018 An Analysis of Annotated Corpora for Emotion Classification in Text
abstract
Several datasets have been annotated and published for classification of emotions. They differ in several ways: (1) the use of different annotation schemata (e. g., discrete label sets, including joy, anger, fear, or sadness or continuous values including valence, or arousal), (2) the domain, and, (3) the file formats. This leads to several research gaps: supervised models often only use a limited set of available resources. Additionally, no previous work has compared emotion corpora in a systematic manner. We aim at contributing to this situation with a survey of the datasets, and aggregate them in a common file format with a common annotation schema. Based on this aggregation, we perform the first cross-corpus classification experiments in the spirit of future research enabled by this paper, in order to gain insight and a better understanding of differences of models inferred from the data. This work also simplifies the choice of the most appropriate resources for developing a model for a novel domain. One result from our analysis is that a subset of corpora is better classified with models trained on a different corpus. For none of the corpora, training on all data altogether is better than using a subselection of the resources. Our unified corpus is available at http://www.ims.uni-stuttgart.de/data/unifyemotion.
Laura Ana Maria Bostan, Roman Klinger
COLING2
2018 Who Feels What and Why? Annotation of a Literature Corpus with Semantic Roles of Emotions
abstract
Most approaches to emotion analysis in fictional texts focus on detecting the emotion expressed in text. We argue that this is a simplification which leads to an overgeneralized interpretation of the results, as it does not take into account who experiences an emotion and why. Emotions play a crucial role in the interaction between characters and the events they are involved in. Until today, no specific corpora that capture such an interaction were available for literature. We aim at filling this gap and present a publicly available corpus based on Project Gutenberg, REMAN (Relational EMotion ANnotation), manually annotated for spans which correspond to emotion trigger phrases and entities/events in the roles of experiencers, targets, and causes of the emotion. We provide baseline results for the automatic prediction of these relational structures and show that emotion lexicons are not able to encompass the high variability of emotion expressions and demonstrate that statistical models benefit from joint modeling of emotions with its roles in all subtasks. The corpus that we provide enables future research on the recognition of emotions and associated entities in text. It supports qualitative literary studies and digital humanities. The corpus is available at http://www.ims.uni-stuttgart.de/data/reman .
Evgeny Kim, Roman Klinger
COLING2
2018 An Empirical Analysis of the Role of Amplifiers, Downtoners, and Negations in Emotion Classification in Microblogs
abstract
The effect of amplifiers, downtoners, and negations has been studied in general and particularly in the context of sentiment analysis. However, there is only limited work which aims at transferring the results and methods to discrete classes of emotions, e.g., joy, anger, fear, sadness, surprise, and disgust. For instance, it is not straight-forward to interpret which emotion the phrase "not happy" expresses. With this paper, we aim at obtaining a better understanding of such modifiers in the context of emotion-bearing words and their impact on document-level emotion classification, namely, microposts on Twitter. We select an appropriate scope detection method for modifiers of emotion words, incorporate it in a document-level emotion classification model as additional bag of words and show that this approach improves the performance of emotion classification. In addition, we build a term weighting approach based on the different modifiers into a lexical model for the analysis of the semantics of modifiers and their impact on emotion meaning. We show that amplifiers separate emotions expressed with an emotion-bearing word more clearly from other secondary connotations. Downtoners have the opposite effect. In addition, we discuss the meaning of negations of emotion-bearing words. For instance we show empirically that "not happy" is closer to sadness than to anger and that fear words in the scope of downtoners often express surprise.
Florian Strohm, Roman Klinger
DSAA2
2018 Assessing the Impact of Single and Pairwise Slot Constraints in a Factor Graph Model for Template-Based Information Extraction
Hendrik ter Horst, Matthias Hartung, Roman Klinger, Nicole Brazda, Hans Werner Müller, Philipp Cimiano
NLDB3
2018 On the Semantic Similarity of Disease Mentions in MEDLINE and Twitter
Camilo Thorne, Roman Klinger
NLDB2
2018 What you use, not what you do: Automatic classification and similarity detection of recipes
Hanna Kicherer, Marcel Dittrich, Lukas Grebe, Christian Scheible, Roman Klinger
Data Knowl. Eng.5
2017 Identifying Right-Wing Extremism in German Twitter Profiles: A Classification Approach
Matthias Hartung, Roman Klinger, Franziska Schmidtke, Lars Vogel
NLDB2
2017 What You Use, Not What You Do: Automatic Classification of Recipes
Hanna Kicherer, Marcel Dittrich, Lukas Grebe, Christian Scheible, Roman Klinger
NLDB5
2017 Does Optical Character Recognition and Caption Generation Improve Emotion Detection in Microblog Posts?
Roman Klinger
NLDB1
2017 Fine-Grained Opinion Mining from Mobile App Reviews with Word Embedding Features
Mario Sänger, Ulf Leser, Roman Klinger
NLDB3
2016 Model Architectures for Quotation Detection
abstract
Quotation detection is the task of locating spans of quoted speech in text.The state of the art treats this problem as a sequence labeling task and employs linear-chain conditional random fields.We question the efficacy of this choice: The Markov assumption in the model prohibits it from making joint decisions about the begin, end, and internal context of a quotation.We perform an extensive analysis with two new model architectures.We find that (a), simple boundary classification combined with a greedy prediction strategy is competitive with the state of the art; (b), a semi-Markov model significantly outperforms all others, by relaxing the Markov assumption.
Christian Scheible, Roman Klinger, Sebastian Padó
ACL (1)2
2016 SCARE ― The Sentiment Corpus of App Reviews with Fine-grained Annotations in German
Mario Sänger, Ulf Leser, Steffen Kemmerer, Peter Adolphs, Roman Klinger
LREC5
2015 Instance Selection Improves Cross-Lingual Model Training for Fine-Grained Sentiment Analysis
abstract
Scarcity of annotated corpora for many languages is a bottleneck for training finegrained sentiment analysis models that can tag aspects and subjective phrases.We propose to exploit statistical machine translation to alleviate the need for training data by projecting annotated data in a source language to a target language such that a supervised fine-grained sentiment analysis system can be trained.To avoid a negative influence of poor-quality translations, we propose a filtering approach based on machine translation quality estimation measures to select only high-quality sentence pairs for projection.We evaluate on the language pair German/English on a corpus of product reviews annotated for both languages and compare to in-target-language training.Projection without any filtering leads to 23 % F 1 in the task of detecting aspect phrases, compared to 41 % F 1 for in-target-language training.Our approach obtains up to 47 % F 1 .Further, we show that the detection of subjective phrases is competitive to in-target-language training without filtering.
Roman Klinger, Philipp Cimiano
CoNLL1
2014 The USAGE review corpus for fine grained multi lingual opinion analysis
Roman Klinger, Philipp Cimiano
LREC1
2013 Orthonormal Explicit Topic Analysis for Cross-Lingual Document Matching
abstract
Cross-lingual topic modelling has applications in machine translation, word sense disambiguation and terminology alignment.Multilingual extensions of approaches based on latent (LSI), generative (LDA, PLSI) as well as explicit (ESA) topic modelling can induce an interlingual topic space allowing documents in different languages to be mapped into the same space and thus to be compared across languages.In this paper, we present a novel approach that combines latent and explicit topic modelling approaches in the sense that it builds on a set of explicitly defined topics, but then computes latent relations between these.Thus, the method combines the benefits of both explicit and latent topic modelling approaches.We show that on a crosslingual mate retrieval task, our model significantly outperforms LDA, LSI, and ESA, as well as a baseline that translates every word in a document into the target language.
John P. McCrae, Philipp Cimiano, Roman Klinger
EMNLP3
2012 E-Government and Policy Simulation in Intelligent Virtual Environments
Fotis Aisopos, Magdalini Kardara, Philipp Senger, Roman Klinger, Athanasios Papaoikonomou, Konstantinos Tserpes, Michael Gardner, Theodora A. Varvarigou
WEBIST4
2011 Challenges in the association of human single nucleotide polymorphism mentions with unique database identifiers
abstract
BACKGROUND: Most information on genomic variations and their associations with phenotypes are covered exclusively in scientific publications rather than in structured databases. These texts commonly describe variations using natural language; database identifiers are seldom mentioned. This complicates the retrieval of variations, associated articles, as well as information extraction, e. g. the search for biological implications. To overcome these challenges, procedures to map textual mentions of variations to database identifiers need to be developed. RESULTS: This article describes a workflow for normalization of variation mentions, i.e. the association of them to unique database identifiers. Common pitfalls in the interpretation of single nucleotide polymorphism (SNP) mentions are highlighted and discussed. The developed normalization procedure achieves a precision of 98.1 % and a recall of 67.5% for unambiguous association of variation mentions with dbSNP identifiers on a text corpus based on 296 MEDLINE abstracts containing 527 mentions of SNPs. The annotated corpus is freely available at http://www.scai.fraunhofer.de/snp-normalization-corpus.html. CONCLUSIONS: Comparable approaches usually focus on variations mentioned on the protein sequence and neglect problems for other SNP mentions. The results presented here indicate that normalizing SNPs described on DNA level is more difficult than the normalization of SNPs described on protein level. The challenges associated with normalization are exemplified with ambiguities and errors, which occur in this corpus.
Philippe Thomas 0002, Roman Klinger, Laura Inés Furlong, Martin Hofmann-Apitius, Christoph M. Friedrich
BMC Bioinform.2
2009 Identification of histone modifications in biomedical text for supporting epigenomic research
abstract
BACKGROUND: Posttranslational modifications of histones influence the structure of chromatine and in such a way take part in the regulation of gene expression. Certain histone modification patterns, distributed over the genome, are connected to cell as well as tissue differentiation and to the adaption of organisms to their environment. Abnormal changes instead influence the development of disease states like cancer. The regulation mechanisms for modifying histones and its functionalities are the subject of epigenomics investigation and are still not completely understood. Text provides a rich resource of knowledge on epigenomics and modifications of histones in particular. It contains information about experimental studies, the conditions used, and results. To our knowledge, no approach has been published so far for identifying histone modifications in text. RESULTS: We have developed an approach for identifying histone modifications in biomedical literature with Conditional Random Fields (CRF) and for resolving the recognized histone modification term variants by term standardization. For the term identification F1 measures of 0.84 by 10-fold cross-validation on the training corpus and 0.81 on an independent test corpus have been obtained. The standardization enabled the correct transformation of 96% of the terms from training and 98% from test the corpus. Due to the lack of terminologies exhaustively covering specific histone modification types, we developed a histone modification term hierarchy for use in a semantic text retrieval system. CONCLUSION: The developed approach highly improves the retrieval of articles describing histone modifications. Since text contains context information about performed studies and experiments, the identification of histone modifications is the basis for supporting literature-based knowledge discovery and hypothesis generation to accelerate epigenomic research.
Corinna Kolárik, Roman Klinger, Martin Hofmann-Apitius
BMC Bioinform.2
2008 Detection of IUPAC and IUPAC-like chemical names
abstract
MOTIVATION: Chemical compounds like small signal molecules or other biological active chemical substances are an important entity class in life science publications and patents. Several representations and nomenclatures for chemicals like SMILES, InChI, IUPAC or trivial names exist. Only SMILES and InChI names allow a direct structure search, but in biomedical texts trivial names and Iupac like names are used more frequent. While trivial names can be found with a dictionary-based approach and in such a way mapped to their corresponding structures, it is not possible to enumerate all IUPAC names. In this work, we present a new machine learning approach based on conditional random fields (CRF) to find mentions of IUPAC and IUPAC-like names in scientific text as well as its evaluation and the conversion rate with available name-to-structure tools. RESULTS: We present an IUPAC name recognizer with an F(1) measure of 85.6% on a MEDLINE corpus. The evaluation of different CRF orders and offset conjunction orders demonstrates the importance of these parameters. An evaluation of hand-selected patent sections containing large enumerations and terms with mixed nomenclature shows a good performance on these cases (F(1) measure 81.5%). Remaining recognition problems are to detect correct borders of the typically long terms, especially when occurring in parentheses or enumerations. We demonstrate the scalability of our implementation by providing results from a full MEDLINE run. AVAILABILITY: We plan to publish the corpora, annotation guideline as well as the conditional random field model as a UIMA component.
Roman Klinger, Corinna Kolárik, Juliane Fluck, Martin Hofmann-Apitius, Christoph M. Friedrich
ISMB1