VLDB 2026 Research / reviewers in the wild / expert
Josef Ruppenhofer
dblp:15/8153
· DBLP profile ↗
42ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-8662-7618ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 9 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Negative Stereotypes - Non-Negative Abusive Utterances about Identity Groups and Their Semantic VariantsabstractWe investigate a specific subtype of implicitly abusive language, focusing on non-negative sentences about identity groups (e.g.Women make good cooks).We introduce a novel dataset comprising such utterances.It not only profiles abusive sentences but also includes various semantic variants of the same characteristic attributed to an identity group, allowing us to systematically examine the impact of different degrees of generalization and perspective framing.Thus we demonstrate that specific variants significantly intensify the perception of abusiveness.By switching identity groups, we highlight how the characteristics often described in stereotypes are not inherently abusive.We also report on classification experiments. Tina Lommel, Elisabeth Eder, Josef Ruppenhofer, Michael Wiegand |
ACL (1) | 3 |
| 2024 | Out of the Mouths of MPs: Speaker Attribution in Parliamentary DebatesabstractThis paper presents GePaDe_SpkAtt , a new corpus for speaker attribution in German parliamentary debates, with more than 7,700 manually annotated events of speech, thought and writing. Our role inventory includes the sources, addressees, messages and topics of the speech event and also two additional roles, medium and evidence. We report baseline results for the automatic prediction of speech events and their roles, with high scores for both, event triggers and roles. Then we apply our model to predict speech events in 20 years of parliamentary debates and investigate the use of factives in the rhetoric of MPs. Ines Rehbein, Josef Ruppenhofer, Annelen Brunner, Simone Paolo Ponzetto |
LREC/COLING | 2 |
| 2024 | Every Verb in Its Right Place? A Roadmap for Operationalizing Developmental Stages in the Acquisition of L2 GermanabstractDevelopmental stages are a linguistic concept claiming that language learning, despite its large inter-individual variance, generally progresses in an ordered, step-like manner. At the core of research has been the acquisition of verb placement by learners, as conceptualized within Processability Theory (Pienemann, 1989). The computational implementation of a system detecting developmental stages is a prerequisite for an automated analysis of L2 language development. However, such an implementation faces two main challenges. The first is the lack of a fully fleshed out, coherent linguistic specification of the stages. The second concerns the translation of the linguistic specification into computational procedures that can extract clauses from learner-produced text and assign them to a developmental stage based on verb placement. Our contribution provides the necessary linguistic specification of the stages as well as detaiiled discussion and recommendations regarding computational implementation. Josef Ruppenhofer, Matthias Schwendemann, Annette Portmann, Katrin Wisniewski, Torsten Zesch |
LREC/COLING | 1 |
| 2024 | Oddballs and Misfits: Detecting Implicit Abuse in Which Identity Groups are Depicted as Deviating from the NormabstractWarning: This paper contains content that may be offensive or upsetting.We address the task of detecting abusive sentences in which identity groups are depicted as deviating from the norm (e.g.Gays sprinkle flour over their gardens for good luck).These abusive utterances need not be stereotypes or negative in sentiment.For this type of abuse, we are the first to present a study on how to detect it.We introduce datasets for this task created via crowdsourcing that include 7 different identity groups.We also report on classification experiments and show that only large language models detect this abuse reliably. Michael Wiegand, Josef Ruppenhofer |
EMNLP | 2 |
| 2024 | Determining sentiment views of verbal multiword expressions using linguistic featuresabstractAbstract We examine the binary classification of sentiment views for verbal multiword expressions (MWEs). Sentiment views denote the perspective of the holder of some opinion. We distinguish between MWEs conveying the view of the speaker of the utterance (e.g., in “The company reinvented the wheel” the holder is the implicit speaker who criticizes the company for creating something already existing) and MWEs conveying the view of explicit entities participating in an opinion event (e.g., in “Peter threw in the towel” the holder is Peter having given up something). The task has so far been examined on unigram opinion words. Since many features found effective for unigrams are not usable for MWEs, we propose novel ones taking into account the internal structure of MWEs, a unigram sentiment-view lexicon and various information from Wiktionary. We also examine distributional methods and show that the corpus on which a representation is induced has a notable impact on the classification. We perform an extrinsic evaluation in the task of opinion holder extraction and show that the learnt knowledge also improves a state-of-the-art classifier trained on BERT. Sentiment-view classification is typically framed as a task in which only little labeled training data are available. As in the case of unigrams, we show that for MWEs a feature-based approach beats state-of-the-art generic methods. Michael Wiegand, Marc Schulder, Josef Ruppenhofer |
Nat. Lang. Eng. | 3 |
| 2023 | Euphemistic Abuse - A New Dataset and Classification Experiments for Implicitly Abusive LanguageabstractWe address the task of identifying euphemistic abuse (e.g.You inspire me to fall asleep) paraphrasing simple explicitly abusive utterances (e.g.You are boring).For this task, we introduce a novel dataset that has been created via crowdsourcing.Special attention has been paid to the generation of appropriate negative (nonabusive) data.We report on classification experiments showing that classifiers trained on previous datasets are less capable of detecting such abuse.Best automatic results are obtained by a classifier that augments training data from our new dataset with automatically-generated GPT-3 completions.We also present a classifier that combines a few manually extracted features that exemplify the major linguistic phenomena constituting euphemistic abuse. Michael Wiegand, Jana Kampfmeier, Elisabeth Eder, Josef Ruppenhofer |
EMNLP | 4 |
| 2022 | Who's in, who's out? Predicting the Inclusiveness or Exclusiveness of Personal Pronouns in Parliamentary DebatesabstractThis paper presents a compositional annotation scheme to capture the clusivity properties of personal pronouns in context, that is their ability to construct and manage in-groups and out-groups by including/excluding the audience and/or non-speech act participants in reference to groups that also include the speaker. We apply and test our schema on pronoun instances in speeches taken from the German parliament. The speeches cover a time period from 2017-2021 and comprise manual annotations for 3,126 sentences. We achieve high inter-annotator agreement for our new schema, with a Cohen’s κ in the range of 89.7-93.2 and a percentage agreement of > 96%. Our exploratory analysis of in/exclusive pronoun use in the parliamentary setting provides some face validity for our new schema. Finally, we present baseline experiments for automatically predicting clusivity in political debates, with promising results for many referential constellations, yielding an overall 84.9% micro F1 for all pronouns. Ines Rehbein, Josef Ruppenhofer |
LREC | 2 |
| 2022 | Identifying Implicitly Abusive Remarks about Identity Groups using a Linguistically Informed ApproachabstractMichael Wiegand, Elisabeth Eder, Josef Ruppenhofer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Michael Wiegand, Elisabeth Eder, Josef Ruppenhofer |
NAACL-HLT | 3 |
| 2021 | Implicitly Abusive Comparisons - A New Dataset and Linguistic AnalysisabstractWe examine the task of detecting implicitly abusive comparisons (e.g.Your hair looks like you have been electrocuted).Implicitly abusive comparisons are abusive comparisons in which abusive words (e.g.dumbass or scum) are absent.We detail the process of creating a novel dataset for this task via crowdsourcing that includes several measures to obtain a sufficiently representative and unbiased set of comparisons.We also present classification experiments that include a range of linguistic features that help us better understand the mechanisms underlying abusive comparisons. Michael Wiegand, Maja Geulig, Josef Ruppenhofer |
EACL | 3 |
| 2021 | Exploiting Emojis for Abusive Language DetectionabstractWe propose to use abusive emojis, such as the middle finger or face vomiting, as a proxy for learning a lexicon of abusive words.Since it represents extralinguistic information, a single emoji can co-occur with different forms of explicitly abusive utterances.We show that our approach generates a lexicon that offers the same performance in cross-domain classification of abusive microposts as the most advanced lexicon induction method.Such an approach, in contrast, is dependent on manually annotated seed words and expensive lexical resources for bootstrapping (e.g.WordNet).We demonstrate that the same emojis can also be effectively used in languages other than English.Finally, we also show that emojis can be exploited for classifying mentions of ambiguous words, such as fuck and bitch, into generally abusive and just profane usages. Michael Wiegand, Josef Ruppenhofer |
EACL | 2 |
| 2021 | Implicitly Abusive Language - What does it actually look like and why are we not getting there?abstractMichael Wiegand, Josef Ruppenhofer, Elisabeth Eder. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Michael Wiegand, Josef Ruppenhofer, Elisabeth Eder |
NAACL-HLT | 2 |
| 2021 | Automatic generation of lexica for sentiment polarity shiftersabstractAbstract Alleviating pain is good and abandoning hope is bad. We instinctively understand how words like alleviate and abandon affect the polarity of a phrase, inverting or weakening it. When these words are content words, such as verbs, nouns, and adjectives, we refer to them as polarity shifters. Shifters are a frequent occurrence in human language and an important part of successfully modeling negation in sentiment analysis; yet research on negation modeling has focused almost exclusively on a small handful of closed-class negation words, such as not, no, and without. A major reason for this is that shifters are far more lexically diverse than negation words, but no resources exist to help identify them. We seek to remedy this lack of shifter resources by introducing a large lexicon of polarity shifters that covers English verbs, nouns, and adjectives. Creating the lexicon entirely by hand would be prohibitively expensive. Instead, we develop a bootstrapping approach that combines automatic classification with human verification to ensure the high quality of our lexicon while reducing annotation costs by over 70%. Our approach leverages a number of linguistic insights; while some features are based on textual patterns, others use semantic resources or syntactic relatedness. The created lexicon is evaluated both on a polarity shifter gold standard and on a polarity classification task. Marc Schulder, Michael Wiegand, Josef Ruppenhofer |
Nat. Lang. Eng. | 3 |
| 2020 | Doctor Who? Framing Through Names and Titles in GermanabstractEntity framing is the selection of aspects of an entity to promote a particular viewpoint towards that entity. We investigate entity framing of political figures through the use of names and titles in German online discourse, enhancing current research in entity framing through titling and naming that concentrates on English only. We collect tweets that mention prominent German politicians and annotate them for stance. We find that the formality of naming in these tweets correlates positively with their stance. This confirms sociolinguistic observations that naming and titling can have a status-indicating function and suggests that this function is dominant in German tweets mentioning political figures. We also find that this status-indicating function is much weaker in tweets from users that are politically left-leaning than in tweets by right-leaning users. This is in line with observations from moral psychology that left-leaning and right-leaning users assign different importance to maintaining social hierarchies. Esther van den Berg, Katharina Korfhage, Josef Ruppenhofer, Michael Wiegand, Katja Markert |
LREC | 3 |
| 2020 | A New Resource for German Causal LanguageabstractWe present a new resource for German causal language, with annotations in context for verbs, nouns and prepositions. Our dataset includes 4,390 annotated instances for more than 150 different triggers. The annotation scheme distinguishes three different types of causal events (CONSEQUENCE , MOTIVATION, PURPOSE). We also provide annotations for semantic roles, i.e. of the cause and effect for the causal event as well as the actor and affected party, if present. In the paper, we present inter-annotator agreement scores for our dataset and discuss problems for annotating causal language. Finally, we present experiments where we frame causal annotation as a sequence labelling problem and report baseline results for the prediciton of causal arguments and for predicting different types of causation. Ines Rehbein, Josef Ruppenhofer |
LREC | 2 |
| 2020 | Improving Sentence Boundary Detection for Spoken Language TranscriptsabstractThis paper presents experiments on sentence boundary detection in transcripts of spoken dialogues. Segmenting spoken language into sentence-like units is a challenging task, due to disfluencies, ungrammatical or fragmented structures and the lack of punctuation. In addition, one of the main bottlenecks for many NLP applications for spoken language is the small size of the training data, as the transcription and annotation of spoken language is by far more time-consuming and labour-intensive than processing written language. We therefore investigate the benefits of data expansion and transfer learning and test different ML architectures for this task. Our results show that data expansion is not straightforward and even data from the same domain does not always improve results. They also highlight the importance of modelling, i.e. of finding the best architecture and data representation for the task at hand. For the detection of boundaries in spoken language transcripts, we achieve a substantial improvement when framing the boundary detection problem assentence pair classification task, as compared to a sequence tagging approach. Ines Rehbein, Josef Ruppenhofer |
LREC | 2 |
| 2020 | Fine-grained Named Entity Annotations for German Biographic InterviewsabstractWe present a fine-grained NER annotations with 30 labels and apply it to German data. Building on the OntoNotes 5.0 NER inventory, our scheme is adapted for a corpus of transcripts of biographic interviews by adding categories for AGE and LAN(guage) and also features extended numeric and temporal categories. Applying the scheme to the spoken data as well as a collection of teaser tweets from newspaper sites, we can confirm its generality for both domains, also achieving good inter-annotator agreement. We also show empirically how our inventory relates to the well-established 4-category NER inventory by re-annotating a subset of the GermEval 2014 NER coarse-grained dataset with our fine label inventory. Finally, we use a BERT-based system to establish some baseline models for NER tagging on our two new datasets. Global results in in-domain testing are quite high on the two datasets, near what was achieved for the coarse inventory on the CoNLLL2003 data. Cross-domain testing produces much lower results due to the severe domain differences. Josef Ruppenhofer, Ines Rehbein, Carolina Flinz |
LREC | 1 |
| 2020 | Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal DependenciesabstractThe paper presents a discussion on the main linguistic phenomena of user-generated texts found in web and social media, and proposes a set of annotation guidelines for their treatment within the Universal Dependencies (UD) framework. Given on the one hand the increasing number of treebanks featuring user-generated content, and its somewhat inconsistent treatment in these resources on the other, the aim of this paper is twofold: (1) to provide a short, though comprehensive, overview of such treebanks - based on available literature - along with their main features and a comparative analysis of their annotation criteria, and (2) to propose a set of tentative UD-based annotation guidelines, to promote consistent treatment of the particular phenomena found in these types of texts. The main goal of this paper is to provide a common framework for those teams interested in developing similar resources in UD, thus enabling cross-linguistic consistency, which is a principle that has always been in the spirit of UD. Manuela Sanguinetti, Cristina Bosco, Lauren Cassidy, Özlem Çetinoglu, Alessandra Teresa Cignarella, Teresa Lynn, Ines Rehbein, Josef Ruppenhofer, Djamé Seddah, Amir Zeldes |
LREC | 8 |
| 2020 | Enhancing a Lexicon of Polarity Shifters through the Supervised Classification of Shifting DirectionsabstractThe sentiment polarity of an expression (whether it is perceived as positive, negative or neutral) can be influenced by a number of phenomena, foremost among them negation. Apart from closed-class negation words like “no”, “not” or “without”, negation can also be caused by so-called polarity shifters. These are content words, such as verbs, nouns or adjectives, that shift polarities in their opposite direction, e.g. “abandoned” in “abandoned hope” or “alleviate” in “alleviate pain”. Many polarity shifters can affect both positive and negative polar expressions, shifting them towards the opposing polarity. However, other shifters are restricted to a single shifting direction. “Recoup” shifts negative to positive in “recoup your losses”, but does not affect the positive polarity of “fortune” in “recoup a fortune”. Existing polarity shifter lexica only specify whether a word can, in general, cause shifting, but they do not specify when this is limited to one shifting direction. To address this issue we introduce a supervised classifier that determines the shifting direction of shifters. This classifier uses both resource-driven features, such as WordNet relations, and data-driven features like in-context polarity conflicts. Using this classifier we enhance the largest available polarity shifter lexicon. Marc Schulder, Michael Wiegand, Josef Ruppenhofer |
LREC | 3 |
| 2018 | Sprucing up the trees - Error detection in treebanksabstractWe present a method for detecting annotation errors in manually and automatically annotated dependency parse trees, based on ensemble parsing in combination with Bayesian inference, guided by active learning. We evaluate our method in different scenarios: (i) for error detection in dependency treebanks and (ii) for improving parsing accuracy on in- and out-of-domain data. Ines Rehbein, Josef Ruppenhofer |
COLING | 2 |
| 2018 | Distinguishing affixoid formations from compoundsabstractWe study German affixoids, a type of morpheme in between affixes and free stems. Several properties have been associated with them – increased productivity; a bleached semantics, which is often evaluative and/or intensifying and thus of relevance to sentiment analysis; and the existence of a free morpheme counterpart – but not been validated empirically. In experiments on a new data set that we make available, we put these key assumptions from the morphological literature to the test and show that despite the fact that affixoids generate many low-frequency formations, we can classify these as affixoid or non-affixoid instances with a best F1-score of 74%. Josef Ruppenhofer, Michael Wiegand, Rebecca Wilm, Katja Markert |
COLING | 1 |
| 2018 | Automatically Creating a Lexicon of Verbal Polarity Shifters: Mono- and Cross-lingual Methods for GermanabstractIn this paper we use methods for creating a large lexicon of verbal polarity shifters and apply them to German. Polarity shifters are content words that can move the polarity of a phrase towards its opposite, such as the verb “abandon” in “abandon all hope”. This is similar to how negation words like “not” can influence polarity. Both shifters and negation are required for high precision sentiment analysis. Lists of negation words are available for many languages, but the only language for which a sizable lexicon of verbal polarity shifters exists is English. This lexicon was created by bootstrapping a sample of annotated verbs with a supervised classifier that uses a set of data- and resource-driven features. We reproduce and adapt this approach to create a German lexicon of verbal polarity shifters. Thereby, we confirm that the approach works for multiple languages. We further improve classification by leveraging cross-lingual information from the English shifter lexicon. Using this improved approach, we bootstrap a large number of German verbal polarity shifters, reducing the annotation effort drastically. The resulting German lexicon of verbal polarity shifters is made publicly available. Marc Schulder, Michael Wiegand, Josef Ruppenhofer |
COLING | 3 |
| 2018 | Introducing a Lexicon of Verbal Polarity Shifters for English
Marc Schulder, Michael Wiegand, Josef Ruppenhofer, Stephanie Köser |
LREC | 3 |
| 2018 | Building a Morphological Treebank for German from a Linguistic Database
Petra Steiner, Josef Ruppenhofer |
LREC | 2 |
| 2018 | Disambiguation of Verbal Shifters
Michael Wiegand, Sylvette Loda, Josef Ruppenhofer |
LREC | 3 |
| 2018 | Inducing a Lexicon of Abusive Words - a Feature-Based ApproachabstractMichael Wiegand, Josef Ruppenhofer, Anna Schmidt, Clayton Greenberg. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Michael Wiegand, Josef Ruppenhofer, Anna Schmidt, Clayton Greenberg |
NAACL-HLT | 2 |
| 2017 | Detecting annotation noise in automatically labelled dataabstractWe introduce a method for error detection in automatically annotated text, aimed at supporting the creation of high-quality language resources at affordable cost.Our method combines an unsupervised generative model with human supervision from active learning.We test our approach on in-domain and out-of-domain data in two languages, in AL simulations and in a real world setting.For all settings, the results show that our method is able to detect annotation errors with high precision and high recall. Ines Rehbein, Josef Ruppenhofer |
ACL (1) | 2 |
| 2017 | Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic FeaturesabstractWe present a major step towards the creation of the first high-coverage lexicon of polarity shifters. In this work, we bootstrap a lexicon of verbs by exploiting various linguistic features. Polarity shifters, such as “abandon”, are similar to negations (e.g. “not”) in that they move the polarity of a phrase towards its inverse, as in “abandon all hope”. While there exist lists of negation words, creating comprehensive lists of polarity shifters is far more challenging due to their sheer number. On a sample of manually annotated verbs we examine a variety of linguistic features for this task. Then we build a supervised classifier to increase coverage. We show that this approach drastically reduces the annotation effort while ensuring a high-precision lexicon. We also show that our acquired knowledge of verbal polarity shifters improves phrase-level sentiment analysis. Marc Schulder, Michael Wiegand, Josef Ruppenhofer, Benjamin Roth 0001 |
IJCNLP(1) | 3 |
| 2016 | Effect Functors for Opinion Inference
Josef Ruppenhofer, Jasper Brandes |
LREC | 1 |
| 2016 | Opinion Holder and Target Extraction on Opinion Compounds - A Linguistic ApproachabstractMichael Wiegand, Christine Bocionek, Josef Ruppenhofer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Michael Wiegand, Christine Bocionek, Josef Ruppenhofer |
HLT-NAACL | 3 |
| 2016 | Separating Actor-View from Speaker-View Opinion Expressions using Linguistic FeaturesabstractWe examine different features and classifiers for the categorization of opinion words into actor and speaker view.To our knowledge, this is the first comprehensive work to address sentiment views on the word level taking into consideration opinion verbs, nouns and adjectives.We consider many high-level features requiring only few labeled training data.A detailed feature analysis produces linguistic insights into the nature of sentiment views.We also examine how far global constraints between different opinion words help to increase classification performance.Finally, we show that our (prior) word-level annotation correlates with contextual sentiment views. Michael Wiegand, Marc Schulder, Josef Ruppenhofer |
HLT-NAACL | 3 |
| 2015 | Opinion Holder and Target Extraction based on the Induction of Verbal CategoriesabstractWe present an approach for opinion role induction for verbal predicates.Our model rests on the assumption that opinion verbs can be divided into three different types where each type is associated with a characteristic mapping between semantic roles and opinion holders and targets.In several experiments, we demonstrate the relevance of those three categories for the task.We show that verbs can easily be categorized with semi-supervised graphbased clustering and some appropriate similarity metric.The seeds are obtained through linguistic diagnostics.We evaluate our approach against a new manuallycompiled opinion role lexicon and perform in-context classification. Michael Wiegand, Josef Ruppenhofer |
CoNLL | 2 |
| 2014 | Comparing methods for deriving intensity scores for adjectivesabstractWe compare several different corpus-based and lexicon-based methods for the scalar ordering of adjectives. Among them, we examine for the first time a low-resource approach based on distinctive-collexeme analysis that just requires a small predefined set of adverbial modi-fiers. While previous work on adjective in-tensity mostly assumes one single scale for all adjectives, we group adjectives into dif-ferent scales which is more faithful to hu-man perception. We also apply the meth-ods to both polar and non-polar adjectives, showing that not all methods are equally suitable for both types of adjectives. 1 Josef Ruppenhofer, Michael Wiegand, Jasper Brandes |
EACL | 1 |
| 2013 | Predicative Adjectives: An Unsupervised Criterion to Extract Subjective Adjectives
Michael Wiegand, Josef Ruppenhofer, Dietrich Klakow |
HLT-NAACL | 2 |
| 2012 | MLSA - A Multi-layered Reference Corpus for German Sentiment Analysis
Simon Clematide, Stefan Gindl, Manfred Klenner, Stefanos Petrakis, Robert Remus, Josef Ruppenhofer, Ulli Waltinger, Michael Wiegand |
LREC | 6 |
| 2012 | Yes we can!? Annotating English modal verbs
Josef Ruppenhofer, Ines Rehbein |
LREC | 1 |
| 2011 | Evaluating the Impact of Coder Errors on Active Learning
Ines Rehbein, Josef Ruppenhofer |
ACL | 2 |
| 2010 | Bringing Active Learning to Life
Ines Rehbein, Josef Ruppenhofer, Alexis Palmer |
COLING | 2 |
| 2010 | There's no Data like More Data? Revisiting the Impact of Data Size on a Classification Task
Ines Rehbein, Josef Ruppenhofer |
LREC | 2 |
| 2010 | Generating FrameNets of Various Granularities: The FrameNet Transformer
Josef Ruppenhofer, Jonas Sunde, Manfred Pinkal |
LREC | 1 |
| 2010 | Speaker Attribution in Cabinet Protocols
Josef Ruppenhofer, Caroline Sporleder, Fabian Shirokov |
LREC | 1 |
| 2008 | Discourse Level Opinion Interpretation
Swapna Somasundaran, Janyce Wiebe, Josef Ruppenhofer |
COLING | 3 |
| 2008 | Finding the Sources and Targets of Subjective Expressions
Josef Ruppenhofer, Swapna Somasundaran, Janyce Wiebe |
LREC | 1 |