Michael Wiegand

dblp:64/4135 · DBLP profile ↗
← Back
39ranked-venue papers
26as first author
13since 2021 · last 2025
0000-0002-5403-1078ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 24 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author
YearPublicationVenuePosition
2025 Beyond Negative Stereotypes - Non-Negative Abusive Utterances about Identity Groups and Their Semantic Variants
abstract
We investigate a specific subtype of implicitly abusive language, focusing on non-negative sentences about identity groups (e.g.Women make good cooks).We introduce a novel dataset comprising such utterances.It not only profiles abusive sentences but also includes various semantic variants of the same characteristic attributed to an identity group, allowing us to systematically examine the impact of different degrees of generalization and perspective framing.Thus we demonstrate that specific variants significantly intensify the perception of abusiveness.By switching identity groups, we highlight how the characteristics often described in stereotypes are not inherently abusive.We also report on classification experiments.
Tina Lommel, Elisabeth Eder, Josef Ruppenhofer, Michael Wiegand
ACL (1)4
2025 Revisiting Implicitly Abusive Language Detection: Evaluating LLMs in Zero-Shot and Few-Shot Settings
abstract
Implicitly abusive language (IAL), unlike its explicit counterpart, lacks overt slurs or unambiguously offensive keywords, such as “bimbo” or “scum”, making it challenging to detect and mitigate. While current research predominantly focuses on explicitly abusive language, the subtler and more covert forms of IAL remain insufficiently studied. The rapid advancement and widespread adoption of large language models (LLMs) have opened new possibilities for various NLP tasks, but their application to IAL detection has been limited. We revisit three very recent challenging datasets of IAL and investigate the potential of LLMs to enhance the detection of IAL in English through zero-shot and few-shot prompting approaches. We evaluate the models’ capabilities in classifying sentences directly as either IAL or benign, and in extracting linguistic features associated with IAL. Our results indicate that classifiers trained on features extracted by advanced LLMs outperform the best previously reported results, achieving near-human performance.
Julia Jaremko, Dagmar Gromann, Michael Wiegand
COLING3
2024 Oddballs and Misfits: Detecting Implicit Abuse in Which Identity Groups are Depicted as Deviating from the Norm
abstract
Warning: This paper contains content that may be offensive or upsetting.We address the task of detecting abusive sentences in which identity groups are depicted as deviating from the norm (e.g.Gays sprinkle flour over their gardens for good luck).These abusive utterances need not be stereotypes or negative in sentiment.For this type of abuse, we are the first to present a study on how to detect it.We introduce datasets for this task created via crowdsourcing that include 7 different identity groups.We also report on classification experiments and show that only large language models detect this abuse reliably.
Michael Wiegand, Josef Ruppenhofer
EMNLP1
2024 Determining sentiment views of verbal multiword expressions using linguistic features
abstract
Abstract We examine the binary classification of sentiment views for verbal multiword expressions (MWEs). Sentiment views denote the perspective of the holder of some opinion. We distinguish between MWEs conveying the view of the speaker of the utterance (e.g., in “The company reinvented the wheel” the holder is the implicit speaker who criticizes the company for creating something already existing) and MWEs conveying the view of explicit entities participating in an opinion event (e.g., in “Peter threw in the towel” the holder is Peter having given up something). The task has so far been examined on unigram opinion words. Since many features found effective for unigrams are not usable for MWEs, we propose novel ones taking into account the internal structure of MWEs, a unigram sentiment-view lexicon and various information from Wiktionary. We also examine distributional methods and show that the corpus on which a representation is induced has a notable impact on the classification. We perform an extrinsic evaluation in the task of opinion holder extraction and show that the learnt knowledge also improves a state-of-the-art classifier trained on BERT. Sentiment-view classification is typically framed as a task in which only little labeled training data are available. As in the case of unigrams, we show that for MWEs a feature-based approach beats state-of-the-art generic methods.
Michael Wiegand, Marc Schulder, Josef Ruppenhofer
Nat. Lang. Eng.1
2023 Euphemistic Abuse - A New Dataset and Classification Experiments for Implicitly Abusive Language
abstract
We address the task of identifying euphemistic abuse (e.g.You inspire me to fall asleep) paraphrasing simple explicitly abusive utterances (e.g.You are boring).For this task, we introduce a novel dataset that has been created via crowdsourcing.Special attention has been paid to the generation of appropriate negative (nonabusive) data.We report on classification experiments showing that classifiers trained on previous datasets are less capable of detecting such abuse.Best automatic results are obtained by a classifier that augments training data from our new dataset with automatically-generated GPT-3 completions.We also present a classifier that combines a few manually extracted features that exemplify the major linguistic phenomena constituting euphemistic abuse.
Michael Wiegand, Jana Kampfmeier, Elisabeth Eder, Josef Ruppenhofer
EMNLP1
2022 Biographically Relevant Tweets - a New Dataset, Linguistic Analysis and Classification Experiments
abstract
We present a new dataset comprising tweets for the novel task of detecting biographically relevant utterances. Biographically relevant utterances are all those utterances that reveal some persistent and non-trivial information about the author of a tweet, e.g. habits, (dis)likes, family status, physical appearance, employment information, health issues etc. Unlike previous research we do not restrict biographical relevance to a small fixed set of pre-defined relations. Next to classification experiments employing state-of-the-art classifiers to establish strong baselines for future work, we carry out a linguistic analysis that compares the predictiveness of various high-level features. We also show that the task is different from established tasks, such as aspectual classification or sentiment analysis.
Michael Wiegand, Rebecca Wilm, Katja Markert
COLING1
2022 "Beste Grüße, Maria Meyer" - Pseudonymization of Privacy-Sensitive Information in Emails
abstract
The exploding amount of user-generated content has spurred NLP research to deal with documents from various digital communication formats (tweets, chats, emails, etc.). Using these texts as language resources implies complying with legal data privacy regulations. To protect the personal data of individuals and preclude their identification, we employ pseudonymization. More precisely, we identify those text spans that carry information revealing an individual’s identity (e.g., names of persons, locations, phone numbers, or dates) and subsequently substitute them with synthetically generated surrogates. Based on CodE Alltag, a German-language email corpus, we address two tasks. The first task is to evaluate various architectures for the automatic recognition of privacy-sensitive entities in raw data. The second task examines the applicability of pseudonymized data as training data for such systems since models learned on original data cannot be published for reasons of privacy protection. As outputs of both tasks, we, first, generate a new pseudonymized version of CodE Alltag compliant with the legal requirements of the General Data Protection Regulation (GDPR). Second, we make accessible a tagger for recognizing privacy-sensitive information in German emails and similar text genres, which is trained on already pseudonymized data.
Elisabeth Eder, Michael Wiegand, Ulrike Krieg-Holz, Udo Hahn
LREC2
2022 Identifying Implicitly Abusive Remarks about Identity Groups using a Linguistically Informed Approach
abstract
Michael Wiegand, Elisabeth Eder, Josef Ruppenhofer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Michael Wiegand, Elisabeth Eder, Josef Ruppenhofer
NAACL-HLT1
2021 Implicitly Abusive Comparisons - A New Dataset and Linguistic Analysis
abstract
We examine the task of detecting implicitly abusive comparisons (e.g.Your hair looks like you have been electrocuted).Implicitly abusive comparisons are abusive comparisons in which abusive words (e.g.dumbass or scum) are absent.We detail the process of creating a novel dataset for this task via crowdsourcing that includes several measures to obtain a sufficiently representative and unbiased set of comparisons.We also present classification experiments that include a range of linguistic features that help us better understand the mechanisms underlying abusive comparisons.
Michael Wiegand, Maja Geulig, Josef Ruppenhofer
EACL1
2021 Exploiting Emojis for Abusive Language Detection
abstract
We propose to use abusive emojis, such as the middle finger or face vomiting, as a proxy for learning a lexicon of abusive words.Since it represents extralinguistic information, a single emoji can co-occur with different forms of explicitly abusive utterances.We show that our approach generates a lexicon that offers the same performance in cross-domain classification of abusive microposts as the most advanced lexicon induction method.Such an approach, in contrast, is dependent on manually annotated seed words and expensive lexical resources for bootstrapping (e.g.WordNet).We demonstrate that the same emojis can also be effectively used in languages other than English.Finally, we also show that emojis can be exploited for classifying mentions of ambiguous words, such as fuck and bitch, into generally abusive and just profane usages.
Michael Wiegand, Josef Ruppenhofer
EACL1
2021 Implicitly Abusive Language - What does it actually look like and why are we not getting there?
abstract
Michael Wiegand, Josef Ruppenhofer, Elisabeth Eder. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Michael Wiegand, Josef Ruppenhofer, Elisabeth Eder
NAACL-HLT1
2021 Python for Linguists
abstract
Teaching programming skills is a hard task.It is even harder if one targets an audience with no or little mathematical background.Although there are books on programming that target such groups, they often fail to raise or maintain interest due to artificial examples that lack reference to the professional issues that the audience typically face.This book fills the gap by addressing linguistics, a profession and academic subject for which basic knowledge of script programming is becoming more and more important.The book Python for Linguists by Michael Hammond is an introductory Python course targeted at linguists with no prior programming background.It succeeds previous books for Perl (Hammond 2008) and Java (Hammond 2002) by the same author, and reflects the current de facto prevalence of Python when it comes to adoption and available packages for natural language processing.We feel it necessary to clarify that the book aims at (general) linguists in the broad sense rather than computational linguists.Its aim is to teach linguists the fundamental concepts of programming using typical examples from linguistics.The book should not be mistaken as a course for learning basic algorithms in computational linguistics.We acknowledge that the author nowhere makes such a claim; however, given the thematic proximity to computational linguistics, one should have the right expectation before working with the book.Chapters 1-5 lay the foundations of the Python programming language, introducing the most important language constructs but deferring object oriented programming to a later part of the book.The focus in Chapters 1 and 2 covers the basic data types (numbers, strings, dictionaries), with a particular emphasis on simple string operations, and introduces some more advanced concepts such as mutability.Chapters 3-5 introduce control structures, input-output operations, and modules.The book goes at great length to visualize the program flow and the state of different variables for different steps in a program execution, which is certainly very helpful for learners with no prior programming experience.The book also guides the learner to understand certain error types that frequently occur in computer programming (but might be unintuitive for beginners).For example, when discussing function calls, much care is devoted to pointing out the unintended consequences stemming from mutability and side effects.
Benjamin Roth 0001, Michael Wiegand
Comput. Linguistics2
2021 Automatic generation of lexica for sentiment polarity shifters
abstract
Abstract Alleviating pain is good and abandoning hope is bad. We instinctively understand how words like alleviate and abandon affect the polarity of a phrase, inverting or weakening it. When these words are content words, such as verbs, nouns, and adjectives, we refer to them as polarity shifters. Shifters are a frequent occurrence in human language and an important part of successfully modeling negation in sentiment analysis; yet research on negation modeling has focused almost exclusively on a small handful of closed-class negation words, such as not, no, and without. A major reason for this is that shifters are far more lexically diverse than negation words, but no resources exist to help identify them. We seek to remedy this lack of shifter resources by introducing a large lexicon of polarity shifters that covers English verbs, nouns, and adjectives. Creating the lexicon entirely by hand would be prohibitively expensive. Instead, we develop a bootstrapping approach that combines automatic classification with human verification to ensure the high quality of our lexicon while reducing annotation costs by over 70%. Our approach leverages a number of linguistic insights; while some features are based on textual patterns, others use semantic resources or syntactic relatedness. The created lexicon is evaluated both on a polarity shifter gold standard and on a polarity classification task.
Marc Schulder, Michael Wiegand, Josef Ruppenhofer
Nat. Lang. Eng.2
2020 Doctor Who? Framing Through Names and Titles in German
abstract
Entity framing is the selection of aspects of an entity to promote a particular viewpoint towards that entity. We investigate entity framing of political figures through the use of names and titles in German online discourse, enhancing current research in entity framing through titling and naming that concentrates on English only. We collect tweets that mention prominent German politicians and annotate them for stance. We find that the formality of naming in these tweets correlates positively with their stance. This confirms sociolinguistic observations that naming and titling can have a status-indicating function and suggests that this function is dominant in German tweets mentioning political figures. We also find that this status-indicating function is much weaker in tweets from users that are politically left-leaning than in tweets by right-leaning users. This is in line with observations from moral psychology that left-leaning and right-leaning users assign different importance to maintaining social hierarchies.
Esther van den Berg, Katharina Korfhage, Josef Ruppenhofer, Michael Wiegand, Katja Markert
LREC4
2020 Enhancing a Lexicon of Polarity Shifters through the Supervised Classification of Shifting Directions
abstract
The sentiment polarity of an expression (whether it is perceived as positive, negative or neutral) can be influenced by a number of phenomena, foremost among them negation. Apart from closed-class negation words like “no”, “not” or “without”, negation can also be caused by so-called polarity shifters. These are content words, such as verbs, nouns or adjectives, that shift polarities in their opposite direction, e.g. “abandoned” in “abandoned hope” or “alleviate” in “alleviate pain”. Many polarity shifters can affect both positive and negative polar expressions, shifting them towards the opposing polarity. However, other shifters are restricted to a single shifting direction. “Recoup” shifts negative to positive in “recoup your losses”, but does not affect the positive polarity of “fortune” in “recoup a fortune”. Existing polarity shifter lexica only specify whether a word can, in general, cause shifting, but they do not specify when this is limited to one shifting direction. To address this issue we introduce a supervised classifier that determines the shifting direction of shifters. This classifier uses both resource-driven features, such as WordNet relations, and data-driven features like in-context polarity conflicts. Using this classifier we enhance the largest available polarity shifter lexicon.
Marc Schulder, Michael Wiegand, Josef Ruppenhofer
LREC2
2018 Distinguishing affixoid formations from compounds
abstract
We study German affixoids, a type of morpheme in between affixes and free stems. Several properties have been associated with them – increased productivity; a bleached semantics, which is often evaluative and/or intensifying and thus of relevance to sentiment analysis; and the existence of a free morpheme counterpart – but not been validated empirically. In experiments on a new data set that we make available, we put these key assumptions from the morphological literature to the test and show that despite the fact that affixoids generate many low-frequency formations, we can classify these as affixoid or non-affixoid instances with a best F1-score of 74%.
Josef Ruppenhofer, Michael Wiegand, Rebecca Wilm, Katja Markert
COLING2
2018 Automatically Creating a Lexicon of Verbal Polarity Shifters: Mono- and Cross-lingual Methods for German
abstract
In this paper we use methods for creating a large lexicon of verbal polarity shifters and apply them to German. Polarity shifters are content words that can move the polarity of a phrase towards its opposite, such as the verb “abandon” in “abandon all hope”. This is similar to how negation words like “not” can influence polarity. Both shifters and negation are required for high precision sentiment analysis. Lists of negation words are available for many languages, but the only language for which a sizable lexicon of verbal polarity shifters exists is English. This lexicon was created by bootstrapping a sample of annotated verbs with a supervised classifier that uses a set of data- and resource-driven features. We reproduce and adapt this approach to create a German lexicon of verbal polarity shifters. Thereby, we confirm that the approach works for multiple languages. We further improve classification by leveraging cross-lingual information from the English shifter lexicon. Using this improved approach, we bootstrap a large number of German verbal polarity shifters, reducing the annotation effort drastically. The resulting German lexicon of verbal polarity shifters is made publicly available.
Marc Schulder, Michael Wiegand, Josef Ruppenhofer
COLING2
2018 Introducing a Lexicon of Verbal Polarity Shifters for English
Marc Schulder, Michael Wiegand, Josef Ruppenhofer, Stephanie Köser
LREC2
2018 Disambiguation of Verbal Shifters
Michael Wiegand, Sylvette Loda, Josef Ruppenhofer
LREC1
2018 Inducing a Lexicon of Abusive Words - a Feature-Based Approach
abstract
Michael Wiegand, Josef Ruppenhofer, Anna Schmidt, Clayton Greenberg. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Michael Wiegand, Josef Ruppenhofer, Anna Schmidt, Clayton Greenberg
NAACL-HLT1
2017 Towards Bootstrapping a Polarity Shifter Lexicon using Linguistic Features
abstract
We present a major step towards the creation of the first high-coverage lexicon of polarity shifters. In this work, we bootstrap a lexicon of verbs by exploiting various linguistic features. Polarity shifters, such as “abandon”, are similar to negations (e.g. “not”) in that they move the polarity of a phrase towards its inverse, as in “abandon all hope”. While there exist lists of negation words, creating comprehensive lists of polarity shifters is far more challenging due to their sheer number. On a sample of manually annotated verbs we examine a variety of linguistic features for this task. Then we build a supervised classifier to increase coverage. We show that this approach drastically reduces the annotation effort while ensuring a high-precision lexicon. We also show that our acquired knowledge of verbal polarity shifters improves phrase-level sentiment analysis.
Marc Schulder, Michael Wiegand, Josef Ruppenhofer, Benjamin Roth 0001
IJCNLP(1)2
2016 Opinion Holder and Target Extraction on Opinion Compounds - A Linguistic Approach
abstract
Michael Wiegand, Christine Bocionek, Josef Ruppenhofer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Michael Wiegand, Christine Bocionek, Josef Ruppenhofer
HLT-NAACL1
2016 Separating Actor-View from Speaker-View Opinion Expressions using Linguistic Features
abstract
We examine different features and classifiers for the categorization of opinion words into actor and speaker view.To our knowledge, this is the first comprehensive work to address sentiment views on the word level taking into consideration opinion verbs, nouns and adjectives.We consider many high-level features requiring only few labeled training data.A detailed feature analysis produces linguistic insights into the nature of sentiment views.We also examine how far global constraints between different opinion words help to increase classification performance.Finally, we show that our (prior) word-level annotation correlates with contextual sentiment views.
Michael Wiegand, Marc Schulder, Josef Ruppenhofer
HLT-NAACL1
2015 Opinion Holder and Target Extraction based on the Induction of Verbal Categories
abstract
We present an approach for opinion role induction for verbal predicates.Our model rests on the assumption that opinion verbs can be divided into three different types where each type is associated with a characteristic mapping between semantic roles and opinion holders and targets.In several experiments, we demonstrate the relevance of those three categories for the task.We show that verbs can easily be categorized with semi-supervised graphbased clustering and some appropriate similarity metric.The seeds are obtained through linguistic diagnostics.We evaluate our approach against a new manuallycompiled opinion role lexicon and perform in-context classification.
Michael Wiegand, Josef Ruppenhofer
CoNLL1
2015 Combining Pattern-Based and Distributional Similarity for Graph-Based Noun Categorization
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow
NLDB1
2014 Separating Brands from Types: an Investigation of Different Features for the Food Domain
Michael Wiegand, Dietrich Klakow
COLING1
2014 Comparing methods for deriving intensity scores for adjectives
abstract
We compare several different corpus-based and lexicon-based methods for the scalar ordering of adjectives. Among them, we examine for the first time a low-resource approach based on distinctive-collexeme analysis that just requires a small predefined set of adverbial modi-fiers. While previous work on adjective in-tensity mostly assumes one single scale for all adjectives, we group adjectives into dif-ferent scales which is more faithful to hu-man perception. We also apply the meth-ods to both polar and non-polar adjectives, showing that not all methods are equally suitable for both types of adjectives. 1
Josef Ruppenhofer, Michael Wiegand, Jasper Brandes
EACL2
2014 Automatic Food Categorization from Large Unlabeled Corpora and Its Impact on Relation Extraction
abstract
We present a weakly-supervised induction method to assign semantic information to food items.We consider two tasks of categorizations being food-type classification and the distinction of whether a food item is composite or not.The categorizations are induced by a graph-based algorithm applied on a large unlabeled domain-specific corpus.We show that the usage of a domain-specific corpus is vital.We do not only outperform a manually designed open-domain ontology but also prove the usefulness of these categorizations in relation extraction, outperforming state-of-the-art features that include syntactic information and Brown clustering.
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow
EACL1
2013 Towards Contextual Healthiness Classification of Food Items - A Linguistic Approach
Michael Wiegand, Dietrich Klakow
IJCNLP1
2013 Predicative Adjectives: An Unsupervised Criterion to Extract Subjective Adjectives
Michael Wiegand, Josef Ruppenhofer, Dietrich Klakow
HLT-NAACL1
2012 Generalization Methods for In-Domain and Cross-Domain Opinion Holder Extraction
Michael Wiegand, Dietrich Klakow
EACL1
2012 MLSA - A Multi-layered Reference Corpus for German Sentiment Analysis
Simon Clematide, Stefan Gindl, Manfred Klenner, Stefanos Petrakis, Robert Remus, Josef Ruppenhofer, Ulli Waltinger, Michael Wiegand
LREC8
2012 A Gold Standard for Relation Extraction in the Food Domain
Michael Wiegand, Benjamin Roth 0001, Eva Lasarcyk, Stephanie Köser, Dietrich Klakow
LREC1
2012 Web-Based Relation Extraction for the Food Domain
Michael Wiegand, Benjamin Roth 0001, Dietrich Klakow
NLDB1
2010 Predictive Features for Detecting Indefinite Polar Sentences
Michael Wiegand, Dietrich Klakow
LREC1
2010 Convolution Kernels for Opinion Holder Extraction
Michael Wiegand, Dietrich Klakow
HLT-NAACL1
2008 Optimizing Language Models for Polarity Classification
Michael Wiegand, Dietrich Klakow
ECIR1
2008 Cost-Sensitive Learning in Answer Extraction
Michael Wiegand, Jochen L. Leidner, Dietrich Klakow
LREC1
2007 Combining term-based and event-based matching for question answering
abstract
In question answering, two main kinds of matching methods for finding answer sentences for a question are term-based approaches -- which are simple, efficient, effective, and yield high recall -- and event-based approaches that take syntactic and semantic information into account. The latter often sacrifice recall for increased precision, but actually capture the meaning of the events denoted by the textual units of a passage or sentence. We propose a robust, data-driven method that learns the mapping between questions and answers using logistic regression and show that combining term-based and event-based approaches significantly outperforms the individual methods.
Michael Wiegand, Jochen L. Leidner, Dietrich Klakow
SIGIR1