VLDB 2026 Research / reviewers in the wild / expert
Henning Wachsmuth
dblp:73/9281
· DBLP profile ↗
72ranked-venue papers
20as first author
32since 2021 · last 2026
0000-0003-2792-621XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 19 first-author · 27 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement LearningabstractEditing human-written text has become a standard use case of large language models (LLMs), for example, to make one's arguments more appropriate for a discussion.Comparing human to LLM-generated edits, however, we observe a mismatch in editing strategies: While LLMs often perform multiple scattered edits and tend to change meaning notably, humans rather encapsulate dependent changes in selfcontained, meaning-preserving edits.In this paper, we present a reinforcement learning approach that teaches LLMs human-like editing to improve the appropriateness of arguments.Our approach produces self-contained sentence-level edit suggestions that can be accepted or rejected independently.We train the approach using group relative policy optimization with a multi-component reward function that jointly optimizes edit-level semantic similarity, fluency, and pattern conformity as well as argument-level appropriateness.In automatic and human evaluation, it outperforms competitive baselines and the state of the art in humanlike editing, with multi-round editing achieving appropriateness close to full rewriting. Timon Ziegenbein, Maja Stahl, Henning Wachsmuth |
ACL (1) | 3 |
| 2026 | Assessing the Persuasive Effect of AI-Generated Image Support of Arguments
Mackwyn Quadras, Manfred Stede, Henning Wachsmuth |
LREC | 3 |
| 2025 | Towards a Perspectivist Turn in Argument Quality AssessmentabstractJulia Romberg, Maximilian Maurer, Henning Wachsmuth, Gabriella Lapesa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Julia Romberg, Maximilian Maurer, Henning Wachsmuth, Gabriella Lapesa |
NAACL (Long Papers) | 3 |
| 2025 | Adaptive Prompting: Ad-hoc Prompt Composition for Social Bias DetectionabstractMaximilian Spliethöver, Tim Knebler, Fabian Fumagalli, Maximilian Muschalik, Barbara Hammer, Eyke Hüllermeier, Henning Wachsmuth. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Maximilian Spliethöver, Tim Knebler, Fabian Fumagalli, Maximilian Muschalik, Barbara Hammer, Eyke Hüllermeier, Henning Wachsmuth |
NAACL (Long Papers) | 7 |
| 2025 | Investigating Co-Constructive Behavior of Large Language Models in Explanation DialoguesabstractThe ability to generate explanations that are understood by explainees is the quintessence of explainable artificial intelligence. Since understanding depends on the explainee’s background and needs, recent research focused on co-constructive explanation dialogues, where an explainer continuously monitors the explainee’s understanding and adapts their explanations dynamically. We investigate the ability of large language models (LLMs) to engage as explainers in co-constructive explanation dialogues. In particular, we present a user study in which explainees interact with an LLM in two settings, one of which involves the LLM being instructed to explain a topic co-constructively. We evaluate the explainees’ understanding before and after the dialogue, as well as their perception of the LLMs’ co-constructive behavior. Our results suggest that LLMs show some co-constructive behaviors, such as asking verification questions, that foster the explainees’ engagement and can improve understanding of a topic. However, their ability to effectively monitor the current understanding and scaffold the explanations accordingly remains limited. Leandra Fichtel, Maximilian Spliethöver, Eyke Hüllermeier, Patricia Jimenez, Nils Oliver Klowait, Stefan Kopp, Axel-Cyrille Ngonga Ngomo, Amelie Sophie Robrecht, Ingrid Scharlau, Lutz Terfloth, Anna-Lisa Vollmer, Henning Wachsmuth |
SIGDIAL | 12 |
| 2024 | LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine FeedbackabstractTimon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning Wachsmuth. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Timon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning Wachsmuth |
ACL (1) | 4 |
| 2024 | Reference-guided Style-Consistent Content TransferabstractIn this paper, we introduce the task of style-consistent content transfer, which concerns modifying a text’s content based on a provided reference statement while preserving its original style. We approach the task by employing multi-task learning to ensure that the modified text meets three important conditions: reference faithfulness, style adherence, and coherence. In particular, we train three independent classifiers for each condition. During inference, these classifiers are used to determine the best modified text variant. Our evaluation, conducted on hotel reviews and news articles, compares our approach with sequence-to-sequence and error correction baselines. The results demonstrate that our approach reasonably generates text satisfying all three conditions. In subsequent analyses, we highlight the strengths and limitations of our approach, providing valuable insights for future research directions. Wei-Fan Chen 0001, Milad Alshomary, Maja Stahl, Khalid Al-Khatib, Benno Stein 0001, Henning Wachsmuth |
LREC/COLING | 6 |
| 2024 | Modeling the Quality of Dialogical ExplanationsabstractExplanations are pervasive in our lives. Mostly, they occur in dialogical form where an explainer discusses a concept or phenomenon of interest with an explainee. Leaving the explainee with a clear understanding is not straightforward due to the knowledge gap between the two participants. Previous research looked at the interaction of explanation moves, dialogue acts, and topics in successful dialogues with expert explainers. However, daily-life explanations often fail, raising the question of what makes a dialogue successful. In this work, we study explanation dialogues in terms of the interactions between the explainer and explainee and how they correlate with the quality of explanations in terms of a successful understanding on the explainee’s side. In particular, we first construct a corpus of 399 dialogues from the Reddit forum Explain Like I am Five and annotate it for interaction flows and explanation quality. We then analyze the interaction flows, comparing them to those appearing in expert dialogues. Finally, we encode the interaction flows using two language models that can handle long inputs, and we provide empirical evidence for the effectiveness boost gained through the encoding in predicting the success of explanation dialogues. Milad Alshomary, Felix Lange 0001, Meisam Booshehri, Meghdut Sengupta, Philipp Cimiano, Henning Wachsmuth |
LREC/COLING | 6 |
| 2024 | The Touché23-ValueEval Dataset for Identifying Human Values behind ArgumentsabstractWhile human values play a crucial role in making arguments persuasive, we currently lack the necessary extensive datasets to develop methods for analyzing the values underlying these arguments on a large scale. To address this gap, we present the Touché23-ValueEval dataset, an expansion of the Webis-ArgValues-22 dataset. We collected and annotated an additional 4780 new arguments, doubling the dataset’s size to 9324 arguments. These arguments were sourced from six diverse sources, covering religious texts, community discussions, free-text arguments, newspaper editorials, and political debates. Each argument is annotated by three crowdworkers for 54 human values, following the methodology established in the original dataset. The Touché23-ValueEval dataset was utilized in the SemEval 2023 Task 4. ValueEval: Identification of Human Values behind Arguments, where an ensemble of transformer models demonstrated state-of-the-art performance. Furthermore, our experiments show that a fine-tuned large language model, Llama-2-7B, achieves comparable results. Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Valentin Barrière, Doratossadat Dastgheib, Omid Ghahroodi, Mohammad Ali Sadraei, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, Benno Stein 0001 |
LREC/COLING | 13 |
| 2024 | Argument Quality Assessment in the Age of Instruction-Following Large Language ModelsabstractThe computational treatment of arguments on controversial issues has been subject to extensive NLP research, due to its envisioned impact on opinion formation, decision making, writing education, and the like. A critical task in any such application is the assessment of an argument’s quality - but it is also particularly challenging. In this position paper, we start from a brief survey of argument quality research, where we identify the diversity of quality notions and the subjectiveness of their perception as the main hurdles towards substantial progress on argument quality assessment. We argue that the capabilities of instruction-following large language models (LLMs) to leverage knowledge across contexts enable a much more reliable assessment. Rather than just fine-tuning LLMs towards leaderboard chasing on assessment tasks, they need to be instructed systematically with argumentation theories and scenarios as well as with ways to solve argument-related problems. We discuss the real-world opportunities and ethical issues emerging thereby. Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, Timon Ziegenbein |
LREC/COLING | 1 |
| 2024 | Overview of Touché 2024: Argumentation Systems
Johannes Kiesel, Çagri Çöltekin, Maximilian Heinrich, Maik Fröbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Münstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 17 |
| 2024 | Analyzing the Use of Metaphors in News Editorials for Political FramingabstractMeghdut Sengupta, Roxanne El Baff, Milad Alshomary, Henning Wachsmuth. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Meghdut Sengupta, Roxanne El Baff, Milad Alshomary, Henning Wachsmuth |
NAACL-HLT | 4 |
| 2024 | A School Student Essay Corpus for Analyzing Interactions of Argumentative Structure and QualityabstractMaja Stahl, Nadine Michel, Sebastian Kilsbach, Julian Schmidtke, Sara Rezat, Henning Wachsmuth. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Maja Stahl, Nadine Michel, Sebastian Kilsbach, Julian Schmidtke, Sara Rezat, Henning Wachsmuth |
NAACL-HLT | 6 |
| 2023 | To Revise or Not to Revise: Learning to Detect Improvable Claims for Argumentative Writing SupportabstractOptimizing the phrasing of argumentative text is crucial in higher education and professional development.However, assessing whether and how the different claims in a text should be revised is a hard task, especially for novice writers.In this work, we explore the main challenges to identifying argumentative claims in need of specific revisions.By learning from collaborative editing behaviors in online debates, we seek to capture implicit revision patterns in order to develop approaches aimed at guiding writers in how to further improve their arguments.We systematically compare the ability of common word embedding models to capture the differences between different versions of the same text, and we analyze their impact on various types of writing issues.To deal with the noisy nature of revision-based corpora, we propose a new sampling strategy based on revision distance.Opposed to approaches from prior work, such sampling can be done without employing additional annotations and judgments.Moreover, we provide evidence that using contextual information and domain knowledge can further improve prediction results.How useful a certain type of context is, depends on the issue the claim is suffering from, though. Gabriella Skitalinskaya, Henning Wachsmuth |
ACL (1) | 2 |
| 2023 | Modeling Appropriate Language in ArgumentationabstractOnline discussion moderators must make adhoc decisions about whether the contributions of discussion participants are appropriate or should be removed to maintain civility.Existing research on offensive language and the resulting tools cover only one aspect among many involved in such decisions.The question of what is considered appropriate in a controversial discussion has not yet been systematically addressed.In this paper, we operationalize appropriate language in argumentation for the first time.In particular, we model appropriateness through the absence of flaws, grounded in research on argument quality assessment, especially in aspects from rhetoric.From these, we derive a new taxonomy of 14 dimensions that determine inappropriate language in online discussions.Building on three argument quality corpora, we then create a corpus of 2191 arguments annotated for the 14 dimensions.Empirical analyses support that the taxonomy covers the concept of appropriateness comprehensively, showing several plausible correlations with argument quality dimensions.Moreover, results of baseline approaches to assessing appropriateness suggest that all dimensions can be modeled computationally on the corpus. Timon Ziegenbein, Shahbaz Syed, Felix Lange 0001, Martin Potthast, Henning Wachsmuth |
ACL (1) | 5 |
| 2023 | Conclusion-based Counter-Argument GenerationabstractIn real-world debates, the most common way to counter an argument is to reason against its main point, that is, its conclusion.Existing work on the automatic generation of natural language counter-arguments does not address the relation to the conclusion, possibly because many arguments leave their conclusion implicit.In this paper, we hypothesize that the key to effective counter-argument generation is to explicitly model the argument's conclusion and to enforce that the stance of the generated counter is opposite to that conclusion.In particular, we propose a multitask approach that jointly learns to generate both the conclusion and the counter of an input argument.The approach employs a stance-based ranking component that selects the counter from a diverse set of generated candidates whose stance best opposes the generated conclusion.In both automatic and manual evaluation, we provide evidence that our approach generates more relevant and stanceadhering counters than strong baselines. Milad Alshomary, Henning Wachsmuth |
EACL | 2 |
| 2023 | Claim Optimization in Computational ArgumentationabstractAn optimal delivery of arguments is key to persuasion in any debate, both for humans and for AI systems.This requires the use of clear and fluent claims relevant to the given debate.Prior work has studied the automatic assessment of argument quality extensively.Yet, no approach actually improves the quality so far.To fill this gap, this paper proposes the task of claim optimization: to rewrite argumentative claims in order to optimize their delivery.As multiple types of optimization are possible, we approach this task by first generating a diverse set of candidate claims using a large language model, such as BART, taking into account contextual information.Then, the best candidate is selected using various quality metrics.In automatic and human evaluation on an English-language corpus, our quality-based candidate selection outperforms several baselines, improving 60% of all claims (worsening 16% only).Follow-up analyses reveal that, beyond copy editing, our approach often specifies claims with details, whereas it adds less evidence than humans do.Moreover, its capabilities generalize well to other domains, such as instructional texts. Gabriella Skitalinskaya, Maximilian Spliethöver, Henning Wachsmuth |
INLG | 3 |
| 2023 | Supporting Requesters in Writing Clear Crowdsourcing Task Descriptions Through Computational Flaw AssessmentabstractQuality control is an, if not the, essential challenge in crowdsourcing. Unsatisfactory responses from crowd workers have been found to particularly result from ambiguous and incomplete task descriptions, often from inexperienced task requesters. However, creating clear task descriptions with sufficient information is a complex process for requesters in crowdsourcing marketplaces. In this paper, we investigate the extent to which requesters can be supported effectively in this process through computational techniques. To this end, we developed a tool that enables requesters to iteratively identify and correct eight common clarity flaws in their task descriptions before deployment on the platform. The tool can be used to write task descriptions from scratch or to assess and improve the clarity of prepared descriptions. It employs machine learning-based natural language processing models trained on real-world task descriptions that score a given task description for the eight clarity flaws. On this basis, the requester can iteratively revise and reassess the task description until it reaches a sufficient level of clarity. In a first user study, we let requesters create task descriptions using the tool and rate the tool’s different aspects of helpfulness thereafter. We then carried out a second user study with crowd workers, as those who are confronted with such descriptions in practice, to rate the clarity of the created task descriptions. According to our results, 65% of the requesters classified the helpfulness of the information provided by the tool high or very high (only 12% as low or very low). The requesters saw some room for improvement though, for example, concerning the display of bad examples. Nevertheless, 76% of the crowd workers believe that the overall clarity of the task descriptions created by the requesters using the tool improves over the initial version. In line with this, the automatically-computed clarity scores of the edited task descriptions were generally higher than those of the initial descriptions, indicating that the tool reliably predicts the clarity of task descriptions in overall terms. Zahra Nouri, Nikhil Prakash, Ujwal Gadiraju, Henning Wachsmuth |
IUI | 4 |
| 2023 | Frame-oriented Summarization of Argumentative DiscussionsabstractOnline discussions on controversial topics with many participants frequently include hundreds of arguments that cover different framings of the topic.But these arguments and frames are often spread across the various branches of the discussion tree structure.This makes it difficult for interested participants to follow the discussion in its entirety as well as to introduce new arguments.In this paper, we present a new rankbased approach to extractive summarization of online discussions focusing on argumentation frames that capture the different aspects of a discussion.Our approach includes three retrieval tasks to find arguments in a discussion that are (1) relevant to a frame of interest, (2) relevant to the topic under discussion, and (3) informative to the reader.Based on a joint ranking by these three criteria for a set of user-selected frames, our approach allows readers to quickly access an ongoing discussion.We evaluate our approach using a test set of 100 controversial Reddit ChangeMyView discussions, for which the relevance of a total of 1871 arguments was manually annotated. Shahbaz Syed, Timon Ziegenbein, Philipp Heinisch, Henning Wachsmuth, Martin Potthast |
SIGDIAL | 4 |
| 2022 | The Moral Debater: A Study on the Computational Generation of Morally Framed ArgumentsabstractAn audience's prior beliefs and morals are strong indicators of how likely they will be affected by a given argument.Utilizing such knowledge can help focus on shared values to bring disagreeing parties towards agreement.In argumentation technology, however, this is barely exploited so far.This paper studies the feasibility of automatically generating morally framed arguments as well as their effect on different audiences.Following the moral foundation theory, we propose a system that effectively generates arguments focusing on different morals.In an in-depth user study, we ask liberals and conservatives to evaluate the impact of these arguments.Our results suggest that, particularly when prior beliefs are challenged, an audience becomes more affected by morally framed arguments. Milad Alshomary, Roxanne El Baff, Timon Ziegenbein, Henning Wachsmuth |
ACL (1) | 4 |
| 2022 | Identifying the Human Values behind ArgumentsabstractJohannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, Benno Stein. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Johannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, Benno Stein 0001 |
ACL (1) | 5 |
| 2022 | "Mama Always Had a Way of Explaining Things So I Could Understand": A Dialogue Corpus for Learning to Construct ExplanationsabstractAs AI is more and more pervasive in everyday life, humans have an increasing demand to understand its behavior and decisions. Most research on explainable AI builds on the premise that there is one ideal explanation to be found. In fact, however, everyday explanations are co-constructed in a dialogue between the person explaining (the explainer) and the specific person being explained to (the explainee). In this paper, we introduce a first corpus of dialogical explanations to enable NLP research on how humans explain as well as on how AI can learn to imitate this process. The corpus consists of 65 transcribed English dialogues from the Wired video series 5 Levels, explaining 13 topics to five explainees of different proficiency. All 1550 dialogue turns have been manually labeled by five independent professionals for the topic discussed as well as for the dialogue act and the explanation move performed. We analyze linguistic patterns of explainers and explainees, and we explore differences across proficiency levels. BERT-based baseline results indicate that sequence information helps predicting topics, acts, and moves effectively. Henning Wachsmuth, Milad Alshomary |
COLING | 1 |
| 2022 | Generating Contrastive Snippets for Argument SearchabstractIn argument search, snippets provide an overview of the aspects discussed by the arguments retrieved for a queried controversial topic. Existing work has focused on generating snippets that are representative of an argument’s content while remaining argumentative. In this work, we argue that the snippets should also be contrastive, that is, they should highlight the aspects that make an argument unique in the context of others. Thereby, aspect diversity is increased and redundancy is reduced. We present and compare two snippet generation approaches that jointly optimize representativeness and contrastiveness. According to our experiments, both approaches have advantages, and one is able to generate representative yet sufficiently contrastive snippets. Milad Alshomary, Jonas Rieskamp, Henning Wachsmuth |
COMMA | 3 |
| 2022 | Overview of Touché 2022: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Shahbaz Syed, Timon Ziegenbein, Meriem Beloucif, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen |
ECIR (2) | 10 |
| 2022 | Scientia Potentia Est - On the Role of Knowledge in Computational ArgumentationabstractAbstract Despite extensive research efforts in recent years, computational argumentation (CA) remains one of the most challenging areas of natural language processing. The reason for this is the inherent complexity of the cognitive processes behind human argumentation, which integrate a plethora of different types of knowledge, ranging from topic-specific facts and common sense to rhetorical knowledge. The integration of knowledge from such a wide range in CA requires modeling capabilities far beyond many other natural language understanding tasks. Existing research on mining, assessing, reasoning over, and generating arguments largely acknowledges that much more knowledge is needed to accurately model argumentation computationally. However, a systematic overview of the types of knowledge introduced in existing CA models is missing, hindering targeted progress in the field. Adopting the operational definition of knowledge as any task-relevant normative information not provided as input, the survey paper at hand fills this gap by (1) proposing a taxonomy of types of knowledge required in CA tasks, (2) systematizing the large body of CA work according to the reliance on and exploitation of these knowledge types for the four main research areas in CA, and (3) outlining and discussing directions for future research efforts in CA. Anne Lauscher, Henning Wachsmuth, Iryna Gurevych, Goran Glavas |
Trans. Assoc. Comput. Linguistics | 2 |
| 2021 | Syntopical Graphs for Computational Argumentation TasksabstractJoe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad Morariu, Varun Manjunatha, Douglas Oard, Philip Resnik, Henning Wachsmuth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Joe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik, Henning Wachsmuth |
ACL/IJCNLP (1) | 9 |
| 2021 | Employing Argumentation Knowledge Graphs for Neural Argument GenerationabstractKhalid Al Khatib, Lukas Trautner, Henning Wachsmuth, Yufang Hou, Benno Stein. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Khalid Al-Khatib, Lukas Trautner, Henning Wachsmuth, Yufang Hou 0001, Benno Stein 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Belief-based Generation of Argumentative ClaimsabstractWhen engaging in argumentative discourse, skilled human debaters tailor claims to the audience's beliefs to construct effective arguments.Recently, the field of computational argumentation witnessed extensive effort to address the automatic generation of arguments.However, existing approaches do not perform any audience-specific adaptation.In this work, we aim to bridge this gap by studying the task of belief-based claim generation: Given a controversial topic and a set of beliefs, generate an argumentative claim tailored to the beliefs.To tackle this task, we model the people's prior beliefs through their stances on controversial topics and extend state-of-the-art text generation models to generate claims conditioned on the beliefs.Our automatic evaluation confirms the ability of our approach to adapt claims to a set of given beliefs.In a manual study, we also evaluate the generated claims in terms of informativeness and their likelihood to be uttered by someone with a respective belief.Our results reveal the limitations of modeling users' beliefs based on their stances.Still, they demonstrate the potential of encoding beliefs into argumentative texts, laying the ground for future exploration of audience reach. Milad Alshomary, Wei-Fan Chen 0001, Timon Ziegenbein, Henning Wachsmuth |
EACL | 4 |
| 2021 | Learning From Revisions: Quality Assessment of Claims in Argumentation at ScaleabstractAssessing the quality of arguments and of the claims the arguments are composed of has become a key task in computational argumentation.However, even if different claims share the same stance on the same topic, their assessment depends on the prior perception and weighting of the different aspects of the topic being discussed.This renders it difficult to learn topic-independent quality indicators.In this paper, we study claim quality assessment irrespective of discussed aspects by comparing different revisions of the same claim.We compile a large-scale corpus with over 377k claim revision pairs of various types from kialo.com, covering diverse topics from politics, ethics, entertainment, and others.We then propose two tasks: (a) assessing which claim of a revision pair is better, and (b) ranking all versions of a claim by quality.Our first experiments with embedding-based logistic regression and transformer-based neural networks show promising results, suggesting that learned indicators generalize well across topics.In a detailed error analysis, we give insights into what quality dimensions of claims can be assessed reliably.We provide the data and scripts needed to reproduce all results. 1 Gabriella Skitalinskaya, Jonas Klaff, Henning Wachsmuth |
EACL | 3 |
| 2021 | Overview of Touché 2021: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Lukas Gienapp, Maik Fröbe, Meriem Beloucif, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen |
ECIR (2) | 9 |
| 2021 | Bias Silhouette Analysis: Towards Assessing the Quality of Bias Metrics for Word Embedding ModelsabstractWord embedding models reflect bias towards genders, ethnicities, and other social groups present in the underlying training data. Metrics such as ECT, RNSB, and WEAT quantify bias in these models based on predefined word lists representing social groups and bias-conveying concepts. How suitable these lists actually are to reveal bias - let alone the bias metrics in general - remains unclear, though. In this paper, we study how to assess the quality of bias metrics for word embedding models. In particular, we present a generic method, Bias Silhouette Analysis (BSA), that quantifies the accuracy and robustness of such a metric and of the word lists used. Given a biased and an unbiased reference embedding model, BSA applies the metric systematically for several subsets of the lists to the models. The variance and rate of convergence of the bias values of each model then entail the robustness of the word lists, whereas the distance between the models' values gives indications of the general accuracy of the metric with the word lists. We demonstrate the behavior of BSA on two standard embedding models for the three mentioned metrics with several word lists from existing research. Maximilian Spliethöver, Henning Wachsmuth |
IJCAI | 2 |
| 2021 | Visual Analysis of Argumentation in EssaysabstractThis paper presents a visual analytics system for exploring, analyzing and comparing argument structures in essay corpora. We provide an overview of the corpus by a list of ArguLines which represent the argument units of each essay by a sequence of glyphs. Each glyph encodes the stance, the depth and the relative position of an argument unit. The overview can be ordered in various ways to reveal patterns and outliers. Subsets of essays can be selected and analyzed in detail using the Argument Unit Occurrence Tree which aggregates the argument structures using hierarchical histograms. This hierarchical view facilitates the estimation of statistics and trends concerning the progression of the argumentation in the essays. It also provides insights into the commonalities and differences between selected subsets. The text view is the necessary textual basis to verify conclusions from the other views and the annotation process. Linking the views and interaction techniques for visual filtering, studying the evolution of stance within a subset of essays and scrutinizing the order of argumentative units enable a deep analysis of essay corpora. Our expert reviews confirmed the utility of the system and revealed detailed and previously unknown information about the argumentation in our sample corpus. Dora Kiesel, Patrick Riehmann, Henning Wachsmuth, Benno Stein 0001, Bernd Fröhlich 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | End-to-End Argumentation Knowledge Graph ConstructionabstractThis paper studies the end-to-end construction of an argumentation knowledge graph that is intended to support argument synthesis, argumentative question answering, or fake news detection, among others. The study is motivated by the proven effectiveness of knowledge graphs for interpretable and controllable text generation and exploratory search. Original in our work is that we propose a model of the knowledge encapsulated in arguments. Based on this model, we build a new corpus that comprises about 16k manual annotations of 4740 claims with instances of the model's elements, and we develop an end-to-end framework that automatically identifies all modeled types of instances. The results of experiments show the potential of the framework for building a web-based argumentation graph that is of high quality and large scale. Khalid Al-Khatib, Yufang Hou 0001, Henning Wachsmuth, Charles Jochim, Francesca Bonin, Benno Stein 0001 |
AAAI | 3 |
| 2020 | Target Inference in Argument Conclusion GenerationabstractIn argumentation, people state premises to reason towards a conclusion.The conclusion conveys a stance towards some target, such as a concept or statement.Often, the conclusion remains implicit, though, since it is self-evident in a discussion or left out for rhetorical reasons.However, the conclusion is key to understanding an argument, and hence, to any application that processes argumentation.We thus study the question to what extent an argument's conclusion can be reconstructed from its premises.In particular, we argue here that a decisive step is to infer a conclusion's target, and we hypothesize that this target is related to the premises' targets.We develop two complementary target inference approaches: one ranks premise targets and selects the top-ranked target as the conclusion target, the other finds a new conclusion target in a learned embedding space using a triplet neural network.Our evaluation on corpora from two domains indicates that a hybrid of both approaches is best, outperforming several strong baselines.According to human annotators, we infer a reasonably adequate conclusion target in 89% of the cases. Milad Alshomary, Shahbaz Syed, Martin Potthast, Henning Wachsmuth |
ACL | 4 |
| 2020 | Analyzing the Persuasive Effect of Style in News Editorial ArgumentationabstractNews editorials argue about political issues in order to challenge or reinforce the stance of readers with different ideologies.Previous research has investigated such persuasive effects for argumentative content.In contrast, this paper studies how important the style of news editorials is to achieve persuasion.To this end, we first compare content-and style-oriented classifiers on editorials from the liberal NYTimes with ideology-specific effect annotations.We find that conservative readers are resistant to NYTimes style, but on liberals, style even has more impact than content.Focusing on liberals, we then cluster the leads, bodies, and endings of editorials, in order to learn about writing style patterns of effective argumentation. Roxanne El Baff, Henning Wachsmuth, Khalid Al-Khatib, Benno Stein 0001 |
ACL | 2 |
| 2020 | Investigating Expectations for Voice-based and Conversational Argument Search on the WebabstractMillions of arguments are shared on the web. Future information systems will be able to exploit this valuable knowledge source and to retrieve arguments relevant and convincing to our specific need---all with an interface as intuitive as asking your friend "Why ...". Although recent advancements in argument mining, conversational search, and voice recognition have put such systems within reach, many questions remain open, especially on the interface side. In this regard the paper at hand presents the first study of argument search behavior. We conduct an online-survey and a focused user study, putting emphasis on what people expect argument search to be like, rather than on what current first-generation systems provide. Our participants expected to use voice-based argument search mostly at home, but also together with others. Moreover, they expect such search systems to provide rich information on retrieved arguments, such as the source, supporting evidence, and background knowledge on entities or events mentioned. In observed interactions with a simulated system we found that the participants adapted their search behavior to different types of tasks, and that up-front categorization of the retrieved arguments is perceived as helpful if this is short. Our findings are directly applicable to the design of argument search systems, not only voice-based ones. Johannes Kiesel, Kevin Lang, Henning Wachsmuth, Eva Hornecker, Benno Stein 0001 |
CHIIR | 3 |
| 2020 | CauseNet: Towards a Causality Graph Extracted from the WebabstractCausal knowledge is seen as one of the key ingredients to advance artificial intelligence. Yet, few knowledge bases comprise causal knowledge to date, possibly due to significant efforts required for validation. Notwithstanding this challenge, we compile CauseNet, a large-scale knowledge base of claimed causal relations between causal concepts. By extraction from different semi- and unstructured web sources, we collect more than 11 million causal relations with an estimated extraction precision of 83% and construct the first large-scale and open-domain causality graph. We analyze the graph to gain insights about causal beliefs expressed on the web and we demonstrate its benefits in basic causal question answering. Future work may use the graph for causal reasoning, computational argumentation, multi-hop question answering, and more. Stefan Heindorf, Yan Scholten, Henning Wachsmuth, Axel-Cyrille Ngonga Ngomo, Martin Potthast |
CIKM | 3 |
| 2020 | Mining Crowdsourcing Problems from Discussion Forums of WorkersabstractCrowdsourcing is used in academia and industry to solve tasks that are easy for humans but hard for computers, in natural language processing mostly to annotate data. The quality of annotations is affected by problems in the task design, task operation, and task evaluation that workers face with requesters in crowdsourcing processes. To learn about the major problems, we provide a short but comprehensive survey based on two complementary studies: (1) a literature review where we collect and organize problems known from interviews with workers, and (2) an empirical data analysis where we use topic modeling to mine workers' complaints from a new English corpus of workers' forum discussions. While literature covers all process phases, problems in the task evaluation are prevalent, including unfair rejections, late payments, and unjustified blockings of workers. According to the data, however, poor task design in terms of malfunctioning environments, bad workload estimation, and privacy violations seems to bother the workers most. Our findings form the basis for future research on how to improve crowdsourcing processes. Zahra Nouri, Henning Wachsmuth, Gregor Engels |
COLING | 2 |
| 2020 | Intrinsic Quality Assessment of ArgumentsabstractSeveral quality dimensions of natural language arguments have been investigated.Some are likely to be reflected in linguistic features (e.g., an argument's arrangement), whereas others depend on context (e.g., relevance) or topic knowledge (e.g., acceptability).In this paper, we study the intrinsic computational assessment of 15 dimensions, i.e., only learning from an argument's text.In systematic experiments with eight feature types on an existing corpus, we observe moderate but significant learning success for most dimensions.Rhetorical quality seems hardest to assess, and subjectivity features turn out strong, although length bias in the corpus impedes full validity.We also find that human assessors differ more clearly to each other than to our approach. Henning Wachsmuth, Till Werner |
COLING | 1 |
| 2020 | Touché: First Shared Task on Argument Retrieval
Alexander Bondarenko 0001, Matthias Hagen, Martin Potthast, Henning Wachsmuth, Meriem Beloucif, Chris Biemann, Alexander Panchenko, Benno Stein 0001 |
ECIR (2) | 4 |
| 2020 | Task Proposal: Abstractive Snippet Generation for Web PagesabstractWe propose a shared task on abstractive snippet generation for web pages, a novel task of generating query-biased abstractive summaries for documents that are to be shown on a search results page.Conventional snippets are extractive in nature, which recently gave rise to copyright claims from news publishers as well as a new copyright legislation being passed in the European Union, limiting the fair use of web page contents for snippets.At the same time, abstractive summarization has matured considerably in recent years, potentially allowing for more personalization of snippets in the future.Taken together, these facts render further research into generating abstractive snippets both timely and promising. Shahbaz Syed, Wei-Fan Chen 0001, Matthias Hagen, Benno Stein 0001, Henning Wachsmuth, Martin Potthast |
INLG | 5 |
| 2020 | Extractive Snippet Generation for ArgumentsabstractSnippets are used in web search to help users assess the relevance of retrieved results to their query. Recently, specialized search engines have arisen that retrieve pro and con arguments on controversial issues. We argue that standard snippet generation is insufficient to represent the core reasoning of an argument. In this paper, we introduce the task of generating a snippet that represents the main claim and reason of an argument. We propose a query-independent extractive summarization approach to this task that uses a variant of PageRank to assess the importance of sentences based on their context and argumentativeness. In both automatic and manual evaluation, our approach outperforms strong baselines. Milad Alshomary, Nick Düsterhus, Henning Wachsmuth |
SIGIR | 3 |
| 2019 | Wikipedia Text Reuse: Within and Without
Milad Alshomary, Michael Völske, Tristan Licht, Henning Wachsmuth, Benno Stein 0001, Matthias Hagen, Martin Potthast |
ECIR (1) | 4 |
| 2019 | Modeling Frames in ArgumentationabstractYamen Ajjour, Milad Alshomary, Henning Wachsmuth, Benno Stein. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yamen Ajjour, Milad Alshomary, Henning Wachsmuth, Benno Stein 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Computational Argumentation Synthesis as a Language Modeling TaskabstractSynthesis approaches in computational argumentation so far are restricted to generating claim-like argument units or short summaries of debates.Ultimately, however, we expect computers to generate whole new arguments for a given stance towards some topic, backing up claims following argumentative and rhetorical considerations.In this paper, we approach such an argumentation synthesis as a language modeling task.In our language model, argumentative discourse units are the "words", and arguments represent the "sentences".Given a pool of units for any unseen topic-stance pair, the model selects a set of unit types according to a basic rhetorical strategy (logos vs. pathos), arranges the structure of the types based on the units' argumentative roles, and finally "phrases" an argument by instantiating the structure with semantically coherent units from the pool.Our evaluation suggests that the model can, to some extent, mimic the human synthesis of strategy-specific arguments. Roxanne El Baff, Henning Wachsmuth, Khalid Al-Khatib, Manfred Stede, Benno Stein 0001 |
INLG | 2 |
| 2019 | Argument Search: Assessing Argument RelevanceabstractWe report on the first user study on assessing argument relevance. Based on a search among more than 300,000 arguments, four standard retrieval models are compared on 40 topics for 20 controversial issues: every issue has one topic with a biased stance and another neutral one. Following TREC, the top results of the different models on a topic were pooled and relevance-judged by one assessor per topic. The assessors also judged the arguments' rhetorical, logical, and dialectical quality, the results of which were cross-referenced with the relevance judgments. Furthermore, the assessors were asked for their personal opinion, and whether it matched the predefined stance of a topic. Among other results, we find that Terrier's implementations of DirichletLM and DPH are on par, significantly outperforming TFIDF and BM25. The judgments of relevance and quality hardly correlate, giving rise to a more diverse set of ranking criteria than relevance alone. We did not measure a significant bias of assessors when their stance is at odds with a topic's stance. Martin Potthast, Lukas Gienapp, Florian Euchner, Nick Heilenkötter, Nico Weidmann, Henning Wachsmuth, Benno Stein 0001, Matthias Hagen |
SIGIR | 6 |
| 2019 | Argumentation Mining
Henning Wachsmuth |
Comput. Linguistics | 1 |
| 2018 | Modeling Deliberative Argumentation Strategies on WikipediaabstractKhalid Al-Khatib, Henning Wachsmuth, Kevin Lang, Jakob Herpel, Matthias Hagen, Benno Stein. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Khalid Al-Khatib, Henning Wachsmuth, Kevin Lang, Jakob Herpel, Matthias Hagen, Benno Stein 0001 |
ACL (1) | 2 |
| 2018 | Retrieval of the Best Counterargument without Prior Topic KnowledgeabstractGiven any argument on any controversial topic, how to counter it?This question implies the challenging retrieval task of finding the best counterargument.Since prior knowledge of a topic cannot be expected in general, we hypothesize the best counterargument to invoke the same aspects as the argument while having the opposite stance.To operationalize our hypothesis, we simultaneously model the similarity and dissimilarity of pairs of arguments, based on the words and embeddings of the arguments' premises and conclusions.A salient property of our model is its independence from the topic at hand, i.e., it applies to arbitrary arguments.We evaluate different model variations on millions of argument pairs derived from the web portal idebate.org.Systematic ranking experiments suggest that our hypothesis is true for many arguments: For 7.6 candidates with opposing stance on average, we rank the best counterargument highest with 60% accuracy.Even among all 2801 test set pairs as candidates, we still find the best one about every third time. Henning Wachsmuth, Shahbaz Syed, Benno Stein 0001 |
ACL (1) | 1 |
| 2018 | Argumentation Synthesis following Rhetorical StrategiesabstractPersuasion is rarely achieved through a loose set of arguments alone. Rather, an effective delivery of arguments follows a rhetorical strategy, combining logical reasoning with appeals to ethics and emotion. We argue that such a strategy means to select, arrange, and phrase a set of argumentative discourse units. In this paper, we model rhetorical strategies for the computational synthesis of effective argumentation. In a study, we let 26 experts synthesize argumentative texts with different strategies for 10 topics. We find that the experts agree in the selection significantly more when following the same strategy. While the texts notably vary for different strategies, especially their arrangement remains stable. The results suggest that our model enables a strategical synthesis. Henning Wachsmuth, Manfred Stede, Roxanne El Baff, Khalid Al-Khatib, Maria Skeppstedt, Benno Stein 0001 |
COLING | 1 |
| 2018 | Challenge or Empower: Revisiting Argumentation Quality in a News Editorial CorpusabstractNews editorials are said to shape public opinion, which makes them a powerful tool and an important source of political argumentation.However, rarely do editorials change anyone's stance on an issue completely, nor do they tend to argue explicitly (but rather follow a subtle rhetorical strategy).So, what does argumentation quality mean for editorials then?We develop the notion that an effective editorial challenges readers with opposing stance, and at the same time empowers the arguing skills of readers that share the editorial's stance -or even challenges both sides.To study argumentation quality based on this notion, we introduce a new corpus with 1000 editorials from the New York Times, annotated for their perceived effect along with the annotators' political orientations.Analyzing the corpus, we find that annotators with different orientation disagree on the effect significantly.While only 1% of all editorials changed anyone's stance, more than 5% meet our notion.We conclude that our corpus serves as a suitable resource for studying the argumentation quality of news editorials. Roxanne El Baff, Henning Wachsmuth, Khalid Al-Khatib, Benno Stein 0001 |
CoNLL | 2 |
| 2018 | Learning to Flip the Bias of News HeadlinesabstractThis paper introduces the task of "flipping" the bias of news articles: Given an article with a political bias (left or right), generate an article with the same topic but opposite bias.To study this task, we create a corpus with bias-labeled articles from allsides.com.As a first step, we analyze the corpus and discuss intrinsic characteristics of bias.They point to the main challenges of bias flipping, which in turn lead to a specific setting in the generation process.The paper in hand narrows down the general bias flipping task to focus on bias flipping for news article headlines.A manual annotation of headlines from each side reveals that they are self-informative in general and often convey bias.We apply an autoencoder incorporating information from an article's content to learn how to automatically flip the bias.From 200 generated headlines, 73 are classified as understandable by annotators, and 83 maintain the topic while having opposite bias.Insights from our analysis shed light on how to solve the main challenges of bias flipping. Wei-Fan Chen 0001, Henning Wachsmuth, Khalid Al-Khatib, Benno Stein 0001 |
INLG | 2 |
| 2018 | Before Name-Calling: Dynamics and Triggers of Ad Hominem Fallacies in Web ArgumentationabstractArguing without committing a fallacy is one of the main requirements of an ideal debate. But even when debating rules are strictly enforced and fallacious arguments punished, arguers often lapse into attacking the opponent by an ad hominem argument. As existing research lacks solid empirical investigation of the typology of ad hominem arguments as well as their potential causes, this paper fills this gap by (1) performing several large-scale annotation studies, (2) experimenting with various neural architectures and validating our working hypotheses, such as controversy or reasonableness, and (3) providing linguistic insights into triggers of ad hominem using explainable neural network architectures. Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, Benno Stein 0001 |
NAACL-HLT | 2 |
| 2018 | The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit WarrantsabstractIvan Habernal, Henning Wachsmuth, Iryna Gurevych, Benno Stein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, Benno Stein 0001 |
NAACL-HLT | 2 |
| 2017 | "PageRank" for Argument RelevanceabstractFuture search engines are expected to deliver pro and con arguments in response to queries on controversial topics.While argument mining is now in the focus of research, the question of how to retrieve the relevant arguments remains open.This paper proposes a radical model to assess relevance objectively at web scale: the relevance of an argument's conclusion is decided by what other arguments reuse it as a premise.We build an argument graph for this model that we analyze with a recursive weighting scheme, adapting key ideas of PageRank.In experiments on a large ground-truth argument graph, the resulting relevance scores correlate with human average judgments.We outline what natural language challenges must be faced at web scale in order to stepwise bring argument relevance to web search engines. Henning Wachsmuth, Benno Stein 0001, Yamen Ajjour |
EACL (1) | 1 |
| 2017 | Computational Argumentation Quality Assessment in Natural LanguageabstractHenning Wachsmuth, Nona Naderi, Yufang Hou, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, Benno Stein. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Henning Wachsmuth, Nona Naderi, Yufang Hou 0001, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, Benno Stein 0001 |
EACL (1) | 1 |
| 2017 | Patterns of Argumentation Strategies across TopicsabstractThis paper presents an analysis of argumentation strategies in news editorials within and across topics.Given nearly 29,000 argumentative editorials from the New York Times, we develop two machine learning models, one for determining an editorial's topic, and one for identifying evidence types in the editorial.Based on the distribution and structure of the identified types, we analyze the usage patterns of argumentation strategies among 12 different topics.We detect several common patterns that provide insights into the manifestation of argumentation strategies.Also, our experiments reveal clear correlations between the topics and the detected patterns. Khalid Al-Khatib, Henning Wachsmuth, Matthias Hagen, Benno Stein 0001 |
EMNLP | 2 |
| 2017 | The Impact of Modeling Overall Argumentation with Tree KernelsabstractSeveral approaches have been proposed to model either the explicit sequential structure of an argumentative text or its implicit hierarchical structure.So far, the adequacy of these models of overall argumentation remains unclear.This paper asks what type of structure is actually important to tackle downstream tasks in computational argumentation.We analyze patterns in the overall argumentation of texts from three corpora.Then, we adapt the idea of positional tree kernels in order to capture sequential and hierarchical argumentative structure together for the first time.In systematic experiments for three text classification tasks, we find strong evidence for the impact of both types of structure.Our results suggest that either of them is necessary while their combination may be beneficial. Henning Wachsmuth, Giovanni Da San Martino, Dora Kiesel, Benno Stein 0001 |
EMNLP | 1 |
| 2017 | A Universal Model for Discourse-Level Argumentation AnalysisabstractThe argumentative structure of texts is increasingly exploited for analysis tasks, for example, for stance classification or the assessment of argumentation quality. Most existing approaches, however, model only the local structure of single arguments. This article considers the question of how to capture the global discourse-level structure of a text for argumentation-related analyses. In particular, we propose to model the global structure as a flow of “task-related rhetorical moves,” such as discourse functions or aspect-based sentiment. By comparing the flow of a text to a set of common flow patterns, we map the text into the feature space of global structures, thus capturing its discourse-level argumentation. We show how to identify different types of flow patterns, and we provide evidence that they generalize well across different domains of texts. In our evaluation for two analysis tasks, the classification of review sentiment and the scoring of essay organization, the features derived from flow patterns prove both effective and more robust than strong baselines. We conclude with a discussion of the universality of modeling flow for discourse-level argumentation analysis. Henning Wachsmuth, Benno Stein 0001 |
ACM Trans. Internet Techn. | 1 |
| 2016 | A News Editorial Corpus for Mining Argumentation StrategiesabstractMany argumentative texts, and news editorials in particular, follow a specific strategy to persuade their readers of some opinion or attitude. This includes decisions such as when to tell an anecdote or where to support an assumption with statistics, which is reflected by the composition of different types of argumentative discourse units in a text. While several argument mining corpora have recently been published, they do not allow the study of argumentation strategies due to incomplete or coarse-grained unit annotations. This paper presents a novel corpus with 300 editorials from three diverse news portals that provides the basis for mining argumentation strategies. Each unit in all editorials has been assigned one of six types by three annotators with a high Fleiss’ Kappa agreement of 0.56. We investigate various challenges of the annotation process and we conduct a first corpus analysis. Our results reveal different strategies across the news portals, exemplifying the benefit of studying editorials—a so far underresourced text genre in argument mining. Khalid Al-Khatib, Henning Wachsmuth, Johannes Kiesel, Matthias Hagen, Benno Stein 0001 |
COLING | 2 |
| 2016 | Using Argument Mining to Assess the Argumentation Quality of EssaysabstractArgument mining aims to determine the argumentative structure of texts. Although it is said to be crucial for future applications such as writing support systems, the benefit of its output has rarely been evaluated. This paper puts the analysis of the output into the focus. In particular, we investigate to what extent the mined structure can be leveraged to assess the argumentation quality of persuasive essays. We find insightful statistical patterns in the structure of essays. From these, we derive novel features that we evaluate in four argumentation-related essay scoring tasks. Our results reveal the benefit of argument mining for assessing argumentation quality. Among others, we improve the state of the art in scoring an essay’s organization and its argument strength. Henning Wachsmuth, Khalid Al-Khatib, Benno Stein 0001 |
COLING | 1 |
| 2016 | Cross-Domain Mining of Argumentative Text through Distant SupervisionabstractKhalid Al-Khatib, Henning Wachsmuth, Matthias Hagen, Jonas Köhler, Benno Stein. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Khalid Al-Khatib, Henning Wachsmuth, Matthias Hagen, Jonas Köhler 0001, Benno Stein 0001 |
HLT-NAACL | 2 |
| 2015 | Sentiment Flow - A General Model of Web Review ArgumentationabstractWeb reviews have been intensively studied in argumentation-related tasks such as sen-timent analysis. However, due to their fo-cus on content-based features, many sen-timent analysis approaches are effective only for reviews from those domains they have been specifically modeled for. This paper puts its focus on domain indepen-dence and asks whether a general model can be found for how people argue in web reviews. Our hypothesis is that people ex-press their global sentiment on a topic with similar sequences of local sentiment inde-pendent of the domain. We model such sentiment flow robustly under uncertainty through abstraction. To test our hypoth-esis, we predict global sentiment based on sentiment flow. In systematic experiments, we improve over the domain independence of strong baselines. Our findings suggest that sentiment flow qualifies as a general model of web review argumentation. 1 Henning Wachsmuth, Johannes Kiesel, Benno Stein 0001 |
EMNLP | 1 |
| 2014 | A Review Corpus for Argumentation Analysis
Henning Wachsmuth, Martin Trenkmann, Benno Stein 0001, Gregor Engels, Tsvetomira Palakarska |
CICLing (2) | 1 |
| 2014 | Modeling Review Argumentation for Robust Sentiment Analysis
Henning Wachsmuth, Martin Trenkmann, Benno Stein 0001, Gregor Engels |
COLING | 1 |
| 2014 | PBlaman: performance blame analysis based on Palladio contractsabstractSUMMARY In performance‐driven software engineering, the performance of a system is evaluated through models before the system is assembled. After assembly, the performance is then validated using performance tests. When a component‐based system fails certain performance requirements during the tests, it is important to find out whether individual components yield performance errors or whether the composition of components is faulty. This task is called performance blame analysis. Existing performance blame analysis approaches and also alternative error analysis approaches are restricted, because they either do not employ expected values, use expected values from regression testing, or use static developer‐set limits. In contrast, this paper describes the new performance blame analysis approach PBlaman that builds upon our previous work and that employs the context‐portable performance contracts of Palladio. PBlaman decides what components to blame by comparing the observed response time data series of each single component operation in a failed test case to the operation's expected response time data series derived from the contracts. System architects are then assisted by a visual presentation of the obtained analysis results. We exemplify the benefits of PBlaman in two case studies, each of which representing applications that follow a particular architectural style. Copyright © 2014 John Wiley & Sons, Ltd. Frank Brüseke, Henning Wachsmuth, Gregor Engels, Steffen Becker 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Automatic Pipeline Construction for Real-Time Annotation
Henning Wachsmuth, Mirko Rose, Gregor Engels |
CICLing (1) | 1 |
| 2013 | Information extraction as a filtering taskabstractInformation extraction is usually approached as an annotation task: Input texts run through several analysis steps of an extraction process in which different semantic concepts are annotated and matched against the slots of templates. We argue that such an approach lacks an efficient control of the input of the analysis steps. In this paper, we hence propose and evaluate a model and a formal approach that consistently put the filtering view in the focus: Before spending annotation effort, filter those portions of the input texts that may contain relevant information for filling a template and discard the others. We model all dependencies between the semantic concepts sought for with a truth maintenance system, which then efficiently infers the portions of text to be annotated in each analysis step. The filtering view enables an information extraction system (1) to annotate only relevant portions of input texts and (2) to easily trade its run-time efficiency for its recall. We provide our approach as an open-source extension of Apache UIMA and we show the potential of our approach in a number of experiments. Henning Wachsmuth, Benno Stein 0001, Gregor Engels |
CIKM | 1 |
| 2013 | Learning Efficient Information Extraction on Heterogeneous Texts
Henning Wachsmuth, Benno Stein 0001, Gregor Engels |
IJCNLP | 1 |
| 2011 | Constructing efficient information extraction pipelinesabstractInformation Extraction (IE) pipelines analyze text through several stages. The pipeline's algorithms determine both its effectiveness and its run-time efficiency. In real-world tasks, however, IE pipelines often fail acceptable run-times because they analyze too much task-irrelevant text. This raises two interesting questions: 1) How much "efficiency potential" depends on the scheduling of a pipeline's algorithms? 2) Is it possible to devise a reliable method to construct efficient IE pipelines? Both questions are addressed in this paper. In particular, we show how to optimize the run-time efficiency of IE pipelines under a given set of algorithms. We evaluate pipelines for three algorithm sets on an industrially relevant task: the extraction of market forecasts from news articles. Using a system-independent measure, we demonstrate that efficiency gains of up to one order of magnitude are possible without compromising a pipeline's original effectiveness. Henning Wachsmuth, Benno Stein 0001, Gregor Engels |
CIKM | 1 |
| 2011 | Back to the Roots of Genres: Text Classification by Language Function
Henning Wachsmuth, Kathrin Bujna |
IJCNLP | 1 |
| 2010 | Efficient Statement Identification for Automatic Market Forecasting
Henning Wachsmuth, Peter Prettenhofer, Benno Stein 0001 |
COLING | 1 |