VLDB 2026 Research / reviewers in the wild / expert
Manfred Stede
dblp:30/5655
· DBLP profile ↗
51ranked-venue papers
15as first author
9since 2021 · last 2026
0000-0001-6819-2043ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 15 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Neural Argumentative Stance Classification in Controversial Topics with Emotion-Lexicon Features
Mohammad Yeghaneh Abkenar, Manfred Stede, Mark A. Finlayson, Davide Picca, Panagiotis Ioannidis 0005 |
LREC | 3 |
| 2026 | Assessing the Persuasive Effect of AI-Generated Image Support of Arguments
Mackwyn Quadras, Manfred Stede, Henning Wachsmuth |
LREC | 2 |
| 2024 | How Diplomats Dispute: The UN Security Council Conflict CorpusabstractWe investigate disputes in the United Nations Security Council (UNSC) by studying the linguistic means of expressing conflicts. As a result, we present the UNSC Conflict Corpus (UNSCon), a collection of 87 UNSC speeches that are annotated for conflicts. We explain and motivate our annotation scheme and report on a series of experiments for automatic conflict classification. Further, we demonstrate the difficulty when dealing with diplomatic language - which is highly complex and often implicit along various dimensions - by providing corpus examples, readability scores, and classification results. Karolina Zaczynska, Peter Bourgonje, Manfred Stede |
LREC/COLING | 3 |
| 2024 | Elaborative Simplification for German-Language TextsabstractThere are many strategies used to simplify texts.In this paper, we focus specifically on the act of inserting information or elaborative simplification.Adding information is done for various reasons, such as providing definitions for concepts, making relations between concepts more explicit, and providing background information that is a prerequisite for the main content.As all of these reasons have the main goal of ensuring coherence, we first conduct a corpus analysis of simplified German-language texts that have been annotated with Rhetorical Structure Theory (RST).We focus specifically on how additional information is incorporated into the RST annotation for a text.We then transfer these insights to automatic simplification using Large Language Models (LLMs), as elaborative simplification is a nuanced task which LLMs still seem to struggle with. Freya Hewett, Hadi Asghari, Manfred Stede |
SIGDIAL | 3 |
| 2024 | Rhetorical Strategies in the UN Security Council: Rhetorical Structure Theory and ConflictsabstractMore and more corpora are being annotated with Rhetorical Structure Theory (RST) trees, often in a multi-layer scenario, as analyzing RST annotations in combination with other layers can lead to a deeper understanding of texts.To date, prior work on RST for the analysis of diplomatic language however, is scarce.We are interested in political speeches and investigate what rhetorical strategies diplomats use to communicate critique or deal with disputes.To this end, we present a new dataset with RST annotations of 82 diplomatic speeches aligned to existing Conflict annotations (UNSC-RST).We explore ways of using rhetorical trees to analyze an annotated multi-layer corpus, looking at both the relation distribution and the tree structure of speeches.In preliminary analyses we already see patterns that are characteristic for particular topics or countries. Karolina Zaczynska, Manfred Stede |
SIGDIAL | 2 |
| 2022 | Extractive Summarisation for German-language Data: A Text-level Approach with Discourse FeaturesabstractWe examine the link between facets of Rhetorical Structure Theory (RST) and the selection of content for extractive summarisation, for German-language texts. For this purpose, we produce a set of extractive summaries for a dataset of German-language newspaper commentaries, a corpus which already has several layers of annotation. We provide an in-depth analysis of the connection between summary sentences and several RST-based features and transfer these insights to various automated summarisation models. Our results show that RST features are informative for the task of extractive summarisation, particularly nuclearity and relations at sentence-level. Freya Hewett, Manfred Stede |
COLING | 2 |
| 2022 | Towards Identifying Alternative-Lexicalization Signals of Discourse RelationsabstractThe task of shallow discourse parsing in the Penn Discourse Treebank (PDTB) framework has traditionally been restricted to identifying those relations that are signaled by a discourse connective (“explicit”) and those that have no signal at all (“implicit”). The third type, the more flexible group of “AltLex” realizations has been neglected because of its small amount of occurrences in the PDTB2 corpus. Their number has grown significantly in the recent PDTB3, and in this paper, we present the first approaches for recognizing these “alternative lexicalizations”. We compare the performance of a pattern-based approach and a sequence labeling model, add an experiment on the pre-classification of candidate sentences, and provide an initial qualitative analysis of the error cases made by both models. René Knaebel, Manfred Stede |
COLING | 2 |
| 2022 | Argument Similarity Assessment in German for Intelligent Tutoring: Crowdsourced Dataset and First ExperimentsabstractNLP technologies such as text similarity assessment, question answering and text classification are increasingly being used to develop intelligent educational applications. The long-term goal of our work is an intelligent tutoring system for German secondary schools, which will support students in a school exercise that requires them to identify arguments in an argumentative source text. The present paper presents our work on a central subtask, viz. the automatic assessment of similarity between a pair of argumentative text snippets in German. In the designated use case, students write out key arguments from a given source text; the tutoring system then evaluates them against a target reference, assessing the similarity level between student work and the reference. We collect a dataset for our similarity assessment task through crowdsourcing as authentic German student data are scarce; we label the collected text pairs with similarity scores on a 5-point scale and run first experiments on the task. We see that a model based on BERT shows promising results, while we also discuss some challenges that we observe. Manfred Stede |
LREC | 2 |
| 2022 | GerCCT: An Annotated Corpus for Mining Arguments in German Tweets on Climate ChangeabstractWhile the field of argument mining has grown notably in the last decade, research on the Twitter medium remains relatively understudied. Given the difficulty of mining arguments in tweets, recent work on creating annotated resources mainly utilized simplified annotation schemes that focus on single argument components, i.e., on claim or evidence. In this paper we strive to fill this research gap by presenting GerCCT, a new corpus of German tweets on climate change, which was annotated for a set of different argument components and properties. Additionally, we labelled sarcasm and toxic language to facilitate the development of tools for filtering out non-argumentative content. This, to the best of our knowledge, renders our corpus the first tweet resource annotated for argumentation, sarcasm and toxic language. We show that a comparatively complex annotation scheme can still yield promising inter-annotator agreement. We further present first good supervised classification results yielded by a fine-tuned BERT architecture. Robin Schaefer, Manfred Stede |
LREC | 2 |
| 2020 | Variation in Coreference Strategies across Genres and Production MediaabstractIn response to (i) inconclusive results in the literature as to the properties of coreference chains in written versus spoken language, and (ii) a general lack of work on automatic coreference resolution on both spoken language and social media, we undertake a corpus study involving the various genre sections of Ontonotes, the Switchboard corpus, and a corpus of Twitter conversations.Using a set of measures that previously have been applied individually to different data sets, we find fairly clear patterns of "behavior" for the different genres/media.Besides their role for psycholinguistic investigation (why do we employ different coreference strategies when we write or speak) and for the placement of Twitter in the spoken-written continuum, we see our results as a contribution to approaching genre-/media-specific coreference resolution. Berfin Aktas, Manfred Stede |
COLING | 2 |
| 2020 | Exploiting a lexical resource for discourse connective disambiguation in GermanabstractIn this paper we focus on connective identification and sense classification for explicit discourse relations in German, as two individual sub-tasks of the overarching Shallow Discourse Parsing task.We successively augment a purely-empirical approach based on contextualised embeddings with linguistic knowledge encoded in a connective lexicon.In this way, we improve over published results for connective identification, achieving a final F 1 -score of 87.93; and we introduce, to the best of our knowledge, first results for German sense classification, achieving an F 1 -score of 87.13.Our approach demonstrates that a connective lexicon can be a valuable resource for those languages that do not have a large PDTB-style-annotated coprus available. Peter Bourgonje, Manfred Stede |
COLING | 2 |
| 2020 | The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse ParsingabstractWe present the Potsdam Commentary Corpus 2.2, a German corpus of news editorials annotated on several different levels. New in the 2.2 version of the corpus are two additional annotation layers for coherence relations following the Penn Discourse TreeBank framework. Specifically, we add relation senses to an already existing layer of discourse connectives and their arguments, and we introduce a new layer with additional coherence relation types, resulting in a German corpus that mirrors the PDTB. The aim of this is to increase usability of the corpus for the task of shallow discourse parsing. In this paper, we provide inter-annotator agreement figures for the new annotations and compare corpus statistics based on the new annotations to the equivalent statistics extracted from the PDTB. Peter Bourgonje, Manfred Stede |
LREC | 2 |
| 2020 | DiMLex-Bangla: A Lexicon of Bangla Discourse ConnectivesabstractWe present DiMLex-Bangla, a newly developed lexicon of discourse connectives in Bangla. The lexicon, upon completion of its first version, contains 123 Bangla connective entries, which are primarily compiled from the linguistic literature and translation of English discourse connectives. The lexicon compilation is later augmented by adding more connectives from a currently developed corpus, called the Bangla RST Discourse Treebank (Das and Stede, 2018). DiMLex-Bangla provides information on syntactic categories of Bangla connectives, their discourse semantics and non-connective uses (if any). It uses the format of the German connective lexicon DiMLex (Stede and Umbach, 1998), which provides a cross-linguistically applicable XML schema. The resource is the first of its kind in Bangla, and is freely available for use in studies on discourse structure and computational applications. Debopam Das, Manfred Stede, Soumya Sankar Ghosh, Lahari Chatterjee |
LREC | 2 |
| 2020 | Semi-Supervised Tri-Training for Explicit Discourse Argument ExpansionabstractThis paper describes a novel application of semi-supervision for shallow discourse parsing. We use a neural approach for sequence tagging and focus on the extraction of explicit discourse arguments. First, additional unlabeled data is prepared for semi-supervised learning. From this data, weak annotations are generated in a first setting and later used in another setting to study performance differences. In our studies, we show an increase in the performance of our models that ranges between 2-10% F1 score. Further, we give some insights to the generated discourse annotations and compare the developed additional relations with the training relations. We release this new dataset of explicit discourse arguments to enable the training of large statistical models. René Knaebel, Manfred Stede |
LREC | 2 |
| 2020 | Shallow Discourse Parsing for Under-Resourced Languages: Combining Machine Translation and Annotation ProjectionabstractShallow Discourse Parsing (SDP), the identification of coherence relations between text spans, relies on large amounts of training data, which so far exists only for English - any other language is in this respect an under-resourced one. For those languages where machine translation from English is available with reasonable quality, MT in conjunction with annotation projection can be an option for producing an SDP resource. In our study, we translate the English Penn Discourse TreeBank into German and experiment with various methods of annotation projection to arrive at the German counterpart of the PDTB. We describe the key characteristics of the corpus as well as some typical sources of errors encountered during its creation. Then we evaluate the GermanPDTB by training components for selected sub-tasks of discourse parsing on this silver data and compare performance to the same components when trained on the gold, original PDTB corpus. Henny Sluyter-Gäthje, Peter Bourgonje, Manfred Stede |
LREC | 3 |
| 2019 | Window-Based Neural Tagging for Shallow Discourse Argument LabelingabstractThis paper describes a novel approach for the task of end-to-end argument labeling in shallow discourse parsing.Our method describes a decomposition of the overall labeling task into subtasks and a general distance-based aggregation procedure.For learning these subtasks, we train a recurrent neural network and gradually replace existing components of our baseline by our model.The model is trained and evaluated on the Penn Discourse Treebank 2 corpus.While it is not as good as knowledge-intensive approaches, it clearly outperforms other models that are also trained without additional linguistic features. René Knaebel, Manfred Stede, Sebastian Stober |
CoNLL | 2 |
| 2019 | Computational Argumentation Synthesis as a Language Modeling TaskabstractSynthesis approaches in computational argumentation so far are restricted to generating claim-like argument units or short summaries of debates.Ultimately, however, we expect computers to generate whole new arguments for a given stance towards some topic, backing up claims following argumentative and rhetorical considerations.In this paper, we approach such an argumentation synthesis as a language modeling task.In our language model, argumentative discourse units are the "words", and arguments represent the "sentences".Given a pool of units for any unseen topic-stance pair, the model selects a set of unit types according to a basic rhetorical strategy (logos vs. pathos), arranges the structure of the types based on the units' argumentative roles, and finally "phrases" an argument by instantiating the structure with semantically coherent units from the pool.Our evaluation suggests that the model can, to some extent, mimic the human synthesis of strategy-specific arguments. Roxanne El Baff, Henning Wachsmuth, Khalid Al-Khatib, Manfred Stede, Benno Stein 0001 |
INLG | 4 |
| 2018 | Argumentation Synthesis following Rhetorical StrategiesabstractPersuasion is rarely achieved through a loose set of arguments alone. Rather, an effective delivery of arguments follows a rhetorical strategy, combining logical reasoning with appeals to ethics and emotion. We argue that such a strategy means to select, arrange, and phrase a set of argumentative discourse units. In this paper, we model rhetorical strategies for the computational synthesis of effective argumentation. In a study, we let 26 experts synthesize argumentative texts with different strategies for 10 topics. We find that the experts agree in the selection significantly more when following the same strategy. While the texts notably vary for different strategies, especially their arrangement remains stable. The results suggest that our model enables a strategical synthesis. Henning Wachsmuth, Manfred Stede, Roxanne El Baff, Khalid Al-Khatib, Maria Skeppstedt, Benno Stein 0001 |
COLING | 2 |
| 2018 | Developing the Bangla RST Discourse Treebank
Debopam Das, Manfred Stede |
LREC | 2 |
| 2018 | A Lexicon of Discourse Markers for Portuguese - LDM-PT
Amália Mendes, Iria del Río Gayo, Manfred Stede, Felix Dombek |
LREC | 3 |
| 2018 | A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations
Elena Musi, Manfred Stede, Leonard Kriese, Smaranda Muresan, Andrea Rocci |
LREC | 2 |
| 2018 | Identifying Explicit Discourse Connectives in GermanabstractWe are working on an end-to-end Shallow Discourse Parsing system for German and in this paper focus on the first subtask: the identification of explicit connectives.Starting with the feature set from an English system and a Random Forest classifier, we evaluate our approach on a (relatively small) German annotated corpus, the Potsdam Commentary Corpus.We introduce new features and experiment with including additional training data obtained through annotation projection and achieve an f-score of 83.89. Peter Bourgonje, Manfred Stede |
SIGDIAL Conference | 2 |
| 2018 | Constructing a Lexicon of English Discourse ConnectivesabstractWe present a new lexicon of English discourse connectives called DiMLex-Eng, built by merging information from two annotated corpora and an additional list of relation signals from the literature.The format follows the German connective lexicon DiMLex, which provides a crosslinguistically applicable XML schema.DiMLex-Eng contains 149 English connectives, and gives information on syntactic categories, discourse semantics and non-connective uses (if any).We report on the development steps and discuss design decisions encountered in the lexicon expansion phase.The resource is freely available for use in studies of discourse structure and computational applications. Debopam Das, Tatjana Scheffler, Peter Bourgonje, Manfred Stede |
SIGDIAL Conference | 4 |
| 2017 | Classifying news versus opinions in newspapers: Linguistic features for domain independenceabstractAbstract Newspaper text can be broadly divided in the classes ‘opinion’ (editorials, commentary, letters to the editor) and ‘neutral’ (reports). We describe a classification system for performing this separation, which uses a set of linguistically motivated features. Working with various English newspaper corpora, we demonstrate that it significantly outperforms bag-of-lemma and PoS-tag models. We conclude that the linguistic features constitute the best method for achieving robustness against change of newspaper or domain. Katarina R. Krüger, Anna Lukowiak, Jonathan Sonntag, Saskia Warzecha, Manfred Stede |
Nat. Lang. Eng. | 5 |
| 2016 | Towards assessing depth of argumentationabstractFor analyzing argumentative text, we propose to study the ‘depth’ of argumentation as one important component, which we distinguish from argument quality. In a pilot study with German newspaper commentary texts, we asked students to rate the degree of argumentativeness, and then looked for correlations with features of the annotated argumentation structure and the rhetorical structure (in terms of RST). The results indicate that the human judgements correlate with our operationalization of depth and with certain structural features of RST trees. Manfred Stede |
COLING | 1 |
| 2016 | Adding Semantic Relations to a Large-Coverage Connective Lexicon of German
Tatjana Scheffler, Manfred Stede |
LREC | 2 |
| 2016 | Parallel Discourse Annotations on a Corpus of Short Texts
Manfred Stede, Stergos D. Afantenos, Andreas Peldszus, Nicholas Asher, Jérémy Perret |
LREC | 1 |
| 2016 | Information structure in the Potsdam Commentary Corpus: Topics
Manfred Stede, Sara Mamprin |
LREC | 1 |
| 2015 | Joint prediction in MST-style discourse parsing for argumentation miningabstractWe introduce a new approach to argumentation mining that we applied to a parallel German/English corpus of short texts annotated with argumentation structure.We focus on structure prediction, which we break into a number of subtasks: relation identification, central claim identification, role classification, and function classification.Our new model jointly predicts different aspects of the structure by combining the different subtask predictions in the edge weights of an evidence graph; we then apply a standard MST decoding algorithm.This model not only outperforms two reasonable baselines and two datadriven models of global argument structure for the difficult subtask of relation identification, but also improves the results for central claim identification and function classification and it compares favorably to a complex mstparser pipeline. Andreas Peldszus, Manfred Stede |
EMNLP | 2 |
| 2014 | Towards Argument Mining from DialogueabstractArgument mining has started to yield early results in automatic analysis of text to produce representations of reason-conclusion structures. This paper addresses for the first time the question of automatically extracting such structures from dialogical settings of argument. More specifically, we introduce theoretical foundations for dialogical argument mining as well as show the initial implementation in a software for dialogue processing, and the application in corpus analysis. We combine analysis of illocutionary structure with structured argumentation frameworks as our scaffolding, and apply a combination of statistical and grammatically based analytical techniques. Katarzyna Budzynska, Mathilde Janier, Juyeon Kang, Chris Reed 0001, Patrick Saint-Dizier, Manfred Stede, Olena Yaskorska-Shah |
COMMA | 6 |
| 2014 | A Model for Processing Illocutionary Structures and Argumentation in Debates
Katarzyna Budzynska, Mathilde Janier, Chris Reed 0001, Patrick Saint-Dizier, Manfred Stede, Olena Yaskorska-Shah |
LREC | 5 |
| 2014 | GraPAT: a Tool for Graph Annotations
Jonathan Sonntag, Manfred Stede |
LREC | 2 |
| 2014 | Potsdam Commentary Corpus 2.0: Annotation for Discourse Research
Manfred Stede, Arne Neumann |
LREC | 1 |
| 2013 | Discourse Processing
Manfred Stede |
HLT-NAACL | 1 |
| 2012 | SemScribe: Natural Language Generation for Medical Reports
Sebastian Varges, Heike Bieler, Manfred Stede, Lukas C. Faulstich, Kristin Irsig, Malik Atalla |
LREC | 3 |
| 2011 | Lexicon-Based Methods for Sentiment AnalysisabstractWe present a lexicon-based approach to extracting sentiment from text. The Semantic Orientation CALculator (SO-CAL) uses dictionaries of words annotated with their semantic orientation (polarity and strength), and incorporates intensification and negation. SO-CAL is applied to the polarity classification task, the process of assigning a positive or negative label to a text that captures the text's opinion towards its main subject matter. We show that SO-CAL's performance is consistent across domains and in completely unseen data. Additionally, we describe the process of dictionary creation, and our use of Mechanical Turk to check dictionaries for consistency and reliability. Maite Taboada, Julian Brooke, Milan Tofiloski, Kimberly D. Voll, Manfred Stede |
Comput. Linguistics | 5 |
| 2009 | Genre-Based Paragraph Classification for Sentiment Analysis
Maite Taboada, Julian Brooke, Manfred Stede |
SIGDIAL Conference | 3 |
| 2006 | SUMMaR: Combining Linguistics and Statistics for Text Summarization
Manfred Stede, Heike Bieler, Stefanie Dipper, Arthit Suriyawongkul |
ECAI | 1 |
| 2004 | Machine-Assisted Rhetorical Structure Annotation
Manfred Stede, Silvan Heintze |
COLING | 1 |
| 2004 | Salience-Driven Text Planning
Christian Chiarcos, Manfred Stede |
INLG | 2 |
| 2003 | Rhetorical Parsing with Underspecification and Forests
Thomas Hanneforth, Silvan Heintze, Manfred Stede |
HLT-NAACL | 3 |
| 2002 | Polibox: Generating Descriptions, Comparisons, and Recommendations from a Database
Manfred Stede |
COLING | 1 |
| 2000 | The hyperonym problem revisited: Conceptual and lexical hierarchies in language generationabstractWhen a lexical item is selected in the language production process, it needs to be explained why none of its superordinates gets selected instead, since their applicability conditions are fulfilled all the same. This question has received much attention in cognitive modelling and not as much in other branches of NLG. This paper describes the various approaches taken, discusses the reasons why they are so different, and argues that production models using symbolic representations should make a distinction between conceptual and lexical hierarchies, which can be organized along fixed levels as studied in (some branches of) lexical semantics. Manfred Stede |
INLG | 1 |
| 2000 | Discourse Particles and Discourse Functions
Manfred Stede, Birte Schmitz |
Mach. Transl. | 1 |
| 1998 | Discourse Marker Choice In Sentence Planning
Brigitte Grote, Manfred Stede |
INLG | 2 |
| 1998 | A Generative Perspective on Verb Alternations
Manfred Stede |
Comput. Linguistics | 1 |
| 1996 | A generative perspecl: ive on verbs and their readingsabstractWe sketch the architecture of a sentence generation module that maps a language-neutral "deep" representation to a language-specific sentence-semantic specification, which is given to a front-end generator.Lexicalizat, ion is tlm main instrument tbr the mapl~ing step, and we examine the role of verb semantics in the process.In particular, we propose a set of rules that derive a range of verb alternations from a single base form, which is one source of lexical paraphrasing in the system. Manfred Stede |
INLG (1) | 1 |
| 1996 | Lexical paraphrases in multilingual sentence generation
Manfred Stede |
Mach. Transl. | 1 |
| 1996 | Scott R. Turner, The Creative Process. A Computer Model of Storytelling and Creativity. Hillsdale, NJ: Lawrence Erlbaum, 1994. ISBN 0-8058-1576-7, £49.95, 298 pp
Manfred Stede |
Nat. Lang. Eng. | 1 |
| 1994 | Generating Multilingual Documents from a Knowledge Base: The TECHDOC Project
Dietmar F. Rösner, Manfred Stede |
COLING | 2 |
| 1993 | Lexical Choice Criteria in Language Generation
Manfred Stede |
EACL | 1 |