VLDB 2026 Research / reviewers in the wild / expert
Annemarie Friedrich
dblp:126/8745
· DBLP profile ↗
22ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-8771-7634ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle SolvingabstractThe rise of large language models (LLMs) has sparked interest in coding assistants. While general-purpose programming languages are well supported, generating code for domain-specific languages remains a challenging problem for LLMs. In this paper, we focus on the LLM-based generation of code for Answer Set Programming (ASP), a particularly effective approach for finding solutions to combinatorial search problems. The effectiveness of LLMs in ASP code generation is currently hindered by the limited number of examples seen during their initial pre-training phase. In this paper, we introduce a novel ASP-solver-in-the-loop approach for solver-guided instruction-tuning of LLMs to addressing the highly complex semantic parsing task inherent in ASP code generation. Our method only requires problem specifications in natural language and their solutions. Specifically, we sample ASP statements for program continuations from LLMs for unriddling logic puzzles. Leveraging the special property of declarative ASP programming that partial encodings increasingly narrow down the solution space, we categorize them into chosen and rejected instances based on solver feedback. We then apply supervised fine-tuning to train LLMs on the curated data and further improve robustness using a solver-guided search that includes best-of-N sampling. Our experiments demonstrate consistent improvements in two distinct prompting settings on two datasets. Timo Pierre Schrader, Lukas Lange, Tobias Kaminski, Simon Razniewski, Annemarie Friedrich |
AAAI | 5 |
| 2026 | Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationabstractTable Question Answering (TQA) aims to answer natural language questions about tabular data, often accompanied by additional contexts such as text passages.The task spans diverse settings, varying in table representation, question/answer complexity, modality involved, and domain.While recent advances in large language models (LLMs) have led to substantial progress in TQA, the field still lacks a systematic organization and understanding of task formulations, core challenges, and methodological trends, particularly in light of emerging research directions such as reinforcement learning.This survey addresses this gap by providing a comprehensive and structured overview of TQA research with a focus on LLM-based methods.We provide a comprehensive categorization of existing benchmarks and task setups.We group current modeling strategies according to the challenges they target, and analyze their strengths and limitations.Furthermore, we highlight underexplored but timely topics that have not been systematically covered in prior research.By unifying disparate research threads and identifying open problems, our survey offers a consolidated foundation for the TQA community, enabling a deeper understanding of the state of the art and guiding future developments in this rapidly evolving area. Wei Zhou 0067, Bolei Ma, Annemarie Friedrich, Mohsen Mesgar |
ACL (1) | 3 |
| 2026 | Annotating Conversational Phases and Communication Techniques: A Corpus of German Teacher-Parent Counseling Conversations
Tobias Hallmen, Kathrin Gietl, Karoline Hillesheim, Annemarie Friedrich, Elisabeth André |
LREC | 4 |
| 2026 | Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval
Valentin Knappich, Anna Hätty, Simon Razniewski, Annemarie Friedrich |
SIGIR | 4 |
| 2025 | Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesabstractUnderstanding pragmatics—the use of language in context—is crucial for developing NLP systems capable of interpreting nuanced language use. Despite recent advances in language technologies, including large language models, evaluating their ability to handle pragmatic phenomena such as implicatures and references remains challenging. To advance pragmatic abilities in models, it is essential to understand current evaluation trends and identify existing limitations. In this survey, we provide a comprehensive review of resources designed for evaluating pragmatic capabilities in NLP, categorizing datasets by the pragmatic phenomena they address. We analyze task designs, data collection methods, evaluation approaches, and their relevance to real-world applications. By examining these resources in the context of modern language models, we highlight emerging trends, challenges, and gaps in existing benchmarks. Our survey aims to clarify the landscape of pragmatic evaluation and guide the development of more comprehensive and targeted benchmarks, ultimately contributing to more nuanced and context-aware NLP models. Bolei Ma, Wei Zhou 0067, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank |
ACL (1) | 7 |
| 2024 | AnnoCTR: A Dataset for Detecting and Linking Entities, Tactics, and Techniques in Cyber Threat ReportsabstractMonitoring the threat landscape to be aware of actual or potential attacks is of utmost importance to cybersecurity professionals. Information about cyber threats is typically distributed using natural language reports. Natural language processing can help with managing this large amount of unstructured information, yet to date, the topic has received little attention. With this paper, we present AnnoCTR, a new CC-BY-SA-licensed dataset of cyber threat reports. The reports have been annotated by a domain expert with named entities, temporal expressions, and cybersecurity-specific concepts including implicitly mentioned techniques and tactics. Entities and concepts are linked to Wikipedia and the MITRE ATT&CK knowledge base, the most widely-used taxonomy for classifying types of attacks. Prior datasets linking to MITRE ATT&CK either provide a single label per document or annotate sentences out-of-context; our dataset annotates entire documents in a much finer-grained way. In an experimental study, we model the annotations of our dataset using state-of-the-art neural models. In our few-shot scenario, we find that for identifying the MITRE ATT&CK concepts that are mentioned explicitly or implicitly in a text, concept descriptions from MITRE ATT&CK are an effective source for training data augmentation. Lukas Lange, Marc Müller, Ghazaleh H. Torbati, Dragan Milchevski, Patrick Grau, Subhash Chandra Pujari, Annemarie Friedrich |
LREC/COLING | 7 |
| 2024 | QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning ScenariosabstractReasoning is key to many decision making processes.It requires consolidating a set of rulelike premises that are often associated with degrees of uncertainty and observations to draw conclusions.In this work, we address both the case where premises are specified as numeric probabilistic rules and situations in which humans state their estimates using words expressing degrees of certainty.Existing probabilistic reasoning datasets simplify the task, e.g., by requiring the model to only rank textual alternatives, by including only binary random variables, or by making use of a limited set of templates that result in less varied text.In this work, we present QUITE, a question answering dataset of real-world Bayesian reasoning scenarios with categorical random variables and complex relationships.QUITE provides high-quality natural language verbalizations of premises together with evidence statements, and expects the answer to a question in the form of an estimated probability.We conduct an extensive set of experiments, finding that logic-based models outperform out-of-the-box large language models on all reasoning types (causal, evidential, and explaining-away).Our results provide evidence that neuro-symbolic models are a promising direction for improving complex reasoning.We release QUITE and code for training and experiments on Github. 1 Timo Pierre Schrader, Lukas Lange, Simon Razniewski, Annemarie Friedrich |
EMNLP | 4 |
| 2024 | FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question AnsweringabstractWei Zhou, Mohsen Mesgar, Heike Adel, Annemarie Friedrich. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wei Zhou 0067, Mohsen Mesgar, Heike Adel, Annemarie Friedrich |
NAACL-HLT | 4 |
| 2024 | SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific DomainabstractIn scientific publications, a substantial part of the information is expressed via figures containing images and diagrams. Hence, the retrieval of relevant figures given a natural language query is an important real-world task. However, due to the lack of training and evaluation data, most existing approaches are either limited to one modality or focus on non-scientific domains, making their application to scientific publications challenging.In this paper, we address this gap by introducing two novel datasets: (1) SciOL, the largest openly-licensed pre-training corpus for multimodal models in the scientific domain, covering multiple sciences including materials science, physics, and computer science, and (2) MuLMS-Img, a high-quality dataset in the materials science domain, manually annotated for various image-text tasks. Our experiments show that pre-training large-scale vision-language models on SciOL increases performance considerably across a broad variety of image-text tasks including figure type classification, optical character recognition, captioning, and figure retrieval. Using MuLMS-Img, we show that integrating text-based features extracted via a fine-tuned model for a specific domain can boost cross-modal scientific figure retrieval performance by up to 50%. Tim Tarsi, Heike Adel, Jan Hendrik Metzen, Matteo Finco, Annemarie Friedrich |
WACV | 6 |
| 2023 | A Kind Introduction to Lexical and Grammatical Aspect, with a Survey of Computational ApproachesabstractAspectual meaning refers to how the internal temporal structure of situations is presented.This includes whether a situation is described as a state or as an event, whether the situation is finished or ongoing, and whether it is viewed as a whole or with a focus on a particular phase.This survey gives an overview of computational approaches to modeling lexical and grammatical aspect along with intuitive explanations of the necessary linguistic concepts and terminology.In particular, we describe the concepts of stativity, telicity, habituality, perfective and imperfective, as well as influential inventories of eventuality and situation types.Aspect is a crucial component of semantics, especially for precise reporting of the temporal structure of situations, and future NLP approaches need to be able to handle and evaluate it systematically. Annemarie Friedrich, Nianwen Xue, Alexis Palmer |
EACL | 1 |
| 2023 | A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language ProcessingabstractMany natural language processing (NLP) tasks are naturally imbalanced, as some target categories occur much more frequently than others in the real world.In such scenarios, current NLP models tend to perform poorly on less frequent classes.Addressing class imbalance in NLP is an active research topic, yet, finding a good approach for a particular task and imbalance scenario is difficult.In this survey, the first overview on class imbalance in deep-learning based NLP, we first discuss various types of controlled and realworld class imbalance.Our survey then covers approaches that have been explicitly proposed for class-imbalanced NLP tasks or, originating in the computer vision community, have been evaluated on them.We organize the methods by whether they are based on sampling, data augmentation, choice of loss function, staged learning, or model design.Finally, we discuss open problems and how to move forward. Sophie Henning, William Beluch, Alexander Fraser 0001, Annemarie Friedrich |
EACL | 4 |
| 2022 | Three Real-World Datasets and Neural Computational Models for Classification Tasks in Patent LandscapingabstractPatent Landscaping, one of the central tasks of intellectual property management, includes selecting and grouping patents according to userdefined technical or application-oriented criteria.While recent transformer-based models have been shown to be effective for classifying patents into taxonomies such as CPC or IPC, there is yet little research on how to support real-world Patent Landscape Studies (PLSs) using natural language processing methods.With this paper, we release three labeled datasets for PLS-oriented classification tasks covering two diverse domains.We provide a qualitative analysis and report detailed corpus statistics.Most research on neural models for patents has been restricted to leveraging titles and abstracts.We compare strong neural and non-neural baselines, proposing a novel model that takes into account textual information from the patents' full texts as well as embeddings created based on the patents' CPC labels.We find that for PLS-oriented classification tasks, going beyond title and abstract is crucial, CPC labels are an effective source of information, and combining all features yields the best results. Subhash Chandra Pujari, Jannik Strötgen, Mark Giereth, Michael Gertz 0001, Annemarie Friedrich |
EMNLP | 5 |
| 2021 | Negation-Instance Based Evaluation of End-to-End Negation ResolutionabstractIn this paper, we revisit the task of negation resolution, which includes the subtasks of cue detection (e.g."not", "never") and scope resolution.In the context of previous shared tasks, a variety of evaluation metrics have been proposed.Subsequent works usually use different subsets of these, including variations and custom implementations, rendering meaningful comparisons between systems difficult.Examining the problem both from a linguistic perspective and from a downstream viewpoint, we here argue for a negation-instance based approach to evaluating negation resolution.Our proposed metrics correspond to expectations over per-instance scores and hence are intuitively interpretable.To render research comparable and to foster future work, we provide results for a set of current state-of-the-art systems for negation resolution on three English corpora, and make our implementation of the evaluation scripts publicly available. Elizaveta Sineva, Stefan Grünewald, Annemarie Friedrich, Jonas Kuhn |
CoNLL | 3 |
| 2021 | Coordinate Constructions in English Enhanced Universal Dependencies: Analysis and Computational ModelingabstractIn this paper, we address the representation of coordinate constructions in Enhanced Universal Dependencies (UD), where relevant dependency links are propagated from conjunction heads to other conjuncts.English treebanks for enhanced UD have been created from gold basic dependencies using a heuristic rule-based converter, which propagates only core arguments.With the aim of determining which set of links should be propagated from a semantic perspective, we create a large-scale dataset of manually edited syntax graphs.We identify several systematic errors in the original data, and propose to also propagate adjuncts.We observe high inter-annotator agreement for this semantic annotation task.Using our new manually verified dataset, we perform the first principled comparison of rule-based and (partially novel) machine-learning based methods for conjunction propagation for English.We show that learning propagation rules is more effective than hand-designing heuristic rules.When using automatic parses, our neural graph-parser based edge predictor outperforms the currently predominant pipelines using a basic-layer tree parser plus converters. Stefan Grünewald, Prisca Piccirilli, Annemarie Friedrich |
EACL | 3 |
| 2021 | A Multi-task Approach to Neural Multi-label Hierarchical Patent Classification Using Transformers
Subhash Chandra Pujari, Annemarie Friedrich, Jannik Strötgen |
ECIR (1) | 2 |
| 2020 | The SOFC-Exp Corpus and Neural Approaches to Information Extraction in the Materials Science DomainabstractAnnemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Marusczyk, Lukas Lange. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Marusczyk, Lukas Lange |
ACL | 1 |
| 2017 | Classification of telicity using cross-linguistic annotation projectionabstractThis paper addresses the automatic recognition of telicity, an aspectual notion.A telic event includes a natural endpoint (she walked home), while an atelic event does not (she walked around).Recognizing this difference is a prerequisite for temporal natural language understanding.In English, this classification task is difficult, as telicity is a covert linguistic category.In contrast, in Slavic languages, aspect is part of a verb's meaning and even available in machine-readable dictionaries.Our contributions are as follows.We successfully leverage additional silver standard training data in the form of projected annotations from parallel English-Czech data as well as context information, improving automatic telicity classification for English significantly compared to previous work.We also create a new data set of English texts manually annotated with telicity. Annemarie Friedrich, Damyana Gateva |
EMNLP | 1 |
| 2016 | Situation entity types: automatic classification of clause-level aspectabstractThis paper describes the first robust approach to automatically labeling clauses with their situation entity type (Smith, 2003), capturing aspectual phenomena at the clause level which are relevant for interpreting both semantics at the clause level and discourse structure.Previous work on this task used a small data set from a limited domain, and relied mainly on words as features, an approach which is impractical in larger settings.We provide a new corpus of texts from 13 genres (40,000 clauses) annotated with situation entity types.We show that our sequence labeling approach using distributional information in the form of Brown clusters, as well as syntactic-semantic features targeted to the task, is robust across genres, reaching accuracies of up to 76%. Annemarie Friedrich, Alexis Palmer, Manfred Pinkal |
ACL (1) | 1 |
| 2015 | Discourse-sensitive Automatic Identification of Generic ExpressionsabstractAnnemarie Friedrich, Manfred Pinkal. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Annemarie Friedrich, Manfred Pinkal |
ACL (1) | 1 |
| 2015 | Automatic recognition of habituals: a three-way classification of clausal aspectabstractThis paper provides the first fully automatic approach for classifying clauses with respect to their aspectual properties as habitual, episodic or static.We bring together two strands of previous work, which address only the related tasks of the episodic-habitual and stative-dynamic distinctions, respectively.Our method combines different sources of information found to be useful for these tasks.We are the first to exhaustively classify all clauses of a text, achieving up to 80% accuracy (baseline 58%) for the three-way classification task, and up to 85% accuracy for related subtasks (baselines 50% and 60%), outperforming previous work.In addition, we provide a new large corpus of Wikipedia texts labeled according to our linguistically motivated guidelines. Annemarie Friedrich, Manfred Pinkal |
EMNLP | 1 |
| 2014 | LQVSumm: A Corpus of Linguistic Quality Violations in Multi-Document Summarization
Annemarie Friedrich, Marina Valeeva, Alexis Palmer |
LREC | 1 |
| 2012 | Suffix Trees as Language Models
Casey Kennington, Martin Kay, Annemarie Friedrich |
LREC | 3 |