EDBT 2026 Demo / reviewers in the wild / expert
Viktor Schlegel
dblp:236/0362
· DBLP profile ↗
23ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-6391-2950ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 18 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware RefinementabstractHao Li, Yizheng Sun, Viktor Schlegel, Kailai Yang, Riza Batista-Navarro, Goran Nenadic. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Li 0074, Yizheng Sun, Viktor Schlegel, Kailai Yang, Riza Theresa Batista-Navarro, Goran Nenadic |
ACL (1) | 3 |
| 2025 | MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic DialoguesabstractAutomatic Speech Recognition (ASR) systems are pivotal in transcribing speech into text, yet the errors they introduce can significantly degrade the performance of downstream tasks like summarization. This issue is particularly pronounced in clinical dialogue summarization, a low-resource domain where supervised data for fine-tuning is scarce, necessitating the use of ASR models as black-box solutions. Employing conventional data augmentation for enhancing the noise robustness of summarization models is not feasible either due to the unavailability of sufficient medical dialogue audio recordings and corresponding ASR transcripts. To address this challenge, we propose MEDSAGE, an approach for generating synthetic samples for data augmentation using Large Language Models (LLMs). Specifically, we leverage the in-context learning capabilities of LLMs and instruct them to generate ASR-like errors based on a few available medical dialogue examples with audio recordings. Experimental results show that LLMs can effectively model ASR noise, and incorporating this noisy data into the training process significantly improves the robustness and accuracy of medical dialogue summarization systems. This approach addresses the challenges of noisy ASR outputs in critical applications, offering a robust solution to enhance the reliability of clinical dialogue summarization. Kuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu, Vijay Prakash Dwivedi, Thanh-Tung Nguyen, Xiaoxue Gao, Nancy F. Chen, Stefan Winkler 0001 |
AAAI | 3 |
| 2025 | uMedSum: A Unified Framework for Clinical Abstractive SummarizationabstractAishik Nagar, Yutong Liu, Andy T. Liu, Viktor Schlegel, Vijay Prakash Dwivedi, Arun-Kumar Kaliya-Perumal, Guna Pratheep Kalanchiam, Yili Tang, Robby T. Tan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Aishik Nagar, Andy T. Liu, Viktor Schlegel, Vijay Prakash Dwivedi, Arun-Kumar Kaliya-Perumal, Guna Pratheep Kalanchiam, Yili Tang, Robby T. Tan |
ACL (1) | 4 |
| 2025 | Evaluating Differentially Private Generation of Domain-Specific TextabstractGenerative AI offers transformative potential for high-stakes domains such as healthcare and finance, yet privacy and regulatory barriers hinder the use of real-world data. To address this, differentially private synthetic data generation has emerged as a promising alternative. In this work, we introduce a unified benchmark to systematically evaluate the utility and fidelity of text datasets generated under formal Differential Privacy (DP) guarantees. Our benchmark addresses key challenges in domain-specific benchmarking, including choice of representative data and realistic privacy budgets, accounting for pre-training and a variety of evaluation metrics. We assess state-of-the-art privacy-preserving generation methods across five domain-specific datasets, revealing significant utility and fidelity degradation compared to real data, especially under strict privacy constraints. These findings underscore the limitations of current approaches, outline the need for advanced privacy-preserving data sharing methods and set a precedent regarding their evaluation in realistic scenarios. Viktor Schlegel, Srinivasan Nandakumar, Iqra Zahid, Yuping Wu 0001, Warren Del-Pinto, Goran Nenadic, Siew-Kei Lam, Jie Zhang 0073, Anil A. Bharath |
CIKM | 2 |
| 2025 | DocDiscNER: Enhanced Document-Level Discontinuous NER via Coordination Ellipses Resolution and Self-Consistency DecodingabstractIdentifying entities in medical text often involves dealing with discontinuous word sequences or entities sharing a common head, which pose significant challenges for traditional Named Entity Recognition (NER) systems. Current state-of-the-art discontinuous NER models typically process each sentence in isolation, overlooking valuable intra-sentence context. However, recent studies have shown that large language models (LLMs) perform exceptionally well when provided such context. In this work, we introduce DocDiscNER, a novel approach to discontinuous NER, which features (i) a context-aware document chunking method that provides contextually related segments as input for LLM-based NER models; (ii) a dataset and approach for coordination ellipses resolution, to address candidate spans sharing common heads and (iii) a self-consistency decoding strategy that uses self-ensembling and a majority voting mechanism to select the most consistent predictions as entity spans. We demonstrate the effectiveness and generalisability of our method on three discontinuous NER benchmarks, achieving new state-of-the-art (SOTA) performance on two of them–CADEC and ShARe-14 (2.48 and 2.2 absolute F1 points gain, respectively); while achieving competitive results on ShARe-13. In addition, our method surpasses previous SOTA performance specifically in recognising discontinuous mentions. A deeper analysis unveils that incorporating semantically relevant context significantly enhances overall NER performance compared to using individual sentences as input. Areej Alhassan, Viktor Schlegel, Rina Carines Cabral, Riza Theresa Batista-Navarro, Soyeon Caren Han, Josiah Poon, Goran Nenadic |
ECAI | 2 |
| 2025 | BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion ModelingabstractTime-series Generation (TSG) is a prominent research area with broad applications in simulations, data augmentation, and counterfactual analysis. While existing methods have shown promise in unconditional single-domain TSG, real-world applications demand for cross-domain approaches capable of controlled generation tailored to domain-specific constraints and instance-level requirements. In this paper, we argue that text can provide semantic insights, domain information and instance-specific temporal patterns, to guide and improve TSG. We introduce “Text-Controlled TSG”, a task focused on generating realistic time series by incorporating textual descriptions. To address data scarcity in this setting, we propose a novel LLM-based Multi-Agent framework that synthesizes diverse, realistic text-to-TS datasets. Furthermore, we introduce Bridge, a hybrid text-controlled TSG framework that integrates semantic prototypes with text description for supporting domain-level guidance. This approach achieves state-of-the-art generation fidelity on 11 of 12 datasets, and improves controllability by up to 12% on MSE and 6% MAE compared to no text input generation, highlighting its potential for generating tailored time-series data. Hao Li 0074, Yu-Hao Huang 0002, Chang Xu 0008, Viktor Schlegel, Renhe Jiang, Riza Theresa Batista-Navarro, Goran Nenadic, Jiang Bian 0002 |
ICML | 4 |
| 2025 | MIRA: Medical Time Series Foundation Model for Real-World Health DataabstractA unified foundation model for medical time series—pretrained on open access and ethically reviewed medical corpora—offers the potential to reduce annotation burdens, minimize model customization, and enable robust transfer across clinical institutions, modalities, and tasks, particularly in data-scarce or privacy-constrained environments. However, existing time series foundation models struggle to handle medical time series data due to its inherent challenges, including irregular intervals, heterogeneous sampling rates, and frequent missingness. To address these challenges, we introduce MIRA, a unified foundation model specifically designed for medical time series forecasting. MIRA incorporates a Continuous-Time Rotary Positional Encoding that enables fine-grained modeling of variable time intervals, a frequency-specific mixture-of-experts layer that routes computation across latent frequency regimes to further promote temporal specialization, and a Continuous Dynamics Extrapolation Block based on Neural ODE that models the continuous trajectory of latent states, enabling accurate forecasting at arbitrary target timestamps. Pretrained on a large-scale and diverse medical corpus comprising over 454 billion time points collect from publicly available datasets, MIRA achieving reductions in forecasting errors by an average of 8% and 6% in out-of-distribution and in-distribution scenarios, respectively. We also introduce a comprehensive benchmark spanning multiple downstream clinical tasks, establishing a foundation for future research in medical time series modeling. Hao Li 0074, Chang Xu 0008, Zhiyuan Feng, Viktor Schlegel, Yu-Hao Huang 0002, Yizheng Sun, Kailai Yang, Yiyao Yu, Jiang Bian 0002 |
NeurIPS | 5 |
| 2025 | Learning to generate and evaluate fact-checking explanations with transformersabstractIn an era increasingly dominated by digital platforms, the spread of misinformation poses a significant challenge, highlighting the need for solutions capable of assessing information veracity. Our research contributes to the field of Explainable Artificial Antelligence (XAI) by developing transformer-based fact-checking models that contextualise and justify their decisions by generating human-accessible explanations. Importantly, we also develop models for automatic evaluation of explanations for fact-checking verdicts across different dimensions such as (self)-contradiction , hallucination , convincingness and overall quality . By introducing human-centred evaluation methods and developing specialised datasets, we emphasise the need for aligning Artificial Intelligence (AI)-generated explanations with human judgements. This approach not only advances theoretical knowledge in XAI but also holds practical implications by enhancing the transparency, reliability and users’ trust in AI-driven fact-checking systems. Furthermore, the development of our metric learning models is a first step towards potentially increasing efficiency and reducing reliance on extensive manual assessment. Based on experimental results, our best performing generative model achieved a Recall-Oriented Understudy for Gisting Evaluation-1 ( ROUGE-1 ) score of 47.77 demonstrating superior performance in generating fact-checking explanations, particularly when provided with high-quality evidence. Additionally, the best performing metric learning model showed a moderately strong correlation with human judgements on objective dimensions such as (self)-contradiction and hallucination , achieving a Matthews Correlation Coefficient (MCC) of around 0.7. • A dataset for fact-checking which includes explanations written by journalists. • Transformer models for generating human-accessible fact-checking explanations. • Multi-dimensional annotations reflecting explanation quality judgements. • A metric learning model scoring explanations aligned with these judgements. Darius Feher, Abdullah Salem Khered, Riza Theresa Batista-Navarro, Viktor Schlegel |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Discontinuous named entities in clinical text: A systematic literature reviewabstractOBJECTIVE: Extracting named entities from clinical free-text presents unique challenges, particularly when dealing with discontinuous entities-mentions that are separated by unrelated words. Traditional NER methods often struggle to accurately identify these entities, prompting the development of specialised computational solutions. This paper systematically reviews and presents the methodologies developed for Discontinuous Named Entity Recognition in clinical texts, highlighting their effectiveness and the challenges they face. METHOD: We conducted a systematic literature review focused on discontinuous named entities, using structured searches across four Computer Science-related and one medical-related electronic database. A combination of search terms, grouped into three synonym categories-problem, entity/approach, and task-yielded 2,442 articles. Guided by our research objectives, we identified five key dimensions to systematically annotate and normalise the data for comprehensive analysis. RESULT: The review included 44 studies which were coded across several key dimensions: the chronological development of approaches, the corpora used, the downstream tasks affected by discontinuous named entities, the methodological approaches proposed to address the issue, and the reported performance outcomes. The discussion section examines the challenges encountered in this area and suggests potential directions for future research. CONCLUSION: Significant progress has been made in discontinuous named entity recognition; however, there remains a need for more adaptable, generalisable solutions that are independent of custom annotation schemes. Exploring various configurations of generative language models presents a promising avenue for advancing this area. Additionally, future research should investigate the impact of precise versus imprecise recognition of discontinuous entities on clinical downstream tasks to better understand its practical implications in healthcare applications. Areej Alhassan, Viktor Schlegel, Monira Aloud, Riza Theresa Batista-Navarro, Goran Nenadic |
J. Biomed. Informatics | 2 |
| 2024 | A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the CHATGPT Era and BeyondabstractAbhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler, See-Kiong Ng, Soujanya Poria. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler 0001, See-Kiong Ng, Soujanya Poria |
EACL (1) | 3 |
| 2024 | Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?abstractState-of-the-art Large Language Models (LLMs) are accredited with an increasing number of different capabilities, ranging from reading comprehension over advanced mathematical and reasoning skills to possessing scientific knowledge.In this paper we focus on multi-hop reasoning-the ability to identify and integrate information from multiple textual sources.Given the concerns with the presence of simplifying cues in existing multi-hop reasoning benchmarks, which allow models to circumvent the reasoning requirement, we set out to investigate whether LLMs are prone to exploiting such simplifying cues.We find evidence that they indeed circumvent the requirement to perform multi-hop reasoning, but they do so in more subtle ways than what was reported about their fine-tuned pre-trained language model (PLM) predecessors.We propose a challenging multi-hop reasoning benchmark by generating seemingly plausible multi-hop reasoning chains that ultimately lead to incorrect answers.We evaluate multiple open and proprietary state-of-the-art LLMs and show that their multi-hop reasoning performance is affected, as indicated by up to 45% relative decrease in F1 score when presented with such seemingly plausible alternatives.We also find that-while LLMs tend to ignore misleading lexical cues-misleading reasoning paths indeed present a significant challenge.The code and data are made available at https: //github.com/zawedcvg/Are-Large- Language-Models-Attentive-Readers. Neeladri Bhuiya, Viktor Schlegel, Stefan Winkler 0001 |
EMNLP | 2 |
| 2023 | Do You Hear The People Sing? Key Point Analysis via Iterative Clustering and Abstractive SummarisationabstractArgument summarisation is a promising but currently under-explored field.Recent work has aimed to provide textual summaries in the form of concise and salient short texts, i.e., key points (KPs), in a task known as Key Point Analysis (KPA).One of the main challenges in KPA is finding high-quality key point candidates from dozens of arguments even in a small corpus.Furthermore, evaluating key points is crucial in ensuring that the automatically generated summaries are useful.Although automatic methods for evaluating summarisation have considerably advanced over the years, they mainly focus on sentence-level comparison, making it difficult to measure the quality of a summary (a set of KPs) as a whole.Aggravating this problem is the fact that human evaluation is costly and unreproducible.To address the above issues, we propose a two-step abstractive summarisation framework based on neural topic modelling with an iterative clustering procedure, to generate key points which are aligned with how humans identify key points.Our experiments show that our framework advances the state of the art in KPA, with performance improvement of up to 14 (absolute) percentage points, in terms of both ROUGE and our own proposed evaluation metrics 1 .Furthermore, we evaluate the generated summaries using a novel set-based evaluation toolkit.Our quantitative analysis demonstrates the effectiveness of our proposed evaluation metrics in assessing the quality of generated KPs.Human evaluation further demonstrates the advantages of our approach and validates that our proposed evaluation metric is more consistent with human judgment than ROUGE scores.Children express themselves through the clothes they wear and should be able to do this at school School uniform is harming the student's self expression School uniforms are expensive [...]Children should be able to dress as they wish, within reason, at school rather than being restricted from expressing themselves through their clothes.Children should be allowed to express themselves School uniform is unaffordable for many single parents and should be abandoned.School uniforms are an expense that many families can't afford.there are plenty of ways to get very cheap clothing, but discounted uniforms are more difficult to obtain.School uniforms are expensive and puts an undue burden on the parents of the students.School uniforms are expensive for the school and take money from other important programs.School unforms stifle freedom of expression.they can be costly and make circumstances difficult for those on a budget [... Hao Li 0074, Viktor Schlegel, Riza Theresa Batista-Navarro, Goran Nenadic |
ACL (1) | 2 |
| 2023 | Few-shot entity linking of food namesabstractEntity linking (EL), the task of automatically matching mentions in text to concepts in a target knowledge base, remains under-explored when it comes to the food domain, despite its many potential applications, e.g., finding the nutritional value of ingredients in databases. In this paper, we describe the creation of new resources supporting the development of EL methods applied to the food domain: the E.Care Knowledge Base (E.Care KB) which contains 664 food concepts and the E.Care dataset, a corpus of 468 cooking recipes where ingredient names have been manually linked to corresponding concepts in the E.Care KB. We developed and evaluated different methods for EL, namely, deep learning-based approaches underpinned by Siamese networks trained under a few-shot learning setting, traditional machine learning-based approaches underpinned by support vector machines (SVMs) and unsupervised approaches based on string matching algorithms. Combining the strengths of each of these approaches, we built a hybrid model for food EL that balances the trade-offs between performance and inference speed. Specifically, our hybrid model obtains 89.40% accuracy and links mentions at an average speed of 0.24 seconds per mention, whereas our best deep learning-based model, SVM model and unsupervised model obtain accuracies of 86.99%, 87.19% and 87.43% at inference speeds of 0.007, 0.66 and 0.02 seconds per mention, respectively. Darius Feher, Faridz Ibrahim, Zhuyan Cheng, Viktor Schlegel, Tom Maidment, James Bagshaw, Riza Theresa Batista-Navarro |
Inf. Process. Manag. | 4 |
| 2023 | Global information-aware argument mining based on a top-down multi-turn QA modelabstractArgument mining (AM) aims to automatically generate a graph that represents the argument structure of a document. Most previous AM models only pay attention to a single argument component (AC) to classify the type of the AC or a pair of ACs to identify and classify the argumentative relation (AR) between the two ACs. These models ignore the impact of global argument structure of the documents, which is important, especially in some highly structured genres such as scientific papers, where the process of argumentation is relatively fixed. Inspired by this, we propose a novel two-stage model which leverages global structure information to support AM. The first stage uses a multi-turn question-answering model to incrementally generate an initial argumentative graph that identifies relations among ACs. At each turn, all ACs related to the query AC are generated simultaneously, such that the sibling global information between the answer ACs is considered. In addition, the partially constructed graph is used as global structure information to support the extension of the graph with additional ACs. After the whole initial graph structure has been determined, the second stage assigns semantic types to both the ACs and ARs among them, leveraging information from this initial graph as global structure information. We test the proposed methods on two scientific datasets (one is the AbstRCT dataset including 659 abstracts about cancer research and the other is the SciARG dataset that consists of 225 computer linguistic abstracts and 285 biomedical abstracts) and a student essay dataset PE with 402 essays. Our experiments show that our model improves the state-of-the-art performance on two scientific datasets for different AM subtasks, with average improvements of 1%, 2.41%, 1.1% for the ACC, ARI and ARC task respectively on the AbstRCT dataset, and 2.36%, 1.84%, 8.87% for the ACC, ARI and ARC task on the SciARG dataset. Our model also achieves comparative results on the PE datasets: 87.7% of F1 scores for the ACC task, 81.4% for the ARI task and 78.8% for the ARC task. Boyang Liu 0002, Viktor Schlegel, Paul Thompson 0002, Riza Theresa Batista-Navarro, Sophia Ananiadou |
Inf. Process. Manag. | 2 |
| 2023 | A survey of methods for revealing and overcoming weaknesses of data-driven Natural Language UnderstandingabstractAbstract Recent years have seen a growing number of publications that analyse Natural Language Understanding (NLU) datasets for superficial cues, whether they undermine the complexity of the tasks underlying those datasets and how they impact those models that are optimised and evaluated on this data. This structured survey provides an overview of the evolving research area by categorising reported weaknesses in models and datasets and the methods proposed to reveal and alleviate those weaknesses for the English language. We summarise and discuss the findings and conclude with a set of recommendations for possible future research directions. We hope that it will be a useful resource for researchers who propose new datasets to assess the suitability and quality of their data to evaluate various phenomena of interest, as well as those who propose novel NLU approaches, to further understand the implications of their improvements with respect to their model’s acquired capabilities. Viktor Schlegel, Goran Nenadic, Riza Theresa Batista-Navarro |
Nat. Lang. Eng. | 1 |
| 2022 | Can Transformers Reason in Fragments of Natural Language?abstractState-of-the-artdeep-learning-based approaches to Natural Language Processing (NLP) are credited with various capabilities that involve reasoning with natural language texts.In this paper we carry out a large-scale empirical study investigating the detection of formally valid inferences in controlled fragments of natural language for which the satisfiability problem becomes increasingly complex.We find that, while transformerbased language models perform surprisingly well in these scenarios, a deeper analysis reveals that they appear to overfit to superficial patterns in the data rather than acquiring the logical principles governing the reasoning in these fragments. Viktor Schlegel, Kamen V. Pavlov, Ian Pratt-Hartmann |
EMNLP | 1 |
| 2022 | 'Am I the Bad One'? Predicting the Moral Judgement of the Crowd Using Pre-trained Language ModelsabstractNatural language processing (NLP) has been shown to perform well in various tasks, such as answering questions, ascertaining natural language inference and anomaly detection. However, there are few NLP-related studies that touch upon the moral context conveyed in text. This paper studies whether state-of-the-art, pre-trained language models are capable of passing moral judgments on posts retrieved from a popular Reddit user board. Reddit is a social discussion website and forum where posts are promoted by users through a voting system. In this work, we construct a dataset that can be used for moral judgement tasks by collecting data from the AITA? (Am I the A*******?) subreddit. To model our task, we harnessed the power of pre-trained language models, including BERT, RoBERTa, RoBERTa-large, ALBERT and Longformer. We then fine-tuned these models and evaluated their ability to predict the correct verdict as judged by users for each post in the datasets. RoBERTa showed relative improvements across the three datasets, exhibiting a rate of 87% accuracy and a Matthews correlation coefficient (MCC) of 0.76, while the use of the Longformer model slightly improved the performance when used with longer sequences, achieving 87% accuracy and 0.77 MCC. Areej Alhassan, Jinkai Zhang, Viktor Schlegel |
LREC | 3 |
| 2022 | RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced LabourabstractForced labour is the most common type of modern slavery, and it is increasingly gaining the attention of the research and social community. Recent studies suggest that artificial intelligence (AI) holds immense potential for augmenting anti-slavery action. However, AI tools need to be developed transparently in cooperation with different stakeholders. Such tools are contingent on the availability and access to domain-specific data, which are scarce due to the near-invisible nature of forced labour. To the best of our knowledge, this paper presents the first openly accessible English corpus annotated for multi-class and multi-label forced labour detection. The corpus consists of 989 news articles retrieved from specialised data sources and annotated according to risk indicators defined by the International Labour Organization (ILO). Each news article was annotated for two aspects: (1) indicators of forced labour as classification labels and (2) snippets of the text that justify labelling decisions. We hope that our data set can help promote research on explainability for multi-class and multi-label text classification. In this work, we explain our process for collecting the data underpinning the proposed corpus, describe our annotation guidelines and present some statistical analysis of its content. Finally, we summarise the results of baseline experiments based on different variants of the Bidirectional Encoder Representation from Transformer (BERT) model. Erick Mendez Guzman, Viktor Schlegel, Riza Theresa Batista-Navarro |
LREC | 2 |
| 2022 | Incorporating Zoning Information into Argument Mining from Biomedical LiteratureabstractThe goal of text zoning is to segment a text into zones (i.e., Background, Conclusion) that serve distinct functions. Argumentative zoning, a specific text zoning scheme for the scientific domain, is considered as the antecedent for argument mining by many researchers. Surprisingly, however, little work is concerned with exploiting zoning information to improve the performance of argument mining models, despite the relatedness of the two tasks. In this paper, we propose two transformer-based models to incorporate zoning information into argumentative component identification and classification tasks. One model is for the sentence-level argument mining task and the other is for the token-level task. In particular, we add the zoning labels predicted by an off-the-shelf model to the beginning of each sentence, inspired by the convention commonly used biomedical abstracts. Moreover, we employ multi-head attention to transfer the sentence-level zoning information to each token in a sentence. Based on experiment results, we find a significant improvement in F1-scores for both sentence- and token-level tasks. It is worth mentioning that these zoning labels can be obtained with high accuracy by utilising readily available automated methods. Thus, existing argument mining models can be improved by incorporating zoning information without any additional annotation cost. Boyang Liu 0002, Viktor Schlegel, Riza Theresa Batista-Navarro, Sophia Ananiadou |
LREC | 2 |
| 2021 | Semantics Altering Modifications for Evaluating Comprehension in Machine ReadingabstractAdvances in NLP have yielded impressive results for the task of machine reading comprehension (MRC), with approaches having been reported to achieve performance comparable to that of humans. In this paper, we investigate whether state-of-the-art MRC models are able to correctly process Semantics Altering Modifications (SAM): linguistically-motivated phenomena that alter the semantics of a sentence while preserving most of its lexical surface form. We present a method to automatically generate and align challenge sets featuring original and altered examples. We further propose a novel evaluation methodology to correctly assess the capability of MRC systems to process these examples independent of the data they were optimised on, by discounting for effects introduced by domain shift. In a large-scale empirical study, we apply the methodology in order to evaluate extractive MRC models with regard to their capability to correctly process SAM-enriched data. We comprehensively cover 12 different state-of-the-art neural architecture configurations and four training datasets and find that -- despite their well-known remarkable performance -- optimised models consistently struggle to correctly process semantically altered data. Viktor Schlegel, Goran Nenadic, Riza Theresa Batista-Navarro |
AAAI | 1 |
| 2021 | Is the Understanding of Explicit Discourse Relations Required in Machine Reading Comprehension?abstractAn in-depth analysis of the level of language understanding required by existing Machine Reading Comprehension (MRC) benchmarks can provide insight into the reading capabilities of machines.In this paper, we propose an ablation-based methodology to assess the extent to which MRC datasets evaluate the understanding of explicit discourse relations.We define seven MRC skills which require the understanding of different discourse relations.We then introduce ablation methods that verify whether these skills are required to succeed on a dataset.By observing the drop in performance of neural MRC models evaluated on the original and the modified dataset, we can measure to what degree the dataset requires these skills, in order to be understood correctly.Experiments on three large-scale datasets with the BERT-base and ALBERT-xxlarge model show that the relative changes for all skills are small (less than 6%).These results imply that most of the answered questions in the examined datasets do not require understanding the discourse structure of the text.To specifically probe for natural language understanding, there is a need to design more challenging benchmarks that can correctly evaluate the intended skills 1 . Viktor Schlegel, Riza Theresa Batista-Navarro |
EACL | 2 |
| 2020 | A Framework for Evaluation of Machine Reading Comprehension Gold StandardsabstractMachine Reading Comprehension (MRC) is the task of answering a question over a paragraph of text. While neural MRC systems gain popularity and achieve noticeable performance, issues are being raised with the methodology used to establish their performance, particularly concerning the data design of gold standards that are used to evaluate them. There is but a limited understanding of the challenges present in this data, which makes it hard to draw comparisons and formulate reliable hypotheses. As a first step towards alleviating the problem, this paper proposes a unifying framework to systematically investigate the present linguistic features, required reasoning and background knowledge and factual correctness on one hand, and the presence of lexical cues as a lower bound for the requirement of understanding on the other hand. We propose a qualitative annotation schema for the first and a set of approximative metrics for the latter. In a first application of the framework, we analyse modern MRC gold standards and present our findings: the absence of features that contribute towards lexical ambiguity, the varying factual correctness of the expected answers and the presence of lexical cues, all of which potentially lower the reading comprehension complexity and quality of the evaluation data. Viktor Schlegel, Marco Valentino, André Freitas, Goran Nenadic, Riza Theresa Batista-Navarro |
LREC | 1 |
| 2019 | Vajra: step-by-step programming with natural languageabstractBuilding natural language programming systems that are geared towards end-users requires the abstraction of formalisms inherently introduced by programming languages, capturing the intent of natural language inputs and mapping it to existing programming language constructs. Viktor Schlegel, Benedikt Lang, Siegfried Handschuh, André Freitas |
IUI | 1 |