Riza Theresa Batista-Navarro

dblp:92/11424 · also Riza Batista-Navarro · DBLP profile ↗
← Back
40ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-6693-7531ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 1 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement
abstract
Hao Li, Yizheng Sun, Viktor Schlegel, Kailai Yang, Riza Batista-Navarro, Goran Nenadic. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Li 0074, Yizheng Sun, Viktor Schlegel, Kailai Yang, Riza Theresa Batista-Navarro, Goran Nenadic
ACL (1)5
2025 Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability Problems
abstract
Tharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann, Riza Batista-Navarro. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Tharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann, Riza Theresa Batista-Navarro
ACL (1)5
2025 DocDiscNER: Enhanced Document-Level Discontinuous NER via Coordination Ellipses Resolution and Self-Consistency Decoding
abstract
Identifying entities in medical text often involves dealing with discontinuous word sequences or entities sharing a common head, which pose significant challenges for traditional Named Entity Recognition (NER) systems. Current state-of-the-art discontinuous NER models typically process each sentence in isolation, overlooking valuable intra-sentence context. However, recent studies have shown that large language models (LLMs) perform exceptionally well when provided such context. In this work, we introduce DocDiscNER, a novel approach to discontinuous NER, which features (i) a context-aware document chunking method that provides contextually related segments as input for LLM-based NER models; (ii) a dataset and approach for coordination ellipses resolution, to address candidate spans sharing common heads and (iii) a self-consistency decoding strategy that uses self-ensembling and a majority voting mechanism to select the most consistent predictions as entity spans. We demonstrate the effectiveness and generalisability of our method on three discontinuous NER benchmarks, achieving new state-of-the-art (SOTA) performance on two of them–CADEC and ShARe-14 (2.48 and 2.2 absolute F1 points gain, respectively); while achieving competitive results on ShARe-13. In addition, our method surpasses previous SOTA performance specifically in recognising discontinuous mentions. A deeper analysis unveils that incorporating semantically relevant context significantly enhances overall NER performance compared to using individual sentences as input.
Areej Alhassan, Viktor Schlegel, Rina Carines Cabral, Riza Theresa Batista-Navarro, Soyeon Caren Han, Josiah Poon, Goran Nenadic
ECAI4
2025 Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study
abstract
Vision-Language Models (VLMs) are powerful yet computationally intensive for widespread practical deployments.To address such challenge without costly re-training, post-training acceleration techniques like quantization and token reduction are extensively explored.However, current acceleration evaluations primarily target minimal overall performance degradation, overlooking a crucial question: does the accelerated model still give the same answers to the same questions as it did before acceleration?This is vital for stability-centered industrial applications where consistently correct answers for specific, known situations are paramount, such as in AI-based disease diagnosis.We systematically investigate this for accelerated VLMs, testing four leading models (LLaVA-1.5,LLaVA-Next, Qwen2-VL, Qwen2.5-VL) with eight acceleration methods on ten multimodal benchmarks.Our findings are stark: despite minimal aggregate performance drops, accelerated models changed original answers up to 20% of the time.Critically, up to 6.5% of these changes converted correct answers to incorrect.Input perturbations magnified these inconsistencies, and the trend is confirmed by case studies with the medical VLM LLaVA-Med.This research reveals a significant oversight in VLM acceleration, stressing an urgent need for instance-level stability checks to ensure trustworthy real-world deployment.
Yizheng Sun, Hao Li 0074, Chang Xu 0008, Hongpeng Zhou, Chenghua Lin 0002, Riza Theresa Batista-Navarro
EMNLP6
2025 BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion Modeling
abstract
Time-series Generation (TSG) is a prominent research area with broad applications in simulations, data augmentation, and counterfactual analysis. While existing methods have shown promise in unconditional single-domain TSG, real-world applications demand for cross-domain approaches capable of controlled generation tailored to domain-specific constraints and instance-level requirements. In this paper, we argue that text can provide semantic insights, domain information and instance-specific temporal patterns, to guide and improve TSG. We introduce “Text-Controlled TSG”, a task focused on generating realistic time series by incorporating textual descriptions. To address data scarcity in this setting, we propose a novel LLM-based Multi-Agent framework that synthesizes diverse, realistic text-to-TS datasets. Furthermore, we introduce Bridge, a hybrid text-controlled TSG framework that integrates semantic prototypes with text description for supporting domain-level guidance. This approach achieves state-of-the-art generation fidelity on 11 of 12 datasets, and improves controllability by up to 12% on MSE and 6% MAE compared to no text input generation, highlighting its potential for generating tailored time-series data.
Hao Li 0074, Yu-Hao Huang 0002, Chang Xu 0008, Viktor Schlegel, Renhe Jiang, Riza Theresa Batista-Navarro, Goran Nenadic, Jiang Bian 0002
ICML6
2025 TriG-NER: Triplet-Grid Framework for Discontinuous Named Entity Recognition
abstract
Discontinuous Named Entity Recognition (DNER) presents a challenging problem where entities may be scattered across multiple non-adjacent tokens, making traditional sequence labelling approaches inadequate. Existing methods predominantly rely on custom tagging schemes to handle these discontinuous entities, resulting in models tightly coupled to specific tagging strategies and lacking generalisability across diverse datasets. To address these challenges, we propose TriG-NER, a novel Triplet-Grid Framework that introduces a generalisable approach to learning robust token-level representations for discontinuous entity extraction. Our framework applies triplet loss at the token level, where similarity is defined by word pairs existing within the same entity, effectively pulling together similar and pushing apart dissimilar ones. This approach enhances entity boundary detection and reduces the dependency on specific tagging schemes by focusing on word-pair relationships within a flexible grid structure. We evaluate TriG-NER on three benchmark DNER datasets and demonstrate significant improvements over existing grid-based architectures. These results underscore our framework's effectiveness in capturing complex entity structures and its adaptability to various tagging schemes, setting a new benchmark for discontinuous entity extraction.
Rina Carines Cabral, Soyeon Caren Han, Areej Alhassan, Riza Theresa Batista-Navarro, Goran Nenadic, Josiah Poon
WWW4
2025 Learning to generate and evaluate fact-checking explanations with transformers
abstract
In an era increasingly dominated by digital platforms, the spread of misinformation poses a significant challenge, highlighting the need for solutions capable of assessing information veracity. Our research contributes to the field of Explainable Artificial Antelligence (XAI) by developing transformer-based fact-checking models that contextualise and justify their decisions by generating human-accessible explanations. Importantly, we also develop models for automatic evaluation of explanations for fact-checking verdicts across different dimensions such as (self)-contradiction , hallucination , convincingness and overall quality . By introducing human-centred evaluation methods and developing specialised datasets, we emphasise the need for aligning Artificial Intelligence (AI)-generated explanations with human judgements. This approach not only advances theoretical knowledge in XAI but also holds practical implications by enhancing the transparency, reliability and users’ trust in AI-driven fact-checking systems. Furthermore, the development of our metric learning models is a first step towards potentially increasing efficiency and reducing reliance on extensive manual assessment. Based on experimental results, our best performing generative model achieved a Recall-Oriented Understudy for Gisting Evaluation-1 ( ROUGE-1 ) score of 47.77 demonstrating superior performance in generating fact-checking explanations, particularly when provided with high-quality evidence. Additionally, the best performing metric learning model showed a moderately strong correlation with human judgements on objective dimensions such as (self)-contradiction and hallucination , achieving a Matthews Correlation Coefficient (MCC) of around 0.7. • A dataset for fact-checking which includes explanations written by journalists. • Transformer models for generating human-accessible fact-checking explanations. • Multi-dimensional annotations reflecting explanation quality judgements. • A metric learning model scoring explanations aligned with these judgements.
Darius Feher, Abdullah Salem Khered, Riza Theresa Batista-Navarro, Viktor Schlegel
Eng. Appl. Artif. Intell.4
2025 Discontinuous named entities in clinical text: A systematic literature review
abstract
OBJECTIVE: Extracting named entities from clinical free-text presents unique challenges, particularly when dealing with discontinuous entities-mentions that are separated by unrelated words. Traditional NER methods often struggle to accurately identify these entities, prompting the development of specialised computational solutions. This paper systematically reviews and presents the methodologies developed for Discontinuous Named Entity Recognition in clinical texts, highlighting their effectiveness and the challenges they face. METHOD: We conducted a systematic literature review focused on discontinuous named entities, using structured searches across four Computer Science-related and one medical-related electronic database. A combination of search terms, grouped into three synonym categories-problem, entity/approach, and task-yielded 2,442 articles. Guided by our research objectives, we identified five key dimensions to systematically annotate and normalise the data for comprehensive analysis. RESULT: The review included 44 studies which were coded across several key dimensions: the chronological development of approaches, the corpora used, the downstream tasks affected by discontinuous named entities, the methodological approaches proposed to address the issue, and the reported performance outcomes. The discussion section examines the challenges encountered in this area and suggests potential directions for future research. CONCLUSION: Significant progress has been made in discontinuous named entity recognition; however, there remains a need for more adaptable, generalisable solutions that are independent of custom annotation schemes. Exploring various configurations of generative language models presents a promising avenue for advancing this area. Additionally, future research should investigate the impact of precise versus imprecise recognition of discontinuous entities on clinical downstream tasks to better understand its practical implications in healthcare applications.
Areej Alhassan, Viktor Schlegel, Monira Aloud, Riza Theresa Batista-Navarro, Goran Nenadic
J. Biomed. Informatics4
2024 Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
abstract
Efforts to apply transformer-based language models (TLMs) to the problem of reasoning in natural language have enjoyed ever-increasing success in recent years.The most fundamental task in this area to which nearly all others can be reduced is that of determining satisfiability.However, from a logical point of view, satisfiability problems vary along various dimensions, which may affect TLMs' ability to learn how to solve them.The problem instances of satisfiability in natural language can belong to different computational complexity classes depending on the language fragment in which they are expressed.Although prior research has explored the problem of natural language satisfiability, the above-mentioned point has not been discussed adequately.Hence, we investigate how problem instances from varying computational complexity classes and having different grammatical constructs impact TLMs' ability to learn rules of inference.Furthermore, to faithfully evaluate TLMs, we conduct an empirical study to explore the distribution of satisfiability problems.
Tharindu Madusanka, Ian Pratt-Hartmann, Riza Theresa Batista-Navarro
ACL (1)3
2024 IDEM: The IDioms with EMotions Dataset for Emotion Recognition
abstract
Idiomatic expressions are used in everyday language and typically convey affect, i.e., emotion. However, very little work investigating the extent to which automated methods can recognise emotions expressed in idiom-containing text has been undertaken. This can be attributed to the lack of emotion-labelled datasets that support the development and evaluation of such methods. In this paper, we present the IDioms with EMotions (IDEM) dataset consisting of a total of 9685 idiom-containing sentences that were generated and labelled with any one of 36 emotion types, with the help of the GPT-4 generative language model. Human validation by two independent annotators showed that more than 51% of the generated sentences are ideal examples, with the annotators reaching an agreement rate of 62% measured in terms of Cohen’s Kappa coefficient. To establish baseline performance on IDEM, various transformer-based emotion recognition approaches were implemented and evaluated. Results show that a RoBERTa model fine-tuned as a sequence classifier obtains a weighted F1-score of 58.73%, when the sequence provided as input specifies the idiom contained in a given sentence, together with its definition. Since this input configuration is based on the assumption that the idiom contained in the given sentence is already known, we also sought to assess the feasibility of automatically identifying the idioms contained in IDEM sentences. To this end, a hybrid idiom identification approach combining a rule-based method and a deep learning-based model was developed, whose performance on IDEM was determined to be 84.99% in terms of F1-score.
Alexander Prochnow, Johannes E. Bendler, Caroline Lange, Foivos Ioannis Tzavellos, Bas Marco Göritzer, Marijn ten Thij, Riza Theresa Batista-Navarro
LREC/COLING7
2024 CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models using Real and Synthetic Back-Translation Data
abstract
Neural Machine Translation (NMT) for low-resource languages remains a challenge for many NLP researchers. In this work, we deploy a standard data augmentation methodology by back-translation to a new language translation direction, i.e., Cantonese-to-English. We present the models we fine-tuned using the limited amount of real data and the synthetic data we generated using back-translation by three models: OpusMT, NLLB, and mBART.We carried out automatic evaluation using a range of different metrics including those that are lexical-based and embedding-based.Furthermore, we create a user-friendly interface for the models we included in this project, CantonMT, and make it available to facilitate Cantonese-to-English MT research. Researchers can add more models to this platform via our open-source CantonMT toolkit, available at https://github.com/kenrickkung/CantoneseTranslation.
Kung Yin Hong, Lifeng Han, Riza Theresa Batista-Navarro, Goran Nenadic
EAMT (1)3
2023 Do You Hear The People Sing? Key Point Analysis via Iterative Clustering and Abstractive Summarisation
abstract
Argument summarisation is a promising but currently under-explored field.Recent work has aimed to provide textual summaries in the form of concise and salient short texts, i.e., key points (KPs), in a task known as Key Point Analysis (KPA).One of the main challenges in KPA is finding high-quality key point candidates from dozens of arguments even in a small corpus.Furthermore, evaluating key points is crucial in ensuring that the automatically generated summaries are useful.Although automatic methods for evaluating summarisation have considerably advanced over the years, they mainly focus on sentence-level comparison, making it difficult to measure the quality of a summary (a set of KPs) as a whole.Aggravating this problem is the fact that human evaluation is costly and unreproducible.To address the above issues, we propose a two-step abstractive summarisation framework based on neural topic modelling with an iterative clustering procedure, to generate key points which are aligned with how humans identify key points.Our experiments show that our framework advances the state of the art in KPA, with performance improvement of up to 14 (absolute) percentage points, in terms of both ROUGE and our own proposed evaluation metrics 1 .Furthermore, we evaluate the generated summaries using a novel set-based evaluation toolkit.Our quantitative analysis demonstrates the effectiveness of our proposed evaluation metrics in assessing the quality of generated KPs.Human evaluation further demonstrates the advantages of our approach and validates that our proposed evaluation metric is more consistent with human judgment than ROUGE scores.Children express themselves through the clothes they wear and should be able to do this at school School uniform is harming the student's self expression School uniforms are expensive [...]Children should be able to dress as they wish, within reason, at school rather than being restricted from expressing themselves through their clothes.Children should be allowed to express themselves School uniform is unaffordable for many single parents and should be abandoned.School uniforms are an expense that many families can't afford.there are plenty of ways to get very cheap clothing, but discounted uniforms are more difficult to obtain.School uniforms are expensive and puts an undue burden on the parents of the students.School uniforms are expensive for the school and take money from other important programs.School unforms stifle freedom of expression.they can be costly and make circumstances difficult for those on a budget [...
Hao Li 0074, Viktor Schlegel, Riza Theresa Batista-Navarro, Goran Nenadic
ACL (1)3
2023 Identifying the limits of transformers when performing model-checking with natural language
abstract
Can transformers learn to comprehend logical semantics in natural language?Although many strands of work on natural language inference have focussed on transformer models' ability to perform reasoning on text, the above question has not been answered adequately.This is primarily because the logical problems that have been studied in the context of natural language inference have their computational complexity vary with the logical and grammatical constructs within the sentences.As such, it is difficult to access whether the difference in accuracy is due to logical semantics or the difference in computational complexity.A problem that is much suited to address this issue is that of the model-checking problem, whose computational complexity remains constant (for fragments derived from first-order logic).However, the model-checking problem remains untouched in natural language inference research.Thus, we investigated the problem of model-checking with natural language to adequately answer the question of how the logical semantics of natural language affects transformers' performance 1 .Our results imply that the language fragment has a significant impact on the performance of transformer models.Furthermore, we hypothesise that a transformer model can at least partially understand the logical semantics in natural language but can not completely learn the rules governing the modelchecking algorithm.
Tharindu Madusanka, Riza Theresa Batista-Navarro, Ian Pratt-Hartmann
EACL2
2023 TIMELINE: Exhaustive Annotation of Temporal Relations Supporting the Automatic Ordering of Events in News Articles
abstract
Temporal relation extraction models have thus far been hindered by a number of issues in existing temporal relation-annotated news datasets, including: (1) low inter-annotator agreement due to the lack of specificity of their annotation guidelines in terms of what counts as a temporal relation; (2) the exclusion of long-distance relations within a given document (those spanning across different paragraphs); and (3) the exclusion of events that are not centred on verbs.This paper aims to alleviate these issues by presenting a new annotation scheme that clearly defines the criteria based on which temporal relations should be annotated.Additionally, the scheme includes events even if they are not expressed as verbs (e.g., nominalised events).Furthermore, we propose a method for annotating all temporal relations-including long-distance ones-which automates the process, hence reducing time and manual effort on the part of annotators.The result is a new dataset, the TIMELINE corpus, in which improved inter-annotator agreement was obtained, in comparison with previously reported temporal relation datasets.We report the results of training and evaluating baseline temporal relation extraction models on the new corpus, and compare them with results obtained on the widely used MATRES corpus.
Sarah Alsayyahi, Riza Theresa Batista-Navarro
EMNLP2
2023 Not all quantifiers are equal: Probing Transformer-based language models' understanding of generalised quantifiers
abstract
How do different generalised quantifiers affect the behaviour of transformer-based language models (TLMs)?The recent popularity of TLMs and the central role generalised quantifiers have traditionally played in linguistics and logic bring this question into particular focus.The current research investigating this subject has not utilised a task defined purely in a logical sense, and thus, has not captured the underlying logical significance of generalised quantifiers.Consequently, they have not answered the aforementioned question faithfully or adequately.Therefore, we investigate how different generalised quantifiers affect TLMs by employing a textual entailment problem defined in a purely logical sense, namely, modelchecking with natural language.Our approach permits the automatic construction of datasets with respect to which we can assess the ability of TLMs to learn the meanings of generalised quantifiers.Our investigation reveals that TLMs generally can comprehend the logical semantics of the most common generalised quantifiers, but that distinct quantifiers influence TLMs in varying ways.
Tharindu Madusanka, Iqra Zahid, Hao Li 0074, Ian Pratt-Hartmann, Riza Theresa Batista-Navarro
EMNLP5
2023 Few-shot entity linking of food names
abstract
Entity linking (EL), the task of automatically matching mentions in text to concepts in a target knowledge base, remains under-explored when it comes to the food domain, despite its many potential applications, e.g., finding the nutritional value of ingredients in databases. In this paper, we describe the creation of new resources supporting the development of EL methods applied to the food domain: the E.Care Knowledge Base (E.Care KB) which contains 664 food concepts and the E.Care dataset, a corpus of 468 cooking recipes where ingredient names have been manually linked to corresponding concepts in the E.Care KB. We developed and evaluated different methods for EL, namely, deep learning-based approaches underpinned by Siamese networks trained under a few-shot learning setting, traditional machine learning-based approaches underpinned by support vector machines (SVMs) and unsupervised approaches based on string matching algorithms. Combining the strengths of each of these approaches, we built a hybrid model for food EL that balances the trade-offs between performance and inference speed. Specifically, our hybrid model obtains 89.40% accuracy and links mentions at an average speed of 0.24 seconds per mention, whereas our best deep learning-based model, SVM model and unsupervised model obtain accuracies of 86.99%, 87.19% and 87.43% at inference speeds of 0.007, 0.66 and 0.02 seconds per mention, respectively.
Darius Feher, Faridz Ibrahim, Zhuyan Cheng, Viktor Schlegel, Tom Maidment, James Bagshaw, Riza Theresa Batista-Navarro
Inf. Process. Manag.7
2023 Global information-aware argument mining based on a top-down multi-turn QA model
abstract
Argument mining (AM) aims to automatically generate a graph that represents the argument structure of a document. Most previous AM models only pay attention to a single argument component (AC) to classify the type of the AC or a pair of ACs to identify and classify the argumentative relation (AR) between the two ACs. These models ignore the impact of global argument structure of the documents, which is important, especially in some highly structured genres such as scientific papers, where the process of argumentation is relatively fixed. Inspired by this, we propose a novel two-stage model which leverages global structure information to support AM. The first stage uses a multi-turn question-answering model to incrementally generate an initial argumentative graph that identifies relations among ACs. At each turn, all ACs related to the query AC are generated simultaneously, such that the sibling global information between the answer ACs is considered. In addition, the partially constructed graph is used as global structure information to support the extension of the graph with additional ACs. After the whole initial graph structure has been determined, the second stage assigns semantic types to both the ACs and ARs among them, leveraging information from this initial graph as global structure information. We test the proposed methods on two scientific datasets (one is the AbstRCT dataset including 659 abstracts about cancer research and the other is the SciARG dataset that consists of 225 computer linguistic abstracts and 285 biomedical abstracts) and a student essay dataset PE with 402 essays. Our experiments show that our model improves the state-of-the-art performance on two scientific datasets for different AM subtasks, with average improvements of 1%, 2.41%, 1.1% for the ACC, ARI and ARC task respectively on the AbstRCT dataset, and 2.36%, 1.84%, 8.87% for the ACC, ARI and ARC task on the SciARG dataset. Our model also achieves comparative results on the PE datasets: 87.7% of F1 scores for the ACC task, 81.4% for the ARI task and 78.8% for the ARC task.
Boyang Liu 0002, Viktor Schlegel, Paul Thompson 0002, Riza Theresa Batista-Navarro, Sophia Ananiadou
Inf. Process. Manag.4
2023 A survey of methods for revealing and overcoming weaknesses of data-driven Natural Language Understanding
abstract
Abstract Recent years have seen a growing number of publications that analyse Natural Language Understanding (NLU) datasets for superficial cues, whether they undermine the complexity of the tasks underlying those datasets and how they impact those models that are optimised and evaluated on this data. This structured survey provides an overview of the evolving research area by categorising reported weaknesses in models and datasets and the methods proposed to reveal and alleviate those weaknesses for the English language. We summarise and discuss the findings and conclude with a set of recommendations for possible future research directions. We hope that it will be a useful resource for researchers who propose new datasets to assess the suitability and quality of their data to evaluate various phenomena of interest, as well as those who propose novel NLU approaches, to further understand the implications of their improvements with respect to their model’s acquired capabilities.
Viktor Schlegel, Goran Nenadic, Riza Theresa Batista-Navarro
Nat. Lang. Eng.3
2022 RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced Labour
abstract
Forced labour is the most common type of modern slavery, and it is increasingly gaining the attention of the research and social community. Recent studies suggest that artificial intelligence (AI) holds immense potential for augmenting anti-slavery action. However, AI tools need to be developed transparently in cooperation with different stakeholders. Such tools are contingent on the availability and access to domain-specific data, which are scarce due to the near-invisible nature of forced labour. To the best of our knowledge, this paper presents the first openly accessible English corpus annotated for multi-class and multi-label forced labour detection. The corpus consists of 989 news articles retrieved from specialised data sources and annotated according to risk indicators defined by the International Labour Organization (ILO). Each news article was annotated for two aspects: (1) indicators of forced labour as classification labels and (2) snippets of the text that justify labelling decisions. We hope that our data set can help promote research on explainability for multi-class and multi-label text classification. In this work, we explain our process for collecting the data underpinning the proposed corpus, describe our annotation guidelines and present some statistical analysis of its content. Finally, we summarise the results of baseline experiments based on different variants of the Bidirectional Encoder Representation from Transformer (BERT) model.
Erick Mendez Guzman, Viktor Schlegel, Riza Theresa Batista-Navarro
LREC3
2022 Incorporating Zoning Information into Argument Mining from Biomedical Literature
abstract
The goal of text zoning is to segment a text into zones (i.e., Background, Conclusion) that serve distinct functions. Argumentative zoning, a specific text zoning scheme for the scientific domain, is considered as the antecedent for argument mining by many researchers. Surprisingly, however, little work is concerned with exploiting zoning information to improve the performance of argument mining models, despite the relatedness of the two tasks. In this paper, we propose two transformer-based models to incorporate zoning information into argumentative component identification and classification tasks. One model is for the sentence-level argument mining task and the other is for the token-level task. In particular, we add the zoning labels predicted by an off-the-shelf model to the beginning of each sentence, inspired by the convention commonly used biomedical abstracts. Moreover, we employ multi-head attention to transfer the sentence-level zoning information to each token in a sentence. Based on experiment results, we find a significant improvement in F1-scores for both sentence- and token-level tasks. It is worth mentioning that these zoning labels can be obtained with high accuracy by utilising readily available automated methods. Thus, existing argument mining models can be improved by incorporating zoning information without any additional annotation cost.
Boyang Liu 0002, Viktor Schlegel, Riza Theresa Batista-Navarro, Sophia Ananiadou
LREC3
2021 Semantics Altering Modifications for Evaluating Comprehension in Machine Reading
abstract
Advances in NLP have yielded impressive results for the task of machine reading comprehension (MRC), with approaches having been reported to achieve performance comparable to that of humans. In this paper, we investigate whether state-of-the-art MRC models are able to correctly process Semantics Altering Modifications (SAM): linguistically-motivated phenomena that alter the semantics of a sentence while preserving most of its lexical surface form. We present a method to automatically generate and align challenge sets featuring original and altered examples. We further propose a novel evaluation methodology to correctly assess the capability of MRC systems to process these examples independent of the data they were optimised on, by discounting for effects introduced by domain shift. In a large-scale empirical study, we apply the methodology in order to evaluate extractive MRC models with regard to their capability to correctly process SAM-enriched data. We comprehensively cover 12 different state-of-the-art neural architecture configurations and four training datasets and find that -- despite their well-known remarkable performance -- optimised models consistently struggle to correctly process semantically altered data.
Viktor Schlegel, Goran Nenadic, Riza Theresa Batista-Navarro
AAAI3
2021 Is the Understanding of Explicit Discourse Relations Required in Machine Reading Comprehension?
abstract
An in-depth analysis of the level of language understanding required by existing Machine Reading Comprehension (MRC) benchmarks can provide insight into the reading capabilities of machines.In this paper, we propose an ablation-based methodology to assess the extent to which MRC datasets evaluate the understanding of explicit discourse relations.We define seven MRC skills which require the understanding of different discourse relations.We then introduce ablation methods that verify whether these skills are required to succeed on a dataset.By observing the drop in performance of neural MRC models evaluated on the original and the modified dataset, we can measure to what degree the dataset requires these skills, in order to be understood correctly.Experiments on three large-scale datasets with the BERT-base and ALBERT-xxlarge model show that the relative changes for all skills are small (less than 6%).These results imply that most of the answered questions in the examined datasets do not require understanding the discourse structure of the text.To specifically probe for natural language understanding, there is a need to design more challenging benchmarks that can correctly evaluate the intended skills 1 .
Viktor Schlegel, Riza Theresa Batista-Navarro
EACL3
2020 ParlVote: A Corpus for Sentiment Analysis of Political Debates
abstract
Debate transcripts from the UK Parliament contain information about the positions taken by politicians towards important topics, but are difficult for people to process manually. While sentiment analysis of debate speeches could facilitate understanding of the speakers’ stated opinions, datasets currently available for this task are small when compared to the benchmark corpora in other domains. We present ParlVote, a new, larger corpus of parliamentary debate speeches for use in the evaluation of sentiment analysis systems for the political domain. We also perform a number of initial experiments on this dataset, testing a variety of approaches to the classification of sentiment polarity in debate speeches. These include a linear classifier as well as a neural network trained using a transformer word embedding model (BERT), and fine-tuned on the parliamentary speeches. We find that in many scenarios, a linear classifier trained on a bag-of-words text representation achieves the best results. However, with the largest dataset, the transformer-based model combined with a neural classifier provides the best performance. We suggest that further experimentation with classification models and observations of the debate content and structure are required, and that there remains much room for improvement in parliamentary sentiment analysis.
Gavin Abercrombie, Riza Theresa Batista-Navarro
LREC2
2020 A Framework for Evaluation of Machine Reading Comprehension Gold Standards
abstract
Machine Reading Comprehension (MRC) is the task of answering a question over a paragraph of text. While neural MRC systems gain popularity and achieve noticeable performance, issues are being raised with the methodology used to establish their performance, particularly concerning the data design of gold standards that are used to evaluate them. There is but a limited understanding of the challenges present in this data, which makes it hard to draw comparisons and formulate reliable hypotheses. As a first step towards alleviating the problem, this paper proposes a unifying framework to systematically investigate the present linguistic features, required reasoning and background knowledge and factual correctness on one hand, and the presence of lexical cues as a lower bound for the requirement of understanding on the other hand. We propose a qualitative annotation schema for the first and a set of approximative metrics for the latter. In a first application of the framework, we analyse modern MRC gold standards and present our findings: the absence of features that contribute towards lexical ambiguity, the varying factual correctness of the expected answers and the presence of lexical cues, all of which potentially lower the reading comprehension complexity and quality of the evaluation data.
Viktor Schlegel, Marco Valentino, André Freitas, Goran Nenadic, Riza Theresa Batista-Navarro
LREC5
2019 Policy Preference Detection in Parliamentary Debate Motions
abstract
Debate motions (proposals) tabled in the UK Parliament contain information about the stated policy preferences of the Members of Parliament who propose them, and are key to the analysis of all subsequent speeches given in response to them.We attempt to automatically label debate motions with codes from a pre-existing coding scheme developed by political scientists for the annotation and analysis of political parties' manifestos.We develop annotation guidelines for the task of applying these codes to debate motions at two levels of granularity and produce a dataset of manually labelled examples.We evaluate the annotation process and the reliability and utility of the labelling scheme, finding that inter-annotator agreement is comparable with that of other studies conducted on manifesto data.Moreover, we test a variety of ways of automatically labelling motions with the codes, ranging from similarity matching to neural classification methods, and evaluate them against the gold standard labels.From these experiments, we note that established supervised baselines are not always able to improve over simple lexical heuristics.At the same time, we detect a clear and evident benefit when employing BERT, a state-of-the-art deep language representation model, even in classification scenarios with over 30 different labels and limited amounts of training data.
Gavin Abercrombie, Federico Nanni, Riza Theresa Batista-Navarro, Simone Paolo Ponzetto
CoNLL3
2019 Topic Modelling vs Distant Supervision: A Comparative Evaluation Based on the Classification of Parliamentary Enquiries
Riza Theresa Batista-Navarro, Oliver Hawkins
TPDL1
2019 Using Prior Knowledge to Facilitate Computational Reading of Arabic Calligraphy
Seetah ALSalamah, Riza Theresa Batista-Navarro, Ross D. King
IDEAL (2)2
2019 Studying the Evolution of the 'Circular Economy' Concept Using Topic Modelling
Sampriti Mahanty, Frank Boons, Julia Handl, Riza Theresa Batista-Navarro
IDEAL (2)4
2019 Whose story is it anyway? Automatic extraction of accounts from news articles
abstract
Narratives are comprised of stories that provide insight into social processes. To facilitate the analysis of narratives in a more efficient manner, natural language processing (NLP) methods have been employed in order to automatically extract information from textual sources, e.g., newspaper articles. Existing work on automatic narrative extraction, however, has ignored the nested character of narratives. In this work, we argue that a narrative may contain multiple accounts given by different actors. Each individual account provides insight into the beliefs and desires underpinning an actor’s actions. We present a pipeline for automatically extracting accounts, consisting of NLP methods for: (1) named entity recognition, (2) event extraction, and (3) attribution extraction. Machine learning-based models for named entity recognition were trained based on a state-of-the-art neural network architecture for sequence labelling. For event extraction, we developed a hybrid approach combining the use of semantic role labelling tools, the FrameNet repository of semantic frames, and a lexicon of event nouns. Meanwhile, attribution extraction was addressed with the aid of a dependency parser and Levin’s verb classes. To facilitate the development and evaluation of these methods, we constructed a new corpus of news articles, in which named entities, events and attributions have been manually marked up following a novel annotation scheme that covers over 20 event types relating to socio-economic phenomena. Evaluation results show that relative to a baseline method underpinned solely by semantic role labelling tools, our event extraction approach optimises recall by 12.22–14.20 percentage points (reaching as high as 92.60% on one data set). Meanwhile, the use of Levin’s verb classes in attribution extraction obtains optimal performance in terms of F-score, outperforming a baseline method by 7.64–11.96 percentage points. Our proposed approach was applied on news articles focused on industrial regeneration cases. This facilitated the generation of accounts of events that are attributed to specific actors.
Frank Boons, Riza Theresa Batista-Navarro
Inf. Process. Manag.3
2018 Using semantic frames to identify related textual requirements: an initial validation
abstract
Identifying relationships between requirements described in natural language (NL) is a difficult task in requirements engineering (RE). This paper presents a novel approach that uses Semantic Frames in FrameNet to find the relationships between requirements. Our initial validation shows that the approach is promising, with an F-Score of 83%. Our next step is to use the approach to identify implicit requirements relationships and finding requirements traceability links.
Waad Alhoshan, Liping Zhao 0001, Riza Theresa Batista-Navarro
ESEM3
2018 'Aye' or 'No'? Speech-level Sentiment Analysis of Hansard UK Parliamentary Debate Transcripts
Gavin Abercrombie, Riza Theresa Batista-Navarro
LREC2
2018 Towards a Corpus of Requirements Documents Enriched with Semantic Frame Annotations
abstract
Software requirements are typically written in natural language, which need to be transformed into a more formal representation. Natural language processing techniques have been applied to aid in this transformation. Semantic parsing, for instance, adds semantic structure to text. It however requires supporting corpora which are still missing in requirements engineering. To address this gap, we developed FN-RE, a corpus of requirements documents, which was annotated based on semantic frames in FrameNet. Each requirement statement was manually labelled by two annotators by selecting suitable semantic frames and related frame elements. We obtained an average agreement of 72.85% between the two annotators, measured by F-score, thus indicating that the annotations provided in our corpus are reliable.
Waad Alhoshan, Riza Theresa Batista-Navarro, Liping Zhao 0001
RE2
2018 LitPathExplorer: a confidence-based visual text analytics tool for exploring literature-enriched pathway models
abstract
Motivation: Pathway models are valuable resources that help us understand the various mechanisms underpinning complex biological processes. Their curation is typically carried out through manual inspection of published scientific literature to find information relevant to a model, which is a laborious and knowledge-intensive task. Furthermore, models curated manually cannot be easily updated and maintained with new evidence extracted from the literature without automated support. Results: We have developed LitPathExplorer, a visual text analytics tool that integrates advanced text mining, semi-supervised learning and interactive visualization, to facilitate the exploration and analysis of pathway models using statements (i.e. events) extracted automatically from the literature and organized according to levels of confidence. LitPathExplorer supports pathway modellers and curators alike by: (i) extracting events from the literature that corroborate existing models with evidence; (ii) discovering new events which can update models; and (iii) providing a confidence value for each event that is automatically computed based on linguistic features and article metadata. Our evaluation of event extraction showed a precision of 89% and a recall of 71%. Evaluation of our confidence measure, when used for ranking sampled events, showed an average precision ranging between 61 and 73%, which can be improved to 95% when the user is involved in the semi-supervised learning process. Qualitative evaluation using pair analytics based on the feedback of three domain experts confirmed the utility of our tool within the context of pathway model exploration. Availability and implementation: LitPathExplorer is available at http://nactem.ac.uk/LitPathExplorer_BI/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Axel J. Soto, Chrysoula Zerva, Riza Theresa Batista-Navarro, Sophia Ananiadou
Bioinform.3
2017 Modelling the Coverage of Dipterocarp Trees in Central Visayas, Philippines
abstract
With the rapid decline of dipterocarp coverage in the Philippines, efforts in restoration would benefit from having a suitability map that would indicate the suitable areas where different dipterocarp species could thrive. Another use of these maps would be for assessing whether particular locations should be considered as Protected Areas in which exploitation would be limited, if not completely prohibited. We obtained data from the Department of Environment and Natural Resources of the Philippines, which consists of dipterocarp-related information gathered over localities in the country's Central Visayas region. Six climate data parameters were then chosen and obtained from the Philippine Atmospheric Geophysical and Astronomical Services Administration. We employed a maximum entropy-based niche modelling approach to generate a suitability map. The model that we produced obtained an area under the curve (AUC) score of 0.955 with minimum temperature being the variable with the highest contribution. The map produced indicates that dipterocarp trees are more likely to appear in the lower regions of Central Visayas.
Adrian Jose Sabado, Geoffrey Solano, Marilou Nicolas, Riza Theresa Batista-Navarro, Roselyn Gabud, Vincent Hilomen
eScience4
2017 Using uncertainty to link and rank evidence from biomedical literature for model curation
abstract
MOTIVATION: In recent years, there has been great progress in the field of automated curation of biomedical networks and models, aided by text mining methods that provide evidence from literature. Such methods must not only extract snippets of text that relate to model interactions, but also be able to contextualize the evidence and provide additional confidence scores for the interaction in question. Although various approaches calculating confidence scores have focused primarily on the quality of the extracted information, there has been little work on exploring the textual uncertainty conveyed by the author. Despite textual uncertainty being acknowledged in biomedical text mining as an attribute of text mined interactions (events), it is significantly understudied as a means of providing a confidence measure for interactions in pathways or other biomedical models. In this work, we focus on improving identification of textual uncertainty for events and explore how it can be used as an additional measure of confidence for biomedical models. RESULTS: We present a novel method for extracting uncertainty from the literature using a hybrid approach that combines rule induction and machine learning. Variations of this hybrid approach are then discussed, alongside their advantages and disadvantages. We use subjective logic theory to combine multiple uncertainty values extracted from different sources for the same interaction. Our approach achieves F-scores of 0.76 and 0.88 based on the BioNLP-ST and Genia-MK corpora, respectively, making considerable improvements over previously published work. Moreover, we evaluate our proposed system on pathways related to two different areas, namely leukemia and melanoma cancer research. AVAILABILITY AND IMPLEMENTATION: The leukemia pathway model used is available in Pathway Studio while the Ras model is available via PathwayCommons. Online demonstration of the uncertainty extraction system is available for research purposes at http://argo.nactem.ac.uk/test. The related code is available on https://github.com/c-zrv/uncertainty_components.git. Details on the above are available in the Supplementary Material. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chrysoula Zerva, Riza Theresa Batista-Navarro, Philip Day, Sophia Ananiadou
Bioinform.2
2016 A Text Mining Framework for Accelerating the Semantic Curation of Literature
Riza Theresa Batista-Navarro, Jennifer Hammock, William Ulate, Sophia Ananiadou
TPDL1
2014 Interoperability and Customisation of Annotation Schemata in Argo
Rafal Rak, Jacob Carter, Andrew Rowley, Riza Theresa Batista-Navarro, Sophia Ananiadou
LREC4
2013 Facilitating the Analysis of Discourse Phenomena in an Interoperable NLP Platform
Riza Theresa Batista-Navarro, Georgios Kontonatsios, Claudiu Mihaila, Paul Thompson 0002, Rafal Rak, Raheel Nawaz, Ioannis Korkontzelos, Sophia Ananiadou
CICLing (1)1
2012 What's in a Name? Entity Type Variation across Two Biomedical Subdomains
Claudiu Mihaila, Riza Theresa Batista-Navarro
EACL2
2011 Detecting experimental techniques and selecting relevant documents for protein-protein interactions from biomedical literature
abstract
BACKGROUND: The selection of relevant articles for curation, and linking those articles to experimental techniques confirming the findings became one of the primary subjects of the recent BioCreative III contest. The contest's Protein-Protein Interaction (PPI) task consisted of two sub-tasks: Article Classification Task (ACT) and Interaction Method Task (IMT). ACT aimed to automatically select relevant documents for PPI curation, whereas the goal of IMT was to recognise the methods used in experiments for identifying the interactions in full-text articles. RESULTS: We proposed and compared several classification-based methods for both tasks, employing rich contextual features as well as features extracted from external knowledge sources. For IMT, a new method that classifies pair-wise relations between every text phrase and candidate interaction method obtained promising results with an F1 score of 64.49%, as tested on the task's development dataset. We also explored ways to combine this new approach and more conventional, multi-label document classification methods. For ACT, our classifiers exploited automatically detected named entities and other linguistic information. The evaluation results on the BioCreative III PPI test datasets showed that our systems were very competitive: one of our IMT methods yielded the best performance among all participants, as measured by F1 score, Matthew's Correlation Coefficient and AUC iP/R; whereas for ACT, our best classifier was ranked second as measured by AUC iP/R, and also competitive according to other metrics. CONCLUSIONS: Our novel approach that converts the multi-class, multi-label classification problem to a binary classification problem showed much promise in IMT. Nevertheless, on the test dataset the best performance was achieved by taking the union of the output of this method and that of a multi-class, multi-label document classifier, which indicates that the two types of systems complement each other in terms of recall. For ACT, our system exploited a rich set of features and also obtained encouraging results. We examined the features with respect to their contributions to the classification results, and concluded that contextual words surrounding named entities, as well as the MeSH headings associated with the documents were among the main contributors to the performance.
Xinglong Wang, Rafal Rak, Angelo C. Restificar, Chikashi Nobata, C. J. Rupp, Riza Theresa Batista-Navarro, Raheel Nawaz, Sophia Ananiadou
BMC Bioinform.6