EDBT 2026 Demo / reviewers in the wild / expert
Maria Liakata
dblp:65/6264
· DBLP profile ↗
52ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-5765-0416ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 4 first-author · 18 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Responsible Evaluation of AI for Mental HealthabstractHiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz 0001, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza del Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych |
ACL (1) | 15 |
| 2025 | Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI EvaluationabstractScientific language models drive research innovation but require extensive fine-tuning on large datasets.This work enhances such models by improving their inference and evaluation capabilities with minimal or no additional training.Focusing on molecule caption generation, we explore post-training synergies between alignment fine-tuning and model merging in a cross-modal setup.We reveal intriguing insights into the behaviour and suitability of such methods while significantly surpassing state-of-the-art models.Moreover, we propose a novel atomic-level evaluation method leveraging off-the-shelf Natural Language Inference (NLI) models for use in the unseen chemical domain.Our experiments demonstrate that our evaluation operates at the right level of granularity, effectively handling multiple content units and subsentence reasoning, while widely adopted NLI methods consistently misalign with assessment criteria. Dimitris Gkoumas, Maria Liakata |
ACL (1) | 2 |
| 2025 | Temporal reasoning for timeline summarisation in social mediaabstractThis paper explores whether enhancing temporal reasoning capabilities in Large Language Models (LLMs) can improve the quality of timeline summarisation, the task of summarising long texts containing sequences of events, such as social media threads.We first introduce NarrativeReason, a novel dataset focused on temporal relations among sequential events within narratives, distinguishing it from existing temporal reasoning datasets that primarily address pair-wise event relations.Our approach then combines temporal reasoning with timeline summarisation through a knowledge distillation framework, where we first fine-tune a teacher model on temporal reasoning tasks and then distill this knowledge into a student model while simultaneously training it for the task of timeline summarisation.Experimental results demonstrate that our model achieves superior performance on out-of-domain mental health-related timeline summarisation tasks, which involve long social media threads with repetitions of events and a mix of emotions, highlighting the importance and generalisability of leveraging temporal reasoning to improve timeline summarisation. Jiayu Song, Mahmud Elahi Akhter, Dana Atzil-Slonim, Maria Liakata |
ACL (1) | 4 |
| 2025 | Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process SupervisionabstractLarge language models (LLMs) have shown strong performance in many reasoning benchmarks. However, recent studies have pointed to memorization, rather than generalization, as one of the leading causes for such performance. LLMs, in fact, are susceptible to content variations, demonstrating a lack of robust planning or symbolic abstractions supporting their reasoning process. To improve reliability, many attempts have been made to combine LLMs with symbolic methods. Nevertheless, existing approaches fail to effectively leverage symbolic representations due to the challenges involved in developing reliable and scalable verification mechanisms. In this paper, we propose to overcome such limitations by synthesizing high-quality symbolic reasoning trajectories with stepwise pseudo-labels at scale via Monte Carlo estimation. A Process Reward Model (PRM) can be efficiently trained based on the synthesized data and then used to select more symbolic trajectories. The trajectories are then employed with Direct Preference Optimization (DPO) and Supervised Fine-Tuning (SFT) to improve logical reasoning and generalization. Our results on benchmarks (i.e., FOLIO and LogicAsker) show the effectiveness of the proposed method with gains on frontier and open-weight models. Moreover, additional experiments on claim verification data reveal that fine-tuning on the generated symbolic reasoning trajectories enhances out-of-domain generalizability, suggesting the potential impact of the proposed method in enhancing planning and logical reasoning. Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Maria Liakata, Nikolaos Aletras |
EMNLP | 4 |
| 2025 | Evaluating Synthetic Data Generation from User Generated TextabstractAbstract User-generated content provides a rich resource to study social and behavioral phenomena. Although its application potential is currently limited by the paucity of expert labels and the privacy risks inherent in personal data, synthetic data can help mitigate this bottleneck. In this work, we introduce an evaluation framework to facilitate research on synthetic language data generation for user-generated text. We define a set of aspects for assessing data quality, namely, style preservation, meaning preservation, and divergence, as a proxy for privacy. We introduce metrics corresponding to each aspect. Moreover, through a set of generation strategies and representative tasks and baselines across domains, we demonstrate the relation between the quality aspects of synthetic user generated content, generation strategies, metrics, and downstream performance. To our knowledge, our work is the first unified evaluation framework for user-generated text in relation to the specified aspects, offering both intrinsic and extrinsic evaluation. We envisage it will facilitate developments towards shareable, high-quality synthetic language data. Jenny Chim, Julia Ive, Maria Liakata |
Comput. Linguistics | 3 |
| 2024 | Knowledge Graphs for Real-World Rumour VerificationabstractDespite recent progress in automated rumour verification, little has been done on evaluating rumours in a real-world setting. We advance the state-of-the-art on the PHEME dataset, which consists of Twitter response threads collected as a rumour was unfolding. We automatically collect evidence relevant to PHEME and use it to construct knowledge graphs in a time-sensitive manner, excluding information post-dating rumour emergence. We identify discrepancies between the evidence retrieved and PHEME’s labels, which are discussed in detail and amended to release an updated dataset. We develop a novel knowledge graph approach which finds paths linking disjoint fragments of evidence. Our rumour verification model which combines evidence from the graph outperforms the state-of-the-art on PHEME and has superior generisability when evaluated on a temporally distant rumour verification dataset. John Dougrez-Lewis, Elena Kochkina, Maria Liakata, Yulan He 0001 |
LREC/COLING | 3 |
| 2024 | A Multi-Task Transformer Model for Fine-grained Labelling of Chest X-Ray ReportsabstractPrecise understanding of free-text radiology reports through localised extraction of clinical findings can enhance medical imaging applications like computer-aided diagnosis. We present a new task, that of segmenting radiology reports into topically meaningful passages (segments) and a transformer-based model that both segments reports into semantically coherent segments and classifies each segment using a set of 37 radiological abnormalities, thus enabling fine-grained analysis. This contrasts with prior work that performs classification on full reports without localisation. Trained on over 2.7 million unlabelled chest X-ray reports and over 28k segmented and labelled reports, our model achieves state-of-the-art performance on report segmentation (0.0442 WinDiff) and multi-label classification (0.84 report-level macro F1) over 37 radiological labels and 8 NLP-specific labels. This work establishes new benchmarks for fine-grained understanding of free-text radiology reports, with precise localisation of semantics unlocking new opportunities to improve computer vision model training and clinical decision support. We open-source our annotation tool, model code and pretrained weights to encourage future research. Yuanyi Zhu, Maria Liakata, Giovanni Montana |
LREC/COLING | 2 |
| 2024 | LongEval: Longitudinal Evaluation of Model Performance at CLEF 2024
Rabab Alkhalifa, Hsuvas Borkakoty, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Tobias Fink, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, David Iommi, Maria Liakata, Harish Tayyar Madabushi, Pablo Medina-Alias, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga |
ECIR (6) | 11 |
| 2024 | TempoFormer: A Transformer for Temporally-aware Representations in Change DetectionabstractDynamic representation learning plays a pivotal role in understanding the evolution of linguistic content over time.On this front both context and time dynamics as well as their interplay are of prime importance.Current approaches model context via pre-trained representations, which are typically temporally agnostic.Previous work on modelling context and temporal dynamics has used recurrent methods, which are slow and prone to overfitting.Here we introduce TempoFormer, the first task-agnostic transformer-based and temporally-aware model for dynamic representation learning.Our approach is jointly trained on inter and intra context dynamics and introduces a novel temporal variation of rotary positional embeddings.The architecture is flexible and can be used as the temporal representation foundation of other models or applied to different transformer-based architectures.We show new SOTA performance on three different realtime change detection tasks. Talia Tseriotou, Adam Tsakalidis, Maria Liakata |
EMNLP | 3 |
| 2023 | Creation and evaluation of timelines for longitudinal user postsabstractThere is increasing interest to work with user generated content in social media, especially textual posts over time.Currently there is no consistent way of segmenting user posts into timelines in a meaningful way that improves the quality and cost of manual annotation.Here we propose a set of methods for segmenting longitudinal user posts into timelines likely to contain interesting moments of change in a user's behaviour, based on their online posting activity.We also propose a novel framework for evaluating timelines and show its applicability in the context of two different social media datasets.Finally, we present a discussion of the linguistic content of highly ranked timelines.1 Anthony Hills, Adam Tsakalidis, Federico Nanni, Ioannis Zachos, Maria Liakata |
EACL | 5 |
| 2023 | LongEval: Longitudinal Evaluation of Model Performance at CLEF 2023
Rabab Alkhalifa, Iman Munire Bilal, Hsuvas Borkakoty, José Camacho-Collados, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, Elena Kochkina, Maria Liakata, Daniel Loureiro, Harish Tayyar Madabushi, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga |
ECIR (3) | 12 |
| 2023 | Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with DementiaabstractDementia is associated with language disorders which impede communication.Here, we automatically learn linguistic disorder patterns by making use of a moderately-sized pre-trained language model and forcing it to focus on reformulated natural language processing (NLP) tasks and associated linguistic patterns.Our experiments show that NLP tasks that encapsulate contextual information and enhance the gradient signal with linguistic patterns benefit performance.We then use the probability estimates from the best model to construct digital linguistic markers measuring the overall quality in communication and the intensity of a variety of language disorders.We investigate how the digital markers characterize dementia speech from a longitudinal perspective.We find that our proposed communication marker is able to robustly and reliably characterize the language of people with dementia, outperforming existing linguistic approaches; and shows external validity via significant correlation with clinical markers of behaviour.Finally, our proposed linguistic disorder markers provide useful insights into gradual language impairment associated with disease progression. Dimitris Gkoumas, Matthew Purver, Maria Liakata |
EMNLP | 3 |
| 2023 | A Digital Language Coherence Marker for Monitoring DementiaabstractThe use of spontaneous language to derive appropriate digital markers has become an emergent, promising and non-intrusive method to diagnose and monitor dementia.Here we propose methods to capture language coherence as a cost-effective, human-interpretable digital marker for monitoring cognitive changes in people with dementia.We introduce a novel task to learn the temporal logical consistency of utterances in short transcribed narratives and investigate a range of neural approaches.We compare such language coherence patterns between people with dementia and healthy controls and conduct a longitudinal evaluation against three clinical bio-markers to investigate the reliability of our proposed digital coherence marker.The coherence marker shows a significant difference between people with mild cognitive impairment, those with Alzheimer's Disease and healthy controls.Moreover our analysis shows high association between the coherence marker and the clinical bio-markers as well as generalisability potential to other related conditions. Dimitris Gkoumas, Adam Tsakalidis, Maria Liakata |
EMNLP | 3 |
| 2023 | Evaluating the generalisability of neural rumour verification modelsabstractResearch on automated social media rumour verification, the task of identifying the veracity of questionable information circulating on social media, has yielded neural models achieving high performance, with accuracy scores that often exceed 90%. However, none of these studies focus on the real-world generalisability of the proposed approaches, that is whether the models perform well on datasets other than those on which they were initially trained and tested. In this work we aim to fill this gap by assessing the generalisability of top performing neural rumour verification models covering a range of different architectures from the perspectives of both topic and temporal robustness. For a more complete evaluation of generalisability, we collect and release COVID-RV, a novel dataset of Twitter conversations revolving around COVID-19 rumours. Unlike other existing COVID-19 datasets, our COVID-RV contains conversations around rumours that follow the format of prominent rumour verification benchmarks, while being different from them in terms of topic and time scale, thus allowing better assessment of the temporal robustness of the models. We evaluate model performance on COVID-RV and three popular rumour verification datasets to understand limitations and advantages of different model architectures, training datasets and evaluation scenarios. We find a dramatic drop in performance when testing models on a different dataset from that used for training. Further, we evaluate the ability of models to generalise in a few-shot learning setup, as well as when word embeddings are updated with the vocabulary of a new, unseen rumour. Drawing upon our experiments we discuss challenges and make recommendations for future research directions in addressing this important problem. Elena Kochkina, Tamanna Hossain, Robert L. Logan IV, Miguel Arana-Catania, Rob Procter, Arkaitz Zubiaga, Sameer Singh 0001, Yulan He 0001, Maria Liakata |
Inf. Process. Manag. | 9 |
| 2022 | Identifying Moments of Change from Longitudinal User TextabstractAdam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim, Jiayu Song, Maria Liakata. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Adam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim, Jiayu Song, Maria Liakata |
ACL (1) | 6 |
| 2022 | Unsupervised Opinion Summarisation in the Wasserstein SpaceabstractOpinion summarisation synthesises opinions expressed in a group of documents discussing the same topic to produce a single summary.Recent work has looked at opinion summarisation of clusters of social media posts.Such posts are noisy and have unpredictable structure, posing additional challenges for the construction of the summary distribution and the preservation of meaning compared to online reviews, which has been so far the focus of opinion summarisation.To address these challenges we present WassOS, an unsupervised abstractive summarization model which makes use of the Wasserstein distance.A Variational Autoencoder is used to get the distribution of documents/posts, and the distributions are disentangled into separate semantic and syntactic spaces.The summary distribution is obtained using the Wasserstein barycenter of the semantic and syntactic distributions.A latent variable sampled from the summary distribution is fed into a GRU decoder with a transformer layer to produce the final summary.Our experiments on multiple datasets including Twitter clusters, Reddit threads, and reviews show that WassOS almost always outperforms the state-of-the-art on ROUGE metrics and consistently produces the best summaries with respect to meaning preservation according to human evaluations. Jiayu Song, Iman Munire Bilal, Adam Tsakalidis, Rob Procter, Maria Liakata |
EMNLP | 5 |
| 2022 | Natural Language Inference with Self-Attention for Veracity Assessment of Pandemic ClaimsabstractMiguel Arana-Catania, Elena Kochkina, Arkaitz Zubiaga, Maria Liakata, Robert Procter, Yulan He. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Miguel Arana-Catania, Elena Kochkina, Arkaitz Zubiaga, Maria Liakata, Rob Procter, Yulan He 0001 |
NAACL-HLT | 4 |
| 2022 | Template-based Abstractive Microblog Opinion SummarisationabstractAbstract We introduce the task of microblog opinion summarization (MOS) and share a dataset of 3100 gold-standard opinion summaries to facilitate research in this domain. The dataset contains summaries of tweets spanning a 2-year period and covers more topics than any other public Twitter summarization dataset. Summaries are abstractive in nature and have been created by journalists skilled in summarizing news articles following a template separating factual information (main story) from author opinions. Our method differs from previous work on generating gold-standard summaries from social media, which usually involves selecting representative posts and thus favors extractive summarization models. To showcase the dataset’s utility and challenges, we benchmark a range of abstractive and extractive state-of-the-art summarization models and achieve good performance, with the former outperforming the latter. We also show that fine-tuning is necessary to improve performance and investigate the benefits of using different sample sizes. Iman Munire Bilal, Bo Wang 0034, Adam Tsakalidis, Dong Nguyen 0002, Rob Procter, Maria Liakata |
Trans. Assoc. Comput. Linguistics | 6 |
| 2021 | Evaluation of Thematic Coherence in MicroblogsabstractIman Munire Bilal, Bo Wang, Maria Liakata, Rob Procter, Adam Tsakalidis. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Iman Munire Bilal, Bo Wang 0034, Maria Liakata, Rob Procter, Adam Tsakalidis |
ACL/IJCNLP (1) | 3 |
| 2021 | Boosting Low-Resource Biomedical QA via Entity-Aware Masking StrategiesabstractBiomedical question-answering (QA) has gained increased attention for its capability to provide users with high-quality information from a vast scientific literature.Although an increasing number of biomedical QA datasets has been recently made available, those resources are still rather limited and expensive to produce.Transfer learning via pre-trained language models (LMs) has been shown as a promising approach to leverage existing general-purpose knowledge.However, finetuning these large models can be costly and time consuming, often yielding limited benefits when adapting to specific themes of specialised domains, such as the COVID-19 literature.To bootstrap further their domain adaptation, we propose a simple yet unexplored approach, which we call biomedical entity-aware masking (BEM).We encourage masked language models to learn entity-centric knowledge based on the pivotal entities characterizing the domain at hand, and employ those entities to drive the LM fine-tuning.The resulting strategy is a downstream process applicable to a wide variety of masked LMs, not requiring additional memory or components in the neural architectures.Experimental results show performance on par with state-of-the-art models on several biomedical QA datasets. Gabriele Pergola, Elena Kochkina, Lin Gui 0003, Maria Liakata, Yulan He 0001 |
EACL | 4 |
| 2021 | CD\^2CR: Co-reference resolution across documents and domainsabstractCross-document co-reference resolution (CDCR) is the task of identifying and linking mentions to entities and concepts across many text documents.Current state-of-the-art models for this task assume that all documents are of the same type (e.g.news articles) or fall under the same theme.However, it is also desirable to perform CDCR across different domains (type or theme).A particular use case we focus on in this paper is the resolution of entities mentioned across scientific work and newspaper articles that discuss them.Identifying the same entities and corresponding concepts in both scientific articles and news can help scientists understand how their work is represented in mainstream media.We propose a new task and English language dataset for cross-document cross-domain co-reference resolution (CD 2 CR).The task aims to identify links between entities across heterogeneous document types.We show that in this cross-domain, cross-document setting, existing CDCR models do not perform well and we provide a baseline model that outperforms current state-of-the-art CDCR models on CD 2 CR.Our data set, annotation tool and guidelines as well as our model for cross-document cross-domain co-reference are all supplied as open access open source resources. James Ravenscroft, Amanda Clare, Arie Cattan, Ido Dagan, Maria Liakata |
EACL | 5 |
| 2021 | Modelling Paralinguistic Properties in Conversational Speech to Detect Bipolar Disorder and Borderline Personality DisorderabstractBipolar disorder (BD) and borderline personality disorder (BPD) are two chronic mental health conditions that clinicians find challenging to distinguish based on clinical interviews, due to their overlapping symptoms. In this work, we investigate the automatic detection of these two conditions by modelling both verbal and non-verbal cues in a set of interviews. We propose a new approach of modelling short-term features with visibility-signature transform, and compare it with widely used high-level statistical functions. We demonstrate the superior performance of our proposed signature-based model. Furthermore, we show the role of different sets of features in characterising BD and BPD. Bo Wang 0034, Yue Wu 0016, Nemanja Vaci, Maria Liakata, Terry J. Lyons, Kate Saunders |
ICASSP | 4 |
| 2020 | Estimating predictive uncertainty for rumour verification modelsabstractThe inability to correctly resolve rumours circulating online can have harmful real-world consequences.We present a method for incorporating model and data uncertainty estimates into natural language processing models for automatic rumour verification.We show that these estimates can be used to filter out model predictions likely to be erroneous, so that these difficult instances can be prioritised by a human fact-checker.We propose two methods for uncertainty-based instance rejection, supervised and unsupervised.We also show how uncertainty estimates can be used to interpret model performance as a rumour unfolds. Elena Kochkina, Maria Liakata |
ACL | 2 |
| 2020 | tBERT: Topic Models and BERT Joining Forces for Semantic Similarity DetectionabstractSemantic similarity detection is a fundamental task in natural language understanding.Adding topic information has been useful for previous feature-engineered semantic similarity models as well as neural models for other tasks.There is currently no standard way of combining topics with pretrained contextual representations such as BERT.We propose a novel topic-informed BERT-based architecture for pairwise semantic similarity detection and show that our model improves performance over strong neural baselines across a variety of English language datasets.We find that the addition of topics to BERT helps particularly with resolving domain-specific cases. Nicole Peinelt, Dong Nguyen 0002, Maria Liakata |
ACL | 3 |
| 2020 | Sequential Modelling of the Evolution of Word Representations for Semantic Change DetectionabstractSemantic change detection concerns the task of identifying words whose meaning has changed over time.Current state-of-the-art approaches operating on neural embeddings detect the level of semantic change in a word by comparing its vector representation in two distinct time periods, without considering its evolution through time.In this work, we propose three variants of sequential models for detecting semantically shifted words, effectively accounting for the changes in the word representations over time.Through extensive experimentation under various settings with synthetic and real data we showcase the importance of sequential modelling of word vectors through time for semantic change detection.Finally, we compare different approaches in a quantitative manner, demonstrating that temporal modelling of word representations yields a clear-cut advantage in performance. Adam Tsakalidis, Maria Liakata |
EMNLP (1) | 2 |
| 2020 | Learning to Detect Bipolar Disorder and Borderline Personality Disorder with Language and Speech in Non-Clinical InterviewsabstractBipolar disorder (BD) and borderline personality disorder (BPD) are both chronic psychiatric disorders.However, their overlapping symptoms and common comorbidity make it challenging for the clinicians to distinguish the two conditions on the basis of a clinical interview.In this work, we first present a new multi-modal dataset containing interviews involving individuals with BD or BPD being interviewed about a non-clinical topic .We investigate the automatic detection of the two conditions, and demonstrate a good linear classifier that can be learnt using a down-selected set of features from the different aspects of the interviews and a novel approach of summarising these features.Finally, we find that different sets of features characterise BD and BPD, thus providing insights into the difference between the automatic screening of the two conditions. Bo Wang 0034, Yue Wu 0016, Niall Taylor, Terry J. Lyons, Maria Liakata, Alejo J. Nevado-Holgado, Kate Saunders |
INTERSPEECH | 5 |
| 2019 | Aiming beyond the Obvious: Identifying Non-Obvious Cases in Semantic Similarity DatasetsabstractExisting datasets for scoring text pairs in terms of semantic similarity contain instances whose resolution differs according to the degree of difficulty.This paper proposes to distinguish obvious from non-obvious text pairs based on superficial lexical overlap and ground-truth labels.We characterise existing datasets in terms of containing difficult cases and find that recently proposed models struggle to capture the non-obvious cases of semantic similarity.We describe metrics that emphasise cases of similarity which require more complex inference and propose that these are used for evaluating systems for semantic similarity. Nicole Peinelt, Maria Liakata, Dong Nguyen 0002 |
ACL (1) | 2 |
| 2019 | A Path Signature Approach for Speech Emotion RecognitionabstractAutomatic speech emotion recognition (SER) remains a \ndifficult task within human-computer interaction, despite increasing interest in the research community. One key challenge is how to effectively integrate short-term characterisation \nof speech segments with long-term information such as temporal variations. Motivated by the numerical approximation theory of stochastic differential equations (SDEs), we propose the \nnovel use of path signatures. The latter provide a pathwise definition to solve SDEs, for the integration of short speech frames. \nFurthermore we propose a hierarchical tree structure of path signatures, to capture both global and local information. A simple tree-based convolutional neural network (TBCNN) is used \nfor learning the structural information stemming from dyadic \npath-tree signatures. Our experimental results on a widely \nused benchmark dataset demonstrate comparable performance \nto complex neural network based systems. Bo Wang 0034, Maria Liakata, Hao Ni 0001, Terry J. Lyons, Alejo J. Nevado-Holgado, Kate Saunders |
INTERSPEECH | 2 |
| 2019 | Gaussian Processes for Rumour Stance Classification in Social MediaabstractSocial media tend to be rife with rumours while new reports are released piecemeal during breaking news. Interestingly, one can mine multiple reactions expressed by social media users in those situations, exploring their stance towards rumours, ultimately enabling the flagging of highly disputed rumours as being potentially false. In this work, we set out to develop an automated, supervised classifier that uses multi-task learning to classify the stance expressed in each individual tweet in a conversation around a rumour as either supporting, denying or questioning the rumour. Using a Gaussian Process classifier, and exploring its effectiveness on two datasets with very different characteristics and varying distributions of stances, we show that our approach consistently outperforms competitive baseline classifiers. Our classifier is especially effective in estimating the distribution of different types of stance associated with a given rumour, which we set forth as a desired characteristic for a rumour-tracking system that will show both ordinary users of Twitter and professional news practitioners how others orient to the disputed veracity of a rumour, with the final aim of establishing its actual truth value. Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Arkaitz Zubiaga, Maria Liakata, Rob Procter |
ACM Trans. Inf. Syst. | 5 |
| 2018 | Nowcasting the Stance of Social Media Users in a Sudden Vote: The Case of the Greek ReferendumabstractModelling user voting intention in social media is an important research area, with applications in analysing electorate behaviour, online political campaigning and advertising. Previous approaches mainly focus on predicting national general elections, which are regularly scheduled and where data of past results and opinion polls are available. However, there is no evidence of how such models would perform during a sudden vote under time-constrained circumstances. That poses a more challenging task compared to traditional elections, due to its spontaneous nature. In this paper, we focus on the 2015 Greek bailout referendum, aiming to nowcast on a daily basis the voting intention of 2,197 Twitter users. We propose a semi-supervised multiple convolution kernel learning approach, leveraging temporally sensitive text and network information. Our evaluation under a real-time simulation framework demonstrates the effectiveness and robustness of our approach against competitive baselines, achieving a significant 20% increase in F-score compared to solely text-based models. Adam Tsakalidis, Nikolaos Aletras, Alexandra I. Cristea, Maria Liakata |
CIKM | 4 |
| 2018 | All-in-one: Multi-task Learning for Rumour VerificationabstractAutomatic resolution of rumours is a challenging task that can be broken down into smaller components that make up a pipeline, including rumour detection, rumour tracking and stance classification, leading to the final outcome of determining the veracity of a rumour. In previous work, these steps in the process of rumour verification have been developed as separate components where the output of one feeds into the next. We propose a multi-task learning approach that allows joint training of the main and auxiliary tasks, improving the performance of rumour verification. We examine the connection between the dataset properties and the outcomes of the multi-task learning models used. Elena Kochkina, Maria Liakata, Arkaitz Zubiaga |
COLING | 2 |
| 2018 | Can We Assess Mental Health Through Social Media and Smart Devices? Addressing Bias in Methodology and Evaluation
Adam Tsakalidis, Maria Liakata, Theodoros Damoulas, Alexandra I. Cristea |
ECML/PKDD (3) | 2 |
| 2018 | Discourse-aware rumour stance classification in social media using sequential classifiers
Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Isabelle Augenstein |
Inf. Process. Manag. | 3 |
| 2018 | Using clinical Natural Language Processing for health outcomes research: Overview and actionable suggestions for future advancesabstractThe importance of incorporating Natural Language Processing (NLP) methods in clinical informatics research has been increasingly recognized over the past years, and has led to transformative advances. Typically, clinical NLP systems are developed and evaluated on word, sentence, or document level annotations that model specific attributes and features, such as document content (e.g., patient status, or report type), document section types (e.g., current medications, past medical history, or discharge summary), named entities and concepts (e.g., diagnoses, symptoms, or treatments) or semantic attributes (e.g., negation, severity, or temporality). From a clinical perspective, on the other hand, research studies are typically modelled and evaluated on a patient- or population-level, such as predicting how a patient group might respond to specific treatments or patient monitoring over time. While some NLP tasks consider predictions at the individual or group user level, these tasks still constitute a minority. Owing to the discrepancy between scientific objectives of each field, and because of differences in methodological evaluation priorities, there is no clear alignment between these evaluation approaches. Here we provide a broad summary and outline of the challenging issues involved in defining appropriate intrinsic and extrinsic evaluation methods for NLP research that is to be used for clinical outcomes research, and vice versa. A particular focus is placed on mental health research, an area still relatively understudied by the clinical NLP research community, but where NLP methods are of notable relevance. Recent advances in clinical NLP method development have been significant, but we propose more emphasis needs to be placed on rigorous evaluation for the field to advance further. To enable this, we provide actionable suggestions, including a minimal protocol that could be used when reporting clinical NLP method development and its evaluation. Sumithra Velupillai, Hanna Suominen, Maria Liakata, Angus Roberts, Anoop D. Shah, Katherine Morley, David Osborn, Joseph Hayes, Robert Stewart 0002, Johnny Downs, Wendy W. Chapman, Rina Dutta |
J. Biomed. Informatics | 3 |
| 2017 | Supporting the Use of User Generated Content in Journalistic PracticeabstractSocial media and user-generated content (UGC) are increasingly important features of journalistic work in a number of different ways. However, their use presents major challenges, not least because information posted on social media is not always reliable and therefore its veracity needs to be checked before it can be considered as fit for use in the reporting of news. We report on the results of a series of in-depth ethnographic studies of journalist work practices undertaken as part of the requirements gathering for a prototype of a social media verification 'dashboard' and its subsequent evaluation. We conclude with some reflections upon the broader implications of our findings for the design of tools to support journalistic work. Peter Tolmie, Rob Procter, Dave W. Randall 0001, Mark Rouncefield, Christian Burger, Geraldine Wong Sak Hoi, Arkaitz Zubiaga, Maria Liakata |
CHI | 8 |
| 2017 | TDParse: Multi-target-specific sentiment recognition on TwitterabstractExisting target-specific sentiment recognition methods consider only a single target per tweet, and have been shown to miss nearly half of the actual targets mentioned.We present a corpus of UK election tweets, with an average of 3.09 entities per tweet and more than one type of sentiment in half of the tweets.This requires a method for multi-target specific sentiment recognition, which we develop by using the context around a target as well as syntactic dependencies involving the target.We present results of our method on both a benchmark corpus of single targets and the multi-target election corpus, showing state-of-the art performance in both corpora and outperforming previous approaches to multi-target sentiment task as well as deep learning models for singletarget sentiment. Bo Wang 0034, Maria Liakata, Arkaitz Zubiaga, Rob Procter |
EACL (1) | 2 |
| 2017 | Towards Real-Time, Country-Level Location Classification of Worldwide TweetsabstractThe increase of interest in using social media as a source for research has motivated tackling the challenge of automatically geolocating tweets, given the lack of explicit location information in the majority of tweets. In contrast to much previous work that has focused on location classification of tweets restricted to a specific country, here we undertake the task in a broader context by classifying global tweets at the country level, which is so far unexplored in a real-time scenario. We analyze the extent to which a tweet's country of origin can be determined by making use of eight tweet-inherent features for classification. Furthermore, we use two datasets, collected a year apart from each other, to analyze the extent to which a model trained from historical tweets can still be leveraged for classification of new tweets. With classification experiments on all 217 countries in our datasets, as well as on the top 25 countries, we offer some insights into the best use of tweet-inherent features for an accurate country-level classification of tweets. We find that the use of a single feature, such as the use of tweet content alone-the most widely used feature in previous work-leaves much to be desired. Choosing an appropriate combination of both tweet content and metadata can actually lead to substantial improvements of between 20 and 50 percent. We observe that tweet content, the user's self-reported location and the user's real name, all of which are inherent in a tweet and available in a real-time scenario, are particularly useful to determine the country of origin. We also experiment on the applicability of a model trained on historical tweets to classify new tweets, finding that the choice of a particular combination of features whose utility does not fade over time can actually lead to comparable performance, avoiding the need to retrain. However, the difficulty of achieving accurate classification increases slightly for countries with multiple commonalities, especially for English and Spanish speaking countries. Arkaitz Zubiaga, Alexander Voß, Rob Procter, Maria Liakata, Bo Wang 0034, Adam Tsakalidis |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | SentiHood: Targeted Aspect Based Sentiment Analysis Dataset for Urban NeighbourhoodsabstractIn this paper, we introduce the task of targeted aspect-based sentiment analysis. The goal is to extract fine-grained information with respect to entities mentioned in user comments. This work extends both aspect-based sentiment analysis – that assumes a single entity per document — and targeted sentiment analysis — that assumes a single sentiment towards a target entity. In particular, we identify the sentiment towards each aspect of one or more entities. As a testbed for this task, we introduce the SentiHood dataset, extracted from a question answering (QA) platform where urban neighbourhoods are discussed by users. In this context units of text often mention several aspects of one or more neighbourhoods. This is the first time that a generic social media platform,i.e. QA, is used for fine-grained opinion mining. Text coming from QA platforms are far less constrained compared to text from review specific platforms which current datasets are based on. We develop several strong baselines, relying on logistic regression and state-of-the-art recurrent neural networks Marzieh Saeidi, Guillaume Bouchard, Maria Liakata, Sebastian Riedel 0001 |
COLING | 3 |
| 2016 | Combining Heterogeneous User Generated Data to Sense Well-beingabstractIn this paper we address a new problem of predicting affect and well-being scales in a real-world setting of heterogeneous, longitudinal and non-synchronous textual as well as non-linguistic data that can be harvested from on-line media and mobile phones. We describe the method for collecting the heterogeneous longitudinal data, how features are extracted to address missing information and differences in temporal alignment, and how the latter are combined to yield promising predictions of affect and well-being on the basis of widely used psychological scales. We achieve a coefficient of determination (R^2) of 0.71-0.76 and a correlation coefficient of 0.68-0.87 which is higher than the state-of-the art in equivalent multi-modal tasks for affect. Adam Tsakalidis, Maria Liakata, Theodoros Damoulas, Brigitte Jellinek, Weisi Guo, Alexandra I. Cristea |
COLING | 2 |
| 2016 | Stance Classification in Rumours as a Sequential Task Exploiting the Tree Structure of Social Media ConversationsabstractRumour stance classification, the task that determines if each tweet in a collection discussing a rumour is supporting, denying, questioning or simply commenting on the rumour, has been attracting substantial interest. Here we introduce a novel approach that makes use of the sequence of transitions observed in tree-structured conversation threads in Twitter. The conversation threads are formed by harvesting users’ replies to one another, which results in a nested tree-like structure. Previous work addressing the stance classification task has treated each tweet as a separate unit. Here we analyse tweets by virtue of their position in a sequence and test two sequential classifiers, Linear-Chain CRF and Tree CRF, each of which makes different assumptions about the conversational structure. We experiment with eight Twitter datasets, collected during breaking news, and show that exploiting the sequential structure of Twitter conversations achieves significant improvements over the non-sequential methods. Our work is the first to model Twitter conversations as a tree structure in this manner, introducing a novel way of tackling NLP tasks on Twitter conversations. Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik |
COLING | 3 |
| 2016 | Applying Core Scientific Concepts to Context-Based Citation Recommendation
Daniel Duma, Maria Liakata, Amanda Clare, James Ravenscroft, Ewan Klein |
LREC | 2 |
| 2016 | Multi-label Annotation in Scientific Articles - The Multi-label Cancer Risk Assessment Corpus
James Ravenscroft, Anika Oellrich, Shyamasree Saha, Maria Liakata |
LREC | 4 |
| 2016 | CRUDE: Combining Resource Usage Data and Error Logs for Accurate Error Detection in Large-Scale Distributed SystemsabstractThe use of console logs for error detection in large scale distributed systems has proven to be useful to system administrators. However, such logs are typically redundant and incomplete, making accurate detection very difficult. In an attempt to increase this accuracy, we complement these incomplete console logs with resource usage data, which captures the resource utilisation of every job in the system. We then develop a novel error detection methodology, the CRUDE approach, that makes use of both the resource usage data and console logs. We thus make the following specific technical contributions: we develop (i) a clustering algorithm to group nodes with similar behaviour, (ii) an anomaly detection algorithm to identify jobs with anomalous resource usage, (iii) an algorithm that links jobs with anomalous resource usage with erroneous nodes. We then evaluate our approach using console logs and resource usage data from the Ranger Supercomputer. Our results are positive: (i) our approach detects errors with a true positive rate of about 80%, and (ii) when compared with the well-known Nodeinfo error detection algorithm, our algorithm provides an average improvement of around 85% over Nodeinfo, with a best-case improvement of 250%. Nentawe Gurumdimma, Arshad Jhumka, Maria Liakata, Edward Chuah, James C. Browne |
SRDS | 3 |
| 2014 | Biological network extraction from scientific literature: state of the art and challengesabstractNetworks of molecular interactions explain complex biological processes, and all known information on molecular events is contained in a number of public repositories including the scientific literature. Metabolic and signalling pathways are often viewed separately, even though both types are composed of interactions involving proteins and other chemical entities. It is necessary to be able to combine data from all available resources to judge the functionality, complexity and completeness of any given network overall, but especially the full integration of relevant information from the scientific literature is still an ongoing and complex task. Currently, the text-mining research community is steadily moving towards processing the full body of the scientific literature by making use of rich linguistic features such as full text parsing, to extract biological interactions. The next step will be to combine these with information from scientific databases to support hypothesis generation for the discovery of new knowledge and the extension of biological networks. The generation of comprehensive networks requires technologies such as entity grounding, coordination resolution and co-reference resolution, which are not fully solved and are required to further improve the quality of results. Here, we analyse the state of the art for the extraction of network information from the scientific literature and the evaluation of extraction methods against reference corpora, discuss challenges involved and identify directions for future research. Chen Li 0011, Maria Liakata, Dietrich Rebholz-Schuhmann |
Briefings Bioinform. | 2 |
| 2013 | A Discourse-Driven Content Model for Summarising Scientific Articles Evaluated in a Complex Question Answering TaskabstractWe present a method which exploits automatically generated scientific discourse annotations to create a content model for the summarisation of scientific articles.Full papers are first automatically annotated using the CoreSC scheme, which captures 11 contentbased concepts such as Hypothesis, Result, Conclusion etc at the sentence level.A content model which follows the sequence of CoreSC categories observed in abstracts is used to provide the skeleton of the summary, making a distinction between dependent and independent categories.Summary creation is also guided by the distribution of CoreSC categories found in the full articles, in order to adequately represent the article content.Finally, we demonstrate the usefulness of the summaries by evaluating them in a complex question answering task.Results are very encouraging as summaries of papers from automatically obtained CoreSCs enable experts to answer 66% of complex content-related questions designed on the basis of paper abstracts.The questions were answered with a precision of 75%, where the upper bound for human summaries (abstracts) was 95%. Maria Liakata, Simon Dobnik, Shyamasree Saha, Colin R. Batchelor, Dietrich Rebholz-Schuhmann |
EMNLP | 1 |
| 2012 | Automatic recognition of conceptualization zones in scientific articles and two life science applicationsabstractMOTIVATION: Scholarly biomedical publications report on the findings of a research investigation. Scientists use a well-established discourse structure to relate their work to the state of the art, express their own motivation and hypotheses and report on their methods, results and conclusions. In previous work, we have proposed ways to explicitly annotate the structure of scientific investigations in scholarly publications. Here we present the means to facilitate automatic access to the scientific discourse of articles by automating the recognition of 11 categories at the sentence level, which we call Core Scientific Concepts (CoreSCs). These include: Hypothesis, Motivation, Goal, Object, Background, Method, Experiment, Model, Observation, Result and Conclusion. CoreSCs provide the structure and context to all statements and relations within an article and their automatic recognition can greatly facilitate biomedical information extraction by characterizing the different types of facts, hypotheses and evidence available in a scientific publication. RESULTS: We have trained and compared machine learning classifiers (support vector machines and conditional random fields) on a corpus of 265 full articles in biochemistry and chemistry to automatically recognize CoreSCs. We have evaluated our automatic classifications against a manually annotated gold standard, and have achieved promising accuracies with 'Experiment', 'Background' and 'Model' being the categories with the highest F1-scores (76%, 62% and 53%, respectively). We have analysed the task of CoreSC annotation both from a sentence classification as well as sequence labelling perspective and we present a detailed feature evaluation. The most discriminative features are local sentence features such as unigrams, bigrams and grammatical dependencies while features encoding the document structure, such as section headings, also play an important role for some of the categories. We discuss the usefulness of automatically generated CoreSCs in two biomedical applications as well as work in progress. AVAILABILITY: A web-based tool for the automatic annotation of articles with CoreSCs and corresponding documentation is available online at http://www.sapientaproject.com/software http://www.sapientaproject.com also contains detailed information pertaining to CoreSC annotation and links to annotation guidelines as well as a corpus of manually annotated articles, which served as our training data. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Maria Liakata, Shyamasree Saha, Simon Dobnik, Colin R. Batchelor, Dietrich Rebholz-Schuhmann |
Bioinform. | 1 |
| 2011 | A comparison and user-based evaluation of models of textual information structure in the context of cancer risk assessmentabstractBACKGROUND: Many practical tasks in biomedicine require accessing specific types of information in scientific literature; e.g. information about the results or conclusions of the study in question. Several schemes have been developed to characterize such information in scientific journal articles. For example, a simple section-based scheme assigns individual sentences in abstracts under sections such as Objective, Methods, Results and Conclusions. Some schemes of textual information structure have proved useful for biomedical text mining (BIO-TM) tasks (e.g. automatic summarization). However, user-centered evaluation in the context of real-life tasks has been lacking. METHODS: We take three schemes of different type and granularity--those based on section names, Argumentative Zones (AZ) and Core Scientific Concepts (CoreSC)--and evaluate their usefulness for a real-life task which focuses on biomedical abstracts: Cancer Risk Assessment (CRA). We annotate a corpus of CRA abstracts according to each scheme, develop classifiers for automatic identification of the schemes in abstracts, and evaluate both the manual and automatic classifications directly as well as in the context of CRA. RESULTS: Our results show that for each scheme, the majority of categories appear in abstracts, although two of the schemes (AZ and CoreSC) were developed originally for full journal articles. All the schemes can be identified in abstracts relatively reliably using machine learning. Moreover, when cancer risk assessors are presented with scheme annotated abstracts, they find relevant information significantly faster than when presented with unannotated abstracts, even when the annotations are produced using an automatic classifier. Interestingly, in this user-based evaluation the coarse-grained scheme based on section names proved nearly as useful for CRA as the finest-grained CoreSC scheme. CONCLUSIONS: We have shown that existing schemes aimed at capturing information structure of scientific documents can be applied to biomedical abstracts and can be identified in them automatically with an accuracy which is high enough to benefit a real-life task in biomedicine. Anna Korhonen, Maria Liakata, Ilona Silins, Johan Högberg, Ulla Stenius |
BMC Bioinform. | 3 |
| 2010 | Corpora for the Conceptualisation and Zoning of Scientific Papers
Maria Liakata, Simone Teufel, Advaith Siddharthan, Colin R. Batchelor |
LREC | 1 |
| 2005 | A Two-level Morphology of Malagasy
Mary Dalrymple, Maria Liakata, Lisa Mackie |
PACLIC | 2 |
| 2004 | Learning theories from text
Maria Liakata, Stephen G. Pulman |
COLING | 1 |
| 2002 | From Trees to Predicate-argument Structures
Maria Liakata, Stephen G. Pulman |
COLING | 1 |
| 2000 | Named Entity Recognition in Greek Texts
Iason Demiros, Sotiris Boutsis, Voula Giouli, Maria Liakata, Harris Papageorgiou, Stelios Piperidis |
LREC | 4 |