Beatrice Alex

dblp:17/2942 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-7279-1476ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 52% Trustworthy machine learning · 47% Speech recognition and synthesis · 2%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 100%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 9 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › fairness
bias identification
0.912025
Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals · CHI 2025
Machine learning › Trustworthy machine learning
fairness and bias
0.912025
Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals · CHI 2025
Medical and health informatics
clinical text processing
0.912025
Adverse Event Extraction from Discharge Summaries: A New Dataset, Annotation Scheme, and Initial Findings · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › text classification
multi-label text classification
0.512021
CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
text classification
0.512021
CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification · EMNLP (1) 2021
Medical and health informatics › clinical informatics
clinical coding
0.112021
CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification · EMNLP (1) 2021
Compilers and program optimization
parsing
0.112007
Using Foreign Inclusion Detection to Improve Parsing Performance · EMNLP-CoNLL 2007
Natural language and speech › Speech recognition and synthesis › speech analysis
language identification
0.112005
An Unsupervised System for Identifying English Inclusions in German Text · ACL 2005
Programming languages and type systems › syntax
grammar
0.012007
Using Foreign Inclusion Detection to Improve Parsing Performance · EMNLP-CoNLL 2007

Methods — techniques the papers use, named apart from their topics

workshop · 1.7machine learning · 1.7dataset construction · 1.7annotation scheme · 1.7hierarchical metrics · 1.0depth-based ontology representation · 1.0mixed-methods · 0.9mixed methods · 0.9unsupervised learning · 0.1n-gram analysis · 0.1foreign inclusion detection · 0.1
YearPublicationVenuePosition
2025 Adverse Event Extraction from Discharge Summaries: A New Dataset, Annotation Scheme, and Initial Findings
abstract
Imane Guellil, Salomé Andres, Atul Anand, Bruce Guthrie, Huayu Zhang, Abul Hasan, Honghan Wu, Beatrice Alex. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Imene Guellil, Salomé Andres, Atul Anand, Bruce Guthrie, Abul Hasan, Honghan Wu, Beatrice Alex
ACL (1)8
2025 Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals
abstract
Despite numerous efforts to mitigate their biases, ML systems continue to harm already-marginalized people. While predominant ML approaches assume bias can be removed and fair models can be created, we show that these are not always possible, nor desirable, goals. We reframe the problem of ML bias by creating models to identify biased language, drawing attention to a dataset's biases rather than trying to remove them. Then, through a workshop, we evaluated the models for a specific use case: workflows of information and heritage professionals. Our findings demonstrate the limitations of ML for identifying bias due to its contextual nature, the way in which approaches to mitigating it can simultaneously privilege and oppress different communities, and its inevitability. We demonstrate the need to expand ML approaches to bias and fairness, providing a mixed-methods approach to investigating the feasibility of removing bias or achieving fairness in a given ML use case.
Lucy Havens, Benjamin Bach, Melissa Terras, Beatrice Alex
CHI4
2025 Perceptions of Edinburgh: Capturing neighbourhood characteristics by clustering geoparsed local news
abstract
The communities that we live in affect our health in ways that are complex and hard to define. Moreover, our understanding of the place-based processes affecting health and inequalities is limited. This undermines the development of robust policy interventions to improve local health and well-being. News media provides social and community information that may be useful in health studies. Here we propose a methodology for characterising neighbourhoods by using local news articles. More specifically, we show how we can use Natural Language Processing (NLP) to unlock further information about neighbourhoods by analysing, geoparsing and clustering news articles. Our work is novel because we combine street-level geoparsing tailored to the locality with clustering of full news articles, enabling a more detailed examination of neighbourhood characteristics. We evaluate our outputs and show via a confluence of evidence, both from a qualitative and a quantitative perspective, that the themes we extract from news articles are sensible and reflect many characteristics of the real world. This is significant because it allows us to better understand the effects of neighbourhoods on health. Our findings on neighbourhood characterisation using news data will support a new generation of place-based research which examines a wider set of spatial processes and how they affect health, enabling new epidemiological research. • Novel methodology for characterising neighbourhoods based on local news. • Natural Language Processing can create meaningful neighbourhood characterisations from local news articles. • Extensive quantitative and qualitative evaluation show analysis is sound.
Andreas Grivas, Claire Grover, Richard Tobin, Clare Llewellyn, Eleojo Oluwaseun Abubakar, Chunyu Zheng, Chris Dibben, David J. Pearce 0001, Beatrice Alex
Inf. Process. Manag.10
2024 Enhancing Natural Language Processing Capabilities in Geriatric Patient Care: An Annotation Scheme and Guidelines
Imene Guellil, Salomé Andres, Bruce Guthrie, Atul Anand, Abul Kalam Hasan, Honghan Wu, Beatrice Alex
NLDB (2)8
2024 Can GPT-3.5 generate and code discharge summaries?
abstract
OBJECTIVES: The aim of this study was to investigate GPT-3.5 in generating and coding medical documents with International Classification of Diseases (ICD)-10 codes for data augmentation on low-resource labels. MATERIALS AND METHODS: Employing GPT-3.5 we generated and coded 9606 discharge summaries based on lists of ICD-10 code descriptions of patients with infrequent (or generation) codes within the MIMIC-IV dataset. Combined with the baseline training set, this formed an augmented training set. Neural coding models were trained on baseline and augmented data and evaluated on an MIMIC-IV test set. We report micro- and macro-F1 scores on the full codeset, generation codes, and their families. Weak Hierarchical Confusion Matrices determined within-family and outside-of-family coding errors in the latter codesets. The coding performance of GPT-3.5 was evaluated on prompt-guided self-generated data and real MIMIC-IV data. Clinicians evaluated the clinical acceptability of the generated documents. RESULTS: Data augmentation results in slightly lower overall model performance but improves performance for the generation candidate codes and their families, including 1 absent from the baseline training data. Augmented models display lower out-of-family error rates. GPT-3.5 identifies ICD-10 codes by their prompted descriptions but underperforms on real data. Evaluators highlight the correctness of generated concepts while suffering in variety, supporting information, and narrative. DISCUSSION AND CONCLUSION: While GPT-3.5 alone given our prompt setting is unsuitable for ICD-10 coding, it supports data augmentation for training neural models. Augmentation positively affects generation code families but mainly benefits codes with existing examples. Augmentation reduces out-of-family errors. Documents generated by GPT-3.5 state prompted concepts correctly but lack variety, and authenticity in narratives.
Matús Falis, Aryo Pradipta Gema, Hang Dong 0002, Luke Daines, Siddharth Basetti, Michael Holder, Rose S. Penfold, Alexandra Birch, Beatrice Alex
J. Am. Medical Informatics Assoc.9
2021 Extending defoe for the Efficient Analysis of Historical Texts at Scale
abstract
This paper presents the new facilities provided in defoe, a parallel toolbox for querying a wealth of digitised newspapers and books at scale. defoe has been extended to work with further Natural Language Processing () tools such as the Edinburgh Geoparser, to store the preprocessed text in several storage facilities and to support different types of queries and analyses. We have also extended the collection of XML schemas supported by defoe, increasing the versatility of the tool for the analysis of digital historical textual data at scale. Finally, we have conducted several studies in which we worked with humanities and social science researchers who posed complex and interested questions to large-scale digital collections. Results shows that defoe allows researchers to conduct their studies and obtain results faster, while all the large-scale text mining complexity is automatically handled by defoe.
Rosa Filgueira, Claire Grover, Vasilios Karaiskos, Beatrice Alex, Sarah Van Eyndhoven, Lisa Gotthard, Melissa Terras
e-Science4
2021 CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification
abstract
Large-Scale Multi-Label Text Classification (LMTC) includes tasks with hierarchical label spaces, such as automatic assignment of ICD-9 codes to discharge summaries.Performance of models in prior art is evaluated with standard precision, recall, and F 1 measures without regard for the rich hierarchical structure.In this work we argue for hierarchical evaluation of the predictions of neural LMTC models.With the example of the ICD-9 ontology we describe a structural issue in the representation of the structured label space in prior art, and propose an alternative representation based on the depth of the ontology.We propose a set of metrics for hierarchical evaluation using the depthbased representation.We compare the evaluation scores from the proposed metrics with previously used metrics on prior art LMTC models for ICD-9 coding in MIMIC-III.We also propose further avenues of research involving the proposed ontological representation.
Matús Falis, Hang Dong 0002, Alexandra Birch, Beatrice Alex
EMNLP (1)4
2021 Comparison of rule-based and neural network models for negation detection in radiology reports
abstract
Abstract Using natural language processing, it is possible to extract structured information from raw text in the electronic health record (EHR) at reasonably high accuracy. However, the accurate distinction between negated and non-negated mentions of clinical terms remains a challenge. EHR text includes cases where diseases are stated not to be present or only hypothesised, meaning a disease can be mentioned in a report when it is not being reported as present. This makes tasks such as document classification and summarisation more difficult. We have developed the rule-based EdIE-R-Neg, part of an existing text mining pipeline called EdIE-R (Edinburgh Information Extraction for Radiology reports), developed to process brain imaging reports, ( https://www.ltg.ed.ac.uk/software/edie-r/ ) and two machine learning approaches; one using a bidirectional long short-term memory network and another using a feedforward neural network. These were developed on data from the Edinburgh Stroke Study (ESS) and tested on data from routine reports from NHS Tayside (Tayside). Both datasets consist of written reports from medical scans. These models are compared with two existing rule-based models: pyConText (Harkema et al. 2009. Journal of Biomedical Informatics42(5), 839–851), a python implementation of a generalisation of NegEx, and NegBio (Peng et al. 2017. NegBio: A high-performance tool for negation and uncertainty detection in radiology reports. arXiv e-prints, p. arXiv:1712.05898 ), which identifies negation scopes through patterns applied to a syntactic representation of the sentence. On both the test set of the dataset from which our models were developed, as well as the largely similar Tayside test set, the neural network models and our custom-built rule-based system outperformed the existing methods. EdIE-R-Neg scored highest on F1 score, particularly on the test set of the Tayside dataset, from which no development data were used in these experiments, showing the power of custom-built rule-based systems for negation detection on datasets of this size. The performance gap of the machine learning models to EdIE-R-Neg on the Tayside test set was reduced through adding development Tayside data into the ESS training set, demonstrating the adaptability of the neural network models.
D. Sykes, Andreas Grivas, Claire Grover, Richard Tobin, Cathie Sudlow, William Whiteley, Andrew M. McIntosh, Heather Whalley, Beatrice Alex
Nat. Lang. Eng.9
2019 Detecting Topic-Oriented Speaker Stance in Conversational Speech
abstract
Being able to detect topics and speaker stances in conversations is a key requirement for developing spoken language understanding systems that are personalized and adaptive. In this work, we explore how topic-oriented speaker stance is expressed in conversational speech. To do this, we present a new set of topic and stance annotations of the CallHome corpus of spontaneous dialogues. Specifically, we focus on six stances-positivity, certainty, surprise, amusement, interest, and comfort-which are useful for characterizing important aspects of a conversation, such as whether a conversation is going well or not. Based on this, we investigate the use of neural network models for automatically detecting speaker stance from speech in multi-turn, multi-speaker contexts. In particular, we examine how performance changes depending on how input feature representations are constructed and how this is related to dialogue structure. Our experiments show that incorporating both lexical and acoustic features is beneficial for stance detection. However, we observe variation in whether using hierarchical models for encoding lexical and acoustic information improves performance, suggesting that some aspects of speaker stance are expressed more locally than others. Overall, our findings highlight the importance of modelling interaction dynamics and non-lexical content for stance detection.
Catherine Lai, Beatrice Alex, Johanna D. Moore, Leimin Tian, Tatsuro Hori, Gianpiero Francesca
INTERSPEECH2
2016 Homing in on Twitter Users: Evaluating an Enhanced Geoparser for User Profile Locations
Beatrice Alex, Clare Llewellyn, Claire Grover, Jon Oberlander, Richard Tobin
LREC1
2015 Extracting a Topic Specific Dataset from a Twitter Archive
Clare Llewellyn, Claire Grover, Beatrice Alex, Jon Oberlander, Richard Tobin
TPDL3
2008 Comparing Corpus-based to Web-based Lookup Techniques for Automatic English Inclusion Detection
Beatrice Alex
LREC1
2008 Exploiting Multiply Annotated Corpora in Biomedical Information Extraction Tasks
Barry Haddow, Beatrice Alex
LREC2
2007 Using Foreign Inclusion Detection to Improve Parsing Performance
Beatrice Alex, Amit Dubey, Frank Keller
EMNLP-CoNLL1
2006 The Impact of Annotation on the Performance of Protein Tagging in Biomedical Text
Beatrice Alex, Malvina Nissim, Claire Grover
LREC1
2005 An Unsupervised System for Identifying English Inclusions in German Text
Beatrice Alex
ACL1
2005 Investigating the Effects of Selective Sampling on the Annotation Task
Ben Hachey, Beatrice Alex, Markus Becker 0002
CoNLL2
2005 Exploring the boundaries: gene and protein identification in biomedical text
abstract
BACKGROUND: Good automatic information extraction tools offer hope for automatic processing of the exploding biomedical literature, and successful named entity recognition is a key component for such tools. METHODS: We present a maximum-entropy based system incorporating a diverse set of features for identifying gene and protein names in biomedical abstracts. RESULTS: This system was entered in the BioCreative comparative evaluation and achieved a precision of 0.83 and recall of 0.84 in the "open" evaluation and a precision of 0.78 and recall of 0.85 in the "closed" evaluation. CONCLUSION: Central contributions are rich use of features derived from the training data at multiple levels of granularity, a focus on correctly identifying entity boundaries, and the innovative use of several external knowledge sources including full MEDLINE abstracts and web searches.
Jenny Rose Finkel, Shipra Dingare, Christopher D. Manning, Malvina Nissim, Beatrice Alex, Claire Grover
BMC Bioinform.5