Paul Rayson

dblp:96/5736 · also Paul Edward Rayson · DBLP profile ↗
← Back
64ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0002-1257-2191ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 17 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021Software engineering, systems software and programming languages · 7Computer networks · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Structured Prompting for Arabic Essay Proficiency: A Trait-Centric Evaluation Approach
Salim Al Mandhari, Hieu Pham Dinh, Mo El-Haj, Paul Rayson
LREC4
2026 Creating a Hybrid Rule and Neural Network Based Semantic Tagger Using Silver Standard Data: The PyMUSAS Framework for Multilingual Semantic Annotation
Andrew Moore 0001, Paul Rayson, Dawn Archer, Tim Czerniak, Dawn Knight, Daisy Monika Lal, Gearóid Ó Donnchadha, Mícheál J. Ó Meachair, Scott Piao, Elaine Uí Dhonnchadha, Johanna Vuorinen, Yan Yabo, Xiaobin Yang
LREC2
2026 FreeTxt-Vi: A Benchmarked Vietnamese-English Toolkit for Segmentation, Sentiment, and Summarisation
abstract
FreeTxt-Vi is a free and open-source web-based toolkit for creating and analysing bilingual Vietnamese-English text collections.Positioned at the intersection of corpus linguistics and natural language processing (NLP), it enables users to build, explore, and interpret free-text data without requiring programming expertise.The system combines established corpus analysis features such as concordancing, keyword analysis, word relation exploration, and interactive visualisation with modern transformer-based NLP components for sentiment analysis and summarisation.A key contribution of this work is the design of a unified bilingual NLP pipeline that integrates a hybrid VnCoreNLP + Byte Pair Encoding (BPE) segmentation strategy, a fine-tuned TabularisAI sentiment classifier, and a fine-tuned Qwen2.5 model for abstractive summarisation.Unlike existing text analysis platforms, FreeTxt-Vi is evaluated as a set of language processing components.We conduct a three-part evaluation covering segmentation, sentiment analysis, and summarisation, and demonstrate that our approach achieves competitive or superior performance compared to widely used baselines in both Vietnamese and English.By reducing technical barriers to multilingual text analysis, FreeTxt-Vi supports reproducible research and promotes the development of language resources for Vietnamese, a widely spoken but underrepresented language in NLP.The toolkit is applicable to a wide range of domains, including education, digital humanities, cultural heritage, and the social sciences, where qualitative text data are common but often difficult to process at scale.
Hung Huy Nguyen, Mo El-Haj, Paul Rayson, Dawn Knight
LREC3
2025 Sinhala Encoder-only Language Models and Evaluation
abstract
Recently, language models (LMs) have produced excellent results in many natural language processing (NLP) tasks. However, their effectiveness is highly dependent on available pre-training resources, which is particularly challenging for low-resource languages such as Sinhala. Furthermore, the scarcity of benchmarks to evaluate LMs is also a major concern for low-resource languages. In this paper, we address these two challenges for Sinhala by (i) collecting the largest monolingual corpus for Sinhala, (ii) training multiple LMs on this corpus and (iii) compiling the first Sinhala NLP benchmark (Sinhala-GLUE) and evaluating LMs on it. We show the Sinhala LMs trained in this paper outperform the popular multilingual LMs, such as XLM-R and existing Sinhala LMs in downstream NLP tasks. All the trained LMs are publicly available. We also make Sinhala-GLUE publicly available as a public leaderboard, and we hope that it will enable further advancements in developing and evaluating LMs for Sinhala.
Tharindu Ranasinghe, Hansi Hettiarachchi, Nadeesha Pathirana, Damith Premasiri, Lasitha Uyangodage, Isuri Anuradha, Alistair Plum, Paul Rayson, Ruslan Mitkov
ACL (1)8
2025 AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
Mo El-Haj, Paul Rayson
IEEE Big Data2
2024 The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment
abstract
The Igbo language is facing a risk of becoming endangered, as indicated by a 2025 UNESCO study. This highlights the need to develop language technologies for Igbo to foster communication, learning and preservation. To create robust, impactful, and widely adopted language technologies for Igbo, it is essential to incorporate the multi-dialectal nature of the language. The primary obstacle in achieving dialectal-aware language technologies is the lack of comprehensive dialectal datasets. In response, we present the IgboAPI dataset, a multi-dialectal Igbo-English dictionary dataset, developed with the aim of enhancing the representation of Igbo dialects. Furthermore, we illustrate the practicality of the IgboAPI dataset through two distinct studies: one focusing on Igbo semantic lexicon and the other on machine translation. In the semantic lexicon project, we successfully establish an initial Igbo semantic lexicon for the Igbo semantic tagger, while in the machine translation study, we demonstrate that by finetuning existing machine translation systems using the IgboAPI dataset, we significantly improve their ability to handle dialectal variations in sentences.
Chris C. Emezue, Ifeoma Okoh, Chinedu E. Mbonu, Chiamaka Ijeoma Chukwuneke, Daisy Monika Lal, Ignatius Ezeani, Paul Rayson, Ijemma Onwuzulike, Chukwuma Onyebuchi Okeke, Gerald Okey Nweya, Bright Ikechukwu Ogbonna, Chukwuebuka Uchenna Oraegbunam, Chidinma A. Nwafor, Akudo Amarachukwu Osuagwu
LREC/COLING7
2024 Annotator Disagreement-Based Analysis for Developing Bias Benchmark Datasets in Resource-Restricted Settings
Vithya Yogarajan, Paul Rayson, Gillian Dobbie, Aaron Keesing, Te Taka Keegan, Diana Benavides-Prado, Michael Witbrock
ICONIP (10)2
2023 Open-Source Thesaurus Development for Under-Resourced Languages: a Welsh Case Study
Nouran Khallaf, Elin Arfon, Mo El-Haj, Jonathan Morris, Dawn Knight, Paul Rayson, Tymaa Hammouda, Mustafa Jarrar
LDK6
2023 FinAraT5: A text to text model for financial Arabic text understanding and generation
Nadhem Zmandar, Mo El-Haj, Paul Rayson
LDK3
2023 A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers
Nadhem Zmandar, Mahmoud El-Haj, Paul Rayson
NLDB3
2023 UNLT: Urdu Natural Language Toolkit
abstract
Abstract This study describes a Natural Language Processing (NLP) toolkit, as the first contribution of a larger project, for an under-resourced language—Urdu. In previous studies, standard NLP toolkits have been developed for English and many other languages. There is also a dire need for standard text processing tools and methods for Urdu, despite it being widely spoken in different parts of the world with a large amount of digital text being readily available. This study presents the first version of the UNLT (Urdu Natural Language Toolkit) which contains three key text processing tools required for an Urdu NLP pipeline; word tokenizer, sentence tokenizer, and part-of-speech (POS) tagger. The UNLT word tokenizer employs a morpheme matching algorithm coupled with a state-of-the-art stochastic n -gram language model with back-off and smoothing characteristics for the space omission problem. The space insertion problem for compound words is tackled using a dictionary look-up technique. The UNLT sentence tokenizer is a combination of various machine learning, rule-based, regular-expressions, and dictionary look-up techniques. Finally, the UNLT POS taggers are based on Hidden Markov Model and Maximum Entropy-based stochastic techniques. In addition, we have developed large gold standard training and testing data sets to improve and evaluate the performance of new techniques for Urdu word tokenization, sentence tokenization, and POS tagging. For comparison purposes, we have compared the proposed approaches with several methods. Our proposed UNLT, the training and testing data sets, and supporting resources are all free and publicly available for academic use.
Jawad Shafi, Hafiz Rizwan Iqbal, Rao Muhammad Adeel Nawab, Paul Rayson
Nat. Lang. Eng.4
2023 Semantic Tagging for the Urdu Language: Annotated Corpus and Multi-Target Classification Methods
abstract
Extracting and analysing meaning-related information from natural language data has attracted the attention of researchers in various fields, such as natural language processing, corpus linguistics, information retrieval, and data science. An important aspect of such automatic information extraction and analysis is the annotation of language data using semantic tagging tools. Different semantic tagging tools have been designed to carry out various levels of semantic analysis, for instance, named entity recognition and disambiguation, sentiment analysis, word sense disambiguation, content analysis, and semantic role labelling. Common to all of these tasks, in the supervised setting, is the requirement for a manually semantically annotated corpus, which acts as a knowledge base from which to train and test potential word and phrase-level sense annotations. Many benchmark corpora have been developed for various semantic tagging tasks, but most are for English and other European languages. There is a dearth of semantically annotated corpora for the Urdu language, which is widely spoken and used around the world. To fill this gap, this study presents a large benchmark corpus and methods for the semantic tagging task for the Urdu language. The proposed corpus contains 8,000 tokens in the following domains or genres: news, social media, Wikipedia, and historical text (each domain having 2K tokens). The corpus has been manually annotated with 21 major semantic fields and 232 sub-fields with the USAS (UCREL Semantic Analysis System) semantic taxonomy which provides a comprehensive set of semantic fields for coarse-grained annotation. Each word in our proposed corpus has been annotated with at least one and up to nine semantic field tags to provide a detailed semantic analysis of the language data, which allowed us to treat the problem of semantic tagging as a supervised multi-target classification task. To demonstrate how our proposed corpus can be used for the development and evaluation of Urdu semantic tagging methods, we extracted local, topical and semantic features from the proposed corpus and applied seven different supervised multi-target classifiers to them. Results show an accuracy of 94% on our proposed corpus which is free and publicly available to download.
Jawad Shafi, Rao Muhammad Adeel Nawab, Paul Rayson
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 Cross-lingual Text Reuse Detection at Document Level for English-Urdu Language Pair
abstract
In recent years, the problem of Cross-Lingual Text Reuse Detection (CLTRD) has gained the interest of the research community due to the availability of large digital repositories and automatic Machine Translation (MT) systems. These systems are readily available and openly accessible, which makes it easier to reuse text across languages but hard to detect. In previous studies, different corpora and methods have been developed for CLTRD at the sentence/passage level for the English-Urdu language pair. However, there is a lack of large standard corpora and methods for CLTRD for the English-Urdu language pair at the document level. To overcome this limitation, the significant contribution of this study is the development of a large benchmark cross-lingual (English-Urdu) text reuse corpus, called the TREU (Text Reuse for English-Urdu) corpus. It contains English to Urdu real cases of text reuse at the document level. The corpus is manually labelled into three categories (Wholly Derived = 672, Partially Derived = 888, and Non Derived = 697) with the source text in English and the derived text in the Urdu language. Another contribution of this study is the evaluation of the TREU corpus using a diversified range of methods to show its usefulness and how it can be utilized in the development of automatic methods for measuring cross-lingual (English-Urdu) text reuse at the document level. The best evaluation results, for both binary ( F 1 = 0.78) and ternary ( F 1 = 0.66) classification tasks, are obtained using a combination of all Translation plus Mono-lingual Analysis (T+MA) based methods. The TREU corpus is publicly available to promote CLTRD research in an under-resourced language, i.e., Urdu.
Muhammad Sharjeel, Iqra Muneer, Sumaira Nosheen, Rao Muhammad Adeel Nawab, Paul Rayson
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2022 IgboBERT Models: Building and Training Transformer Models for the Igbo Language
abstract
This work presents a standard Igbo named entity recognition (IgboNER) dataset as well as the results from training and fine-tuning state-of-the-art transformer IgboNER models. We discuss the process of our dataset creation - data collection and annotation and quality checking. We also present experimental processes involved in building an IgboBERT language model from scratch as well as fine-tuning it along with other non-Igbo pre-trained models for the downstream IgboNER task. Our results show that, although the IgboNER task benefited hugely from fine-tuning large transformer model, fine-tuning a transformer model built from scratch with comparatively little Igbo text data seems to yield quite decent results for the IgboNER task. This work will contribute immensely to IgboNLP in particular as well as the wider African and low-resource NLP efforts Keywords: Igbo, named entity recognition, BERT models, under-resourced, dataset
Chiamaka Ijeoma Chukwuneke, Ignatius Ezeani, Paul Rayson, Mahmoud El-Haj
LREC3
2022 CoFiF Plus: A French Financial Narrative Summarisation Corpus
abstract
Natural Language Processing is increasingly being applied in the finance and business industry to analyse the text of many different types of financial documents. Given the increasing growth of firms around the world, the volume of financial disclosures and financial texts in different languages and forms is increasing sharply and therefore the study of language technology methods that automatically summarise content has grown rapidly into a major research area. Corpora for financial narrative summarisation exists in English, but there is a significant lack of financial text resources in the French language. To remedy this, we present CoFiF Plus, the first French financial narrative summarisation dataset providing a comprehensive set of financial text written in French. The dataset has been extracted from French financial reports published in PDF file format. It is composed of 1,703 reports from the most capitalised companies in France (Euronext Paris) covering a time frame from 1995 to 2021. This paper describes the collection, annotation and validation of the financial reports and their summaries. It also describes the dataset and gives the results of some baseline summarisers. Our datasets will be openly available upon the acceptance of the paper.
Nadhem Zmandar, Tobias Daudert, Sina Ahmadi, Mahmoud El-Haj, Paul Rayson
LREC5
2021 Multilingual Financial Word Embeddings for Arabic, English and French
abstract
Natural Language Processing is increasingly being applied to analyse the text of many different types of financial documents. For many tasks, it has been shown that standard language models and tools need to be adapted to the financial domain in order to properly represent domain specific vocabulary, styles and meanings. Previous work has almost exclusively focused on English financial text, so in this paper we describe the creation of novel financial word embeddings for three languages: English, French and Arabic. In order to evaluate the effectiveness of the embeddings, we started by evaluating the English embeddings on a sentiment analysis classification task using the existing FinancialPhrase dataset and show improved performance over a standard GloVe based model using convolutional neural networks.
Nadhem Zmandar, Mahmoud El-Haj, Paul Rayson
IEEE BigData3
2021 GIBBON: General-purpose Information-Based Bayesian Optimisation
abstract
This paper describes a general-purpose extension of max-value entropy search, a popular approach for Bayesian Optimisation (BO). A novel approximation is proposed for the information gain -- an information-theoretic quantity central to solving a range of BO problems, including noisy, multi-fidelity and batch optimisations across both continuous and highly-structured discrete spaces. Previously, these problems have been tackled separately within information-theoretic BO, each requiring a different sophisticated approximation scheme, except for batch BO, for which no computationally-lightweight information-theoretic approach has previously been proposed. GIBBON (General-purpose Information-Based Bayesian OptimisatioN) provides a single principled framework suitable for all the above, out-performing existing approaches whilst incurring substantially lower computational overheads. In addition, GIBBON does not require the problem's search space to be Euclidean and so is the first high-performance yet computationally light-weight acquisition function that supports batch BO over general highly structured input spaces like molecular search and gene design. Moreover, our principled derivation of GIBBON yields a natural interpretation of a popular batch BO heuristic based on determinantal point processes. Finally, we analyse GIBBON across a suite of synthetic benchmark tasks, a molecular search loop, and as part of a challenging batch multi-fidelity framework for problems with controllable experimental noise.
Henry B. Moss, David S. Leslie, Javier González 0002, Paul Rayson
J. Mach. Learn. Res.4
2021 MasakhaNER: Named Entity Recognition for African Languages
abstract
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition (NER) in ten African languages. We detail the characteristics of these languages to help researchers and practitioners better understand the challenges they pose for NER tasks. We analyze our datasets and conduct an extensive empirical evaluation of state- of-the-art methods across both supervised and transfer learning settings. Finally, we release the data, code, and models to inspire future research on African NLP.1
David Ifeoluwa Adelani, Jade Z. Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew 0002, Israel Abebe Azime, Shamsuddeen Hassan Muhammad, Chris C. Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba O. Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin P. Adewumi, Paul Rayson, Mofe Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Ijeoma Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane Mboup, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing K. Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima Diop, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, Salomey Osei
Trans. Assoc. Comput. Linguistics31
2020 Developing an Arabic Infectious Disease Ontology to Include Non-Standard Terminology
abstract
Building ontologies is a crucial part of the semantic web endeavour. In recent years, research interest has grown rapidly in supporting languages such as Arabic in NLP in general but there has been very little research on medical ontologies for Arabic. We present a new Arabic ontology in the infectious disease domain to support various important applications including the monitoring of infectious disease spread via social media. This ontology meaningfully integrates the scientific vocabularies of infectious diseases with their informal equivalents. We use ontology learning strategies with manual checking to build the ontology. We applied three statistical methods for term extraction from selected Arabic infectious diseases articles: TF-IDF, C-value, and YAKE. We also conducted a study, by consulting around 100 individuals, to discover the informal terms related to infectious diseases in Arabic. In future work, we will automatically extract the relations for infectious disease concepts but for now these are manually created. We report two complementary experiments to evaluate the ontology. First, a quantitative evaluation of the term extraction results and an additional qualitative evaluation by a domain expert.
Lama Alsudias, Paul Rayson
LREC2
2020 LexiDB: Patterns & Methods for Corpus Linguistic Database Management
abstract
LexiDB is a tool for storing, managing and querying corpus data. In contrast to other database management systems (DBMSs), it is designed specifically for text corpora. It improves on other corpus management systems (CMSs) because data can be added and deleted from corpora on the fly with the ability to add live data to existing corpora. LexiDB sits between these two categories of DBMSs and CMSs, more specialised to language data than a general purpose DBMS but more flexible than a traditional static corpus management system. Previous work has demonstrated the scalability of LexiDB in response to the growing need to be able to scale out for ever growing corpus datasets. Here, we present the patterns and methods developed in LexiDB for storage, retrieval and querying of multi-level annotated corpus data. These techniques are evaluated and compared to an existing CMS (Corpus Workbench CWB - CQP) and indexer (Lucene). We find that LexiDB consistently outperforms existing tools for corpus queries. This is particularly apparent with large corpora and when handling queries with large result sets
Matthew Coole, Paul Rayson, John A. Mariani
LREC2
2020 Infrastructure for Semantic Annotation in the Genomics Domain
abstract
We describe a novel super-infrastructure for biomedical text mining which incorporates an end-to-end pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature, combining NLP and corpus linguistics methods. The infrastructure permits extreme-scale research on the open access PubMed Central archive. It combines an updatable Gene Ontology Semantic Tagger (GOST) for entity identification and semantic markup in the literature, with a NLP pipeline scheduler (Buster) to collect and process the corpus, and a bespoke columnar corpus database (LexiDB) for indexing. The corpus database is distributed to permit fast indexing, and provides a simple web front-end with corpus linguistics methods for sub-corpus comparison and retrieval. GOST is also connected as a service in the Language Application (LAPPS) Grid, in which context it is interoperable with other NLP tools and data in the Grid and can be combined with them in more complex workflows. In a literature based discovery setting, we have created an annotated corpus of 9,776 papers with 5,481,543 words.
Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani, Sheryl Prentice, Nancy Ide, Jo Knight, Scott Piao, John A. Mariani, Paul Rayson, Keith Suderman
LREC10
2020 BOSS: Bayesian Optimization over String Spaces
abstract
This article develops a Bayesian optimization (BO) method which acts directly over raw strings, proposing the first uses of string kernels and genetic algorithms within BO loops. Recent applications of BO over strings have been hindered by the need to map inputs into a smooth and unconstrained latent space. Learning this projection is computationally and data-intensive. Our approach instead builds a powerful Gaussian process surrogate model based on string kernels, naturally supporting variable length inputs, and performs efficient acquisition function maximization for spaces with syntactic constraints. Experiments demonstrate considerably improved optimization over existing approaches across a broad range of constraints, including the popular setting where syntax is governed by a context-free grammar.
Henry B. Moss, David S. Leslie, Daniel Beck, Javier González 0002, Paul Rayson
NeurIPS5
2020 MUMBO: MUlti-task Max-Value Bayesian Optimization
Henry B. Moss, David S. Leslie, Paul Rayson
ECML/PKDD (3)3
2020 Known and unknown requirements in healthcare
abstract
We report experience in requirements elicitation of domain knowledge from experts in clinical and cognitive neurosciences. The elicitation target was a causal model for early signs of dementia indicated by changes in user behaviour and errors apparent in logs of computer activity. A Delphi-style process consisting of workshops with experts followed by a questionnaire was adopted. The paper describes how the elicitation process had to be adapted to deal with problems encountered in terminology and limited consensus among the experts. In spite of the difficulties encountered, a partial causal model of user behavioural pathologies and errors was elicited. This informed requirements for configuring data- and text-mining tools to search for the specific data patterns. Lessons learned for elicitation from experts are presented, and the implications for requirements are discussed as “unknown unknowns”, as well as configuration requirements for directing data-/text-mining tools towards refining awareness requirements in healthcare applications.
Alistair G. Sutcliffe, Peter Sawyer, Gemma Stringer, Samuel Couth, Laura J. E. Brown, Ann Gledson, Christopher Bull 0001, Paul Rayson, John A. Keane, Xiaojun Zeng, Iracema Leroi
Requir. Eng.8
2019 FIESTA: Fast IdEntification of State-of-The-Art models using adaptive bandit algorithms
abstract
We present FIESTA, a model selection approach that significantly reduces the computational resources required to reliably identify state-of-the-art performance from large collections of candidate models.Despite being known to produce unreliable comparisons, it is still common practice to compare model evaluations based on single choices of random seeds.We show that reliable model selection also requires evaluations based on multiple train-test splits (contrary to common practice in many shared tasks).Using bandit theory from the statistics literature, we are able to adaptively determine appropriate numbers of data splits and random seeds used to evaluate each model, focusing computational resources on the evaluation of promising models whilst avoiding wasting evaluations on models with lower performance.Furthermore, our userfriendly Python implementation produces confidence guarantees of correctly selecting the optimal model.We evaluate our algorithms by selecting between 8 target-dependent sentiment analysis methods using dramatically fewer model evaluations than current model selection approaches.
Henry B. Moss, Andrew Moore 0001, David S. Leslie, Paul Rayson
ACL (1)4
2019 CLEU - A Cross-language english-urdu corpus and benchmark for text reuse experiments
abstract
Text reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi‐lingual content on the Web has increased cross‐language text reuse to an unprecedented scale. Although researchers have proposed methods to detect it, one major drawback is the unavailability of large‐scale gold standard evaluation resources built on real cases. To overcome this problem, we propose a cross‐language sentence/passage level text reuse corpus for the English‐Urdu language pair. The Cross‐Language English‐Urdu Corpus (CLEU) has source text in English whereas the derived text is in Urdu. It contains in total 3,235 sentence/passage pairs manually tagged into three categories that is near copy, paraphrased copy, and independently written. Further, as a second contribution, we evaluate the Translation plus Mono‐lingual Analysis method using three sets of experiments on the proposed dataset to highlight its usefulness. Evaluation results (f1=0.732 binary, f1=0.552 ternary classification) indicate that it is harder to detect cross‐language real cases of text reuse, especially when the language pairs have unrelated scripts. The corpus is a useful benchmark resource for the future development and assessment of cross‐language text reuse detection systems for the English‐Urdu language pair.
Iqra Muneer, Muhammad Sharjeel, Muntaha Iqbal, Rao Muhammad Adeel Nawab, Paul Rayson
J. Assoc. Inf. Sci. Technol.5
2019 A Sense Annotated Corpus for All-Words Urdu Word Sense Disambiguation
abstract
Word Sense Disambiguation (WSD) aims to automatically predict the correct sense of a word used in a given context. All human languages exhibit word sense ambiguity, and resolving this ambiguity can be difficult. Standard benchmark resources are required to develop, compare, and evaluate WSD techniques. These are available for many languages, but not for Urdu, despite this being a language with more than 300 million speakers and large volumes of text available digitally. To fill this gap, this study proposes a novel benchmark corpus for the Urdu All-Words WSD task. The corpus contains 5,042 words of Urdu running text in which all ambiguous words (856 instances) are manually tagged with senses from the Urdu Lughat dictionary. A range of baseline WSD models based on n -gram are applied to the corpus, and the best performance (accuracy of 57.71%) is achieved using word 4-gram. The corpus is freely available to the research community to encourage further WSD research in Urdu.
Ali Saeed, Rao Muhammad Adeel Nawab, Mark Stevenson 0001, Paul Rayson
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2018 Bringing replication and reproduction together with generalisability in NLP: Three reproduction studies for Target Dependent Sentiment Analysis
abstract
Lack of repeatability and generalisability are two significant threats to continuing scientific development in Natural Language Processing. Language models and learning methods are so complex that scientific conference papers no longer contain enough space for the technical depth required for replication or reproduction. Taking Target Dependent Sentiment Analysis as a case study, we show how recent work in the field has not consistently released code, or described settings for learning methods in enough detail, and lacks comparability and generalisability in train, test or validation data. To investigate generalisability and to enable state of the art comparative evaluations, we carry out the first reproduction studies of three groups of complementary methods and perform the first large-scale mass evaluation on six different English datasets. Reflecting on our experiences, we recommend that future replication or reproduction experiments should always consider a variety of datasets alongside documenting and releasing their methods and published code in order to minimise the barriers to both repeatability and generalisability. We have released our code with a model zoo on GitHub with Jupyter Notebooks to aid understanding and full documentation, and we recommend that others do the same with their papers at submission time through an anonymised GitHub account.
Andrew Moore 0001, Paul Rayson
COLING2
2018 Using J-K-fold Cross Validation To Reduce Variance When Tuning NLP Models
abstract
K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so our performance estimates are in fact stochastic, with variability that can be substantial for natural language processing tasks. We demonstrate that these unstable estimates cannot be relied upon for effective parameter tuning. The resulting tuned parameters are highly sensitive to how our data is partitioned, meaning that we often select sub-optimal parameter choices and have serious reproducibility issues. Instead, we propose to use the less variable J-K-fold CV, in which J independent K-fold cross validations are used to assess performance. Our main contributions are extending J-K-fold CV from performance estimation to parameter tuning and investigating how to choose J and K. We argue that variability is more important than bias for effective tuning and so advocate lower choices of K than are typically seen in the NLP literature and instead use the saved computation to increase J. To demonstrate the generality of our recommendations we investigate a wide range of case-studies: sentiment classification (both general and target-specific), part-of-speech tagging and document classification.
Henry B. Moss, David S. Leslie, Paul Rayson
COLING3
2018 Arabic Dialect Identification in the Context of Bivalency and Code-Switching
Mahmoud El-Haj, Paul Rayson, Mariam Aboelezz
LREC2
2018 Profiling Medical Journal Articles Using a Gene Ontology Semantic Tagger
Mahmoud El-Haj, Paul Rayson, Scott Piao, Jo Knight
LREC2
2018 Towards a Welsh Semantic Annotation System
Scott Piao, Paul Rayson, Dawn Knight, Gareth Watkins
LREC2
2017 A time-sensitive historical thesaurus-based semantic tagger for deep semantic annotation
abstract
Automatic extraction and analysis of meaning-related information from natural language data has been an important issue in a number of research areas, such as natural language processing (NLP), text mining, corpus linguistics, and data science. An important aspect of such information extraction and analysis is the semantic annotation of language data using a semantic tagger. In practice, various semantic annotation tools have been designed to carry out different levels of semantic annotation, such as topics of documents, semantic role labeling, named entities or events. Currently, the majority of existing semantic annotation tools identify and tag partial core semantic information in language data, but they tend to be applicable only for modern language corpora. While such semantic analyzers have proven useful for various purposes, a semantic annotation tool that is capable of annotating deep semantic senses of all lexical units, or all-words tagging, is still desirable for a deep, comprehensive semantic analysis of language data. With large-scale digitization efforts underway, delivering historical corpora with texts dating from the last 400 years, a particularly challenging aspect is the need to adapt the annotation in the face of significant word meaning change over time. In this paper, we report on the development of a new semantic tagger (the Historical Thesaurus Semantic Tagger), and discuss challenging issues we faced in this work. This new semantic tagger is built on existing NLP tools and incorporates a large-scale historical English thesaurus linked to the Oxford English Dictionary. Employing contextual disambiguation algorithms, this tool is capable of annotating lexical units with a historically-valid highly fine-grained semantic categorization scheme that contains about 225,000 semantic concepts and 4,033 thematic semantic categories. In terms of novelty, it is adapted for processing historical English data, with rich information about historical usage of words and a spelling variant normalizer for historical forms of English. Furthermore, it is able to make use of knowledge about the publication date of a text to adapt its output. In our evaluation, the system achieved encouraging accuracies ranging from 77.12% to 91.08% on individual test texts. Applying time-sensitive methods improved results by as much as 3.54% and by 1.72% on average.
Scott Piao, Fraser Dallachy, Alistair Baron, Jane Demmen, Steve Wattam, Philip Durkin, James McCracken, Paul Rayson, Marc Alexander
Comput. Speech Lang.8
2016 lexiDB: A scalable corpus database management system
abstract
lexiDB is a scalable corpus database management system designed to fulfill corpus linguistics retrieval queries on multi-billion-word multiply-annotated corpora. It is based on a distributed architecture that allows the system to scale out to support ever larger text collections. This paper presents an overview of the architecture behind lexiDB as well as a demonstration of its functionality. We present lexiDB's performance metrics based on the AWS (Amazon Web Services) infrastructure with two part-of-speech and semantically tagged billion word corpora: Historical Hansard and EEBO (Early English Books Online).
Matthew Coole, Paul Rayson, John A. Mariani
IEEE BigData2
2016 Sampling labelled profile data for identity resolution
abstract
Identity resolution capability for social networking profiles is important for a range of purposes, from open-source intelligence applications to forming semantic web connections. Yet replication of research in this area is hampered by the lack of access to ground-truth data linking the identities of profiles from different networks. Almost all data sources previously used by researchers are no longer available, and historic datasets are both of decreasing relevance to the modern social networking landscape and ethically troublesome regarding the preservation and publication of personal data. We present and evaluate a method which provides researchers in identity resolution with easy access to a realistically-challenging labelled dataset of online profiles, drawing on four of the currently largest and most influential online social networks. We validate the comparability of samples drawn through this method and discuss the implications of this mechanism for researchers as well as potential alternatives and extensions.
Matthew Edwards 0001, Stephen Wattam, Paul Rayson, Awais Rashid
IEEE BigData3
2016 OSMAN ― A Novel Arabic Readability Metric
Mahmoud El-Haj, Paul Rayson
LREC2
2016 Learning Tone and Attribution for Financial Text Mining
Mahmoud El-Haj, Paul Rayson, Steven Young 0001, Andrew Moore 0001, Martin Walker, Thomas Schleicher, Vasiliki Athanasakou
LREC2
2016 Lexical Coverage Evaluation of Large-scale Multilingual Semantic Lexicons for Twelve Languages
Scott Piao, Paul Rayson, Dawn Archer, Francesca Bianchi, Carmen Dayrell, Mahmoud El-Haj, Ricardo-María Jiménez, Dawn Knight, Michal Kren, Laura Löfberg, Rao Muhammad Adeel Nawab, Jawad Shafi, Phoey Lee Teh, Olga Mudraya
LREC2
2016 UPPC - Urdu Paraphrase Plagiarism Corpus
Muhammad Sharjeel, Paul Rayson, Rao Muhammad Adeel Nawab
LREC2
2016 Reversing the Polarity with Emoticons
Phoey Lee Teh, Paul Rayson, Irina Pak, Scott Piao, Seow Mei Yeng
NLDB2
2015 Scaling out for extreme scale corpus data
abstract
Much of the previous work in Big Data has focussed on numerical sources of information. However, with the `narrative turn' in many disciplines gathering pace and commercial organisations beginning to realise the value of their textual assets, natural language data is fast catching up as an exploitable source of information for decision making. With vast quantities of unstructured textual data on the web, in social media, and in newly digitised historical document archives, the 5Vs (Volume, Velocity, Variety, Value and Veracity) apply equally well, if not more so, to big textual data. Corpus linguistics, the computer-aided study of large collections of naturally occurring language data, has been dealing with big data for fifty years. Corpus linguistics methods impose complex requirements on the retrieval, annotation and analysis of text in terms of displaying narrow contexts for each occurrence of a word or linguistic feature being studied and counting co-occurrences with other words or features to determine significant patterns in language. This, coupled with the distribution of language features in accordance with Zipf's Law, poses complex challenges for data models and corpus software dealing with extreme scale language data. A related issue is the non-random nature of language and the `burstiness' of word occurrences, or what we might put in Big Data terms as a sixth `V' called Viscosity. We report experiments to examine and compare the capabilities of two No-SQL databases in clustered configurations for the indexing, retrieval and analysis of billion-word corpora, since this size is the current state-of-the-art in corpus linguistics. We find that modern DBMSs (Database Management Systems) are capable of handling this extreme scale corpus data set for simple queries but are limited when querying for more frequent words or more complex queries.
Matthew Coole, Paul Rayson, John A. Mariani
IEEE BigData2
2015 Dementia and Social Sustainability: Challenges for Software Engineering
abstract
Dementia is a serious threat to social sustainability. As life expectancy increases, more people are developing dementia. At the same time, demographic change is reducing the economically active part of the population. Care of people with dementia imposes great emotional and financial strain on sufferers, their families and society at large. In response, significant research resources are being focused on dementia. One research thread is focused on using computer technology to monitor people in at-risk groups to improve rates of early diagnosis. In this paper we provide an overview of dementia monitoring research and identify a set of scientific challenges for the engineering of dementia-monitoring software, with implications for other mental health self-management systems.
Peter Sawyer, Alistair G. Sutcliffe, Paul Rayson, Christopher Bull 0001
ICSE (2)3
2015 Sentiment analysis tools should take account of the number of exclamation marks!!!
abstract
There are various factors that affect the sentiment level expressed in textual comments. Capitalization of letters tends to mark something for attention and repeating of letters tends to strengthen the emotion. Emoticons are used to help visualize facial expressions which can affect understanding of text. In this paper, we show the effect of the number of exclamation marks used, via testing with twelve online sentiment tools. We present opinions gathered from 500 respondents towards "like" and "dislike" values, with a varying number of exclamation marks. Results show that only 20% of the online sentiment tools tested considered the number of exclamation marks in their returned scores. However, results from our human raters show that the more exclamation marks used for positive comments, the more they have higher "like" values than the same comments with fewer exclamations marks. Similarly, adding more exclamation marks for negative comments, results in a higher "dislike".
Phoey Lee Teh, Paul Rayson, Irina Pak, Scott Piao
iiWAS2
2015 Development of the Multilingual Semantic Annotation System
abstract
Scott Piao, Francesca Bianchi, Carmen Dayrell, Angela D’Egidio, Paul Rayson. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Scott Piao, Francesca Bianchi, Carmen Dayrell, Angela D'Egidio, Paul Rayson
HLT-NAACL5
2014 Dealing with heterogeneous big data when geoparsing historical corpora
abstract
It has long been known that `variety' is one of the key challenges and opportunities of big data. This is especially true when we consider the variety of content in historical corpora resulting from large-scale digitisation activities. Collections such as Early English Books Online (EEBO) and the British Library 19th Century Newspapers are extremely large and heterogeneous data sources containing a variety of content in terms of time, location, topic, style and quality. The range of geographical locations referenced in these corpora poses a difficult challenge for state of the art geoparsing tools. In the context of our work on Spatial Humanities analyses, we present our solution for dealing with the variety and scale of these corpora.
C. J. Rupp, Paul Rayson, Ian N. Gregory, Andrew Hardie, Amelia Joulain, Daniel Hartmann
IEEE BigData2
2014 Detecting Document Structure in a Very Large Corpus of UK Financial Reports
Mahmoud El-Haj, Paul Rayson, Steven Young 0001, Martin Walker
LREC2
2014 Experiences with Parallelisation of an Existing NLP Pipeline: Tagging Hansard
Stephen Wattam, Paul Rayson, Marc Alexander, Jean Anderson
LREC2
2014 Discovering affect-laden requirements to achieve system acceptance
abstract
Novel envisioned systems face the risk of rejection by their target user community and the requirements engineer must be sensitive to the factors that will determine acceptance or rejection. Conventionally, technology acceptance is determined by perceived usefulness and ease-of-use, but in some domains other factors play an important role. In healthcare systems, particularly, ethical and emotional factors can be crucial. In this paper we describe an approach to requirements discovery that we developed for such systems. We describe how we have applied our approach to a novel system to passively monitor users for signs of cognitive decline consistent with the onset of dementia. A key challenge was eliciting users' reactions to emotionally charged events never before experienced by them at first hand. Our goal was to understand the range of users' emotional responses and their values and motivations, and from these formulate requirements that would maximise the likelihood of acceptance of the system. The problem was heightened by the fact that the key stakeholders were elderly people who represent a poorly studied user constituency. We discuss the elicitation and analysis methodologies used, and our experience with tool support. We conclude by reflecting on the affect issues for RE and for technology acceptance.
Alistair G. Sutcliffe, Paul Rayson, Christopher Bull 0001, Peter Sawyer
RE2
2013 Customising geoparsing and georeferencing for historical texts
abstract
In order to better support the text mining of historical texts, we propose a combination of complementary techniques from Geographical Information Systems, computational and corpus linguistics. In previous work, we have described this as `visual gisting' to extract important themes from text and locate those themes on a map representing geographical information contained in the text. Here, we describe the steps that were found necessary to apply standard analysis and resolution tools to identify place names in a specific corpus of historical texts. This task is seen as an initial and prerequisite step for further analysis and comparison by combining the information we extract from a corpus with information from other sources, including other text corpora. The process is intended to support close reading of historical texts on a much larger scale by highlighting using exploratory and data-driven approaches which parts of the corpus warrant further close analysis. Our case study presented here is from a corpus of Lake District travel literature. We discuss the customisations that we have to make to existing tools to extract placename information and visualise it on a map.
C. J. Rupp, Paul Rayson, Alistair Baron, Christopher Donaldson, Ian N. Gregory, Andrew Hardie, Patricia Murrieta-Flores
IEEE BigData2
2012 Understanding Actionable Knowledge in Social Media: BBC Question Time and Twitter, a Case Study
Maria Angela Ferrario, William Simm, Jon Whittle 0001, Paul Rayson, Maria Terzi, Jane M. Binner
ICWSM4
2012 Document Attrition in Web Corpora: an Exploration
Stephen Wattam, Paul Rayson, Damon Berridge
LREC2
2008 An Exploratory Study of Information Retrieval Techniques in Domain Analysis
abstract
Domain analysis involves not only looking at standard requirements documents (e.g., use case specifications) but also at customer information packs, market analyses, etc. Looking across all these documents and deriving, in a practical and scalable way, a feature model that is comprised of coherent abstractions is a fundamental and non-trivial challenge. We conduct an exploratory study to investigate the suitability of Information Retrieval (IR) techniques for scalable identification of commonalities and variabilities in requirement specifications for software product lines. Accordingly, based on observations derived from industrial experience and on state-of-the-art research and practice, we also propose an initial framework, leveraging IR to systematically abstract requirements from existing specifications of a given domain into a feature model. We evaluate this framework, present a roadmap for its further extension, and formulate hypotheses to guide future work in exploring IR techniques for domain analysis.
Vander Alves, Christa Schwanninger, Luciano Barbosa, Awais Rashid, Peter Sawyer, Paul Rayson, Christoph Pohl, Andreas Rummler
SPLC6
2008 A framework for P2P application development
James Walkerdine, Danny Hughes 0001, Paul Rayson, John Simms, Kiel Mark Gilleade, John A. Mariani, Ian Sommerville
Comput. Commun.3
2008 A flexible framework to experiment with ontology learning techniques
Ricardo Gacitúa, Peter Sawyer, Paul Rayson
Knowl. Based Syst.3
2006 ASSIST: Automated Semantic Assistance for Translators
Serge Sharoff, Bogdan Babych, Paul Rayson, Olga Mudraya, Scott Piao
EACL3
2005 EA-Miner: a tool for automating aspect-oriented requirements identification
abstract
Aspect-Oriented requirements engineering helps to achieve early separation of concerns by supporting systematic analysis of broadly-scoped properties such as security, real-time constraints, etc. The early identification and separation of aspects and base abstractions crosscut by them helps to avoid costly refactorings at later stages such as design and code. However, if not handled effectively, the aspect identification task can become a bottleneck requiring a significant effort due to the large amount of, often poorly structured or imprecise, information available to a requirements engineer. In this paper, we describe a tool, EA-Miner, that provides effective automated support for identifying and separating aspectual and non-aspectual concerns as well as their crosscutting relationships at the requirements level. The tool utilises natural language processing techniques to reason about the properties of the concerns and model their structure and relationships.
Américo Sampaio, Ruzanna Chitchyan, Paul Rayson
ASE3
2005 Early-AIM: An Approach for Identifying Aspects in Requirements
abstract
Identifying aspects at an early stage helps to achieve separation of crosscutting concerns in the initial system analysis, instead of deferring such decisions to later stages of design and code, and thus, having to perform costly refactorings. This paper describes the early-AIM (early aspects identification method) approach that utilises corpus-based natural language processing (NLP) techniques to effectively enable the identification and modelling of early aspects in a semi-automated way.
Américo Sampaio, Awais Rashid, Paul Rayson
RE3
2005 Comparing and combining a semantic tagger and a statistical tool for MWE extraction
Scott Piao, Paul Rayson, Dawn Archer, Tony McEnery
Comput. Speech Lang.2
2005 Shallow Knowledge as an Aid to Deep Understanding in Early Phase Requirements Engineering
abstract
Requirements engineering's continuing dependence on natural language description has made it the focus of several efforts to apply language engineering techniques. The raw textual material that forms an input to early phase requirements engineering and which informs the subsequent formulation of the requirements is inevitably uncontrolled and this makes its processing very hard. Nevertheless, sufficiently robust techniques do exist that can be used to aid the requirements engineer provided that the scope of what can be achieved is understood. In this paper, we show how combinations of lexical and shallow semantic analysis techniques developed from corpus linguistics can help human analysts acquire the deep understanding needed as the first step towards the synthesis of requirements.
Peter Sawyer, Paul Rayson, Kenneth Cosh
IEEE Trans. Software Eng.2
2004 Evaluating Lexical Resources for a Semantic Tagger
Scott Piao, Paul Rayson, Dawn Archer, Tony McEnery
LREC2
2004 Language Resources and Tools for Supporting the System Engineering Process
Victor Onditi, Paul Rayson, B. Ransom, Devina Ramduny-Ellis, Ian Sommerville, Alan J. Dix
NLDB2
2004 P2P-4-DL: Digital Library over Peer-to-Peer
abstract
The P2P-4-DL project aims to investigate and build a DL system that would operate over a P2P structure. Rather than storing digital objects centrally they remain the responsibility of the individual peers that provide them. This allows the system to utilise network resources more efficiently as well as providing users with a greater sense of control over the digital objects they share. Our prototype also draws upon natural language processing (NLP) techniques in an attempt to increase the usability of the system. Other related work within this area includes EDUTELLA, a RDF based P2P infrastructure that can support the development of DL's.
James Walkerdine, Paul Rayson
Peer-to-Peer Computing2
2000 The REVERE Project: Experiments with the Application of Probabilistic NLP to Systems Engineering
Paul Rayson, Luke Emmet, Roger Garside, Peter Sawyer
NLDB1
1998 Supporting Information Evolution on the WWW
Ian Sommerville, Tom Rodden, Paul Rayson, Andrew Kirby, Alan J. Dix
World Wide Web3