Paul Rayson

dblp:96/5736 · also Paul Edward Rayson · DBLP profile ↗
← Back
17ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0002-1257-2191ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 7 (1 first)Big Data, Cloud & Distributed Data Systems · 7Knowledge Engineering, Semantic Web & Information Systems · 2Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2025 AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
Mo El-Haj, Paul Rayson
IEEE Big Data2
2023 Open-Source Thesaurus Development for Under-Resourced Languages: a Welsh Case Study
Nouran Khallaf, Elin Arfon, Mo El-Haj, Jonathan Morris, Dawn Knight, Paul Rayson, Tymaa Hammouda, Mustafa Jarrar
LDK6
2023 FinAraT5: A text to text model for financial Arabic text understanding and generation
Nadhem Zmandar, Mo El-Haj, Paul Rayson
LDK3
2023 A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers
Nadhem Zmandar, Mahmoud El-Haj, Paul Rayson
NLDB3
2021 Multilingual Financial Word Embeddings for Arabic, English and French
abstract
Natural Language Processing is increasingly being applied to analyse the text of many different types of financial documents. For many tasks, it has been shown that standard language models and tools need to be adapted to the financial domain in order to properly represent domain specific vocabulary, styles and meanings. Previous work has almost exclusively focused on English financial text, so in this paper we describe the creation of novel financial word embeddings for three languages: English, French and Arabic. In order to evaluate the effectiveness of the embeddings, we started by evaluating the English embeddings on a sentiment analysis classification task using the existing FinancialPhrase dataset and show improved performance over a standard GloVe based model using convolutional neural networks.
Nadhem Zmandar, Mahmoud El-Haj, Paul Rayson
IEEE BigData3
2020 MUMBO: MUlti-task Max-Value Bayesian Optimization
Henry B. Moss, David S. Leslie, Paul Rayson
ECML/PKDD (3)3
2019 CLEU - A Cross-language english-urdu corpus and benchmark for text reuse experiments
abstract
Text reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi‐lingual content on the Web has increased cross‐language text reuse to an unprecedented scale. Although researchers have proposed methods to detect it, one major drawback is the unavailability of large‐scale gold standard evaluation resources built on real cases. To overcome this problem, we propose a cross‐language sentence/passage level text reuse corpus for the English‐Urdu language pair. The Cross‐Language English‐Urdu Corpus (CLEU) has source text in English whereas the derived text is in Urdu. It contains in total 3,235 sentence/passage pairs manually tagged into three categories that is near copy, paraphrased copy, and independently written. Further, as a second contribution, we evaluate the Translation plus Mono‐lingual Analysis method using three sets of experiments on the proposed dataset to highlight its usefulness. Evaluation results (f1=0.732 binary, f1=0.552 ternary classification) indicate that it is harder to detect cross‐language real cases of text reuse, especially when the language pairs have unrelated scripts. The corpus is a useful benchmark resource for the future development and assessment of cross‐language text reuse detection systems for the English‐Urdu language pair.
Iqra Muneer, Muhammad Sharjeel, Muntaha Iqbal, Rao Muhammad Adeel Nawab, Paul Rayson
J. Assoc. Inf. Sci. Technol.5
2016 lexiDB: A scalable corpus database management system
abstract
lexiDB is a scalable corpus database management system designed to fulfill corpus linguistics retrieval queries on multi-billion-word multiply-annotated corpora. It is based on a distributed architecture that allows the system to scale out to support ever larger text collections. This paper presents an overview of the architecture behind lexiDB as well as a demonstration of its functionality. We present lexiDB's performance metrics based on the AWS (Amazon Web Services) infrastructure with two part-of-speech and semantically tagged billion word corpora: Historical Hansard and EEBO (Early English Books Online).
Matthew Coole, Paul Rayson, John A. Mariani
IEEE BigData2
2016 Sampling labelled profile data for identity resolution
abstract
Identity resolution capability for social networking profiles is important for a range of purposes, from open-source intelligence applications to forming semantic web connections. Yet replication of research in this area is hampered by the lack of access to ground-truth data linking the identities of profiles from different networks. Almost all data sources previously used by researchers are no longer available, and historic datasets are both of decreasing relevance to the modern social networking landscape and ethically troublesome regarding the preservation and publication of personal data. We present and evaluate a method which provides researchers in identity resolution with easy access to a realistically-challenging labelled dataset of online profiles, drawing on four of the currently largest and most influential online social networks. We validate the comparability of samples drawn through this method and discuss the implications of this mechanism for researchers as well as potential alternatives and extensions.
Matthew Edwards 0001, Stephen Wattam, Paul Rayson, Awais Rashid
IEEE BigData3
2016 Reversing the Polarity with Emoticons
Phoey Lee Teh, Paul Rayson, Irina Pak, Scott Piao, Seow Mei Yeng
NLDB2
2015 Scaling out for extreme scale corpus data
abstract
Much of the previous work in Big Data has focussed on numerical sources of information. However, with the `narrative turn' in many disciplines gathering pace and commercial organisations beginning to realise the value of their textual assets, natural language data is fast catching up as an exploitable source of information for decision making. With vast quantities of unstructured textual data on the web, in social media, and in newly digitised historical document archives, the 5Vs (Volume, Velocity, Variety, Value and Veracity) apply equally well, if not more so, to big textual data. Corpus linguistics, the computer-aided study of large collections of naturally occurring language data, has been dealing with big data for fifty years. Corpus linguistics methods impose complex requirements on the retrieval, annotation and analysis of text in terms of displaying narrow contexts for each occurrence of a word or linguistic feature being studied and counting co-occurrences with other words or features to determine significant patterns in language. This, coupled with the distribution of language features in accordance with Zipf's Law, poses complex challenges for data models and corpus software dealing with extreme scale language data. A related issue is the non-random nature of language and the `burstiness' of word occurrences, or what we might put in Big Data terms as a sixth `V' called Viscosity. We report experiments to examine and compare the capabilities of two No-SQL databases in clustered configurations for the indexing, retrieval and analysis of billion-word corpora, since this size is the current state-of-the-art in corpus linguistics. We find that modern DBMSs (Database Management Systems) are capable of handling this extreme scale corpus data set for simple queries but are limited when querying for more frequent words or more complex queries.
Matthew Coole, Paul Rayson, John A. Mariani
IEEE BigData2
2015 Sentiment analysis tools should take account of the number of exclamation marks!!!
abstract
There are various factors that affect the sentiment level expressed in textual comments. Capitalization of letters tends to mark something for attention and repeating of letters tends to strengthen the emotion. Emoticons are used to help visualize facial expressions which can affect understanding of text. In this paper, we show the effect of the number of exclamation marks used, via testing with twelve online sentiment tools. We present opinions gathered from 500 respondents towards "like" and "dislike" values, with a varying number of exclamation marks. Results show that only 20% of the online sentiment tools tested considered the number of exclamation marks in their returned scores. However, results from our human raters show that the more exclamation marks used for positive comments, the more they have higher "like" values than the same comments with fewer exclamations marks. Similarly, adding more exclamation marks for negative comments, results in a higher "dislike".
Phoey Lee Teh, Paul Rayson, Irina Pak, Scott Piao
iiWAS2
2014 Dealing with heterogeneous big data when geoparsing historical corpora
abstract
It has long been known that `variety' is one of the key challenges and opportunities of big data. This is especially true when we consider the variety of content in historical corpora resulting from large-scale digitisation activities. Collections such as Early English Books Online (EEBO) and the British Library 19th Century Newspapers are extremely large and heterogeneous data sources containing a variety of content in terms of time, location, topic, style and quality. The range of geographical locations referenced in these corpora poses a difficult challenge for state of the art geoparsing tools. In the context of our work on Spatial Humanities analyses, we present our solution for dealing with the variety and scale of these corpora.
C. J. Rupp, Paul Rayson, Ian N. Gregory, Andrew Hardie, Amelia Joulain, Daniel Hartmann
IEEE BigData2
2013 Customising geoparsing and georeferencing for historical texts
abstract
In order to better support the text mining of historical texts, we propose a combination of complementary techniques from Geographical Information Systems, computational and corpus linguistics. In previous work, we have described this as `visual gisting' to extract important themes from text and locate those themes on a map representing geographical information contained in the text. Here, we describe the steps that were found necessary to apply standard analysis and resolution tools to identify place names in a specific corpus of historical texts. This task is seen as an initial and prerequisite step for further analysis and comparison by combining the information we extract from a corpus with information from other sources, including other text corpora. The process is intended to support close reading of historical texts on a much larger scale by highlighting using exploratory and data-driven approaches which parts of the corpus warrant further close analysis. Our case study presented here is from a corpus of Lake District travel literature. We discuss the customisations that we have to make to existing tools to extract placename information and visualise it on a map.
C. J. Rupp, Paul Rayson, Alistair Baron, Christopher Donaldson, Ian N. Gregory, Andrew Hardie, Patricia Murrieta-Flores
IEEE BigData2
2012 Understanding Actionable Knowledge in Social Media: BBC Question Time and Twitter, a Case Study
Maria Angela Ferrario, William Simm, Jon Whittle 0001, Paul Rayson, Maria Terzi, Jane M. Binner
ICWSM4
2004 Language Resources and Tools for Supporting the System Engineering Process
Victor Onditi, Paul Rayson, B. Ransom, Devina Ramduny-Ellis, Ian Sommerville, Alan J. Dix
NLDB2
2000 The REVERE Project: Experiments with the Application of Probabilistic NLP to Systems Engineering
Paul Rayson, Luke Emmet, Roger Garside, Peter Sawyer
NLDB1