EDBT 2026 Demo / reviewers in the wild / expert
Paul Rayson
dblp:96/5736 · also Paul Edward Rayson
· DBLP profile ↗
17ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0002-1257-2191ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 7 (1 first)Big Data, Cloud & Distributed Data Systems · 7Knowledge Engineering, Semantic Web & Information Systems · 2Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
Mo El-Haj, Paul Rayson |
IEEE Big Data | 2 |
| 2023 | Open-Source Thesaurus Development for Under-Resourced Languages: a Welsh Case Study
Nouran Khallaf, Elin Arfon, Mo El-Haj, Jonathan Morris, Dawn Knight, Paul Rayson, Tymaa Hammouda, Mustafa Jarrar |
LDK | 6 |
| 2023 | FinAraT5: A text to text model for financial Arabic text understanding and generation
Nadhem Zmandar, Mo El-Haj, Paul Rayson |
LDK | 3 |
| 2023 | A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers
Nadhem Zmandar, Mahmoud El-Haj, Paul Rayson |
NLDB | 3 |
| 2021 | Multilingual Financial Word Embeddings for Arabic, English and FrenchabstractNatural Language Processing is increasingly being applied to analyse the text of many different types of financial documents. For many tasks, it has been shown that standard language models and tools need to be adapted to the financial domain in order to properly represent domain specific vocabulary, styles and meanings. Previous work has almost exclusively focused on English financial text, so in this paper we describe the creation of novel financial word embeddings for three languages: English, French and Arabic. In order to evaluate the effectiveness of the embeddings, we started by evaluating the English embeddings on a sentiment analysis classification task using the existing FinancialPhrase dataset and show improved performance over a standard GloVe based model using convolutional neural networks. Nadhem Zmandar, Mahmoud El-Haj, Paul Rayson |
IEEE BigData | 3 |
| 2020 | MUMBO: MUlti-task Max-Value Bayesian Optimization
Henry B. Moss, David S. Leslie, Paul Rayson |
ECML/PKDD (3) | 3 |
| 2019 | CLEU - A Cross-language english-urdu corpus and benchmark for text reuse experimentsabstractText reuse is becoming a serious issue in many fields and research shows that it is much harder to detect when it occurs across languages. The recent rise in multi‐lingual content on the Web has increased cross‐language text reuse to an unprecedented scale. Although researchers have proposed methods to detect it, one major drawback is the unavailability of large‐scale gold standard evaluation resources built on real cases. To overcome this problem, we propose a cross‐language sentence/passage level text reuse corpus for the English‐Urdu language pair. The Cross‐Language English‐Urdu Corpus (CLEU) has source text in English whereas the derived text is in Urdu. It contains in total 3,235 sentence/passage pairs manually tagged into three categories that is near copy, paraphrased copy, and independently written. Further, as a second contribution, we evaluate the Translation plus Mono‐lingual Analysis method using three sets of experiments on the proposed dataset to highlight its usefulness. Evaluation results (f1=0.732 binary, f1=0.552 ternary classification) indicate that it is harder to detect cross‐language real cases of text reuse, especially when the language pairs have unrelated scripts. The corpus is a useful benchmark resource for the future development and assessment of cross‐language text reuse detection systems for the English‐Urdu language pair. Iqra Muneer, Muhammad Sharjeel, Muntaha Iqbal, Rao Muhammad Adeel Nawab, Paul Rayson |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2016 | lexiDB: A scalable corpus database management systemabstractlexiDB is a scalable corpus database management system designed to fulfill corpus linguistics retrieval queries on multi-billion-word multiply-annotated corpora. It is based on a distributed architecture that allows the system to scale out to support ever larger text collections. This paper presents an overview of the architecture behind lexiDB as well as a demonstration of its functionality. We present lexiDB's performance metrics based on the AWS (Amazon Web Services) infrastructure with two part-of-speech and semantically tagged billion word corpora: Historical Hansard and EEBO (Early English Books Online). Matthew Coole, Paul Rayson, John A. Mariani |
IEEE BigData | 2 |
| 2016 | Sampling labelled profile data for identity resolutionabstractIdentity resolution capability for social networking profiles is important for a range of purposes, from open-source intelligence applications to forming semantic web connections. Yet replication of research in this area is hampered by the lack of access to ground-truth data linking the identities of profiles from different networks. Almost all data sources previously used by researchers are no longer available, and historic datasets are both of decreasing relevance to the modern social networking landscape and ethically troublesome regarding the preservation and publication of personal data. We present and evaluate a method which provides researchers in identity resolution with easy access to a realistically-challenging labelled dataset of online profiles, drawing on four of the currently largest and most influential online social networks. We validate the comparability of samples drawn through this method and discuss the implications of this mechanism for researchers as well as potential alternatives and extensions. Matthew Edwards 0001, Stephen Wattam, Paul Rayson, Awais Rashid |
IEEE BigData | 3 |
| 2016 | Reversing the Polarity with Emoticons
Phoey Lee Teh, Paul Rayson, Irina Pak, Scott Piao, Seow Mei Yeng |
NLDB | 2 |
| 2015 | Scaling out for extreme scale corpus dataabstractMuch of the previous work in Big Data has focussed on numerical sources of information. However, with the `narrative turn' in many disciplines gathering pace and commercial organisations beginning to realise the value of their textual assets, natural language data is fast catching up as an exploitable source of information for decision making. With vast quantities of unstructured textual data on the web, in social media, and in newly digitised historical document archives, the 5Vs (Volume, Velocity, Variety, Value and Veracity) apply equally well, if not more so, to big textual data. Corpus linguistics, the computer-aided study of large collections of naturally occurring language data, has been dealing with big data for fifty years. Corpus linguistics methods impose complex requirements on the retrieval, annotation and analysis of text in terms of displaying narrow contexts for each occurrence of a word or linguistic feature being studied and counting co-occurrences with other words or features to determine significant patterns in language. This, coupled with the distribution of language features in accordance with Zipf's Law, poses complex challenges for data models and corpus software dealing with extreme scale language data. A related issue is the non-random nature of language and the `burstiness' of word occurrences, or what we might put in Big Data terms as a sixth `V' called Viscosity. We report experiments to examine and compare the capabilities of two No-SQL databases in clustered configurations for the indexing, retrieval and analysis of billion-word corpora, since this size is the current state-of-the-art in corpus linguistics. We find that modern DBMSs (Database Management Systems) are capable of handling this extreme scale corpus data set for simple queries but are limited when querying for more frequent words or more complex queries. Matthew Coole, Paul Rayson, John A. Mariani |
IEEE BigData | 2 |
| 2015 | Sentiment analysis tools should take account of the number of exclamation marks!!!abstractThere are various factors that affect the sentiment level expressed in textual comments. Capitalization of letters tends to mark something for attention and repeating of letters tends to strengthen the emotion. Emoticons are used to help visualize facial expressions which can affect understanding of text. In this paper, we show the effect of the number of exclamation marks used, via testing with twelve online sentiment tools. We present opinions gathered from 500 respondents towards "like" and "dislike" values, with a varying number of exclamation marks. Results show that only 20% of the online sentiment tools tested considered the number of exclamation marks in their returned scores. However, results from our human raters show that the more exclamation marks used for positive comments, the more they have higher "like" values than the same comments with fewer exclamations marks. Similarly, adding more exclamation marks for negative comments, results in a higher "dislike". Phoey Lee Teh, Paul Rayson, Irina Pak, Scott Piao |
iiWAS | 2 |
| 2014 | Dealing with heterogeneous big data when geoparsing historical corporaabstractIt has long been known that `variety' is one of the key challenges and opportunities of big data. This is especially true when we consider the variety of content in historical corpora resulting from large-scale digitisation activities. Collections such as Early English Books Online (EEBO) and the British Library 19th Century Newspapers are extremely large and heterogeneous data sources containing a variety of content in terms of time, location, topic, style and quality. The range of geographical locations referenced in these corpora poses a difficult challenge for state of the art geoparsing tools. In the context of our work on Spatial Humanities analyses, we present our solution for dealing with the variety and scale of these corpora. C. J. Rupp, Paul Rayson, Ian N. Gregory, Andrew Hardie, Amelia Joulain, Daniel Hartmann |
IEEE BigData | 2 |
| 2013 | Customising geoparsing and georeferencing for historical textsabstractIn order to better support the text mining of historical texts, we propose a combination of complementary techniques from Geographical Information Systems, computational and corpus linguistics. In previous work, we have described this as `visual gisting' to extract important themes from text and locate those themes on a map representing geographical information contained in the text. Here, we describe the steps that were found necessary to apply standard analysis and resolution tools to identify place names in a specific corpus of historical texts. This task is seen as an initial and prerequisite step for further analysis and comparison by combining the information we extract from a corpus with information from other sources, including other text corpora. The process is intended to support close reading of historical texts on a much larger scale by highlighting using exploratory and data-driven approaches which parts of the corpus warrant further close analysis. Our case study presented here is from a corpus of Lake District travel literature. We discuss the customisations that we have to make to existing tools to extract placename information and visualise it on a map. C. J. Rupp, Paul Rayson, Alistair Baron, Christopher Donaldson, Ian N. Gregory, Andrew Hardie, Patricia Murrieta-Flores |
IEEE BigData | 2 |
| 2012 | Understanding Actionable Knowledge in Social Media: BBC Question Time and Twitter, a Case Study
Maria Angela Ferrario, William Simm, Jon Whittle 0001, Paul Rayson, Maria Terzi, Jane M. Binner |
ICWSM | 4 |
| 2004 | Language Resources and Tools for Supporting the System Engineering Process
Victor Onditi, Paul Rayson, B. Ransom, Devina Ramduny-Ellis, Ian Sommerville, Alan J. Dix |
NLDB | 2 |
| 2000 | The REVERE Project: Experiments with the Application of Probabilistic NLP to Systems Engineering
Paul Rayson, Luke Emmet, Roger Garside, Peter Sawyer |
NLDB | 1 |