EDBT 2026 Demo / reviewers in the wild / expert
Paul Cook
dblp:61/7223
· DBLP profile ↗
35ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 5 first-author · 2 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multilingual Phishing Email Detection Using Lightweight Federated LearningabstractGiven the escalating global threat of phishing emails, it is imperative to develop effective solutions to mitigate their potentially devastating impacts on society. This study endeavours to construct a federated multilingual spam detection system employing logistic regression, specifically targeting English, French, and Russian emails. This is the first work to the best of our knowledge which considers a non-deep learning setting for federated learning, and combines federated learning with multilingual phishing detection. Evaluation of the models is based on accuracy metrics which are compared with a most frequent class baseline. Our findings indicate that an optimal configuration comprises 10 clients undergoing 100 epochs of training with 100 rounds of federated learning, resulting in superior performance. Notably, this approach significantly outperforms the baseline, achieving an accuracy of $89.46 \%$ compared to $70 \%$. Dakota Staples, Hung Cao, Saqib Hakak, Paul Cook |
PST | 4 |
| 2024 | WaCadie: Towards an Acadian French CorpusabstractCorpora are important assets within the natural language processing (NLP) and linguistics communities, as they allow the training of models and corpus-based studies of languages. However, corpora do not exist for many languages and language varieties, such as Acadian French. In this paper, we first show that off-the-shelf NLP systems perform more poorly on Acadian French than on standard French. An Acadian French corpus could, therefore, potentially be used to improve NLP models for this dialect. Then, leveraging web-as-corpus methodologies, specifically BootCaT, domain crawling, and social media scraping, we create three corpora of Acadian French. To evaluate these corpora, drawing on the linguistic literature on Acadian French, we propose 22 statistical corpus-based measures of the extent to which a corpus is Acadian French. We use these measures to compare these newly built corpora to known Acadian French text and find that all three corpora include some traces of Acadian French. Jeremy Robichaud, Paul Cook |
LREC/COLING | 2 |
| 2023 | A Comparison of Machine Learning Algorithms for Multilingual Phishing DetectionabstractWith phishing emails being a major problem worldwide which is only getting larger by the year, there needs to exist solutions to combat them as they can cause tremendous harm to society. This research aims to compare numerous machine learning models and transformers for multilingual spam detection using English, French, and Russian emails. We evaluate the models using accuracy as two of the three experiments have nearly balanced test data. Our results show that, on average, XLM-Roberta performs the best out of all of the tested models in terms of accuracy. Dakota Staples, Saqib Hakak, Paul Cook |
PST | 3 |
| 2022 | Leveraging a Bilingual Dictionary to Learn Wolastoqey Word RepresentationsabstractWord embeddings (Mikolov et al., 2013; Pennington et al., 2014) have been used to bolster the performance of natural language processing systems in a wide variety of tasks, including information retrieval (Roy et al., 2018) and machine translation (Qi et al., 2018). However, approaches to learning word embeddings typically require large corpora of running text to learn high quality representations. For many languages, such resources are unavailable. This is the case for Wolastoqey, also known as Passamaquoddy-Maliseet, an endangered low-resource Indigenous language. As there exist no large corpora of running text for Wolastoqey, in this paper, we leverage a bilingual dictionary to learn Wolastoqey word embeddings by encoding their corresponding English definitions into vector representations using pretrained English word and sequence representation models. Specifically, we consider representations based on pretrained word2vec (Mikolov et al., 2013), RoBERTa (Liu et al., 2019) and sentence-BERT (Reimers and Gurevych, 2019) models. We evaluate these embeddings in word prediction tasks focused on part-of-speech, animacy, and transitivity; semantic clustering; and reverse dictionary search. In all evaluations we demonstrate that approaches using these embeddings outperform task-specific baselines, without requiring any language-specific training or fine-tuning. Diego Bear, Paul Cook |
LREC | 2 |
| 2020 | Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi'kmaq Language ModellingabstractMi’kmaq is an Indigenous language spoken primarily in Eastern Canada. It is polysynthetic and low-resource. In this paper we consider a range of n-gram and RNN language models for Mi’kmaq. We find that an RNN language model, initialized with pre-trained fastText embeddings, performs best, highlighting the importance of sub-word information for Mi’kmaq language modelling. We further consider approaches to language modelling that incorporate cross-lingual word embeddings, but do not see improvements with these models. Finally we consider language models that operate over segmentations produced by SentencePiece — which include sub-word units as tokens — as opposed to word-level models. We see improvements for this approach over word-level language models, again indicating that sub-word modelling is important for Mi’kmaq language modelling. Jeremie Boudreau, Akankshya Patra, Ashima Suvarna, Paul Cook |
LREC | 4 |
| 2020 | Evaluating Approaches to Personalizing Language ModelsabstractIn this work, we consider the problem of personalizing language models, that is, building language models that are tailored to the writing style of an individual. Because training language models requires a large amount of text, and individuals do not necessarily possess a large corpus of their writing that could be used for training, approaches to personalizing language models must be able to rely on only a small amount of text from any one user. In this work, we compare three approaches to personalizing a language model that was trained on a large background corpus using a relatively small amount of text from an individual user. We evaluate these approaches using perplexity, as well as two measures based on next word prediction for smartphone soft keyboards. Our results show that when only a small amount of user-specific text is available, an approach based on priming gives the most improvement, while when larger amounts of user-specific text are available, an approach based on language model interpolation performs best. We carry out further experiments to show that these approaches to personalization outperform language model adaptation based on demographic factors. Milton King, Paul Cook |
LREC | 2 |
| 2020 | Evaluating Sub-word Embeddings in Cross-lingual ModelsabstractCross-lingual word embeddings create a shared space for embeddings in two languages, and enable knowledge to be transferred between languages for tasks such as bilingual lexicon induction. One problem, however, is out-of-vocabulary (OOV) words, for which no embeddings are available. This is particularly problematic for low-resource and morphologically-rich languages, which often have relatively high OOV rates. Approaches to learning sub-word embeddings have been proposed to address the problem of OOV words, but most prior work has not considered sub-word embeddings in cross-lingual models. In this paper, we consider whether sub-word embeddings can be leveraged to form cross-lingual embeddings for OOV words. Specifically, we consider a novel bilingual lexicon induction task focused on OOV words, for language pairs covering several language families. Our results indicate that cross-lingual representations for OOV words can indeed be formed from sub-word embeddings, including in the case of a truly low-resource morphologically-rich language. Ali Hakimi Parizi, Paul Cook |
LREC | 2 |
| 2019 | Building Personalized Language Models Through Language Model Interpolation
Milton King, Paul Cook |
CICLing (1) | 2 |
| 2018 | Android authorship attribution through string analysisabstractWith the rising popularity of Android mobile devices, the amount of malicious applications targeting the Android platform has been increasing tremendously. To mitigate the risk of malicious apps, there is a need for an automated system to detect these applications. Current detection techniques rely on the signatures of well-documented malware, and hence may not be able to detect new malware samples. Instead of generating signatures for malware samples themselves, in this work, we propose to develop a lightweight system that can generate signatures of malware writers by leveraging the string components present in their Android binaries. Using these author signatures, we can effectively detect a wide range of existing, as well as any new, malware samples generated by particular authors. The proposed system achieved 98%, 96%, and 71% accuracy over datasets of 1559 benign, 262 malicious, and 96 obfuscated Android applications, respectively. The string-based approach achieved 71% of accuracy compared to only 50% obtained with the existing Ding and Samadzadeh's system. Vaibhavi Kalgutkar, Natalia Stakhanova, Paul Cook, Alina Matyukhina |
ARES | 3 |
| 2018 | Towards Language Technology for Mi'kmaq
Anant Maheshwari, Léo Bouscarrat, Paul Cook |
LREC | 3 |
| 2016 | Determining the Multiword Expression Inventory of a Surprise LanguageabstractMuch previous research on multiword expressions (MWEs) has focused on the token- and type-level tasks of MWE identification and extraction, respectively. Such studies typically target known prevalent MWE types in a given language. This paper describes the first attempt to learn the MWE inventory of a “surprise” language for which we have no explicit prior knowledge of MWE patterns, certainly no annotated MWE data, and not even a parallel corpus. Our proposed model is trained on a treebank with MWE relations of a source language, and can be applied to the monolingual corpus of the surprise language to identify its MWE construction types. Bahar Salehi, Paul Cook, Timothy Baldwin |
COLING | 2 |
| 2016 | Evaluating a Topic Modelling Approach to Measuring Corpus Similarity
Richard Fothergill, Paul Cook, Timothy Baldwin |
LREC | 2 |
| 2016 | Classifying Out-of-vocabulary Terms in a Domain-Specific Social Media Corpus
SoHyun Park, Afsaneh Fazly, Annie Lee, Brandon Seibel, Wenjie Zi, Paul Cook |
LREC | 6 |
| 2015 | Cross-lingual Transfer for Unsupervised Dependency Parsing Without Parallel DataabstractCross-lingual transfer has been shown to produce good results for dependency parsing of resource-poor languages.Although this avoids the need for a target language treebank, most approaches have still used large parallel corpora.However, parallel data is scarce for low-resource languages, and we report a new method that does not need parallel data.Our method learns syntactic word embeddings that generalise over the syntactic contexts of a bilingual vocabulary, and incorporates these into a neural network parser.We show empirical improvements over a baseline delexicalised parser on both the CoNLL and Universal Dependency Treebank datasets.We analyse the importance of the source languages, and show that combining multiple source-languages leads to a substantial improvement. Long Duong, Trevor Cohn, Steven Bird, Paul Cook |
CoNLL | 4 |
| 2015 | A Neural Network Model for Low-Resource Universal Dependency ParsingabstractAccurate dependency parsing requires large treebanks, which are only available for a few languages. We propose a method that takes advantage of shared structure across languages to build a mature parser using less training data. We propose a model for learning a shared “univer-sal ” parser that operates over an inter-lingual continuous representation of lan-guage, along with language-specific map-ping components. Compared with super-vised learning, our methods give a con-sistent 8-10 % improvement across several treebanks in low-resource simulations. 1 Long Duong, Trevor Cohn, Steven Bird, Paul Cook |
EMNLP | 4 |
| 2015 | A Word Embedding Approach to Predicting the Compositionality of Multiword ExpressionsabstractThis paper presents the first attempt to use word embeddings to predict the compositionality of multiword expressions.We consider both single-and multi-prototype word embeddings.Experimental results show that, in combination with a back-off method based on string similarity, word embeddings outperform a method using count-based distributional similarity.Our best results are competitive with, or superior to, state-of-the-art methods over three standard compositionality datasets, which include two types of multiword expressions and two languages. Bahar Salehi, Paul Cook, Timothy Baldwin |
HLT-NAACL | 2 |
| 2014 | Learning Word Sense Distributions, Detecting Unattested Senses and Identifying Novel Senses Using Topic ModelsabstractUnsupervised word sense disambiguation (WSD) methods are an attractive approach to all-words WSD due to their non-reliance on expensive annotated data.Unsupervised estimates of sense frequency have been shown to be very useful for WSD due to the skewed nature of word sense distributions.This paper presents a fully unsupervised topic modelling-based approach to sense frequency estimation, which is highly portable to different corpora and sense inventories, in being applicable to any part of speech, and not requiring a hierarchical sense inventory, parsing or parallel text.We demonstrate the effectiveness of the method over the tasks of predominant sense learning and sense distribution acquisition, and also the novel tasks of detecting senses which aren't attested in the corpus, and identifying novel senses in the corpus which aren't captured in the sense inventory. Jey Han Lau, Paul Cook, Diana McCarthy, Spandana Gella, Timothy Baldwin |
ACL (1) | 2 |
| 2014 | Novel Word-sense Identification
Paul Cook, Jey Han Lau, Diana McCarthy, Timothy Baldwin |
COLING | 1 |
| 2014 | One Sense per Tweeter ... and Other Lexical Semantic Tales of TwitterabstractIn recent years, microblogs such as Twitter have emerged as a new communication channel.Twitter in particular has become the target of a myriad of content-based applications including trend analysis and event detection, but there has been little fundamental work on the analysis of word usage patterns in this text type.In this paper -inspired by the one-sense-perdiscourse heuristic of Gale et al. ( 1992) -we investigate user-level sense distributions, and detect strong support for "one sense per tweeter".As part of this, we construct a novel sense-tagged lexical sample dataset based on Twitter and a web corpus. Spandana Gella, Paul Cook, Timothy Baldwin |
EACL | 2 |
| 2014 | Using Distributional Similarity of Multi-way Translations to Predict Multiword Expression CompositionalityabstractWe predict the compositionality of multiword expressions using distributional similarity between each component word and the overall expression, based on translations into multiple languages.We evaluate the method over English noun compounds, English verb particle constructions and German noun compounds.We show that the estimation of compositionality is improved when using translations into multiple languages, as compared to simply using distributional similarity in the source language.We further find that string similarity complements distributional similarity. Bahar Salehi, Paul Cook, Timothy Baldwin |
EACL | 2 |
| 2014 | What Can We Get From 1000 Tokens? A Case Study of Multilingual POS Tagging For Resource-Poor Languagesabstractby stating that they required an external tag dictionary.We have corrected these inaccuracies to reflect their modest data requirements. Long Duong, Trevor Cohn, Karin Verspoor, Steven Bird, Paul Cook |
EMNLP | 5 |
| 2014 | Detecting Non-compositional MWE Components using WiktionaryabstractWe propose a simple unsupervised ap-proach to detecting non-compositional components in multiword expressions based on Wiktionary. The approach makes use of the definitions, synonyms and trans-lations in Wiktionary, and is applicable to any type of MWE in any language, assum-ing the MWE is contained in Wiktionary. Our experiments show that the proposed approach achieves higher F-score than state-of-the-art methods. 1 Bahar Salehi, Paul Cook, Timothy Baldwin |
EMNLP | 2 |
| 2014 | Text-Based Twitter User Geolocation PredictionabstractGeographical location is vital to geospatial applications like local search and event detection. In this paper, we investigate and improve on the task of text-based geolocation prediction of Twitter users. Previous studies on this topic have typically assumed that geographical references (e.g., gazetteer terms, dialectal words) in a text are indicative of its authors location. However, these references are often buried in informal, ungrammatical, and multilingual data, and are therefore non-trivial to identify and exploit. We present an integrated geolocation prediction framework and investigate what factors impact on prediction accuracy. First, we evaluate a range of feature selection methods to obtain location indicative words. We then evaluate the impact of non-geotagged tweets, language, and user-declared metadata on geolocation prediction. In addition, we evaluate the impact of temporal variance on model generalisation, and discuss how users differ in terms of their geolocatability. We achieve state-of-the-art results for the text-based Twitter user geolocation task, and also provide the most extensive exploration of the task to date. Our findings provide valuable insights into the design of robust, practical text-based geolocation prediction systems. Bo Han 0002, Paul Cook, Timothy Baldwin |
J. Artif. Intell. Res. | 2 |
| 2013 | How Noisy Social Media Text, How Diffrnt Social Media Sources?
Timothy Baldwin, Paul Cook, Marco Lui, Andrew MacKinlay |
IJCNLP | 2 |
| 2013 | Increasing the Quality and Quantity of Source Language Data for Unsupervised Cross-Lingual POS Tagging
Long Duong, Paul Cook, Steven Bird, Pavel Pecina |
IJCNLP | 2 |
| 2013 | Lexical normalization for social media textabstractTwitter provides access to large volumes of data in real time, but is notoriously noisy, hampering its utility for NLP. In this article, we target out-of-vocabulary words in short text messages and propose a method for identifying and normalizing lexical variants. Our method uses a classifier to detect lexical variants, and generates correction candidates based on morphophonemic similarity. Both word similarity and context are then exploited to select the most probable correction candidate for the word. The proposed method doesn't require any annotations, and achieves state-of-the-art performance over an SMS corpus and a novel dataset based on Twitter. Bo Han 0002, Paul Cook, Timothy Baldwin |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Geolocation Prediction in Social Media Data by Finding Location Indicative Words
Bo Han 0002, Paul Cook, Timothy Baldwin |
COLING | 2 |
| 2012 | A Support Platform for Event Detection using Social Intelligence
Timothy Baldwin, Paul Cook, Bo Han 0002, Aaron Harwood, Shanika Karunasekera, Masud Moshtaghi |
EACL | 2 |
| 2012 | Word Sense Induction for Novel Sense Detection
Jey Han Lau, Paul Cook, Diana McCarthy, David Newman 0001, Timothy Baldwin |
EACL | 2 |
| 2012 | Automatically Constructing a Normalisation Dictionary for Microblogs
Bo Han 0002, Paul Cook, Timothy Baldwin |
EMNLP-CoNLL | 2 |
| 2011 | Automatic identification of words with novel but infrequent senses
Paul Cook, Graeme Hirst |
PACLIC | 1 |
| 2011 | A Way with Words: Recent Advances in Lexical Theory and Analysis: A Festschrift for Patrick Hanks Gilles-Maurice de Schryver (editor) (Ghent University and University of the Western Cape)Kampala: Menha Publishers, 2010, vii+375 pp; ISBN 978-9970-10-101-6, €59.95abstractComputing Lexical RelationsChurch kicks off Part II by responding to an earlier position piece by Kilgarriff (2007).Church argues that we should not abandon the use of large, but noisy and unbalanced, corpora, and discusses tasks to which such corpora are better suited than cleaner and more balanced, but smaller, corpora. Paul Cook |
Comput. Linguistics | 1 |
| 2010 | Automatically Identifying Changes in the Semantic Orientation of Words
Paul Cook, Suzanne Stevenson |
LREC | 1 |
| 2010 | Automatically Identifying the Source Words of Lexical Blends in EnglishabstractNewly coined words pose problems for natural language processing systems because they are not in a system's lexicon, and therefore no lexical information is available for such words. A common way to form new words is lexical blending, as in cosmeceutical, a blend of cosmetic and pharmaceutical. We propose a statistical model for inferring a blend's source words drawing on observed linguistic properties of blends; these properties are largely based on the recognizability of the source words in a blend. We annotate a set of 1,186 recently coined expressions which includes 515 blends, and evaluate our methods on a 324-item subset. In this first study of novel blends we achieve an accuracy of 40% on the task of inferring a blend's source words, which corresponds to a reduction in error rate of 39% over an informed baseline. We also give preliminary results showing that our features for source word identification can be used to distinguish blends from other kinds of novel words. Paul Cook, Suzanne Stevenson |
Comput. Linguistics | 1 |
| 2009 | Unsupervised Type and Token Identification of Idiomatic ExpressionsabstractIdiomatic expressions are plentiful in everyday language, yet they remain mysterious, as it is not clear exactly how people learn and understand them. They are of special interest to linguists, psycholinguists, and lexicographers, mainly because of their syntactic and semantic idiosyncrasies as well as their unclear lexical status. Despite a great deal of research on the properties of idioms in the linguistics literature, there is not much agreement on which properties are characteristic of these expressions. Because of their peculiarities, idiomatic expressions have mostly been overlooked by researchers in computational linguistics. In this article, we look into the usefulness of some of the identified linguistic properties of idioms for their automatic recognition. Specifically, we develop statistical measures that each model a specific property of idiomatic expressions by looking at their actual usage patterns in text. We use these statistical measures in a type-based classification task where we automatically separate idiomatic expressions (expressions with a possible idiomatic interpretation) from similar-on-the-surface literal phrases (for which no idiomatic interpretation is possible). In addition, we use some of the measures in a token identification task where we distinguish idiomatic and literal usages of potentially idiomatic expressions in context. Afsaneh Fazly, Paul Cook, Suzanne Stevenson |
Comput. Linguistics | 2 |