VLDB 2026 Research / reviewers in the wild / expert
Ian D. Wood
dblp:170/2706 · also Ian David Wood
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2024
0000-0002-6094-0358ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ConvoCache: Smart Re-Use of Chatbot Responses
Conor Atkins, Ian D. Wood, Mohamed Ali Kâafar, Hassan Jameel Asghar, Nardine Basta, Michal Kepkowski |
INTERSPEECH | 2 |
| 2023 | Those Aren't Your Memories, They're Somebody Else's: Seeding Misinformation in Chat Bot Memories
Conor Atkins, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Ian D. Wood, Mohamed Ali Kâafar |
ACNS (1) | 4 |
| 2023 | Exploring the Distinctive Tweeting Patterns of Toxic Twitter UsersabstractIn the pursuit of bolstering user safety, social media platforms deploy active moderation strategies, including content removal and user suspension. These measures target users engaged in discussions marked by hate speech or toxicity, often linked to specific keywords or hashtags. Nonetheless, the increasing prevalence of toxicity indicates that certain users adeptly circumvent these measures.This study examines consistently toxic users on Twitter (rebranded as X) Rather than relying on traditional methods based on specific topics or hashtags, we employ a novel approach based on patterns of toxic tweets, yielding deeper insights into their behavior.We analyzed 38 million tweets from the timelines of 12,148 Twitter users and identified the top 1,457 users who consistently exhibit toxic behavior, relying on metrics like the Gini index and Toxicity score. By comparing their posting patterns to those of non-consistently toxic users, we have uncovered distinctive temporal patterns, including contiguous activity spans, inter-tweet intervals (referred to as “Burstiness”), and churn analysis. These findings provide strong evidence for the existence of a unique tweeting pattern associated with toxic behavior on Twitter.Crucially, our methodology transcends Twitter and can be adapted to various social media platforms, facilitating the identification of consistently toxic users based on their posting behavior. This research contributes to ongoing efforts to combat online toxicity and offers insights for refining moderation strategies in the digital realm. We are committed to open research and will provide our code and data to the research community. Hina Qayyum, Muhammad Ikram 0001, Benjamin Zi Hao Zhao, Ian D. Wood, Nicolas Kourtellis, Mohamed Ali Kâafar |
IEEE Big Data | 4 |
| 2023 | On mission Twitter Profiles: A Study of Selective Toxic BehaviorabstractThe argument for persistent social media influence campaigns, often funded by malicious entities, is gaining traction. These entities utilize instrumented profiles to disseminate divisive content and disinformation, shaping public perception. Despite ample evidence of these instrumented profiles, few identification methods exist to locate them in the wild. To evade detection and appear genuine, small clusters of instrumented profiles engage in unrelated discussions, diverting attention from their true goals [34]. This strategic thematic diversity conceals their selective polarity towards certain topics and fosters public trust [49]. This study aims to characterize profiles potentially used for influence operations, termed “on-mission profiles,” relying solely on thematic content diversity within unlabeled data. Distinguishing this work is its focus on content volume and toxicity towards specific themes. Longitudinal data from 138K Twitter (rebranded as X) profiles and 293M tweets enables profiling based on theme diversity. High thematic diversity groups predominantly produce toxic content concerning specific themes, like politics, health, and news—classifying them as “on-mission” profiles. Using the identified on-mission” profiles, we design a classifier for unseen, unlabeled data. Employing a linear SVM model, we train and test it on an 80/20% split of the most diverse profiles. The classifier achieves a flawless 100% accuracy, facilitating the discovery of previously unknown “on-mission” profiles in the wild. Hina Qayyum, Muhammad Ikram 0001, Benjamin Zi Hao Zhao, Ian D. Wood, Nicolas Kourtellis, Mohamed Ali Kâafar |
IEEE Big Data | 4 |
| 2023 | Unintended Memorization and Timing Attacks in Named Entity Recognition ModelsabstractNamed entity recognition models (NER), are widely used for identifying named entities (e.g., individuals, locations, and other information) in text documents. Machine learning based NER models are increasingly being applied in privacy-sensitive applications that need automatic and scalable identification of sensitive information to redact text for data sharing. In this paper, we study the setting when NER models are available as a black-box service for identifying sensitive information in user documents and show that these models are vulnerable to membership inference on their training datasets. With updated pre-trained NER models from spaCy, we demonstrate two distinct membership attacks on these models. Our first attack capitalizes on unintended memorization in the NER's underlying neural network, a phenomenon NNs are known to be vulnerable to. Our second attack leverages a timing side-channel to target NER models that maintain vocabularies constructed from the training data. We show that different functional paths of words within the training dataset in contrast to words not previously seen have measurable differences in execution time. Revealing membership status of training samples has clear privacy implications. For example, in text redaction, sensitive words or phrases to be found and removed, are at risk of being detected in the training dataset. Our experimental evaluation includes the redaction of both password and health data, presenting both security risks and a privacy/regulatory issues. This is exacerbated by results that indicate memorization after only a single phrase. We achieved a 70% AUC in our first attack on a text redaction use-case. We also show overwhelming success in the second timing attack with an 99.23% AUC. Finally we discuss potential mitigation approaches to realize the safe use of NER models in light of the presented privacy and security implications of membership inference attacks. Rana Salal Ali, Benjamin Zi Hao Zhao, Hassan Jameel Asghar, Tham Nguyen, Ian D. Wood, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 5 |
| 2022 | How Not to Handle Keys: Timing Attacks on FIDO Authenticator PrivacyabstractThis paper presents a timing attack on the FIDO2 (Fast IDentity Online) authentication protocol that allows attackers to link user accounts stored in vulnerable authenticators, a serious privacy concern. FIDO2 is a new standard specified by the FIDO industry alliance for secure token online authentication. It complements the W3C WebAuthn specification by providing means to use a USB token or other authenticator (which holds the secret authenticating material and implements FIDO protocols) as a second factor during the authentication process. From a cryptographic perspective, the protocol is a simple challenge-response where the elliptic curve digital signature algorithm is used to sign challenges. To protect the privacy of the user the token uses unique key pairs per service. To accommodate for small memory, tokens use various techniques that make use of a special parameter called a key handle sent by the service to the token with which the token can securely produce an authentication key (through generation or decryption). We identify and analyse a vulnerability in the way the processing of key handles is implemented that allows attackers to remotely link user accounts on multiple services. We show that for vulnerable authenticators there is a difference between the time it takes to process a key handle for a different service but correct authenticator, and for a different authenticator but correct service. This difference can be used to perform a timing attack allowing an adversary to link user’s accounts across services. We present several real world examples of adversaries that are in a position to execute our attack and can benefit from linking accounts. We found that two of the eight hardware authenticators we tested were vulnerable despite FIDO level 1 certification, indicating a not insignificant problem. This vulnerability cannot be easily mitigated on authenticators because, for security reasons, they usually do not allow firmware updates. In addition, we show that due to the way existing browsers implement the WebAuthn standard, the attack can be executed remotely. However, we discuss countermeasures that can be implemented by browser providers to mitigate the remote form of the attack. Michal Kepkowski, Lucjan Hanzlik, Ian D. Wood, Mohamed Ali Kâafar |
Proc. Priv. Enhancing Technol. | 3 |
| 2021 | Mention Flags (MF): Constraining Transformer-based Text GeneratorsabstractYufei Wang, Ian Wood, Stephen Wan, Mark Dras, Mark Johnson. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yufei Wang 0003, Ian D. Wood, Stephen Wan 0001, Mark Dras, Mark Johnson 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | ECOL-R: Encouraging Copying in Novel Object Captioning with Reinforcement LearningabstractNovel Object Captioning is a zero-shot Image Captioning task requiring describing objects not seen in the training captions, but for which information is available from external object detectors.The key challenge is to select and describe all salient detected novel objects in the input images.In this paper, we focus on this challenge and propose the ECOL-R model (Encouraging Copying of Object Labels with Reinforced Learning), a copy-augmented transformer model that is encouraged to accurately describe the novel object labels.This is achieved via a specialised reward function in the SCST reinforcement learning framework (Rennie et al., 2017) that encourages novel object mentions while maintaining the caption quality.We further restrict the SCST training to the images where detected objects are mentioned in reference captions to train the ECOL-R model.We additionally improve our copy mechanism via Abstract Labels, which transfer knowledge from known to novel object types, and a Morphological Selector, which determines the appropriate inflected forms of novel object labels.The resulting model sets new state-of-the-art on the nocaps (Agrawal et al., 2019) and held-out COCO (Hendricks et al., 2016) benchmarks. Yufei Wang 0003, Ian D. Wood, Stephen Wan 0001, Mark Johnson 0001 |
EACL | 2 |
| 2021 | Integrating Lexical Information into Entity Neighbourhood Representations for Relation PredictionabstractRelation prediction informed from a combination of text corpora and curated knowledge bases, combining knowledge graph completion with relation extraction, is a relatively little studied task.A system that can perform this task has the ability to extend an arbitrary set of relational database tables with information extracted from a document corpus.OpenKi (Zhang et al., 2019) addresses this task through extraction of named entities and predicates via OpenIE tools then learning relation embeddings from the resulting entityrelation graph for relation prediction, outperforming previous approaches.We present an extension of OpenKi that incorporates embeddings of text-based representations of the entities and the relations.We demonstrate that this results in a substantial performance increase over a system without this information. Ian D. Wood, Mark Johnson 0001, Stephen Wan 0001 |
NAACL-HLT | 1 |
| 2019 | Incorporate User Representation for Personal Question Answer Selection Using Siamese NetworkabstractMany natural language questions are inherently subjective. They can not be answered properly if we do not know the personal preferences of the answerer. For example, "Do you like cats?" There is no "the only correct answer" to this question. To answer it, the model has to be able to capture the persona of the answerers. However, the users usually do not answer different questions with equal chance. Instead, while some are answered with a high frequency, others are hardly answered by anyone. To deal with this imbalanced sparsity in data, we first introduce a Siamese Network to capture the preferences patterns of the users. Then the model is ensembled with an additional dense layer to predict the answers of the users. Applying to an online dating dataset, our approach achieves a high accuracy of 78.7%. Zihao Qi, Dario Bertero, Ian D. Wood, Pascale Fung |
ICASSP | 3 |
| 2018 | A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set
Ian D. Wood, John P. McCrae, Vladimir Andryushechkin, Paul Buitelaar |
LREC | 1 |
| 2018 | Towards a Crowd-Sourced WordNet for Colloquial EnglishabstractPrinceton WordNet is one of the most widely-used resources for natural language processing, but is updated only infrequently and cannot keep up with the fast-changing usage of the English language on social media platforms such as Twitter.The Colloquial WordNet aims to provide an open platform whereby anyone can contribute, while still following the structure of WordNet.Many crowdsourced lexical resources often have significant quality issues, and as such care must be taken in the design of the interface to ensure quality.In this paper, we present the development of a platform that can be opened on the Web to any lexicographer who wishes to contribute to this resource and the lexicographic methodology applied by this interface. John P. McCrae, Ian D. Wood, Amanda Hicks |
GWC | 2 |
| 2018 | MixedEmotions: An Open-Source Toolbox for Multimodal Emotion AnalysisabstractRecently, there is an increasing tendency to embed functionalities for recognizing emotions from user-generated media content in automated systems such as call-centre operations, recommendations, and assistive technologies, providing richer and more informative user and content profiles. However, to date, adding these functionalities was a tedious, costly, and time-consuming effort, requiring identification and integration of diverse tools with diverse interfaces as required by the use case at hand. The MixedEmotions Toolbox leverages the need for such functionalities by providing tools for text, audio, video, and linked data processing within an easily integrable plug-and-play platform. These functionalities include: 1) for text processing: emotion and sentiment recognition; 2) for audio processing: emotion, age, and gender recognition; 3) for video processing: face detection and tracking, emotion recognition, facial landmark localization, head pose estimation, face alignment, and body pose estimation; and 4) for linked data: knowledge graph integration. Moreover, the MixedEmotions Toolbox is open-source and free. In this paper, we present this toolbox in the context of the existing landscape, and provide a range of detailed benchmarks on standard test-beds showing its state-of-the-art performance. Furthermore, three real-world use cases show its effectiveness, namely, emotion-driven smart TV, call center monitoring, and brand reputation analysis. Paul Buitelaar, Ian D. Wood, Sapna Negi, Mihael Arcan, John P. McCrae, Andrejs Abele, Cécile Robin, Vladimir Andryushechkin, Housam Ziad, Hesam Sagha, Maximilian Schmitt, Björn W. Schuller, J. Fernando Sánchez-Rada, Carlos Angel Iglesias, Carlos Navarro, Andreas Giefer, Nicolaus Heise, Vincenzo Masucci, Francesco A. Danza, Ciro Caterino, Pavel Smrz, Michal Hradis, Filip Povolný, Marek Klimes, Pavel Matejka, Giovanni Tummarello |
IEEE Trans. Multim. | 2 |
| 2017 | The Colloquial WordNet: Extending Princeton WordNet with Neologisms
John P. McCrae, Ian D. Wood, Amanda Hicks |
LDK | 2 |