Stephen Wan 0001

dblp:w/StephenWan · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0001-7505-1417ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 PEQQS: a Dataset for Probing Extractive Quantity-focused Question Answering from Scientific Literature
abstract
Question Answering (QA) and Information Retrieval (IR) play a crucial role in information-seeking pipelines implemented in many emerging AI research assistant applications. Large Language Models (LLMs) have demonstrated exceptional effectiveness on QA tasks, with Retrieval Augmented Generation (RAG) techniques often boosting the results. However, in many of those emerging applications, the onus of conducting the actual literature search falls on the user, i.e. the user searches for the relevant literature and the LLM-based assistant extracts the solicited answers from each of the user-supplied documents. The interplay between the quality of the user-conducted search and the quality of the final results remains understudied.
Maciej Rybinski, Necva Bölücü, Huichen Yang, Stephen Wan 0001
CIKM4
2025 A Position Paper on the Automatic Generation of Machine Learning Leaderboards
abstract
An important task in machine learning (ML) research is comparing prior work, which is often performed via ML leaderboards: a tabular overview of experiments with comparable conditions (e.g., same task, dataset, and metric).However, the growing volume of literature creates challenges in creating and maintaining these leaderboards.To ease this burden, researchers have developed methods to extract leaderboard entries from research papers for automated leaderboard curation.Yet, prior work varies in problem framing, complicating comparisons and limiting real-world applicability.In this position paper, we present the first overview of Automatic Leaderboard Generation (ALG) research, identifying fundamental differences in assumptions, scope, and output formats.We propose an ALG unified conceptual framework to standardise how the ALG task is defined.We offer ALG benchmarking guidelines, including recommendations for datasets and metrics that promote fair, reproducible evaluation.Lastly, we outline challenges and new directions for ALG, such as, advocating for broader coverage by including all reported results and richer metadata.
Roelien C. Timmer, Yufang Hou 0001, Stephen Wan 0001
EMNLP3
2025 Abductive Computational Systems: Creative Abduction and Future Directions
Abhinav Sood, Kazjon Grace, Stephen Wan 0001, Cécile Paris
ICCC3
2024 Detecting Online Community Practices with Large Language Models: A Case Study of Pro-Ukrainian Publics on Twitter
abstract
Communities on social media display distinct patterns of linguistic expression and behaviour, collectively referred to as practices.These practices can be traced in textual exchanges, and reflect the intentions, knowledge, values, and norms of users and communities.This paper introduces a comprehensive methodological workflow for computational identification of such practices within social media texts.By focusing on supporters of Ukraine during the Russia-Ukraine war in (1) the activist collective NAFO and (2) the Eurovision Twitter community, we present a gold-standard data set capturing their unique practices.Using this corpus, we perform practice prediction experiments with both open-source baseline models and OpenAI's large language models.Our results demonstrate that closed-source models, especially GPT-4, achieve superior performance, particularly with prompts that incorporate salient features of practices, or utilize Chain-of-Thought prompting.This study provides a detailed error analysis and offers valuable insights into improving the precision of practice identification, thereby supporting context-sensitive moderation and advancing the understanding of online community dynamics.1 communities.In
Kateryna Kasianenko, Shima Khanehzar, Stephen Wan 0001, Ehsan Dehghan, Axel Bruns
EMNLP3
2024 What Causes the Failure of Explicit to Implicit Discourse Relation Recognition?
abstract
Wei Liu, Stephen Wan, Michael Strube. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wei Liu 0145, Stephen Wan 0001, Michael Strube 0001
NAACL-HLT2
2024 An adaptive approach to noisy annotations in scientific information extraction
Necva Bölücü, Maciej Rybinski, Xiang Dai 0001, Stephen Wan 0001
Inf. Process. Manag.4
2023 Rethinking the Role of Entity Type in Relation Classification
abstract
Xiang Dai, Sarvnaz Karimi, Stephen Wan. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xiang Dai 0001, Sarvnaz Karimi, Stephen Wan 0001
IJCNLP (1)3
2023 SciHarvester: Searching Scientific Documents for Numerical Values
abstract
A challenge for search technologies is to support scientific literature surveys that present overviews of the reported numerical values documented for specific physical properties. We present SciHarvester, a system tailored to address this problem for agronomic science. It provides an interface to search PubAg documents, allowing complex queries involving restrictions on numerical values. SciHarvester identifies relevant documents and generates overview of reported parameter values. The system allows interrogation of the results to explain the system's performance. Our evaluations demonstrate the promise of incorporating information extraction techniques with the use of neural scoring mechanisms.
Maciej Rybinski, Stephen Wan 0001, Sarvnaz Karimi, Cécile Paris, Brian Jin, Neil I. Huth, Peter J. Thorburn, Dean P. Holzworth
SIGIR2
2021 Mention Flags (MF): Constraining Transformer-based Text Generators
abstract
Yufei Wang, Ian Wood, Stephen Wan, Mark Dras, Mark Johnson. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yufei Wang 0003, Ian D. Wood, Stephen Wan 0001, Mark Dras, Mark Johnson 0001
ACL/IJCNLP (1)3
2021 ECOL-R: Encouraging Copying in Novel Object Captioning with Reinforcement Learning
abstract
Novel Object Captioning is a zero-shot Image Captioning task requiring describing objects not seen in the training captions, but for which information is available from external object detectors.The key challenge is to select and describe all salient detected novel objects in the input images.In this paper, we focus on this challenge and propose the ECOL-R model (Encouraging Copying of Object Labels with Reinforced Learning), a copy-augmented transformer model that is encouraged to accurately describe the novel object labels.This is achieved via a specialised reward function in the SCST reinforcement learning framework (Rennie et al., 2017) that encourages novel object mentions while maintaining the caption quality.We further restrict the SCST training to the images where detected objects are mentioned in reference captions to train the ECOL-R model.We additionally improve our copy mechanism via Abstract Labels, which transfer knowledge from known to novel object types, and a Morphological Selector, which determines the appropriate inflected forms of novel object labels.The resulting model sets new state-of-the-art on the nocaps (Agrawal et al., 2019) and held-out COCO (Hendricks et al., 2016) benchmarks.
Yufei Wang 0003, Ian D. Wood, Stephen Wan 0001, Mark Johnson 0001
EACL3
2021 Integrating Lexical Information into Entity Neighbourhood Representations for Relation Prediction
abstract
Relation prediction informed from a combination of text corpora and curated knowledge bases, combining knowledge graph completion with relation extraction, is a relatively little studied task.A system that can perform this task has the ability to extend an arbitrary set of relational database tables with information extracted from a document corpus.OpenKi (Zhang et al., 2019) addresses this task through extraction of named entities and predicates via OpenIE tools then learning relation embeddings from the resulting entityrelation graph for relation prediction, outperforming previous approaches.We present an extension of OpenKi that incorporates embeddings of text-based representations of the entities and the relations.We demonstrate that this results in a substantial performance increase over a system without this information.
Ian D. Wood, Mark Johnson 0001, Stephen Wan 0001
NAACL-HLT3
2021 Neural Rule-Execution Tracking Machine For Transformer-Based Text Generation
abstract
Sequence-to-Sequence (Seq2Seq) neural text generation models, especially the pre-trained ones (e.g., BART and T5), have exhibited compelling performance on various natural language generation tasks. However, the black-box nature of these models limits their application in tasks where specific rules (e.g., controllable constraints, prior knowledge) need to be executed. Previous works either design specific model structures (e.g., Copy Mechanism corresponding to the rule "the generated output should include certain words in the source input'') or implement specialized inference algorithms (e.g., Constrained Beam Search) to execute particular rules through the text generation. These methods require the careful design case-by-case and are difficult to support multiple rules concurrently. In this paper, we propose a novel module named Neural Rule-Execution Tracking Machine (NRETM) that can be equipped into various transformer-based generators to leverage multiple rules simultaneously to guide the neural generation model for superior generation performance in an unified and scalable way. Extensive experiments on several benchmarks verify the effectiveness of our proposed model in both controllable and general text generation tasks.
Yufei Wang 0003, Can Xu 0002, Huang Hu, Chongyang Tao, Stephen Wan 0001, Mark Dras, Mark Johnson 0001, Daxin Jiang
NeurIPS5
2020 'Watch the Flu': A Tweet Monitoring Tool for Epidemic Intelligence of Influenza in Australia
abstract
‘Watch The Flu’ is a tool that monitors tweets posted in Australia for symptoms of influenza. The tool is a unique combination of two areas of artificial intelligence: natural language processing and time series monitoring, in order to assist public health surveillance. Using a real-time data pipeline, it deploys a web-based dashboard for visual analysis, and sends out emails to a set of users when an outbreak is detected. We expect that the tool will assist public health experts with their decision-making for disease outbreaks, by providing them insights from social media.
Brian Jin, Aditya Joshi 0001, Ross Sparks, Stephen Wan 0001, Cécile Paris, C. Raina MacIntyre
AAAI4
2020 Social Media Relevance Filtering Using Perplexity-Based Positive-Unlabelled Learning
Sunghwan Mac Kim, Stephen Wan 0001, Cécile Paris, Andreas Dünser
ICWSM2
2020 Image Captioning using Facial Expression and Attention
abstract
Benefiting from advances in machine vision and natural language processing techniques, current image captioning systems are able to generate detailed visual descriptions. For the most part, these descriptions represent an objective characterisation of the image, although some models do incorporate subjective aspects related to the observer’s view of the image, such as sentiment; current models, however, usually do not consider the emotional content of images during the caption generation process. This paper addresses this issue by proposing novel image captioning models which use facial expression features to generate image captions. The models generate image captions using long short-term memory networks applying facial features in addition to other visual features at different time steps. We compare a comprehensive collection of image captioning models with and without facial features using all standard evaluation metrics. The evaluation metrics indicate that applying facial features with an attention mechanism achieves the best performance, showing more expressive and more correlated image captions, on an image caption dataset extracted from the standard Flickr 30K dataset, consisting of around 11K images containing faces. An analysis of the generated captions finds that, perhaps unexpectedly, the improvement in caption quality appears to come not from the addition of adjectives linked to emotional aspects of the images, but from more variety in the actions described in the captions.
Omid Mohamad Nezami, Mark Dras, Stephen Wan 0001, Cécile Paris
J. Artif. Intell. Res.3
2019 How to Best Use Syntax in Semantic Role Labelling
abstract
There are many different ways in which external information might be used in an NLP task.This paper investigates how external syntactic information can be used most effectively in the Semantic Role Labeling (SRL) task.We evaluate three different ways of encoding syntactic parses and three different ways of injecting them into a state-of-the-art neural ELMo-based SRL sequence labelling model.We show that using a constituency representation as input features improves performance the most, achieving a new state-of-the-art for non-ensemble SRL models on the in-domain CoNLL'05 and CoNLL'12 benchmarks. 1
Yufei Wang 0003, Mark Johnson 0001, Stephen Wan 0001, Yifang Sun, Wei Wang 0011
ACL (1)3
2019 Automatic Recognition of Student Engagement Using Deep Learning and Facial Expression
Omid Mohamad Nezami, Mark Dras, Leonard G. C. Hamey, Debbie Richards 0001, Stephen Wan 0001, Cécile Paris
ECML/PKDD (3)5
2019 Towards Generating Stylized Image Captions via Adversarial Training
Omid Mohamad Nezami, Mark Dras, Stephen Wan 0001, Cécile Paris, Leonard G. C. Hamey
PRICAI (1)3
2015 Understanding Public Emotional Reactions on Twitter
Stephen Wan 0001, Cécile Paris
ICWSM1
2014 Improving government services with social media feedback
abstract
Social media is an invaluable source of feedback not just about consumer products and services but also about the effectiveness of government services. Our aim is to help analysts identify how government services can be improved based on citizen-contributed feedback found in publicly available social media. We present ongoing research for a social media monitoring interactive prototype with federated search and text analysis functionality. The prototype, developed to fit the workflow of social media monitors in the government sector, collects, analyses, and provides overviews of social media content. It facilitates relevance judgements on specific social media posts to decide whether or not to engage online. Our user log analysis validates the original design requirements and indicates ongoing utility to our federated search approach.
Stephen Wan 0001, Cécile Paris
IUI1
2012 Differences in Language and Style Between Two Social Media Communities
Cécile Paris, Paul Thomas 0001, Stephen Wan 0001
ICWSM3
2010 Focused and aggregated search: a perspective from natural language generation
Cécile Paris, Stephen Wan 0001, Paul Thomas 0001
Inf. Retr.2
2010 Supporting browsing-specific information needs: Introducing the Citation-Sensitive In-Browser Summariser
Stephen Wan 0001, Cécile Paris, Robert Dale
J. Web Semant.1
2009 Improving Grammaticality in Statistical Sentence Generation: Introducing a Dependency Spanning Tree Algorithm with an Argument Satisfaction Model
Stephen Wan 0001, Mark Dras, Robert Dale, Cécile Paris
EACL1
2009 Capturing the User's Reading Context for Tailoring Summaries
Cécile Paris, Stephen Wan 0001
UMAP2
2008 Seed and Grow: Augmenting Statistically Generated Summary Sentences using Schematic Word Patterns
Stephen Wan 0001, Robert Dale, Mark Dras, Cécile Paris
EMNLP1
2007 GLEU: Automatic Evaluation of Sentence-Level Fluency
Andrew Mutton, Mark Dras, Stephen Wan 0001, Robert Dale
ACL3
2004 Generating Overview Summaries of Ongoing Email Thread Discussions
Stephen Wan 0001, Kathy McKeown
COLING1