Jey Han Lau

dblp:32/9014 · DBLP profile ↗
← Back
64ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0002-1647-4628ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 12 first-author · 30 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics
abstract
Measuring the quality of public deliberation requires evaluating not only civility or argument structure, but also the informational progress of a conversation.We introduce a framework for Conversational Information Gain (CIG) that evaluates each utterance in terms of how it advances collective understanding of the target topic.To operationalize CIG, we model an evolving semantic memory of the discussion: the system extracts atomic claims from utterances and incrementally consolidates them into a structured memory state.Using this memory, we score each utterance along three interpretable dimensions: Novelty, Relevance, and Implication Scope.We annotate 80 segments from two moderated deliberative settings (TV debates and community discussions) with these dimensions and show that memory-derived dynamics (e.g., the number of claim updates) correlate more strongly with human-perceived CIG than traditional heuristics such as utterance length or TF-IDF.We develop effective LLM-based CIG predictors paving the way for information-focused conversation quality analysis in dialogues and deliberative success. 1
Ming-Bin Chen, Jey Han Lau, Lea Frermann
ACL (1)2
2026 Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
abstract
Yanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau, Lea Frermann, Biaoyan Fang, Fajri Koto. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau, Lea Frermann, Biaoyan Fang, Fajri Koto
ACL (1)4
2025 Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases
abstract
Rena Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu, Zheng Yuan, Jey Han Lau. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Rena Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu 0001, Jey Han Lau
ACL (1)9
2025 WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks
abstract
Embeddings-as-a-Service (EaaS) is a service offered by large language model (LLM) developers to supply embeddings generated by LLMs.Previous research suggests that EaaS is prone to imitation attacks-attacks that clone the underlying EaaS model by training another model on the queried embeddings.As a result, EaaS watermarks are introduced to protect the intellectual property of EaaS providers.In this paper, we first show that existing EaaS watermarks can be removed by paraphrasing when attackers clone the model.Subsequently, we propose a novel watermarking technique that involves linearly transforming the embeddings, and show that it is empirically and theoretically robust against paraphrasing.1 Training Dataset ` Verification Dataset 0.4 0.6 0 0 0.4 0.6 0.6 0 0.4 0.57 -0.85 1.28 1.28 0.57 -0.85 -0.85 1.28 0.57 Secret Linear TransformationInverse Linear Transformation
Anudeex Shetty, Qiongkai Xu, Jey Han Lau
ACL (1)3
2025 Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations
abstract
We present an evaluation framework for interactive dialogue assessment in the context of English as a Second Language (ESL) speakers. Our framework collects dialogue-level interactivity labels (e.g., topic management; 4 labels in total) and micro-level span features (e.g., backchannels; 17 features in total). Given our annotated data, we study how the micro-level features influence the (higher level) interactivity quality of ESL dialogues by constructing various machine learning-based models. Our results demonstrate that certain micro-level features strongly correlate with interactivity quality, like reference words (e.g., she, her, he), revealing new insights about the interaction between higher-level dialogue quality and lower-level fundamental linguistic signals. Our framework also provides a means to assess ESL communication, which is useful for language assessment.
Rena Gao, Carsten Roever, Jey Han Lau
COLING3
2025 Factual Dialogue Summarization via Learning from Large Language Models
abstract
Factual consistency is an important quality in dialogue summarization. Large language model (LLM)-based automatic text summarization models generate more factually consistent summaries compared to those by smaller pretrained language models, but they face deployment challenges in real-world applications due to privacy or resource constraints. In this paper, we investigate the use of symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization. We employ zero-shot learning to extract symbolic knowledge from LLMs, generating both factually consistent (positive) and inconsistent (negative) summaries. We then apply two contrastive learning objectives on these summaries to enhance smaller summarization models. Experiments with BART, PEGASUS, and Flan-T5 indicate that our approach surpasses strong baselines that rely on complex data augmentation strategies. Our approach demonstrates improved factual consistency while preserving coherence, fluency, and relevance, as verified by both automatic evaluation metrics and human assessments. We provide access to the data and code to facilitate future research.
Rongxin Zhu, Jey Han Lau, Jianzhong Qi 0001
COLING2
2025 Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation
abstract
Knowledge Base Question Answering (KBQA) aims to answer user questions in natural language using rich human knowledge stored in large KBs.As current KBQA methods struggle with unseen knowledge base elements and their novel compositions at test time, we introduce SG-KBQA -a novel model that injects schema contexts into entity retrieval and logical form generation to tackle this issue.It exploits information about the semantics and structure of the knowledge base provided by schema contexts to enhance generalizability.We show that SG-KBQA achieves strong generalizability, outperforming state-of-the-art models on three commonly used benchmark datasets across a variety of test settings.Our source code is available at https://github. com/gaosx2000/SG_KBQA.
Shengxiang Gao, Jey Han Lau, Jianzhong Qi 0001
EMNLP2
2025 Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
abstract
Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance across various objective multimodal perception tasks, yet their application to subjective, emotionally nuanced domains, such as psychological analysis, remains largely unexplored. In this paper, we introduce PICK, a multi-step framework designed for Psychoanalytical Image Comprehension through hierarchical analysis and Knowledge injection with MLLMs, specifically focusing on the House-Tree-Person (HTP) Test, a psychological assessment test. First, we decompose drawings containing multiple instances into semantically meaningful sub-drawings, constructing a hierarchical representation that captures spatial structure and content across three levels: single-object level, multi-object level, and whole level. Next, we analyze these sub-drawings at each level with a targeted focus, extracting psychological or emotional insights from their visual cues. We also introduce an HTP knowledge base and design a feature extraction module, trained with reinforcement learning, to generate a psychological profile for single-object level analysis. This profile captures both holistic stylistic features and dynamic object-specific features (such as those of the house, tree, or person), correlating them with psychological states. Finally, we integrate these multi-faceted information to produce a well-informed assessment that aligns with expert-level reasoning. Our approach bridges the gap between MLLMs and specialized expert domains, offering a structured and interpretable framework for understanding human mental states through visual expression. Experimental results demonstrate that the proposed PICK significantly enhances the capability of MLLMs in psychological analysis. It is further validated as a general framework through extensions to emotion understanding tasks. Codes are released at https://github.com/YanbeiJiang/PICK.
Xueqi Ma, Yanbei Jiang, Sarah M. Erfani, James Bailey 0001, Weifeng Liu 0001, Krista A. Ehinger, Jey Han Lau
ACM Multimedia7
2025 WHoW: A Cross-domain Approach for Analysing Conversation Moderation
abstract
Ming-Bin Chen, Lea Frermann, Jey Han Lau. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ming-Bin Chen, Lea Frermann, Jey Han Lau
NAACL (Long Papers)3
2025 An Interpretable and Crosslingual Method for Evaluating Second-Language Dialogues
abstract
Rena Gao, Jingxuan Wu, Xuetong Wu, Carsten Roever, Jing Wu, Long Lv, Jey Han Lau. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Rena Gao, Xuetong Wu, Carsten Roever, Long Lv, Jingxuan Wu, Jey Han Lau
NAACL (Long Papers)7
2025 Evaluating Evidence Attribution in Generated Fact Checking Explanations
abstract
Rui Xing, Timothy Baldwin, Jey Han Lau. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Rui Xing 0002, Timothy Baldwin, Jey Han Lau
NAACL (Long Papers)3
2024 A Sentiment Consolidation Framework for Meta-Review Generation
abstract
Modern natural language generation systems with Large Language Models (LLMs) exhibit the capability to generate a plausible summary of multiple documents; however, it is uncertain if they truly possess the capability of information consolidation to generate summaries, especially on documents with opinionated information.We focus on meta-review generation, a form of sentiment summarisation for the scientific domain.To make scientific sentiment summarization more grounded, we hypothesize that human meta-reviewers follow a three-layer framework of sentiment consolidation to write meta-reviews.Based on the framework, we propose novel prompting methods for LLMs to generate meta-reviews and evaluation metrics to assess the quality of generated meta-reviews.Our framework is validated empirically as we find that prompting LLMs based on the framework -compared with prompting them with simple instructions -generates better metareviews.11 The code and annotated data are accessible at https: //github.com/oaimli/MetaReviewingLogic.
Jey Han Lau, Eduard H. Hovy
ACL (1)2
2024 KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
Yanbei Jiang, Krista A. Ehinger, Jey Han Lau
IJCAI3
2023 Compressed Heterogeneous Graph for Abstractive Multi-Document Summarization
abstract
Multi-document summarization (MDS) aims to generate a summary for a number of related documents. We propose HGSum — an MDS model that extends an encoder-decoder architecture to incorporate a heterogeneous graph to represent different semantic units (e.g., words and sentences) of the documents. This contrasts with existing MDS models which do not consider different edge types of graphs and as such do not capture the diversity of relationships in the documents. To preserve only key information and relationships of the documents in the heterogeneous graph, HGSum uses graph pooling to compress the input graph. And to guide HGSum to learn the compression, we introduce an additional objective that maximizes the similarity between the compressed graph and the graph constructed from the ground-truth summary during training. HGSum is trained end-to-end with the graph similarity and standard cross-entropy objectives. Experimental results over Multi-News, WCEP-100, and Arxiv show that HGSum outperforms state-of-the-art MDS models. The code for our model and experiments is available at: https://github.com/oaimli/HGSum.
Jianzhong Qi 0001, Jey Han Lau
AAAI3
2023 Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization
abstract
A series of datasets and models have been proposed for summaries generated for wellformatted documents such as news articles.Dialogue summaries, however, have been under explored.In this paper, we present the first dataset with fine-grained factual error annotations named DIASUMFACT.We define finegrained factual error detection as a sentencelevel multi-label classification problem, and we evaluate two state-of-the-art (SOTA) models on our dataset.Both models yield sub-optimal results, with a macro-averaged F1 score of around 0.25 over 6 error classes.We further propose an unsupervised model ENDERANKER via candidate ranking using pretrained encoder-decoder models.Our model performs on par with the SOTA models while requiring fewer resources.These observations confirm the challenges in detecting factual errors from dialogue summaries, which call for further studies, for which our dataset and results offer a solid foundation. 1Lilly: Wanna go out tonight? Marshall: can't :( money's low Lilly: my treat :) Marshall:
Rongxin Zhu, Jianzhong Qi 0001, Jey Han Lau
ACL (1)3
2023 Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective
abstract
Zijian Zhang, Chang Shu, Ya Xiao, Yuan Shen, Di Zhu, Youxin Chen, Jing Xiao, Jey Han Lau, Qian Zhang, Zheng Lu. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Ya Xiao 0006, Youxin Chen, Jing Xiao 0006, Jey Han Lau, Qian Zhang 0018, Zheng Lu 0002
EACL8
2023 The Next Chapter: A Study of Large Language Models in Storytelling
abstract
To enhance the quality of generated stories, recent story generation models have been investigating the utilization of higher-level attributes like plots or commonsense knowledge.The application of prompt-based learning with large language models (LLMs), exemplified by GPT-3, has exhibited remarkable performance in diverse natural language processing (NLP) tasks.This paper conducts a comprehensive investigation, utilizing both automatic and human evaluation, to compare the story generation capacity of LLMs with recent models across three datasets with variations in style, register, and length of stories.The results demonstrate that LLMs generate stories of significantly higher quality compared to other story generation models.Moreover, they exhibit a level of performance that competes with human authors, albeit with the preliminary observation that they tend to replicate real stories in situations involving world knowledge, resembling a form of plagiarism.
Zhuohan Xie, Trevor Cohn, Jey Han Lau
INLG3
2023 MetaTroll: Few-shot Detection of State-Sponsored Trolls with Transformer Adapters
abstract
State-sponsored trolls are the main actors of influence campaigns on social media and automatic troll detection is important to combat misinformation at scale. Existing troll detection models are developed based on training data for known campaigns (e.g. the influence campaign by Russia’s Internet Research Agency on the 2016 US Election), and they fall short when dealing with novel campaigns with new targets. We propose MetaTroll, a text-based troll detection model based on the meta-learning framework that enables high portability and parameter-efficient adaptation to new campaigns using only a handful of labelled samples for few-shot transfer. We introduce campaign-specific transformer adapters to MetaTroll to “memorise” campaign-specific knowledge so as to tackle catastrophic forgetting, where a model “forgets” how to detect trolls from older campaigns due to continual adaptation. Our experiments demonstrate that MetaTroll substantially outperforms baselines and state-of-the-art few-shot text classification models. Lastly, we explore simple approaches to extend MetaTroll to multilingual and multimodal detection. Source code for MetaTroll is available at: https://github.com/ltian678/metatroll-code.git
Xiuzhen Zhang 0001, Jey Han Lau
WWW3
2022 The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature
abstract
Although multi-document summarization (MDS) of the biomedical literature is a highly valuable task that has recently attracted substantial interest, evaluation of the quality of biomedical summaries lacks consistency and transparency.In this paper, using systematic reviews as an example of biomedical MDS, we examine the summaries generated by two current models in order to understand the deficiencies of existing evaluation approaches in the context of the challenges that arise in the MDS task.Based on this analysis, we propose a new approach to human evaluation and identify several challenges that must be overcome to develop effective biomedical MDS systems.
Yulia Otmakhova 0001, Karin Verspoor, Timothy Baldwin, Jey Han Lau
ACL (1)4
2022 One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
abstract
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, Sebastian Ruder. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, Sebastian Ruder
ACL (1)11
2022 An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue Generation
abstract
We study the interpretability issue of taskoriented dialogue systems in this paper.Previously, most neural-based task-oriented dialogue systems employ an implicit reasoning strategy that makes the model predictions uninterpretable to humans.To obtain a transparent reasoning process, we introduce neurosymbolic to perform explicit reasoning that justifies model decisions by reasoning chains.Since deriving reasoning chains requires multihop reasoning for task-oriented dialogues, existing neuro-symbolic approaches would induce error propagation due to the one-phase design.To overcome this, we propose a twophase approach that consists of a hypothesis generator and a reasoner.We first obtain multiple hypotheses, i.e., potential operations to perform the desired task, through the hypothesis generator.Each hypothesis is then verified by the reasoner, and the valid one is selected to conduct the final prediction.The whole system is trained by exploiting raw textual dialogues without using any reasoning chain annotations.Experimental studies on two public benchmark datasets demonstrate that the proposed approach not only achieves better results, but also introduces an interpretable decision process.
Shiquan Yang, Rui Zhang 0003, Sarah M. Erfani, Jey Han Lau
ACL (1)4
2022 Unsupervised Lexical Substitution with Decontextualised Embeddings
abstract
We propose a new unsupervised method for lexical substitution using pre-trained language models. Compared to previous approaches that use the generative capability of language models to predict substitutes, our method retrieves substitutes based on the similarity of contextualised and decontextualised word embeddings, i.e. the average contextual representation of a word in multiple contexts. We conduct experiments in English and Italian, and show that our method substantially outperforms strong baselines and establishes a new state-of-the-art without any explicit supervision or fine-tuning. We further show that our method performs particularly well at predicting low-frequency substitutes, and also generates a diverse list of substitute candidates, reducing morphophonetic or morphosyntactic biases induced by article-noun agreement.
Takashi Wada 0001, Timothy Baldwin, Yuji Matsumoto 0001, Jey Han Lau
COLING4
2022 LipKey: A Large-Scale News Dataset for Absent Keyphrases Generation and Abstractive Summarization
abstract
Summaries, keyphrases, and titles are different ways of concisely capturing the content of a document. While most previous work has released the datasets of keyphrases and summarization separately, in this work, we introduce LipKey, the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. We jointly use the three elements via multi-task training and training as joint structured inputs, in the context of document summarization. We find that including absent keyphrases and titles as additional context to the source document improves transformer-based summarization models.
Fajri Koto, Timothy Baldwin, Jey Han Lau
COLING3
2022 DUCK: Rumour Detection on Social Media by Modelling User and Comment Propagation Networks
abstract
Social media rumours, a form of misinformation, can mislead the public and cause significant economic and social disruption.Motivated by the observation that the user network -which captures who engages with a storyand the comment network -which captures how they react to it -provide complementary signals for rumour detection.In this paper, we propose DUCK (rumour detection with user and comment networks) for rumour detection on social media.We study how to leverage transformers and graph attention networks to jointly model the contents and the structure of social media conversations, as well as the network of users who engage in these conversations.Over four widely used benchmark rumour datasets in English and Chinese, we show that DUCK produces superior performance for detecting rumours, creating a new state-of-the-art.
Xiuzhen Zhang 0001, Jey Han Lau
NAACL-HLT3
2022 FFCI: A Framework for Interpretable Automatic Evaluation of Summarization
abstract
In this paper, we propose FFCI, a framework for fine-grained summarization evaluation that comprises four elements: faithfulness (degree of factual consistency with the source), focus (precision of summary content relative to the reference), coverage (recall of summary content relative to the reference), and inter-sentential coherence (document fluency between adjacent sentences). We construct a novel dataset for focus, coverage, and inter-sentential coherence, and develop automatic methods for evaluating each of the four dimensions of FFCI based on cross-comparison of evaluation metrics and model-based evaluation methods, including question answering (QA) approaches, semantic textual similarity (STS), next-sentence prediction (NSP), and scores derived from 19 pre-trained language models. We then apply the developed metrics in evaluating a broad range of summarization models across two datasets, with some surprising findings.
Fajri Koto, Timothy Baldwin, Jey Han Lau
J. Artif. Intell. Res.3
2021 Top-down Discourse Parsing via Sequence Labelling
abstract
We introduce a top-down approach to discourse parsing that is conceptually simpler than its predecessors (Kobayashi et al., 2020;Zhang et al., 2020).By framing the task as a sequence labelling problem where the goal is to iteratively segment a document into individual discourse units, we are able to eliminate the decoder and reduce the search space for splitting points.We explore both traditional recurrent models and modern pre-trained transformer models for the task, and additionally introduce a novel dynamic oracle for top-down parsing.Based on the Full metric, our proposed LSTM model sets a new state-of-the-art for RST parsing. 1
Fajri Koto, Jey Han Lau, Timothy Baldwin
EACL2
2021 Brief Description of COVID-SEE: The Scientific Evidence Explorer for COVID-19 Related Research
Karin Verspoor, Simon Suster, Yulia Otmakhova 0001, Shevon Mendis, Zenan Zhai, Biaoyan Fang, Jey Han Lau, Timothy Baldwin, Antonio Jimeno-Yepes, David Martínez 0001
ECIR (2)7
2021 IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization
abstract
We present INDOBERTWEET, the first largescale pretrained model for Indonesian Twitter that is trained by extending a monolinguallytrained Indonesian BERT model with additive domain-specific vocabulary.We focus in particular on efficient model adaptation under vocabulary mismatch, and benchmark different ways of initializing the BERT embedding layer for new word types.We find that initializing with the average BERT subword embedding makes pretraining five times faster, and is more effective than proposed methods for vocabulary adaptation in terms of extrinsic evaluation over seven Twitter-based datasets.1
Fajri Koto, Jey Han Lau, Timothy Baldwin
EMNLP (1)2
2021 UniMF: A Unified Framework to Incorporate Multimodal Knowledge Bases intoEnd-to-End Task-Oriented Dialogue Systems
abstract
Knowledge bases (KBs) are usually essential for building practical dialogue systems. Recently we have seen rapidly growing interest in integrating knowledge bases into dialogue systems. However, existing approaches mostly deal with knowledge bases of a single modality, typically textual information. As today's knowledge bases become abundant with multimodal information such as images, audios and videos, the limitation of existing approaches greatly hinders the development of dialogue systems. In this paper, we focus on task-oriented dialogue systems and address this limitation by proposing a novel model that integrates external multimodal KB reasoning with pre-trained language models. We further enhance the model via a novel multi-granularity fusion mechanism to capture multi-grained semantics in the dialogue history. To validate the effectiveness of the proposed model, we collect a new large-scale (14K) dialogue dataset MMDialKB, built upon multimodal KB. Both automatic and human evaluation results on MMDialKB demonstrate the superiority of our proposed framework over strong baselines.
Shiquan Yang, Rui Zhang 0003, Sarah M. Erfani, Jey Han Lau
IJCAI4
2021 Automatic Classification of Neutralization Techniques in the Narrative of Climate Change Scepticism
abstract
Neutralisation techniques, e.g.denial of responsibility and denial of victim, are used in the narrative of climate change scepticism to justify lack of action or to promote an alternative view.We collect manual annotations of neutralised techniques used in these texts, and explore semi-supervised models to automatically classify them.
Shraey Bhatia, Jey Han Lau, Timothy Baldwin
NAACL-HLT2
2021 Discourse Probing of Pretrained Language Models
abstract
Existing work on probing of pretrained language models (LMs) has predominantly focused on sentence-level syntactic tasks.In this paper, we introduce document-level discourse probing to evaluate the ability of pretrained LMs to capture document-level relations.We experiment with 7 pretrained LMs, 4 languages, and 7 discourse probing tasks, and find BART to be overall the best model at capturing discourse -but only in its encoder, with BERT performing surprisingly well as the baseline model.Across the different models, there are substantial differences in which layers best capture discourse information, and large disparities between models.
Fajri Koto, Jey Han Lau, Timothy Baldwin
NAACL-HLT2
2021 Grey-box Adversarial Attack And Defence For Sentiment Classification
abstract
Ying Xu, Xu Zhong, Antonio Jimeno Yepes, Jey Han Lau. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xu Zhong, Antonio Jimeno-Yepes, Jey Han Lau
NAACL-HLT4
2021 Rumour Detection via Zero-Shot Cross-Lingual Transfer Learning
Xiuzhen Zhang 0001, Jey Han Lau
ECML/PKDD (1)3
2020 Give Me Convenience and Give Her Death: Who Should Decide What Uses of NLP are Appropriate, and on What Basis?
abstract
As part of growing NLP capabilities, coupled with an awareness of the ethical dimensions of research, questions have been raised about whether particular datasets and tasks should be deemed off-limits for NLP research.We examine this question with respect to a paper on automatic legal sentencing from EMNLP 2019 which was a source of some debate, in asking whether the paper should have been allowed to be published, who should have been charged with making such a decision, and on what basis.We focus in particular on the role of data statements in ethically assessing research, but also discuss the topic of dual use, and examine the outcomes of similar debates in other scientific disciplines.
Kobi Leins, Jey Han Lau, Timothy Baldwin
ACL2
2020 IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
abstract
Although the Indonesian language is spoken by almost 200 million people and the 10th mostspoken language in the world, 1 it is under-represented in NLP research.Previous work on Indonesian has been hampered by a lack of annotated datasets, a sparsity of language resources, and a lack of resource standardization.In this work, we release the INDOLEM dataset comprising seven tasks for the Indonesian language, spanning morpho-syntax, semantics, and discourse.We additionally release INDOBERT, a new pre-trained language model for Indonesian, and evaluate it over INDOLEM, in addition to benchmarking it against existing resources.Our experiments show that INDOBERT achieves state-of-the-art performance over most of the tasks in INDOLEM.
Fajri Koto, Afshin Rahimi 0001, Jey Han Lau, Timothy Baldwin
COLING3
2020 Forget Me Not: Reducing Catastrophic Forgetting for Domain Adaptation in Reading Comprehension
abstract
The creation of large-scale open domain reading comprehension data sets in recent years has enabled the development of end-to-end neural comprehension models with promising results. To use these models for domains with limited training data, one of the most effective approach is to first pre-train them on large out-of-domain source data and then fine-tune them with the limited target data. The caveat of this is that after fine-tuning the comprehension models tend to perform poorly in the source domain, a phenomenon known as catastrophic forgetting. In this paper, we explore methods that reduce catastrophic forgetting during fine-tuning without assuming access to data from the source domain. We introduce new auxiliary penalty terms and observe the best performance when a combination of auxiliary penalty terms is used to regularise the fine-tuning process for adapting comprehension models. To test our methods, we develop and release 6 narrow domain data sets that can potentially be used as reading comprehension benchmarks.
Xu Zhong, Antonio Jimeno-Yepes, Jey Han Lau
IJCNN4
2020 Less Is More: Rejecting Unreliable Reviews for Product Question Answering
Xiuzhen Zhang 0001, Jey Han Lau, Jeffrey Chan, Cécile Paris
ECML/PKDD (3)3
2020 How Furiously Can Colourless Green Ideas Sleep? Sentence Acceptability in Context
abstract
We study the influence of context on sentence acceptability. First we compare the acceptability ratings of sentences judged in isolation, with a relevant context, and with an irrelevant context. Our results show that context induces a cognitive load for humans, which compresses the distribution of ratings. Moreover, in relevant contexts we observe a discourse coherence effect that uniformly raises acceptability. Next, we test unidirectional and bidirectional language models in their ability to predict acceptability ratings. The bidirectional models show very promising results, with the best model achieving a new state-of-the-art for unsupervised acceptability prediction. The two sets of experiments provide insights into the cognitive aspects of sentence processing and central issues in the computational modeling of text and discourse.
Jey Han Lau, Carlos Santos Armendariz, Matthew Purver, Shalom Lappin
Trans. Assoc. Comput. Linguistics1
2019 Discovering Relevant Reviews for Answering Product-Related Queries
abstract
With the increasing popularity of e-commerce, the number of product-related queries generated by customers is growing. Answering these queries manually in real time is infeasible, and so automatic question-answering systems can be immensely helpful. Product queries are, however, very different from open-domain questions: they tend to be product-specific and the answers they demand can be very subjective. Previous research suggests that reviews are a valuable resource for answering product queries, but a key challenge is the language mismatch between user queries and reviews. To address this, we propose two neural models that discover relevant reviews for answering product queries. We demonstrate that our best model produces strong performance, outperforming state-of-the-art systems by consistently finding the most relevant reviews for product queries.
Jey Han Lau, Xiuzhen Zhang 0001, Jeffrey Chan, Cécile Paris
ICDM2
2018 Deep-speare: A joint neural model of poetic language, meter and rhyme
abstract
In this paper, we propose a joint architecture that captures language, rhyme and meter for sonnet modelling.We assess the quality of generated poems using crowd and expert judgements.The stress and rhyme models perform very well, as generated poems are largely indistinguishable from human-written poems.Expert evaluation, however, reveals that a vanilla language model captures meter implicitly, and that machine-generated poems still underperform in terms of readability and emotion.Our research shows the importance expert evaluation for poetry generation, and that future research should look beyond rhyme/meter and focus on poetic language.
Jey Han Lau, Trevor Cohn, Timothy Baldwin, Julian Brooke, Adam Hammond
ACL (1)1
2018 Document Chunking and Learning Objective Generation for Instruction Design
Khoi-Nguyen Tran, Jey Han Lau, Danish Contractor, Bikram Sengupta, Christopher J. Butler, Mukesh K. Mohania
EDM2
2018 Topic Intrusion for Automatic Topic Model Evaluation
abstract
Topic coherence is increasingly being used to evaluate topic models and filter topics for enduser applications.Topic coherence measures how well topic words relate to each other, but offers little insight into the utility of the topics in describing the documents.In this paper, we explore the topic intrusion task -the task of guessing an outlier topic given a document and a set of topics -and propose a method to automate it.We improve upon the state-of-the-art substantially, demonstrating its viability as an alternative method for topic model evaluation.
Shraey Bhatia, Jey Han Lau, Timothy Baldwin
EMNLP2
2018 Duplicate Detection in Programming Question Answering Communities
abstract
Community-based Question Answering (CQA) websites are attracting increasing numbers of users and contributors in recent years. However, duplicate questions frequently occur in CQA websites and are currently manually identified by the moderators. Automatic duplicate detection, on one hand, alleviates this laborious effort for moderators before taking close actions, and, on the other hand, helps question issuers quickly find answers. A number of studies have looked into related problems, but very limited works target Duplicate Detection in Programming CQA (PCQA), a branch of CQA that is dedicated to programmers. Existing works framed the task as a supervised learning problem on the question pairs and relied on only textual features. Moreover, the issue of selecting candidate duplicates from large volumes of historical questions is often un-addressed. To tackle these issues, we model duplicate detection as a two-stage “ranking-classification” problem over question pairs. In the first stage, we rank the historical questions according to their similarities to the newly issued question and select the top ranked ones as candidates to reduce the search space. In the second stage, we develop novel features that capture both textual similarity and latent semantics on question pairs, leveraging techniques in deep learning and information retrieval literature. Experiments on real-world questions about multiple programming languages demonstrate that our method works very well; in some cases, up to 25% improvement compared to the state-of-the-art benchmarks.
Wei Zhang 0098, Quan Z. Sheng, Jey Han Lau, Ermyas Abebe, Wenjie Ruan
ACM Trans. Internet Techn.3
2017 Topically Driven Neural Language Model
abstract
Language models are typically applied at the sentence level, without access to the broader document context.We present a neural language model that incorporates document context in the form of a topic model-like architecture, thus providing a succinct representation of the broader document context outside of the current sentence.Experiments over a range of datasets demonstrate that our model outperforms a pure sentence-based model in terms of language model perplexity, and leads to topics that are potentially more coherent than those produced by a standard LDA topic model.Our model also has the ability to generate related sentences for a topic, providing another way to interpret topics.
Jey Han Lau, Timothy Baldwin, Trevor Cohn
ACL (1)1
2017 An Automatic Approach for Document-level Topic Model Evaluation
abstract
Topic models jointly learn topics and document-level topic distribution.Extrinsic evaluation of topic models tends to focus exclusively on topic-level evaluation, e.g. by assessing the coherence of topics.We demonstrate that there can be large discrepancies between topic-and documentlevel model quality, and that basing model evaluation on topic-level analysis can be highly misleading.We propose a method for automatically predicting topic model quality based on analysis of documentlevel topic allocations, and provide empirical evidence for its robustness.
Shraey Bhatia, Jey Han Lau, Timothy Baldwin
CoNLL2
2017 End-to-end Network for Twitter Geolocation Prediction and Hashing
abstract
We propose an end-to-end neural network to predict the geolocation of a tweet. The network takes as input a number of raw Twitter metadata such as the tweet message and associated user account information. Our model is language independent, and despite minimal feature engineering, it is interpretable and capable of learning location indicative words and timing patterns. Compared to state-of-the-art systems, our model outperforms them by 2%-6%. Additionally, we propose extensions to the model to compress representation learnt by the network into binary codes. Experiments show that it produces compact codes compared to benchmark hashing algorithms. An implementation of the model is released publicly.
Jey Han Lau, Lianhua Chi, Khoi-Nguyen Tran, Trevor Cohn
IJCNLP(1)1
2017 Detecting Duplicate Posts in Programming QA Communities via Latent Semantics and Association Rules
abstract
Programming community-based question-answering (PCQA) websites such as Stack Overflow enable programmers to find working solutions to their questions. Despite detailed posting guidelines, duplicate questions that have been answered are frequently created. To tackle this problem, Stack Overflow provides a mechanism for reputable users to manually mark duplicate questions. This is a laborious effort, and leads to many duplicate questions remain undetected. Existing duplicate detection methodologies from traditional community based question-answering (CQA) websites are difficult to be adopted directly to PCQA, as PCQA posts often contain source code which is linguistically very different from natural languages. In this paper, we propose a methodology designed for the PCQA domain to detect duplicate questions. We model the detection as a classification problem over question pairs. To extract features for question pairs, our methodology leverages continuous word vectors from the deep learning literature, topic model features and phrases pairs that co-occur frequently in duplicate questions mined using machine translation systems. These features capture semantic similarities between questions and produce a strong performance for duplicate detection. Experiments on a range of real-world datasets demonstrate that our method works very well; in some cases over 30% improvement compared to state-of-the-art benchmarks. As a product of one of the proposed features, the association score feature, we have mined a set of associated phrases from duplicate questions on Stack Overflow and open the dataset to the public.
Wei Zhang 0098, Quan Z. Sheng, Jey Han Lau, Ermyas Abebe
WWW3
2017 Evaluating topic representations for exploring document collections
abstract
Topic models have been shown to be a useful way of representing the content of large document collections, for example, via visualization interfaces (topic browsers). These systems enable users to explore collections by way of latent topics. A standard way to represent a topic is using a term list; that is the top‐n words with highest conditional probability within the topic. Other topic representations such as textual and image labels also have been proposed. However, there has been no comparison of these alternative representations. In this article, we compare 3 different topic representations in a document retrieval task. Participants were asked to retrieve relevant documents based on predefined queries within a fixed time limit, presenting topics in one of the following modalities: (a) lists of terms, (b) textual phrase labels, and (c) image labels. Results show that textual labels are easier for users to interpret than are term lists and image labels. Moreover, the precision of retrieved documents for textual and image labels is comparable to the precision achieved by representing topics using term lists, demonstrating that labeling methods are an effective alternative topic representation.
Nikolaos Aletras, Timothy Baldwin, Jey Han Lau, Mark Stevenson 0001
J. Assoc. Inf. Sci. Technol.3
2016 LexSemTm: A Semantic Dataset Based on All-words Unsupervised Sense Distribution Learning
Andrew Bennett, Timothy Baldwin, Jey Han Lau, Diana McCarthy, Francis Bond
ACL (1)3
2016 Automatic Labelling of Topics with Neural Embeddings
abstract
Topics generated by topic models are typically represented as list of terms. To reduce the cognitive overhead of interpreting these topics for end-users, we propose labelling a topic with a succinct phrase that summarises its theme or idea. Using Wikipedia document titles as label candidates, we compute neural embeddings for documents and words to select the most relevant labels for topics. Comparing to a state-of-the-art topic labelling system, our methodology is simpler, more efficient and finds better topic labels.
Shraey Bhatia, Jey Han Lau, Timothy Baldwin
COLING2
2016 The Sensitivity of Topic Coherence Evaluation to Topic Cardinality
abstract
©2016 Association for Computational Linguistics. When evaluating the quality of topics generated by a topic model, the convention is to score topic coherence - either manually or automatically - using the top-N topic words. This hyper-parameter N, or the cardinality of the topic, is often overlooked and selected arbitrarily. In this paper, we investigate the impact of this cardinality hyper-parameter on topic coherence evaluation. For two automatic topic coherence methodologies, we observe that the correlation with human ratings decreases systematically as the cardinality increases. More interestingly, we find that performance can be improved if the system scores and human ratings are aggregated over several topic cardinalities before computing the correlation. In contrast to the standard practice of using a fixed value of N (e.g. N = 5 or N = 10), our results suggest that calculating topic coherence over several different cardinalities and averaging results in a substantially more stable and robust evaluation. We release the code and the datasets used in this research, for reproducibility.1
Jey Han Lau, Timothy Baldwin
HLT-NAACL1
2015 Unsupervised Prediction of Acceptability Judgements
abstract
Jey Han Lau, Alexander Clark, Shalom Lappin. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jey Han Lau, Alexander Clark, Shalom Lappin
ACL (1)1
2015 TM 2015 - Topic Models: Post-Processing and Applications Workshop
abstract
The main objective of the workshop is to bring together researchers who are interested in applications of topic models and improving their output. Our goal is to create a broad platform for researchers to share ideas that could improve the usability and interpretation of topic models. We expect this will promote topic model applications in other research areas, making their use more effective.
Nikolaos Aletras, Jey Han Lau, Timothy Baldwin, Mark Stevenson 0001
CIKM2
2014 Learning Word Sense Distributions, Detecting Unattested Senses and Identifying Novel Senses Using Topic Models
abstract
Unsupervised word sense disambiguation (WSD) methods are an attractive approach to all-words WSD due to their non-reliance on expensive annotated data.Unsupervised estimates of sense frequency have been shown to be very useful for WSD due to the skewed nature of word sense distributions.This paper presents a fully unsupervised topic modelling-based approach to sense frequency estimation, which is highly portable to different corpora and sense inventories, in being applicable to any part of speech, and not requiring a hierarchical sense inventory, parsing or parallel text.We demonstrate the effectiveness of the method over the tasks of predominant sense learning and sense distribution acquisition, and also the novel tasks of detecting senses which aren't attested in the corpus, and identifying novel senses in the corpus which aren't captured in the sense inventory.
Jey Han Lau, Paul Cook, Diana McCarthy, Spandana Gella, Timothy Baldwin
ACL (1)1
2014 Measuring Gradience in Speakers' Grammaticality Judgements
Jey Han Lau, Alexander Clark, Shalom Lappin
CogSci1
2014 Novel Word-sense Identification
Paul Cook, Jey Han Lau, Diana McCarthy, Timothy Baldwin
COLING2
2014 Machine Reading Tea Leaves: Automatically Evaluating Topic Coherence and Topic Model Quality
abstract
Topic models based on latent Dirichlet al-location and related methods are used in a range of user-focused tasks including doc-ument navigation and trend analysis, but evaluation of the intrinsic quality of the topic model and topics remains an open research area. In this work, we explore the two tasks of automatic evaluation of single topics and automatic evaluation of whole topic models, and provide recom-mendations on the best strategy for per-forming the two tasks, in addition to pro-viding an open-source toolkit for topic and topic model evaluation. 1
Jey Han Lau, David Newman 0001, Timothy Baldwin
EACL1
2014 Automatic Detection and Language Identification of Multilingual Documents
abstract
Language identification is the task of automatically detecting the language(s) present in a document based on the content of the document. In this work, we address the problem of detecting documents that contain text from more than one language ( multilingual documents). We introduce a method that is able to detect that a document is multilingual, identify the languages present, and estimate their relative proportions. We demonstrate the effectiveness of our method over synthetic data, as well as real-world multilingual documents collected from the web.
Marco Lui, Jey Han Lau, Timothy Baldwin
Trans. Assoc. Comput. Linguistics2
2013 Unsupervised Word Class Induction for Under-resourced Languages: A Case Study on Indonesian
Meladel Mistica, Jey Han Lau, Timothy Baldwin
IJCNLP2
2012 On-line Trend Analysis with Topic Models: \#twitter Trends Detection Topic Model Online
Jey Han Lau, Nigel Collier, Timothy Baldwin
COLING1
2012 Bayesian Text Segmentation for Index Term Identification and Keyphrase Extraction
David Newman 0001, Nagendra Koilada, Jey Han Lau, Timothy Baldwin
COLING3
2012 Word Sense Induction for Novel Sense Detection
Jey Han Lau, Paul Cook, Diana McCarthy, David Newman 0001, Timothy Baldwin
EACL1
2011 Automatic Labelling of Topic Models
Jey Han Lau, Karl Grieser, David Newman 0001, Timothy Baldwin
ACL1
2010 Automatic Evaluation of Topic Coherence
David Newman 0001, Jey Han Lau, Karl Grieser, Timothy Baldwin
HLT-NAACL2