Danushka Bollegala

dblp:b/DanushkaBollegala · DBLP profile ↗
← Back
125ranked-venue papers
43as first author
41since 2021 · last 2026
0000-0003-4476-7003ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 101 · 29 first-author · 37 since 2021Databases, data management, data science and information retrieval · 24 · 14 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 11 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 2 since 2021Systems, architecture and hardware · 4 · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A Multilingual Social Bias Benchmark Incorporating Thinking Processes
abstract
Large Language Models (LLMs) can learn both useful knowledge and harmful stereotypes, making bias evaluation essential.Existing frameworks fall into two types: those considering reasoning steps (Thinking Process-Aware Evaluation, TPAE) and those focusing only on final outputs (Straight-to-the-Answer Evaluation, SAE).Prior TPAE studies showed effectiveness in assessing gender bias but relied on template-based, word-counting prompts, limiting generalization to other bias types, languages, and reasoning-based methods.In this study, we introduce MBTP 1 , a multilingual social bias benchmark that incorporates humangenerated pro-and anti-stereotype reasoning as part of the thinking process, and propose a few-shot meta-evaluation method that enables scalable bias assessment without model finetuning.From experiments evaluating 13 social bias categories across 8 languages, we find that human-generated thinking consistently yields higher-quality evaluations than LLM-generated or template-based approaches.Furthermore, TPAE demonstrates superior performance over SAE, highlighting the importance of considering reasoning processes in bias evaluation.Warning: This paper may contain offensive language or harmful content.
Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin
ACL (1)2
2026 Map of Encoders - Mapping Sentence Encoders using Quantum Relative Entropy
abstract
We propose a method to compare and visualise sentence encoders at scale by creating a map of encoders where each sentence encoder is represented in relation to the other sentence encoders. Specifically, we first represent a sentence encoder using an embedding matrix of a sentence set, where each row corresponds to the embedding of a sentence. Next, we compute the PIP matrix for a sentence encoder using its embedding matrix. Finally, we create a feature vector for each sentence encoder that reflects its QRE with respect to a unit base encoder. We construct a map of encoders covering 1101 publicly available sentence encoders, providing a new perspective of the landscape of the pre-trained sentence encoders. Our map accurately reflects various relationships between encoders, where encoders with similar attributes are proximally located on the map. Moreover, our encoder feature vectors can be used to accurately infer downstream task performance of the encoders, such as in retrieval and clustering tasks, demonstrating the correctness of our map.
Gaifan Zhang, Danushka Bollegala
ACL (1)2
2026 Synthetic Data Generation for Training Diversified Commonsense Reasoning Models
abstract
Conversational agents are required to respond to their users not only with high quality (i.e.commonsense-bearing) responses, but also considering multiple plausible alternative scenarios, reflecting the diversity in their responses.Despite the growing need to train diverse commonsense generators, the progress of this line of work has been significantly hindered by the lack of large-scale high-quality diverse commonsense training datasets.Due to the high annotation costs, existing Generative Commonsense Reasoning (GCR) datasets are created using a small number of human annotators, covering only a narrow set of commonsense scenarios.To address this training resource gap, we propose a two-stage method to create CommonSyn, the first-ever synthetic dataset for diversified GCR.Large Language Models (LLMs) fine-tuned on CommonSyn show simultaneous improvements in both generation diversity and quality compared with vanilla models and models fine-tuned on manually annotated datasets.1
Tianhui Zhang, Danushka Bollegala
ACL (1)3
2025 Evaluating the Evaluation of Diversity in Commonsense Generation
abstract
In commonsense generation, given a set of input concepts, a model must generate a response that is not only commonsense bearing, but also capturing multiple diverse viewpoints. Numerous evaluation metrics based on form- and content-level overlap have been proposed in prior work for evaluating the diversity of a commonsense generation model. However, it remains unclear as to which metrics are best suited for evaluating the diversity in commonsense generation. To address this gap, we conduct a systematic meta-evaluation of diversity metrics for commonsense generation. We find that form-based diversity metrics tend to consistently overestimate the diversity in sentence sets, where even randomly generated sentences are assigned overly high diversity scores. We then use an Large Language Model (LLM) to create a novel dataset annotated for the diversity of sentences generated for a commonsense generation task, and use it to conduct a meta-evaluation of the existing diversity evaluation metrics. Our experimental results show that content-based diversity evaluation metrics consistently outperform the form-based counterparts, showing high correlations with the LLM-based ratings. We recommend that future work on commonsense generation should use content-based metrics for evaluating the diversity of their outputs.
Tianhui Zhang, Danushka Bollegala
ACL (1)3
2025 Investigating the Contextualised Word Embedding Dimensions Specified for Contextual and Temporal Semantic Changes
abstract
The sense-aware contextualised word embeddings (SCWEs) encode semantic changes of words within the contextualised word embedding (CWE) spaces. Despite the superior performance of (SCWE) in contextual/temporal semantic change detection (SCD) benchmarks, it remains unclear as to how the meaning changes are encoded in the embedding space. To study this, we compare pre-trained CWEs and their fine-tuned versions on contextual and temporal semantic change benchmarks under Principal Component Analysis (PCA) and Independent Component Analysis (ICA) transformations. Our experimental results reveal (a) although there exist a smaller number of axes that are specific to semantic changes of words in the pre-trained CWE space, this information gets distributed across all dimensions when fine-tuned, and (b) in contrast to prior work studying the geometry of CWEs, we find that PCA to better represent semantic changes than ICA within the top 10% of axes. These findings encourage the development of more efficient SCD methods with a small number of SCD-aware dimensions.
Taichi Aida, Danushka Bollegala
COLING2
2025 The Gaps between Fine Tuning and In-context Learning in Bias Evaluation and Debiasing
abstract
The output tendencies of PLMs vary markedly before and after FT due to the updates to the model parameters. These divergences in output tendencies result in a gap in the social biases of PLMs. For example, there exits a low correlation between intrinsic bias scores of a PLM and its extrinsic bias scores under FT-based debiasing methods. Additionally, applying FT-based debiasing methods to a PLM leads to a decline in performance in downstream tasks. On the other hand, PLMs trained on large datasets can learn without parameter updates via ICL using prompts. ICL induces smaller changes to PLMs compared to FT-based debiasing methods. Therefore, we hypothesize that the gap observed in pre-trained and FT models does not hold true for debiasing methods that use ICL. In this study, we demonstrate that ICL-based debiasing methods show a higher correlation between intrinsic and extrinsic bias scores compared to FT-based methods. Moreover, the performance degradation due to debiasing is also lower in the ICL case compared to that in the FT case.
Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin
COLING2
2025 Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
abstract
Large language models (LLMs) exhibit social biases, prompting the development of various debiasing methods.However, debiasing methods may degrade the capabilities of LLMs.Previous research has evaluated the impact of bias mitigation primarily through tasks measuring general language understanding, which are often unrelated to social biases.In contrast, cultural commonsense is closely related to social biases, as both are rooted in social norms and values.The impact of bias mitigation on cultural commonsense in LLMs has not been well investigated.Considering this gap, we propose SOBACO (SOcial BiAs and Cultural cOmmonsense benchmark), a Japanese benchmark designed to evaluate social biases and cultural commonsense in LLMs in a unified format.We evaluate several LLMs on SOBACO to examine how debiasing methods affect cultural commonsense in LLMs.Our results reveal that the debiasing methods degrade the performance of the LLMs on the cultural commonsense task (up to 75% accuracy deterioration).These results highlight the importance of developing debiasing methods that consider the trade-off with cultural commonsense to improve fairness and utility of LLMs.Warning: This paper contains examples of social biases that can be offensive.
Taisei Yamamoto, Ryoma Kumon, Danushka Bollegala, Hitomi Yanaka
EMNLP3
2025 Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models
abstract
Semantic similarity between two sentences depends on the aspects considered between those sentences.To study this phenomenon, Deshpande et al. (2023) proposed the Conditional Semantic Textual Similarity (C-STS) task and annotated a human-rated similarity dataset containing pairs of sentences compared under two different conditions.However, Tu et al. (2024)found various annotation issues in this dataset and showed that manually re-annotating a small portion of it leads to more accurate C-STS models.Despite these pioneering efforts, the lack of large and accurately annotated C-STS datasets remains a blocker for making progress on this task as evidenced by the subpar performance of the C-STS models.To address this training data need, we resort to Large Language Models (LLMs) to correct the condition statements and similarity ratings in the original dataset proposed by Deshpande et al. (2023).Our proposed method is able to reannotate a large training dataset for the C-STS task with minimal manual effort.Importantly, by training a supervised C-STS model on our cleaned and re-annotated dataset, we achieve a 5.4% statistically significant improvement in Spearman correlation.The re-annotated dataset is available at https://LivNLP.github. io/CSTS-reannotation.
Gaifan Zhang, Yi Zhou 0019, Danushka Bollegala
EMNLP3
2025 Improving Unsupervised Constituency Parsing via Maximizing Semantic Information
abstract
Unsupervised constituency parsers organize phrases within a sentence into a tree-shaped syntactic constituent structure that reflects the organization of sentence semantics. However, the traditional objective of maximizing sentence log-likelihood (LL) does not explicitly account for the close relationship between the constituent structure and the semantics, resulting in a weak correlation between LL values and parsing accuracy. In this paper, we introduce a novel objective that trains parsers by maximizing SemInfo, the semantic information encoded in constituent structures. We introduce a bag-of-substrings model to represent the semantics and estimate the SemInfo value using the probability-weighted information metric. We apply the SemInfo maximization objective to training Probabilistic Context-Free Grammar (PCFG) parsers and develop a Tree Conditional Random Field (TreeCRF)-based model to facilitate the training. Experiments show that SemInfo correlates more strongly with parsing accuracy than LL, establishing SemInfo as a better unsupervised parsing objective. As a result, our algorithm significantly improves parsing accuracy by an average of 7.85 sentence-F1 scores across five PCFG variants and in four languages, achieving state-of-the-art level results in three of the four languages.
Xiangheng He, Yusuke Miyao, Danushka Bollegala
ICLR4
2025 An Ethical Dataset from Real-World Interactions Between Users and Large Language Models
abstract
Recent studies have demonstrated that Large Language Models (LLMs) have ethical-related problems such as social biases, lack of moral reasoning, and generation of offensive content. The existing evaluation metrics and methods to address these ethical challenges use datasets intentionally created by instructing humans to create instances including ethical problems. Therefore, the data does not sufficiently include comprehensive prompts that users actually provide when using LLM services in everyday contexts and outputs that LLMs generate. There may be different tendencies between unethical instances intentionally created by humans and actual user interactions with LLM services, which could result in a lack of comprehensive evaluation. To investigate the difference, we create Eagle datasets extracted from actual interactions between ChatGPT and users that exhibit social biases, opinion biases, toxicity, and immoral problems. Our experiments show that Eagle captures complementary aspects, not covered by existing datasets proposed for evaluation and mitigation. We argue that using both existing and proposed datasets leads to a more comprehensive assessment of the ethics.
Masahiro Kaneko, Danushka Bollegala, Timothy Baldwin
IJCAI2
2025 A Metric Differential Privacy Mechanism for Sentence Embeddings
abstract
Sentence embeddings represent the meaning of a given sentence using a fixed dimensional vector. Different approaches have been proposed in the Natural Language Processing (NLP) community for learning encoders that can produce accurate sentence embeddings that perform well for diverse downstream tasks requiring sentence representations. Despite prior work focusing mainly on creating accurate sentence embeddings, how to keep private the sensitive information contained in the sentences remains an unexplored research problem. In this article, we propose Covering Metric Analytic Gaussian (CMAG), a covering metric Differential Privacy (DP) mechanism for sentence embeddings such that minimal random noise is added to a set of sentence embeddings produced by an encoder to protect the private information expressed in those sentences. Given a sentence embedding s , CMAG considers the Mahalanobis distance between s and the other sentence embeddings s ’ in the local neighbourhood of s to determine the minimal amount of random noise that must be added to s to obtain provable metric DP guarantees. Experimental results show that the proposed DP mechanism protects private information better than previously proposed DP mechanisms while reporting good performance in a broad range of downstream NLP tasks.
Danushka Bollegala, Shuichi Otake, Tomoya Machide, Ken-ichi Kawarabayashi
ACM Trans. Priv. Secur.1
2024 Evaluating Unsupervised Dimensionality Reduction Methods for Pretrained Sentence Embeddings
abstract
Sentence embeddings produced by Pretrained Language Models (PLMs) have received wide attention from the NLP community due to their superior performance when representing texts in numerous downstream applications. However, the high dimensionality of the sentence embeddings produced by PLMs is problematic when representing large numbers of sentences in memory- or compute-constrained devices. As a solution, we evaluate unsupervised dimensionality reduction methods to reduce the dimensionality of sentence embeddings produced by PLMs. Our experimental results show that simple methods such as Principal Component Analysis (PCA) can reduce the dimensionality of sentence embeddings by almost 50%, without incurring a significant loss in performance in multiple downstream tasks. Surprisingly, reducing the dimensionality further improves performance over the original high dimensional versions for the sentence embeddings produced by some PLMs in some tasks.
Gaifan Zhang, Yi Zhou 0019, Danushka Bollegala
LREC/COLING3
2024 Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language Models
abstract
Social biases such as gender or racial biases have been reported in language models (LMs), including Masked Language Models (MLMs).Given that MLMs are continuously trained with increasing amounts of additional data collected over time, an important yet unanswered question is how the social biases encoded with MLMs vary over time.In particular, the number of social media users continues to grow at an exponential rate, and it is a valid concern for the MLMs trained specifically on social media data whether their social biases (if any) would also amplify over time.To empirically analyse this problem, we use a series of MLMs pretrained on chronologically ordered temporal snapshots of corpora.Our analysis reveals that, although social biases are present in all MLMs, most types of social bias remain relatively stable over time (with a few exceptions).To further understand the mechanisms that influence social biases in MLMs, we analyse the temporal corpora used to train the MLMs.Our findings show that some demographic groups, such as male, obtain higher preference over the other, such as female on the training corpora constantly.1
Yi Zhou 0019, Danushka Bollegala, José Camacho-Collados
EMNLP2
2024 Improving Pre-trained Language Model Sensitivity via Mask Specific losses: A case study on Biomedical NER
abstract
Micheal Abaho, Danushka Bollegala, Gary Leeming, Dan Joyce, Iain Buchan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Micheal Abaho, Danushka Bollegala, Gary Leeming, Dan W. Joyce, Iain E. Buchan
NAACL-HLT2
2024 Generating Character Relationship Maps for a Story
Taichi Uchino, Danushka Bollegala, Naiwala P. Chandrasiri
PACLIC2
2024 Community knowledge graph abstraction for enhanced link prediction: A study on PubMed knowledge graph
abstract
OBJECTIVE: As new knowledge is produced at a rapid pace in the biomedical field, existing biomedical Knowledge Graphs (KGs) cannot be manually updated in a timely manner. Previous work in Natural Language Processing (NLP) has leveraged link prediction to infer the missing knowledge in general-purpose KGs. Inspired by this, we propose to apply link prediction to existing biomedical KGs to infer missing knowledge. Although Knowledge Graph Embedding (KGE) methods are effective in link prediction tasks, they are less capable of capturing relations between communities of entities with specific attributes (Fanourakis et al., 2023). METHODS: To address this challenge, we proposed an entity distance-based method for abstracting a Community Knowledge Graph (CKG) from a simplified version of the pre-existing PubMed Knowledge Graph (PKG) (Xu et al., 2020). For link prediction on the abstracted CKG, we proposed an extension approach for the existing KGE models by linking the information in the PKG to the abstracted CKG. The applicability of this extension was proved by employing six well-known KGE models: TransE, TransH, DistMult, ComplEx, SimplE, and RotatE. Evaluation metrics including Mean Rank (MR), Mean Reciprocal Rank (MRR), and Hits@k were used to assess the link prediction performance. In addition, we presented a backtracking process that traces the results of CKG link prediction back to the PKG scale for further comparison. RESULTS: Six different CKGs were abstracted from the PKG by using embeddings of the six KGE methods. The results of link prediction in these abstracted CKGs indicate that our proposed extension can improve the existing KGE methods, achieving a top-10 accuracy of 0.69 compared to 0.5 for TransE, 0.7 compared to 0.54 for TransH, 0.67 compared to 0.6 for DistMult, 0.73 compared to 0.57 for ComplEx, 0.73 compared to 0.63 for SimplE, and 0.85 compared to 0.76 for RotatE on their CKGs, respectively. These improved performances also highlight the wide applicability of the extension approach. CONCLUSION: This study proposed novel insights into abstracting CKGs from the PKG. The extension approach indicated enhanced performance of the existing KGE methods and has applicability. As an interesting future extension, we plan to conduct link prediction for entities that are newly introduced to the PKG.
Danushka Bollegala, Shunsuke Hirose, Yingzi Jin, Tomotake Kozu
J. Biomed. Informatics2
2023 Learning Dynamic Contextualised Word Embeddings via Template-based Temporal Adaptation
abstract
Dynamic contextualised word embeddings (DCWEs) represent the temporal semantic variations of words.We propose a method for learning DCWEs by time-adapting a pretrained Masked Language Model (MLM) using timesensitive templates.Given two snapshots C 1 and C 2 of a corpus taken respectively at two distinct timestamps T 1 and T 2 , we first propose an unsupervised method to select (a) pivot terms related to both C 1 and C 2 , and (b) anchor terms that are associated with a specific pivot term in each individual snapshot.We then generate prompts by filling manually compiled templates using the extracted pivot and anchor terms.Moreover, we propose an automatic method to learn time-sensitive templates from C 1 and C 2 , without requiring any human supervision.Next, we use the generated prompts to adapt a pretrained MLM to T 2 by fine-tuning using those prompts.Multiple experiments show that our proposed method reduces the perplexity of test sentences in C 2 , outperforming the current state-of-the-art.
Xiaohang Tang, Yi Zhou 0019, Danushka Bollegala
ACL (1)3
2023 Evaluating the Robustness of Discrete Prompts
abstract
Discrete prompts have been used for finetuning Pre-trained Language Models for diverse NLP tasks.In particular, automatic methods that generate discrete prompts from a small set of training instances have reported superior performance.However, a closer look at the learnt prompts reveals that they contain noisy and counter-intuitive lexical constructs that would not be encountered in manuallywritten prompts.This raises an important yet understudied question regarding the robustness of automatically learnt discrete prompts when used in downstream tasks.To address this question, we conduct a systematic study of the robustness of discrete prompts by applying carefully designed perturbations into an application using AutoPrompt and then measure their performance in two Natural Language Inference (NLI) datasets.Our experimental results show that although the discrete prompt-based method remains relatively robust against perturbations to NLI inputs, they are highly sensitive to other types of perturbations such as shuffling and deletion of prompt tokens.Moreover, they generalize poorly across different NLI datasets.We hope our findings will inspire future work on robust discrete prompt learning.1
Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, Satoshi Nakamura 0001
EACL2
2023 Comparing Intrinsic Gender Bias Evaluation Measures without using Human Annotated Examples
abstract
Numerous types of social biases have been identified in pre-trained language models (PLMs), and various intrinsic bias evaluation measures have been proposed for quantifying those social biases.Prior works have relied on human annotated examples to compare existing intrinsic bias evaluation measures.However, this approach is not easily adaptable to different languages nor amenable to large scale evaluations due to the costs and difficulties when recruiting human annotators.To overcome this limitation, we propose a method to compare intrinsic gender bias evaluation measures without relying on human-annotated examples.Specifically, we create multiple bias-controlled versions of PLMs using varying amounts of male vs. female gendered sentences, mined automatically from an unannotated corpus using genderrelated word lists.Next, each bias-controlled PLM is evaluated using an intrinsic bias evaluation measure, and the rank correlation between the computed bias scores and the gender proportions used to fine-tune the PLMs is computed.Experiments on multiple corpora and PLMs repeatedly show that the correlations reported by our proposed method that does not require human annotated examples are comparable to those computed using human annotated examples in prior work.
Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki
EACL2
2023 A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models
abstract
Various types of social biases have been reported with pretrained Masked Language Models (MLMs) in prior work.However, multiple underlying factors are associated with an MLM such as its model size, size of the training data, training objectives, the domain from which pretraining data is sampled, tokenization, and languages present in the pretrained corpora, to name a few.It remains unclear as to which of those factors influence social biases that are learned by MLMs.To study the relationship between model factors and the social biases learned by an MLM, as well as the downstream task performance of the model, we conduct a comprehensive study over 39 pretrained MLMs covering different model sizes, training objectives, tokenization methods, training data domains and languages.Our results shed light on important factors often neglected in prior literature, such as tokenization or model objectives.
Yi Zhou 0019, José Camacho-Collados, Danushka Bollegala
EMNLP3
2023 Vis2Hap: Vision-based Haptic Rendering by Cross-modal Generation
abstract
To assist robots in teleoperation tasks, haptic rendering which allows human operators access a virtual touch feeling has been developed in recent years. Most previous haptic rendering methods strongly rely on data collected by tactile sensors. However, tactile data is not widely available for robots due to their limited reachable space and the restrictions of tactile sensors. To eliminate the need for tactile data, in this paper we propose a novel method named as Vis2Hap to generate haptic rendering from visual inputs that can be obtained from a distance without physical interaction. We take the surface texture of objects as key cues to be conveyed to the human operator. To this end, a generative model is designed to simulate the roughness and slipperiness of the object's surface. To embed haptic cues in Vis2Hap, we use height maps from tactile sensors and spectrograms from friction coefficients as the intermediate outputs of the generative model. Once Vis2Hap is trained, it can be used to generate height maps and spectrograms of new surface textures, from which a friction image can be obtained and displayed on a haptic display. The user study demonstrates that our proposed Vis2Hap method enables users to access a realistic haptic feeling similar to that of physical objects. The proposed vision-based haptic rendering has the potential to enhance human operators' perception of the remote environment and facilitate robotic manipulation.
Guanqun Cao, Ningtao Mao, Danushka Bollegala, Min Li 0003, Shan Luo 0001
ICRA4
2023 Learn from Incomplete Tactile Data: Tactile Representation Learning with Masked Autoencoders
abstract
The missing signal caused by the objects being occluded or an unstable sensor is a common challenge during data collection. Such missing signals will adversely affect the results obtained from the data, and this issue is observed more frequently in robotic tactile perception. In tactile perception, due to the limited working space and the dynamic environment, the contact between the tactile sensor and the object is frequently insufficient and unstable, which causes the partial loss of signals, thus leading to incomplete tactile data. The tactile data will therefore contain fewer tactile cues with low information density. In this paper, we propose a tactile representation learning method, named TacMAE, based on Masked Autoencoder to address the problem of incomplete tactile data in tactile perception. In our framework, a portion of the tactile image is masked out to simulate the missing contact regions. By reconstructing the missing signals in the tactile image, the trained model can achieve a high-level understanding of surface geometry and tactile properties from limited tactile cues. The experimental results of tactile texture recognition show that TacMAE can achieve a high recognition accuracy of 71.4% in the zero-shot transfer and 85.8% after fine-tuning, which are 15.2% and 8.2% higher than the results without using masked modeling. The extensive experiments on YCB objects demonstrate the knowledge transferability of our proposed method and the potential to improve efficiency in tactile exploration.
Guanqun Cao, Danushka Bollegala, Shan Luo 0001
IROS3
2022 Unmasking the Mask - Evaluating Social Biases in Masked Language Models
abstract
Masked Language Models (MLMs) have shown superior performances in numerous downstream Natural Language Processing (NLP) tasks. Unfortunately, MLMs also demonstrate significantly worrying levels of social biases. We show that the previously proposed evaluation metrics for quantifying the social biases in MLMs are problematic due to the following reasons: (1) prediction accuracy of the masked tokens itself tend to be low in some MLMs, which leads to unreliable evaluation metrics, and (2) in most downstream NLP tasks, masks are not used; therefore prediction of the mask is not directly related to them, and (3) high-frequency words in the training data are masked more often, introducing noise due to this selection bias in the test cases. Therefore, we propose All Unmasked Likelihood (AUL), a bias evaluation measure that predicts all tokens in a test case given the MLM embedding of the unmasked input and AUL with Attention weights (AULA) to evaluate tokens based on their importance in a sentence. Our experimental results show that the proposed bias evaluation measures accurately detect different types of biases in MLMs, and unlike AUL and AULA, previously proposed measures for MLMs systematically overestimate the measured biases and are heavily influenced by the unmasked tokens in the context.
Masahiro Kaneko, Danushka Bollegala
AAAI2
2022 Sense Embeddings are also Biased - Evaluating Social Biases in Static and Contextualised Sense Embeddings
abstract
Sense embedding learning methods learn different embeddings for the different senses of an ambiguous word.One sense of an ambiguous word might be socially biased while its other senses remain unbiased.In comparison to the numerous prior work evaluating the social biases in pretrained word embeddings, the biases in sense embeddings have been relatively understudied.We create a benchmark dataset for evaluating the social biases in sense embeddings and propose novel sense-specific bias evaluation measures.We conduct an extensive evaluation of multiple static and contextualised sense embeddings for various types of social biases using the proposed measures.Our experimental results show that even in cases where no biases are found at word-level, there still exist worrying levels of social biases at senselevel, which are often ignored by the word-level bias evaluation measures.1 * Danushka Bollegala holds concurrent appointments as a Professor at University of Liverpool and as an Amazon Scholar.This paper describes work performed at the University of Liverpool and is not associated with Amazon.1 The dataset and evaluation scripts are available at github.com/LivNLP/bias-sense.
Yi Zhou 0019, Masahiro Kaneko, Danushka Bollegala
ACL (1)3
2022 Debiasing Isn't Enough! - on the Effectiveness of Debiasing MLMs and Their Social Biases in Downstream Tasks
abstract
We study the relationship between task-agnostic intrinsic and task-specific extrinsic social bias evaluation measures for MLMs, and find that there exists only a weak correlation between these two types of evaluation measures. Moreover, we find that MLMs debiased using different methods still re-learn social biases during fine-tuning on downstream tasks. We identify the social biases in both training instances as well as their assigned labels as reasons for the discrepancy between intrinsic and extrinsic bias evaluation measurements. Overall, our findings highlight the limitations of existing MLM bias evaluation measures and raise concerns on the deployment of MLMs in downstream applications using those measures.
Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki
COLING2
2022 Learning Meta Word Embeddings by Unsupervised Weighted Concatenation of Source Embeddings
abstract
Given multiple source word embeddings learnt using diverse algorithms and lexical resources, meta word embedding learning methods attempt to learn more accurate and wide-coverage word embeddings. Prior work on meta-embedding has repeatedly discovered that simple vector concatenation of the source embeddings to be a competitive baseline. However, it remains unclear as to why and when simple vector concatenation can produce accurate meta-embeddings. We show that weighted concatenation can be seen as a spectrum matching operation between each source embedding and the meta-embedding, minimising the pairwise inner-product loss. Following this theoretical analysis, we propose two \emph{unsupervised} methods to learn the optimal concatenation weights for creating meta-embeddings from a given set of source embeddings. Experimental results on multiple benchmark datasets show that the proposed weighted concatenated meta-embedding methods outperform previously proposed meta-embedding learning methods.
Danushka Bollegala
IJCAI1
2022 A Survey on Word Meta-Embedding Learning
abstract
Meta-embedding (ME) learning is an emerging approach that attempts to learn more accurate word embeddings given existing (source) word embeddings as the sole input. Due to their ability to incorporate semantics from multiple source embeddings in a compact manner with superior performance, ME learning has gained popularity among practitioners in NLP. To the best of our knowledge, there exist no prior systematic survey on ME learning and this paper attempts to fill this need. We classify ME learning methods according to multiple factors such as whether they (a) operate on static or contextualised embeddings, (b) trained in an unsupervised manner or (c) fine-tuned for a particular task/domain. Moreover, we discuss the limitations of existing ME learning methods and highlight potential future research directions.
Danushka Bollegala, James O'Neill
IJCAI1
2022 Query Obfuscation by Semantic Decomposition
abstract
We propose a method to protect the privacy of search engine users by decomposing the queries using semantically related and unrelated distractor terms. Instead of a single query, the search engine receives multiple decomposed query terms. Next, we reconstruct the search results relevant to the original query term by aggregating the search results retrieved for the decomposed query terms. We show that the word embeddings learnt using a distributed representation learning method can be used to find semantically related and distractor query terms. We derive the relationship between the obfuscity achieved through the proposed query anonymisation method and the reconstructability of the original search results using the decomposed queries. We analytically study the risk of discovering the search engine users’ information intents under the proposed query obfuscation method, and empirically evaluate its robustness against clustering-based attacks. Our experimental results show that the proposed method can accurately reconstruct the search results for user queries, without compromising the privacy of the search engine users.
Danushka Bollegala, Tomoya Machide, Ken-ichi Kawarabayashi
LREC1
2022 Unsupervised Attention-based Sentence-Level Meta-Embeddings from Contextualised Language Models
abstract
A variety of contextualised language models have been proposed in the NLP community, which are trained on diverse corpora to produce numerous Neural Language Models (NLMs). However, different NLMs have reported different levels of performances in downstream NLP applications when used as text representations. We propose a sentence-level meta-embedding learning method that takes independently trained contextualised word embedding models and learns a sentence embedding that preserves the complementary strengths of the input source NLMs. Our proposed method is unsupervised and is not tied to a particular downstream task, which makes the learnt meta-embeddings in principle applicable to different tasks that require sentence representations. Specifically, we first project the token-level embeddings obtained by the individual NLMs and learn attention weights that indicate the contributions of source embeddings towards their token-level meta-embeddings. Next, we apply mean and max pooling to produce sentence-level meta-embeddings from token-level meta-embeddings. Experimental results on semantic textual similarity benchmarks show that our proposed unsupervised sentence-level meta-embedding method outperforms previously proposed sentence-level meta-embedding methods as well as a supervised baseline.
Keigo Takahashi, Danushka Bollegala
LREC2
2022 Learning to Borrow- Relation Representation for Without-Mention Entity-Pairs for Knowledge Graph Completion
abstract
Huda Hakami, Mona Hakami, Angrosh Mandya, Danushka Bollegala. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Huda Hakami, Mona Hakami, Angrosh Mandya, Danushka Bollegala
NAACL-HLT4
2022 Gender Bias in Masked Language Models for Multiple Languages
abstract
Masahiro Kaneko, Aizhan Imankulova, Danushka Bollegala, Naoaki Okazaki. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Masahiro Kaneko, Aizhan Imankulova, Danushka Bollegala, Naoaki Okazaki
NAACL-HLT3
2022 Improvement of intervention information detection for automated clinical literature screening during systematic review
abstract
Systematic literature review (SLR) is a crucial method for clinicians and policymakers to make their decisions in a flood of new clinical studies. Because manual literature screening in SLR is a highly laborious task, its automation by natural language processing (NLP) has been welcomed. Although intervention is a key information for literature screening, NLP models for its detection in previous works have not shown adequate performance. In this work, we first design an algorithm for automated construction of high-quality intervention labels by utilizing information retrieved from a clinical trial database. We then design another algorithm for improving model's recall and F1 score by imposing adaptive weights on training instances in the loss function. The intervention detection model trained on the weighted datasets is tested with the Evidence-Based Medicine NLP (EBM-NLP) corpus, and shows 9.7% and 4.0% improvements respectively in recall and F1 score compared to the previous state-of-the-art model on the corpus. The proposed algorithms can boost automation of literature screening during SLR in the clinical domain.
Tadashi Tsubota, Danushka Bollegala, Yingzi Jin, Tomotake Kozu
J. Biomed. Informatics2
2021 Document Ranking for Curated Document Databases Using BERT and Knowledge Graph Embeddings: Introducing GRAB-Rank
Iqra Muhammad, Danushka Bollegala, Frans Coenen, Carrol Gamble, Anna Kearney, Paula R. Williamson
DaWaK2
2021 RelWalk - A Latent Variable Model Approach to Knowledge Graph Embedding
abstract
Embedding entities and relations of a knowledge graph in a low-dimensional space has shown impressive performance in predicting missing links between entities.Although progresses have been achieved, existing methods are heuristically motivated and theoretical understanding of such embeddings is comparatively underdeveloped.This paper extends the random walk model (Arora et al., 2016a) of word embeddings to Knowledge Graph Embeddings (KGEs) to derive a scoring function that evaluates the strength of a relation R between two entities h (head) and t (tail).Moreover, we show that marginal loss minimisation, a popular objective used in much prior work in KGE, follows naturally from the loglikelihood ratio maximisation under the probabilities estimated from the KGEs according to our theoretical relationship.We propose a learning objective motivated by the theoretical analysis to learn KGEs from a given knowledge graph.Using the derived objective, accurate KGEs are learnt from FB15K237 and WN18RR benchmark datasets, providing empirical evidence in support of the theory. *Danushka Bollegala holds concurrent appointments as a Professor at University of Liverpool and as an Amazon Scholar.This paper describes work performed at the University of Liverpool and is not associated with Amazon.
Danushka Bollegala, Huda Hakami, Yuichi Yoshida, Ken-ichi Kawarabayashi
EACL1
2021 Dictionary-based Debiasing of Pre-trained Word Embeddings
abstract
Word embeddings trained on large corpora have shown to encode high levels of unfair discriminatory gender, racial, religious and ethnic biases.In contrast, human-written dictionaries describe the meanings of words in a concise, objective and an unbiased manner.We propose a method for debiasing pre-trained word embeddings using dictionaries, without requiring access to the original training resources or any knowledge regarding the word embedding algorithms used.Unlike prior work, our proposed method does not require the types of biases to be pre-defined in the form of word lists, and learns the constraints that must be satisfied by unbiased word embeddings automatically from dictionary definitions of the words.Specifically, we learn an encoder to generate a debiased version of an input word embedding such that it (a) retains the semantics of the pre-trained word embeddings, (b) agrees with the unbiased definition of the word according to the dictionary, and (c) remains orthogonal to the vector space spanned by any biased basis vectors in the pre-trained word embedding space.Experimental results on standard benchmark datasets show that the proposed method can accurately remove unfair biases encoded in pre-trained word embeddings, while preserving useful semantics.* Danushka Bollegala holds concurrent appointments as a Professor at University of Liverpool and as an Amazon Scholar.This paper describes work performed at the University of Liverpool and is not associated with Amazon.
Masahiro Kaneko, Danushka Bollegala
EACL2
2021 Debiasing Pre-trained Contextualised Embeddings
abstract
In comparison to the numerous debiasing methods proposed for the static noncontextualised word embeddings, the discriminative biases in contextualised embeddings have received relatively little attention.We propose a fine-tuning method that can be applied at token-or sentence-levels to debias pre-trained contextualised embeddings.Our proposed method can be applied to any pretrained contextualised embedding model, without requiring to retrain those models.Using gender bias as an illustrative example, we then conduct a systematic study using several state-of-the-art (SoTA) contextualised representations on multiple benchmark datasets to evaluate the level of biases encoded in different contextualised embeddings before and after debiasing using the proposed method.We find that applying token-level debiasing for all tokens and across all layers of a contextualised embedding model produces the best performance.Interestingly, we observe that there is a trade-off between creating an accurate vs. unbiased contextualised embedding model, and different contextualised embedding models respond differently to this trade-off.* Danushka Bollegala holds concurrent appointments as a Professor at University of Liverpool and as an Amazon Scholar.This paper describes work performed at the University of Liverpool and is not associated with Amazon.
Masahiro Kaneko, Danushka Bollegala
EACL2
2021 Detect and Classify - Joint Span Detection and Classification for Health Outcomes
abstract
A health outcome is a measurement or an observation used to capture and assess the effect of a treatment.Automatic detection of health outcomes from text would undoubtedly speed up access to evidence necessary in healthcare decision making.Prior work on outcome detection has modelled this task as either (a) a sequence labelling task, where the goal is to detect which text spans describe health outcomes, or (b) a classification task, where the goal is to classify a text into a pre-defined set of categories depending on an outcome that is mentioned somewhere in that text.However, this decoupling of span detection and classification is problematic from a modelling perspective and ignores global structural correspondences between sentence-level and wordlevel information present in a given text.To address this, we propose a method that uses both word-level and sentence-level information to simultaneously perform outcome span detection and outcome type classification.In addition to injecting contextual information to hidden vectors, we use label attention to appropriately weight both word and sentence level information.Experimental results on several benchmark datasets for health outcome detection show that our proposed method consistently outperforms decoupled methods, reporting competitive results.
Micheal Abaho, Danushka Bollegala, Paula R. Williamson, Susanna Dodd
EMNLP (1)2
2021 I Wish I Would Have Loved This One, But I Didn't - A Multilingual Dataset for Counterfactual Detection in Product Review
abstract
Counterfactual statements describe events that did not or cannot take place.We consider the problem of counterfactual detection (CFD) in product reviews.For this purpose, we annotate a multilingual CFD dataset from Amazon product reviews covering counterfactual statements written in English, German, and Japanese languages.The dataset is unique as it contains counterfactuals in multiple languages, covers a new application area of ecommerce reviews, and provides high quality professional annotations.We train CFD models using different text representation methods and classifiers.We find that these models are robust against the selectional biases introduced due to cue phrase-based sentence selection.Moreover, our CFD dataset is compatible with prior datasets and can be merged to learn accurate CFD models.Applying machine translation on English counterfactual examples to create multilingual data performs poorly, demonstrating the language-specificity of this problem, which has been ignored so far.
James O'Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, Danushka Bollegala
EMNLP (1)5
2021 Learning Sense-Specific Static Embeddings using Contextualised Word Embeddings as a Proxy
Danushka Bollegala, Yi Zhou 0019
PACLIC1
2021 Backretrieval: An Image-Pivoted Evaluation Metric for Cross-Lingual Text Representations Without Parallel Corpora
abstract
Cross-lingual text representations have gained popularity lately and act as the backbone of many tasks such as unsupervised machine translation and cross-lingual information retrieval, to name a few. However, evaluation of such representations is difficult in the domains beyond standard benchmarks due to the necessity of obtaining domain-specific parallel language data across different pairs of languages. In this paper, we propose an automatic metric for evaluating the quality of cross-lingual textual representations using images as a proxy in a paired image-text evaluation dataset. Experimentally, Backretrieval is shown to highly correlate with ground truth metrics on annotated datasets, and our analysis shows statistically significant improvements over baselines. Our experiments conclude with a case study on a recipe dataset without parallel cross-lingual data. We illustrate how to judge cross-lingual embedding quality with Backretrieval, and validate the outcome with a small human study.
Mikhail Fain, Niall Twomey, Danushka Bollegala
SIGIR3
2021 Unsupervised Abstractive Opinion Summarization by Generating Sentences with Tree-Structured Topic Guidance
abstract
Abstract This paper presents a novel unsupervised abstractive summarization method for opinionated texts. While the basic variational autoencoder-based models assume a unimodal Gaussian prior for the latent code of sentences, we alternate it with a recursive Gaussian mixture, where each mixture component corresponds to the latent code of a topic sentence and is mixed by a tree-structured topic distribution. By decoding each Gaussian component, we generate sentences with tree-structured topic guidance, where the root sentence conveys generic content, and the leaf sentences describe specific topics. Experimental results demonstrate that the generated topic sentences are appropriate as a summary of opinionated texts, which are more informative and cover more input contents than those generated by the recent unsupervised summarization model (Bražinskas et al., 2020). Furthermore, we demonstrate that the variance of latent Gaussians represents the granularity of sentences, analogous to Gaussian word embedding (Vilnis and McCallum, 2015).
Masaru Isonuma, Junichiro Mori, Danushka Bollegala, Ichiro Sakata
Trans. Assoc. Comput. Linguistics3
2020 Tree-Structured Neural Topic Model
abstract
This paper presents a tree-structured neural topic model, which has a topic distribution over a tree with an infinite number of branches.Our model parameterizes an unbounded ancestral and fraternal topic distribution by applying doubly-recurrent neural networks.With the help of autoencoding variational Bayes, our model improves data scalability and achieves competitive performance when inducing latent topics and tree structures, as compared to a prior tree-structured topic model (Blei et al., 2010).This work extends the tree-structured topic model such that it can be incorporated with neural models for downstream tasks.
Masaru Isonuma, Junichiro Mori, Danushka Bollegala, Ichiro Sakata
ACL3
2020 Autoencoding Improves Pre-trained Word Embeddings
abstract
Prior work investigating the geometry of pre-trained word embeddings have shown that word embeddings to be distributed in a narrow cone and by centering and projecting using principal component vectors one can increase the accuracy of a given set of pre-trained word embeddings.However, theoretically this post-processing step is equivalent to applying a linear autoencoder to minimise the squared ℓ 2 reconstruction error.This result contradicts prior work (Mu and Viswanath, 2018) that proposed to remove the top principal components from pre-trained embeddings.We experimentally verify our theoretical claims and show that retaining the top principal components is indeed useful for improving pre-trained word embeddings, without requiring access to additional linguistic resources or labeled data.
Masahiro Kaneko, Danushka Bollegala
COLING2
2020 Graph Convolution over Multiple Dependency Sub-graphs for Relation Extraction
abstract
We propose in this paper a contextualised graph convolution network over multiple dependency sub-graphs for relation extraction.A novel method to construct multiple sub-graphs using words in shortest dependency path and words linked to entities in the dependency graph is proposed.Graph convolution operation is performed over the resulting multiple sub-graphs to obtain more informative features useful for relation extraction.Our experimental results show that the proposed method achieves superior performance over existing GCN-based models achieving stateof-the-art performance on cross-sentence n-ary relation extraction and SemEval 2010 Task 8 sentence-level relation extraction task.Our model also achieves a comparable performance to the SoTA on the TACRED dataset.
Angrosh Mandya, Danushka Bollegala, Frans Coenen
COLING2
2020 Meta-Embedding as Auxiliary Task Regularization
James O'Neill, Danushka Bollegala
ECAI2
2020 Spatio-temporal Attention Model for Tactile Texture Recognition
abstract
Recently, tactile sensing has attracted great interest in robotics, especially for facilitating exploration of unstructured environments and effective manipulation. A detailed understanding of the surface textures via tactile sensing is essential for many of these tasks. Previous works on texture recognition using camera based tactile sensors have been limited to treating all regions in one tactile image or all samples in one tactile sequence equally, which includes much irrelevant or redundant information. In this paper, we propose a novel Spatio-Temporal Attention Model (STAM) for tactile texture recognition, which is the very first of its kind to our best knowledge. The proposed STAM pays attention to both spatial focus of each single tactile texture and the temporal correlation of a tactile sequence. In the experiments to discriminate 100 different fabric textures, the spatially and temporally selective attention has resulted in a significant improvement of the recognition accuracy, by up to 18.8%, compared to the non-attention based models. Specifically, after introducing noisy data that is collected before the contact happens, our proposed STAM can learn the salient features efficiently and the accuracy can increase by 15.23% on average compared with the CNN based baseline approach. The improved tactile texture perception can be applied to facilitate robot tasks like grasping and manipulation.
Guanqun Cao, Yi Zhou 0019, Danushka Bollegala, Shan Luo 0001
IROS3
2020 Language-Independent Tokenisation Rivals Language-Specific Tokenisation for Word Similarity Prediction
abstract
Language-independent tokenisation (LIT) methods that do not require labelled language resources or lexicons have recently gained popularity because of their applicability in resource-poor languages. Moreover, they compactly represent a language using a fixed size vocabulary and can efficiently handle unseen or rare words. On the other hand, language-specific tokenisation (LST) methods have a long and established history, and are developed using carefully created lexicons and training resources. Unlike subtokens produced by LIT methods, LST methods produce valid morphological subwords. Despite the contrasting trade-offs between LIT vs. LST methods, their performance on downstream NLP tasks remain unclear. In this paper, we empirically compare the two approaches using semantic similarity measurement as an evaluation task across a diverse set of languages. Our experimental results covering eight languages show that LST consistently outperforms LIT when the vocabulary size is large, but LIT can produce comparable or better results than LST in many languages with comparatively smaller (i.e. less than 100K words) vocabulary sizes, encouraging the use of LIT when language-specific resources are unavailable, incomplete or a smaller model is required. Moreover, we find that smoothed inverse frequency (SIF) to be an accurate method to create word embeddings from subword embeddings for multilingual semantic similarity prediction tasks. Further analysis of the nearest neighbours of tokens show that semantically and syntactically related tokens are closely embedded in subword embedding spaces.
Danushka Bollegala, Ryuichi Kiryo, Kosuke Tsujino, Haruki Yukawa
LREC1
2020 Do not let the history haunt you: Mitigating Compounding Errors in Conversational Question Answering
abstract
The Conversational Question Answering (CoQA) task involves answering a sequence of inter-related conversational questions about a contextual paragraph. Although existing approaches employ human-written ground-truth answers for answering conversational questions at test time, in a realistic scenario, the CoQA model will not have any access to ground-truth answers for the previous questions, compelling the model to rely upon its own previously predicted answers for answering the subsequent questions. In this paper, we find that compounding errors occur when using previously predicted answers at test time, significantly lowering the performance of CoQA systems. To solve this problem, we propose a sampling strategy that dynamically selects between target answers and model predictions during training, thereby closely simulating the situation at test time. Further, we analyse the severity of this phenomena as a function of the question type, conversation length and domain type.
Angrosh Mandya, James O'Neill, Danushka Bollegala, Frans Coenen
LREC3
2020 Explanation in AI and law: Past, present and future
Katie Atkinson, Trevor J. M. Bench-Capon, Danushka Bollegala
Artif. Intell.3
2019 Gender-preserving Debiasing for Pre-trained Word Embeddings
abstract
Word embeddings learnt from massive text collections have demonstrated significant levels of discriminative biases such as gender, racial or ethnic biases, which in turn bias the down-stream NLP applications that use those word embeddings. Taking gender-bias as a working example, we propose a debiasing method that preserves non-discriminative gender-related information, while removing stereotypical discriminative gender biases from pre-trained word embeddings. Specifically, we consider four types of information: feminine, masculine, gender-neutral and stereotypical, which represent the relationship between gender vs. bias, and propose a debiasing method that (a) preserves the gender-related information in feminine and masculine words, (b) preserves the neutrality in gender-neutral words, and (c) removes the biases from stereotypical words. Experimental results on several previously proposed benchmark datasets show that our proposed method can debias pre-trained word embeddings better than existing SoTA methods proposed for debiasing word embeddings while preserving gender-related but non-discriminative information.
Masahiro Kaneko, Danushka Bollegala
ACL (1)2
2019 Predicting the Quality of Translations Without an Oracle
Yi Zhou 0019, Danushka Bollegala
IC3K2
2019 Automated Bundle Pagination Using Machine Learning
abstract
Coherent division of legal document bundles, whether this is done in the context of court bundles, briefs or some other application, is a time consuming and challenging task. We propose an approach whereby this process can be automated. Two variations are considered. The first addresses the scenario where the topic labelling is pre-defined and adopts a supervised learning approach. The second addresses the scenario where the topic labelling, for whatever reason, is not specified in advance and adopts an unsupervised learning approach. This paper reports on an investigation of both mechanisms using accident claims bundles. The evaluation results indicate that the proposed approaches can be successfully applied to divide legal document bundles.
Alessandro Torrisi, Robert Bevan, Katie Atkinson, Danushka Bollegala, Frans Coenen
ICAIL4
2019 "Touching to See" and "Seeing to Feel": Robotic Cross-modal Sensory Data Generation for Visual-Tactile Perception
abstract
The integration of visual-tactile stimulus is common while humans performing daily tasks. In contrast, using unimodal visual or tactile perception limits the perceivable dimensionality of a subject. However, it remains a challenge to integrate the visual and tactile perception to facilitate robotic tasks. In this paper, we propose a novel framework for the cross-modal sensory data generation for visual and tactile perception. Taking texture perception as an example, we apply conditional generative adversarial networks to generate pseudo visual images or tactile outputs from data of the other modality. Extensive experiments on the ViTac dataset of cloth textures show that the proposed method can produce realistic outputs from other sensory inputs. We adopt the structural similarity index to evaluate similarity of the generated output and real data and results show that realistic data have been generated. Classification evaluation has also been performed to show that the inclusion of generated data can improve the perception performance. The proposed framework has potential to expand datasets for classification tasks, generate sensory outputs that are not easy to access, and also advance integrated visual-tactile perception.
Jet-Tsyn Lee, Danushka Bollegala, Shan Luo 0001
ICRA2
2019 Combining Textual and Visual Information for Typed and Handwritten Text Separation in Legal Documents
Alessandro Torrisi, Robert Bevan, Katie Atkinson, Danushka Bollegala, Frans Coenen
JURIX4
2018 Using k-Way Co-Occurrences for Learning Word Embeddings
abstract
Co-occurrences between two words provide useful insights into the semantics of those words.Consequently, numerous prior work on word embedding learning has used co-occurrences between two wordsas the training signal for learning word embeddings.However, in natural language texts it is common for multiple words to be related and co-occurring in the same context.We extend the notion of co-occurrences to cover k(≥2)-way co-occurrences among a set of k-words.Specifically, we prove a theoretical relationship between the joint probability of k(≥2) words, and the sum of l_2 norms of their embeddings. Next, we propose a learning objective motivated by our theoretical resultthat utilises k-way co-occurrences for learning word embeddings.Our experimental results show that the derived theoretical relationship does indeed hold empirically, anddespite data sparsity, for some smaller k(≤5) values, k-way embeddings perform comparably or better than 2-way embeddings in a range of tasks.
Danushka Bollegala, Yuichi Yoshida, Ken-ichi Kawarabayashi
AAAI1
2018 Learning Word Meta-Embeddings by Autoencoding
abstract
Distributed word embeddings have shown superior performances in numerous Natural Language Processing (NLP) tasks. However, their performances vary significantly across different tasks, implying that the word embeddings learnt by those methods capture complementary aspects of lexical semantics. Therefore, we believe that it is important to combine the existing word embeddings to produce more accurate and complete meta-embeddings of words. We model the meta-embedding learning problem as an autoencoding problem, where we would like to learn a meta-embedding space that can accurately reconstruct all source embeddings simultaneously. Thereby, the meta-embedding space is enforced to capture complementary information in different source embeddings via a coherent common embedding space. We propose three flavours of autoencoded meta-embeddings motivated by different requirements that must be satisfied by a meta-embedding. Our experimental results on a series of benchmark evaluations show that the proposed autoencoded meta-embeddings outperform the existing state-of-the-art meta-embeddings in multiple tasks.
Danushka Bollegala, Cong Bao
COLING1
2018 Why does PairDiff work? - A Mathematical Analysis of Bilinear Relational Compositional Operators for Analogy Detection
abstract
Representing the semantic relations that exist between two given words (or entities) is an important first step in a wide-range of NLP applications such as analogical reasoning, knowledge base completion and relational information retrieval. A simple, yet surprisingly accurate method for representing a relation between two words is to compute the vector offset (PairDiff) between their corresponding word embeddings. Despite the empirical success, it remains unclear as to whether PairDiff is the best operator for obtaining a relational representation from word embeddings. We conduct a theoretical analysis of generalised bilinear operators that can be used to measure the l2 relational distance between two word-pairs. We show that, if the word embed- dings are standardised and uncorrelated, such an operator will be independent of bilinear terms, and can be simplified to a linear form, where PairDiff is a special case. For numerous word embedding types, we empirically verify the uncorrelation assumption, demonstrating the general applicability of our theoretical result. Moreover, we experimentally discover PairDiff from the bilinear relational compositional operator on several benchmark analogy datasets.
Huda Hakami, Kohei Hayashi, Danushka Bollegala
COLING3
2018 An Empirical Study on Fine-Grained Named Entity Recognition
abstract
Named entity recognition (NER) has attracted a substantial amount of research. Recently, several neural network-based models have been proposed and achieved high performance. However, there is little research on fine-grained NER (FG-NER), in which hundreds of named entity categories must be recognized, especially for non-English languages. It is still an open question whether there is a model that is robust across various settings or the proper model varies depending on the language, the number of named entity categories, and the size of training datasets. This paper first presents an empirical comparison of FG-NER models for English and Japanese and demonstrates that LSTM+CNN+CRF (Ma and Hovy, 2016), one of the state-of-the-art methods for English NER, also works well for English FG-NER but does not work well for Japanese, a language that has a large number of character types. To tackle this problem, we propose a method to improve the neural network-based Japanese FG-NER performance by removing the CNN layer and utilizing dictionary and category embeddings. Experiment results show that the proposed method improves Japanese FG-NER F-score from 66.76% to 75.18%.
Khai Mai, Thai-Hoang Pham, Minh Trung Nguyen, Nguyen Tuan Duc, Danushka Bollegala, Ryohei Sasano, Satoshi Sekine
COLING5
2018 Spectral Analysis of Keystroke Streams: Towards Effective Real-time Continuous User Authentication
abstract
Copyright © 2018 by SCITEPRESS – Science and Technology Publications, Lda. All rights reserved. Continuous authentication using keystroke dynamics is significant for applications where continuous monitoring of a user’s identity is desirable, for example in the context of the online assessments and examinations frequently encountered in eLearning environments. In this paper, a novel approach to realtime keystroke continuous authentication is proposed that is founded on a sinusoidal signal based approach that takes into consideration the sequencing of keystrokes. Three alternative time series representations are considered and compared: Keystroke Time Series (KTS), Discrete Fourier Transform (DFT) and Discrete Wavelet Transform (DWT). The proposed process is fully described and analysed using three keystroke dynamics datasets. The evaluation also includes a comparison with the established Feature Vector Representation (FVR) approach. The reported evaluation demonstrates that the proposed method, coupled with the DWT representation, outperforms other approaches to keystroke continuous authentication with a best overall accuracy of 98.24%; a clear indicator that the proposed keystroke continuous authentication using time series analysis has significant potential.
Abdullah Alshehri 0001, Frans Coenen, Danushka Bollegala
ICISSP3
2018 Think Globally, Embed Locally - Locally Linear Meta-embedding of Words
abstract
Distributed word embeddings have shown superior performances in numerous Natural Language Processing (NLP) tasks. However, their performances vary significantly across different tasks, implying that the word embeddings learnt by those methods capture complementary aspects of lexical semantics. Therefore, we believe that it is important to combine the existing word embeddings to produce more accurate and complete meta-embeddings of words. For this purpose, we propose an unsupervised locally linear meta-embedding learning method that takes pre-trained word embeddings as the input, and produces more accurate meta embeddings. Unlike previously proposed meta-embedding learning methods that learn a global projection over all words in a vocabulary, our proposed method is sensitive to the differences in local neighbourhoods of the individual source word embeddings. Moreover, we show that vector concatenation, a previously proposed highly competitive baseline approach for integrating word embeddings, can be derived as a special case of the proposed method. Experimental results on semantic similarity, word analogy, relation classification, and short-text classification tasks show that our meta-embeddings to significantly outperform prior methods in several benchmark datasets, establishing a new state of the art for meta-embeddings.
Danushka Bollegala, Kohei Hayashi, Ken-ichi Kawarabayashi
IJCAI1
2018 Efficient and Effective Case Reject-Accept Filtering: A Study Using Machine Learning
abstract
The decision whether to accept or reject a new case is a well established task undertaken in legal work. This task frequently necessitates domain knowledge and is consequently resource expensive. In this paper it is proposed that early rejection/acceptance of at least a proportion of new cases can be effectively achieved without requiring significant human intervention. The paper proposes, and evaluates, five different AI techniques whereby early case reject-accept can be achieved. The results suggest it is possible for at least a proportion of cases to be processed in this way.
Robert Bevan, Alessandro Torrisi, Katie Atkinson, Danushka Bollegala, Frans Coenen
JURIX4
2018 Joint Learning of Sense and Word Embeddings
Mohammed Alsuhaibani, Danushka Bollegala
LREC2
2018 A Dataset for Inter-Sentence Relation Extraction using Distant Supervision
Angrosh Mandya, Danushka Bollegala, Frans Coenen, Katie Atkinson
LREC2
2018 Sentiment-Stance-Specificity (SSS) Dataset: Identifying Support-based Entailment among Opinions
Pavithra Rajendran, Danushka Bollegala, Simon Parsons
LREC2
2018 ClassiNet - Predicting Missing Features for Short-Text Classification
abstract
Short and sparse texts such as tweets, search engine snippets, product reviews, and chat messages are abundant on the Web. Classifying such short-texts into a pre-defined set of categories is a common problem that arises in various contexts, such as sentiment classification, spam detection, and information recommendation. The fundamental problem in short-text classification is feature sparseness -- the lack of feature overlap between a trained model and a test instance to be classified. We propose ClassiNet -- a network of classifiers trained for predicting missing features in a given instance, to overcome the feature sparseness problem. Using a set of unlabeled training instances, we first learn binary classifiers as feature predictors for predicting whether a particular feature occurs in a given instance. Next, each feature predictor is represented as a vertex v i in the ClassiNet, where a one-to-one correspondence exists between feature predictors and vertices. The weight of the directed edge e ij connecting a vertex v i to a vertex v j represents the conditional probability that given v i exists in an instance, v j also exists in the same instance. We show that ClassiNets generalize word co-occurrence graphs by considering implicit co-occurrences between features. We extract numerous features from the trained ClassiNet to overcome feature sparseness. In particular, for a given instance x , we find similar features from ClassiNet that did not appear in x , and append those features in the representation of x . Moreover, we propose a method based on graph propagation to find features that are indirectly related to a given short-text. We evaluate ClassiNets on several benchmark datasets for short-text classification. Our experimental results show that by using ClassiNet, we can statistically significantly improve the accuracy in short-text classification tasks, without having to use any external resources such as thesauri for finding related features.
Danushka Bollegala, Vincent Atanasov, Takanori Maehara, Ken-ichi Kawarabayashi
ACM Trans. Knowl. Discov. Data1
2017 Classifier-Based Pattern Selection Approach for Relation Instance Extraction
Angrosh Mandya, Danushka Bollegala, Frans Coenen, Katie Atkinson
CICLing (1)2
2017 Behavioural Biometric Continuous User Authentication Using Multivariate Keystroke Streams in the Spectral Domain
Abdullah Alshehri 0001, Frans Coenen, Danushka Bollegala
IC3K3
2017 CLIEL: context-based information extraction from commercial law documents
abstract
The effectiveness of document Information Extraction (IE) is greatly affected by the structure and layout of the documents being considered. In the case of legal documents relating to commercial law, an additional challenge is the many different and varied formats, structures and layouts used. In this paper, we present work on a flexible and scalable IE environment, the CLIEL (Commercial Law Information Extraction based on Layout) environment, for application to commercial law documentation that allows layout rules to be derived and then utilised to support IE. The proposed CLIEL environment operates using NLP (Natural Language Processing) techniques, JAPE (Java Annotation Patterns Engine) rules and some GATE (General Architecture for Text Engineering) modules. The system is fully described and evaluated using a commercial law document corpus. The results demonstrate that considering the layout is beneficial for extracting data point instances from legal document collections.
Matias Garcia-Constantino, Katie Atkinson, Danushka Bollegala, Karl Chapman, Frans Coenen, Claire Roberts, Katy Robson
ICAIL3
2017 TSP: Learning Task-Specific Pivots for Unsupervised Domain Adaptation
Xia Cui 0001, Frans Coenen, Danushka Bollegala
ECML/PKDD (2)3
2017 Dynamic feature scaling for online learning of binary classifiers
Danushka Bollegala
Knowl. Based Syst.1
2017 Compositional approaches for representing relations between words: A comparative study
Huda Hakami, Danushka Bollegala
Knowl. Based Syst.2
2017 A classification approach for detecting cross-lingual biomedical term translations
abstract
Abstract Finding translations for technical terms is an important problem in machine translation. In particular, in highly specialized domains such as biology or medicine, it is difficult to find bilingual experts to annotate sufficient cross-lingual texts in order to train machine translation systems. Moreover, new terms are constantly being generated in the biomedical community, which makes it difficult to keep the translation dictionaries up to date for all language pairs of interest. Given a biomedical term in one language (source language), we propose a method for detecting its translations in a different language (target language). Specifically, we train a binary classifier to determine whether two biomedical terms written in two languages are translations. Training such a classifier is often complicated due to the lack of common features between the source and target languages. We propose several feature space concatenation methods to successfully overcome this problem. Moreover, we study the effectiveness of contextual and character n-gram features for detecting term translations. Experiments conducted using a standard dataset for biomedical term translation show that the proposed method outperforms several competitive baseline methods in terms of mean average precision and top-k translation accuracy.
Huda Hakami, Danushka Bollegala
Nat. Lang. Eng.2
2016 Joint Word Representation Learning Using a Corpus and a Semantic Lexicon
abstract
Methods for learning word representations using large text corpora have received much attention lately due to their impressive performancein numerous natural language processing (NLP) tasks such as, semantic similarity measurement, and word analogy detection.Despite their success, these data-driven word representation learning methods do not considerthe rich semantic relational structure between words in a co-occurring context. On the other hand, already much manual effort has gone into the construction of semantic lexicons such as the WordNetthat represent the meanings of words by defining the various relationships that exist among the words in a language.We consider the question, can we improve the word representations learnt using a corpora by integrating theknowledge from semantic lexicons?. For this purpose, we propose a joint word representation learning method that simultaneously predictsthe co-occurrences of two words in a sentence subject to the relational constrains given by the semantic lexicon.We use relations that exist between words in the lexicon to regularize the word representations learnt from the corpus.Our proposed method statistically significantly outperforms previously proposed methods for incorporating semantic lexicons into wordrepresentations on several benchmark datasets for semantic similarity and word analogy.
Danushka Bollegala, Mohammed Alsuhaibani, Takanori Maehara, Ken-ichi Kawarabayashi
AAAI1
2016 Assessing Weight of Opinion by Aggregating Coalitions of Arguments
abstract
Argument mining promises to be able to extract information from unstructured text that can help us to understand that text. This paper suggests a novel way to use such information once it has been extracted. Attack and support relations between arguments from a set of test texts are identified, the strength of the arguments is computed based on the relations, and arguments are grouped into coalitions. The resulting set of arguments is then used to predict the weight of opinion in new text, by identifying arguments whose weight has been computed, and aggregating these weights. Our approach is evaluated on a corpus of hotel reviews, and compared with an existing method of predicting the sentiment of reviews.
Pavithra Rajendran, Danushka Bollegala, Simon Parsons
COMMA2
2016 Keyboard Usage Authentication Using Time Series Analysis
Abdullah Alshehri 0001, Frans Coenen, Danushka Bollegala
DaWaK3
2016 Cross-Domain Sentiment Classification Using Sentiment Sensitive Embeddings
abstract
Unsupervised Cross-domain Sentiment Classification is the task of adapting a sentiment classifier trained on a particular domain (source domain), to a different domain (target domain), without requiring any labeled data for the target domain. By adapting an existing sentiment classifier to previously unseen target domains, we can avoid the cost for manual data annotation for the target domain. We model this problem as embedding learning, and construct three objective functions that capture: (a) distributional properties ofpivots(i.e., common features that appear in both source and target domains), (b) label constraints in the source domain documents, and (c) geometric properties in the unlabeled documents in both source and target domains. Unlike prior proposals that first learn a lower-dimensional embedding independent of the source domain sentiment labels, and next a sentiment classifier in this embedding, our joint optimisation method learns embeddings that are sensitive to sentiment classification. Experimental results on a benchmark dataset show that by jointly optimising the three objectives we can obtain better performances in comparison to optimising each objective function separately, thereby demonstrating the importance of task-specific embedding learning for cross-domain sentiment classification. Among the individual objective functions, the best performance is obtained by (c). Moreover, the proposed method reports cross-domain sentiment classification accuracies that are statistically comparable to the current state-of-the-art embedding learning methods for cross-domain sentiment classification.
Danushka Bollegala, Tingting Mu, John Yannis Goulermas
IEEE Trans. Knowl. Data Eng.1
2015 Learning Word Representations from Relational Graphs
abstract
Attributes of words and relations between two words are central to numerous tasks in Artificial Intelligence such as knowledge representation, similarity measurement, and analogy detection. Often when two words share one or more attributes in common, they are con- nected by some semantic relations. On the other hand, if there are numerous semantic relations between two words, we can expect some of the attributes of one of the words to be inherited by the other. Motivated by this close connection between attributes and relations, given a relational graph in which words are inter-connected via numerous semantic relations, we propose a method to learn a latent representation for the individual words. The proposed method considers not only the co-occurrences of words as done by existing approaches for word representation learning, but also the semantic relations in which two words co-occur. To evaluate the accuracy of the word representations learnt using the proposed method, we use the learnt word representa- tions to solve semantic word analogy problems. Our experimental results show that it is possible to learn better word representations by using semantic semantics between words.
Danushka Bollegala, Takanori Maehara, Yuichi Yoshida, Ken-ichi Kawarabayashi
AAAI1
2015 Unsupervised Cross-Domain Word Representation Learning
abstract
Danushka Bollegala, Takanori Maehara, Ken-ichi Kawarabayashi. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Danushka Bollegala, Takanori Maehara, Ken-ichi Kawarabayashi
ACL (1)1
2015 A Discourse Search Engine Based on Rhetorical Structure Theory
Pascal Kuyten, Danushka Bollegala, Bernd Hollerit, Helmut Prendinger, Kiyoharu Aizawa
ECIR2
2015 Embedding Semantic Relations into Word Representations
Danushka Bollegala, Takanori Maehara, Ken-ichi Kawarabayashi
IJCAI1
2014 Learning to Predict Distributions of Words Across Domains
abstract
Although the distributional hypothesis has been applied successfully in many natural language processing tasks, systems using distributional information have been limited to a single domain because the distribution of a word can vary between domains as the word’s predominant meaning changes. However, if it were possible to predict how the distribution of a word changes from one domain to another, the predictions could be used to adapt a system trained in one domain to work in another. We propose an unsupervised method to predict the distribution of a word in one domain, given its distribution in another domain. We evaluate our method on two tasks: cross-domain part-of-speech tagging and cross-domain sentiment classification. In both tasks, our method significantly outperforms competitive baselines and returns results that are statistically comparable to current state-of-the-art methods, while requiring no task-specific customisations.
Danushka Bollegala, David J. Weir, John Carroll 0001
ACL (1)1
2013 Learning non-linear ranking functions for web search using probabilistic model building GP
abstract
Ranking the set of search results according to their relevance to a user query is an important task in an Information Retrieval (IR) systems such as a Web Search Engine. Learning the optimal ranking function for this task is a challenging problem because one must consider complex non-linear interactions between numerous factors such as the novelty, authority, contextual similarity, etc. of thousands of documents that contain the user query. We model this task as a non-linear ranking problem, for which we propose Rank-PMBGP, an efficient algorithm to learn an optimal non-linear ranking function using Probabilistic Model Building Genetic Programming. We evaluate the proposed method using the LETOR dataset, a standard benchmark dataset for training and evaluating ranking functions for IR. In our experiments, the proposed method obtains a Mean Average Precision (MAP) score of 0.291, thereby significantly outperforming a non-linear baseline approach that uses Genetic Programming.
Danushka Bollegala, Yoshihiko Hasegawa, Hitoshi Iba
IEEE Congress on Evolutionary Computation2
2013 Mining for Analogous Tuples from an Entity-Relation Graph
Danushka Bollegala, Mitsuru Kusumoto, Yuichi Yoshida, Ken-ichi Kawarabayashi
IJCAI1
2013 Improving relational similarity measurement using symmetries in proportional word analogies
Danushka Bollegala, Tomokazu Goto, Nguyen Tuan Duc, Mitsuru Ishizuka
Inf. Process. Manag.1
2013 Minimally Supervised Novel Relation Extraction Using a Latent Relational Mapping
abstract
The World Wide Web includes semantic relations of numerous types that exist among different entities. Extracting the relations that exist between two entities is an important step in various Web-related tasks such as information retrieval (IR), information extraction, and social network extraction. A supervised relation extraction system that is trained to extract a particular relation type (source relation) might not accurately extract a new type of a relation (target relation) for which it has not been trained. However, it is costly to create training data manually for every new relation type that one might want to extract. We propose a method to adapt an existing relation extraction system to extract new relation types with minimum supervision. Our proposed method comprises two stages: learning a lower dimensional projection between different relations, and learning a relational classifier for the target relation type with instance sampling. First, to represent a semantic relation that exists between two entities, we extract lexical and syntactic patterns from contexts in which those two entities co-occur. Then, we construct a bipartite graph between relation-specific (RS) and relation-independent (RI) patterns. Spectral clustering is performed on the bipartite graph to compute a lower dimensional projection. Second, we train a classifier for the target relation type using a small number of labeled instances. To account for the lack of target relation training instances, we present a one-sided under sampling method. We evaluate the proposed method using a data set that contains 2,000 instances for 20 different relation types. Our experimental results show that the proposed method achieves a statistically significant macroaverage F-score of 62.77. Moreover, the proposed method outperforms numerous baselines and a previously proposed weakly supervised relation extraction method.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IEEE Trans. Knowl. Data Eng.1
2013 Cross-Domain Sentiment Classification Using a Sentiment Sensitive Thesaurus
abstract
Automatic classification of sentiment is important for numerous applications such as opinion mining, opinion summarization, contextual advertising, and market analysis. Typically, sentiment classification has been modeled as the problem of training a binary classifier using reviews annotated for positive or negative sentiment. However, sentiment is expressed differently in different domains, and annotating corpora for every possible domain of interest is costly. Applying a sentiment classifier trained using labeled data for a particular domain to classify sentiment of user reviews on a different domain often results in poor performance because words that occur in the train (source) domain might not appear in the test (target) domain. We propose a method to overcome this problem in cross-domain sentiment classification. First, we create a sentiment sensitive distributional thesaurus using labeled data for the source domains and unlabeled data for both source and target domains. Sentiment sensitivity is achieved in the thesaurus by incorporating document level sentiment labels in the context vectors used as the basis for measuring the distributional similarity between words. Next, we use the created thesaurus to expand feature vectors during train and test times in a binary classifier. The proposed method significantly outperforms numerous baselines and returns results that are comparable with previously proposed cross-domain sentiment classification methods on a benchmark data set containing Amazon user reviews for different types of products. We conduct an extensive empirical analysis of the proposed method on single- and multisource domain adaptation, unsupervised and supervised domain adaptation, and numerous similarity measures for creating the sentiment sensitive thesaurus. Moreover, our comparisons against the SentiWordNet, a lexical resource for word polarity, show that the created sentiment-sensitive thesaurus accurately captures words that express similar sentiments.
Danushka Bollegala, David J. Weir, John Carroll 0001
IEEE Trans. Knowl. Data Eng.1
2012 Multinomial Relation Prediction in Social Data: A Dimension Reduction Approach
abstract
The recent popularization of social web services has made them one of the primary uses of the World Wide Web. An important concept in social web services is social actions such as making connections and communicating with others and adding annotations to web resources. Predicting social actions would improve many fundamental web applications, such as recommendations and web searches. One remarkable characteristic of social actions is that they involve multiple and heterogeneous objects such as users, documents, keywords, and locations. However, the high-dimensional property of such multinomial relations poses one fundamental challenge, that is, predicting multinomial relations with only a limited amount of data. In this paper, we propose a new multinomial relation prediction method, which is robust to data sparsity. We transform each instance of a multinomial relation into a set of binomial relations between the objects and the multinomial relation of the involved objects. We then apply an extension of a low-dimensional embedding technique to these binomial relations, which results in a generalized eigenvalue problem guaranteeing global optimal solutions. We also incorporate attribute information as side information to address the “cold start” problem in multinomial relation prediction. Experiments with various real-world social web service datasets demonstrate that the proposed method is more robust against data sparseness as compared to several existing methods, which can only find sub-optimal solutions.
Nozomi Nori, Danushka Bollegala, Hisashi Kashima
AAAI2
2012 Similarity Is Not Entailment - Jointly Learning Similarity Transformation for Textual Entailment
abstract
Predicting entailment between two given texts is an important task upon which the performance of numerous NLP tasks depend on such as question answering, text summarization, and information extraction. The degree to which two texts are similar has been used extensively as a key feature in much previous work in predicting entailment. However, using similarity scores directly, without proper transformations, results in suboptimal performance. Given a set of lexical similarity measures, we propose a method that jointly learns both (a) a set of non-linear transformation functions for those similarity measures and, (b) the optimal non-linear combination of those transformation functions to predict textual entailment. Our method consistently outperforms numerous baselines, reporting a micro-averaged F-score of 46.48 on the RTE- 7 benchmark dataset. The proposed method is ranked 2-nd among 33 entailment systems participated in RTE-7, demonstrating its competitiveness over numerous other entailment approaches. Although our method is statistically comparable to the current state-of-the-art, we require less external knowledge resources.
Ken-Ichi Yokote, Danushka Bollegala, Mitsuru Ishizuka
AAAI2
2012 Probabilistic model building GP with Belief propagation
abstract
Estimation of distribution algorithms (EDAs) which deal with tree structures as GP are called as probabilistic model building GPs (PMBGPs), and they show better search performance than GP in many problems. A problem of prototype tree-based method, a type of PMBGPs, is that samplings do not always generate the most probable solution, which is the individual with the highest probability and reflects a learned distribution most. This problem wastes a part of learning and increases the number of evaluations to get an optimum solution. In order to overcome this difficulty, this paper proposes a hybrid approach using Belief propagation (BP) in sampling process. BP is an inference algorithm on graphical models and can generate the most probable solution. By applying our approach to benchmark tests, we show that the proposed method is more effective than PLS alone.
Yoshihiko Hasegawa, Danushka Bollegala, Hitoshi Iba
IEEE Congress on Evolutionary Computation3
2012 Automatic Annotation of Ambiguous Personal Names on the Web
abstract
Personal name disambiguation is an important task in social network extraction, evaluation and integration of ontologies, information retrieval, cross‐document coreference resolution and word sense disambiguation. We propose an unsupervised method to automatically annotate people with ambiguous names on the Web using automatically extracted keywords. Given an ambiguous personal name, first, we download text snippets for the given name from a Web search engine. We then represent each instance of the ambiguous name by a term‐entity model (TEM), a model that we propose to represent the Web appearance of an individual. A TEM of a person captures named entities and attribute values that are useful to disambiguate that person from his or her namesakes (i.e., different people who share the same name). We then use group average agglomerative clustering to identify the instances of an ambiguous name that belong to the same person. Ideally, each cluster must represent a different namesake. However, in practice it is not possible to know the number of namesakes for a given ambiguous personal name in advance. To circumvent this problem, we propose a novel normalized cuts‐based cluster stopping criterion to determine the different people on the Web for a given ambiguous name. Finally, we annotate each person with an ambiguous name using keywords selected from the clusters. We evaluate the proposed method on a data set of over 2500 documents covering 200 different people for 20 ambiguous names. Experimental results show that the proposed method outperforms numerous baselines and previously proposed name disambiguation methods. Moreover, the extracted keywords reduce ambiguity of a name in an information retrieval task, which underscores the usefulness of the proposed method in real‐world scenarios.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
Comput. Intell.1
2012 A preference learning approach to sentence ordering for multi-document summarization
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka
Inf. Sci.1
2012 Cross-Language Latent Relational Search between Japanese and English Languages Using a Web Corpus
abstract
Latent relational search is a novel entity retrieval paradigm based on the proportional analogy between two entity pairs. Given a latent relational search query {(Japan, Tokyo), (France, ?)}, a latent relational search engine is expected to retrieve and rank the entity “Paris” as the first answer in the result list. A latent relational search engine extracts entities and relations between those entities from a corpus, such as the Web. Moreover, from some supporting sentences in the corpus, (e.g., “Tokyo is the capital of Japan” and “Paris is the capital and biggest city of France”), the search engine must recognize the relational similarity between the two entity pairs. In cross-language latent relational search, the entity pairs as well as the supporting sentences of the first entity pair and of the second entity pair are in different languages. Therefore, the search engine must recognize similar semantic relations across languages. In this article, we study the problem of cross-language latent relational search between Japanese and English using Web data. To perform cross-language latent relational search in high speed, we propose a multi-lingual indexing method for storing entities and lexical patterns that represent the semantic relations extracted from Web corpora. We then propose a hybrid lexical pattern clustering algorithm to capture the semantic similarity between lexical patterns across languages. Using this algorithm, we can precisely measure the relational similarity between entity pairs across languages, thereby achieving high precision in the task of cross-language latent relational search. Experiments show that the proposed method achieves an MRR of 0.605 on Japanese-English cross-language latent relational search query sets and it also achieves a reasonable performance on the INEX Entity Ranking task.
Nguyen Tuan Duc, Danushka Bollegala, Mitsuru Ishizuka
ACM Trans. Asian Lang. Inf. Process.2
2011 Cross-Language Latent Relational Search: Mapping Knowledge across Languages
abstract
Latent relational search (LRS) is a novel approach for mapping knowledge across two domains. Given a source domain knowledge concerning the Moon, "The Moon is a satellite of the Earth," one can form a question {(Moon, Earth), (Ganymede, ?)} to query an LRS engine for new knowledge in the target domain concerning the Ganymede. An LRS engine relies on some supporting sentences such as ``Ganymede is a natural satellite of Jupiter.'' to retrieve and rank "Jupiter" as the first answer. This paper proposes cross-language latent relational search (CLRS) to extend the knowledge mapping capability of LRS from cross-domain knowledge mapping to cross-domain and cross-language knowledge mapping. In CLRS, the supporting sentences for the source pair might be in a different language with that of the target pair. We represent the relation between two entities in an entity pair by lexical patterns of the context surrounding the two entities. We then propose a novel hybrid lexical pattern clustering algorithm to capture the semantic similarity between paraphrased lexical patterns across languages. Experiments on Japanese-English datasets show that the proposed method achieves an MRR of 0.579 for CLRS task, which is comparable to the MRR of an existing monolingual LRS engine.
Nguyen Tuan Duc, Danushka Bollegala, Mitsuru Ishizuka
AAAI2
2011 Using Multiple Sources to Construct a Sentiment Sensitive Thesaurus for Cross-Domain Sentiment Classification
Danushka Bollegala, David J. Weir, John Carroll 0001
ACL1
2011 An adaptive differential evolution algorithm
abstract
The performance of Differential Evolution (DE) algorithm is significantly affected by its parameter setting. But the choice of parameters is heavily dependent on the problem characteristics. Therefore, recently a couple of adaptation schemes that automatically adjust DE parameters have been proposed. The current work presents another adaptation scheme for DE parameters namely amplification factor and crossover rate. We systematically analyze the effectiveness of the proposed adaptation scheme for DE parameters using a standard benchmark suite consisting of ten functions. The undertaken empirical study shows that the proposed adaptive DE (aDE) algorithm exhibits an overall better performance compared to other prominent adaptive DE algorithms as well as canonical DE.
Nasimul Noman, Danushka Bollegala, Hitoshi Iba
IEEE Congress on Evolutionary Computation2
2011 Semi-supervised Discourse Relation Classification with Structural Learning
Hugo Hernault, Danushka Bollegala, Mitsuru Ishizuka
CICLing (1)2
2011 Using Graph Based Method to Improve Bootstrapping Relation Extraction
Haibo Li 0002, Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
CICLing (2)2
2011 Collaborative exploratory search in real-world context
abstract
We propose Collaborative Exploratory Search (CES), which is an integration of dialog analysis and web search that involves multiparty collaboration to accomplish an exploratory information retrieval goal. Given a real-time dialog between users on a single topic; we define CES as the task of automatically detecting the topic of the dialog and retrieving task-relevant web pages to support the dialog. To recognize the task of the dialog, we apply the Author--Topic model as a topic model. Then, attribute extraction is applied to the dialog to obtain the attributes of the tasks. Finally, a specific search query is generated to identify the task-relevant information. We implement and evaluate the CES system for a commercial in-vehicle conversation. We also develop an iPad application that listens to conversations among users and continuously retrieves relevant web pages. Our experimental results reveal that the proposed method outperforms existing methods, which demonstrates the potential usefulness of collaborative exploratory search with practically usable accuracy levels.
Naoki Tani, Danushka Bollegala, Naiwala P. Chandrasiri, Keisuke Okamoto, Kazunari Nawa, Shuhei Iitsuka, Yutaka Matsuo
CIKM2
2011 RankDE: learning a ranking function for information retrieval using differential evolution
abstract
Learning a ranking function is important for numerous tasks such as information retrieval (IR), question answering, and product recommendation. For example, in information retrieval, a Web search engine is required to rank and return a set of documents relevant to a query issued by a user. We propose RankDE, a ranking method that uses differential evolution (DE) to learn a ranking function to rank a list of documents retrieved by a Web search engine. To the best of our knowledge, the proposed method is the first DE-based approach to learn a ranking function for IR. We evaluate the proposed method using LETOR dataset, a standard benchmark dataset for training and evaluating ranking functions for IR. In our experiments, the proposed method significantly outperforms previously proposed rank learning methods that use evolutionary computation algorithms such as Particle Swam Optimization (PSO) and Genetic Programming (GP), achieving a statistically significant mean average precision (MAP) of 0.339 on TD2003 dataset and 0.430 on the TD2004 dataset. Moreover, the proposed method shows comparable results to the state-of-the-art non-evolutionary computational approaches on this benchmark dataset. We analyze the feature weights learnt by the proposed method to better understand the salient features for the task of learning to rank for information retrieval.
Danushka Bollegala, Nasimul Noman, Hitoshi Iba
GECCO1
2011 Differential evolution with self adaptive local search
abstract
The performance of a memetic algorithm (MA) largely depends on the synergy between its global and local search counterparts. The amount of global exploration and local exploitation to be carried out, for optimal performance, varies with problem type. Therefore, an algorithm should intelligently allocate its computational efforts between genetic search and local search. In this work we propose an adaptive local search method that adjusts the effort for local tuning of individuals, taking feedback from the search. We implemented an MA hybridizing this adaptive local search method with differential evolution algorithm. Experimenting with a standard benchmark suite it was found that the proposed MA can utilize its global and local search components adaptively. The proposed algorithm also exhibited very competitive performance with other existing algorithms.
Nasimul Noman, Danushka Bollegala, Hitoshi Iba
GECCO2
2011 Exploiting User Interest on Social Media for Aggregating Diverse Data and Predicting Interest
Nozomi Nori, Danushka Bollegala, Mitsuru Ishizuka
ICWSM2
2011 Relation Adaptation: Learning to Extract Novel Relations with Minimum Supervision
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IJCAI1
2011 Interest Prediction on Multinomial, Time-Evolving Social Graph
Nozomi Nori, Danushka Bollegala, Mitsuru Ishizuka
IJCAI2
2011 Automatic Discovery of Personal Name Aliases from the Web
abstract
An individual is typically referred by numerous name aliases on the web. Accurate identification of aliases of a given person name is useful in various web related tasks such as information retrieval, sentiment analysis, personal name disambiguation, and relation extraction. We propose a method to extract aliases of a given personal name from the web. Given a personal name, the proposed method first extracts a set of candidate aliases. Second, we rank the extracted candidates according to the likelihood of a candidate being a correct alias of the given name. We propose a novel, automatically extracted lexical pattern-based approach to efficiently extract a large set of candidate aliases from snippets retrieved from a web search engine. We define numerous ranking scores to evaluate candidate aliases using three approaches: lexical pattern frequency, word co-occurrences in an anchor text graph, and page counts on the web. To construct a robust alias detection system, we integrate the different ranking scores into a single ranking function using ranking support vector machines. We evaluate the proposed method on three data sets: an English personal names data set, an English place names data set, and a Japanese personal names data set. The proposed method outperforms numerous baselines and previously proposed name alias extraction methods, achieving a statistically significant mean reciprocal rank (MRR) of 0.67. Experiments carried out using location names and Japanese personal names suggest the possibility of extending the proposed method to extract aliases for different types of named entities, and for different languages. Moreover, the aliases extracted using the proposed method are successfully utilized in an information retrieval task and improve recall by 20 percent in a relation-detection task.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IEEE Trans. Knowl. Data Eng.1
2011 A Web Search Engine-Based Approach to Measure Semantic Similarity between Words
abstract
Measuring the semantic similarity between words is an important component in various tasks on the web such as relation extraction, community mining, document clustering, and automatic metadata extraction. Despite the usefulness of semantic similarity measures in these applications, accurately measuring semantic similarity between two words (or entities) remains a challenging task. We propose an empirical method to estimate semantic similarity using page counts and text snippets retrieved from a web search engine for two words. Specifically, we define various word co-occurrence measures using page counts and integrate those with lexical patterns extracted from text snippets. To identify the numerous semantic relations that exist between two given words, we propose a novel pattern extraction algorithm and a pattern clustering algorithm. The optimal combination of page counts-based co-occurrence measures and lexical pattern clusters is learned using support vector machines. The proposed method outperforms various baselines and previously proposed web-based semantic similarity measures on three benchmark data sets showing a high correlation with human ratings. Moreover, the proposed method significantly improves the accuracy in a community mining task.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IEEE Trans. Knowl. Data Eng.1
2010 A Sequential Model for Discourse Segmentation
Hugo Hernault, Danushka Bollegala, Mitsuru Ishizuka
CICLing2
2010 A Semi-Supervised Approach to Improve Classification of Infrequent Discourse Relations Using Feature Vector Extension
Hugo Hernault, Danushka Bollegala, Mitsuru Ishizuka
EMNLP2
2010 Exploiting Symmetry in Relational Similarity for Ranking Relational Search Results
Tomokazu Goto, Nguyen Tuan Duc, Danushka Bollegala, Mitsuru Ishizuka
PRICAI3
2010 Towards Semi-Supervised Classification of Discourse Relations using Feature Correlations
Hugo Hernault, Danushka Bollegala, Mitsuru Ishizuka
SIGDIAL Conference2
2010 Using Relational Similarity between Word Pairs for Latent Relational Search on the Web
abstract
Latent relational search is a new search paradigm based on the degree of analogy between two word pairs. A latent relational search engine is expected to return the word Paris as an answer to the question mark (?) in the query {(Japan, Tokyo), (France, ?)} because the relation between Japan and Tokyo is highly similar to that between France and Paris. We propose an approach for exploring and indexing word pairs to efficiently retrieve candidate answers for a latent relational search query. Representing relations between two words in a word pair by lexical patterns allows our search engine to achieve a high MRR and high precision for the top 1 ranked result. When evaluating with a Web corpus, the proposed method achieves an MRR of 0.963 and it retrieves correct answer in the top 1 for 95.0% of queries.
Nguyen Tuan Duc, Danushka Bollegala, Mitsuru Ishizuka
Web Intelligence2
2010 Relational duality: unsupervised extraction of semantic relations between entities on the web
abstract
Extracting semantic relations among entities is an important first step in various tasks in Web mining and natural language processing such as information extraction, relation detection, and social network mining. A relation can be expressed extensionally by stating all the instances of that relation or intensionally by defining all the paraphrases of that relation. For example, consider the ACQUISITION relation between two companies. An extensional definition of ACQUISITION contains all pairs of companies in which one company is acquired by another (e.g. (YouTube, Google) or (Powerset, Microsoft)). On the other hand we can intensionally define ACQUISITION as the relation described by lexical patterns such as X is acquired by Y, or Y purchased X, where X and Y denote two companies. We use this dual representation of semantic relations to propose a novel sequential co-clustering algorithm that can extract numerous relations efficiently from unlabeled data. We provide an efficient heuristic to find the parameters of the proposed coclustering algorithm. Using the clusters produced by the algorithm, we train an L1 regularized logistic regression model to identify the representative patterns that describe the relation expressed by each cluster. We evaluate the proposed method in three different tasks: measuring relational similarity between entity pairs, open information extraction (Open IE), and classifying relations in a social network system. Experiments conducted using a benchmark dataset show that the proposed method improves existing relational similarity measures. Moreover, the proposed method significantly outperforms the current state-of-the-art Open IE systems in terms of both precision and recall. The proposed method correctly classifies 53 relation types in an online social network containing 470; 671 nodes and 35; 652; 475 edges, thereby demonstrating its efficacy in real-world relation detection tasks.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WWW1
2010 A bottom-up approach to sentence ordering for multi-document summarization
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka
Inf. Process. Manag.1
2009 A Relational Model of Semantic Similarity between Words using Automatically Extracted Lexical Pattern Clusters from the Web
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
EMNLP1
2009 Measuring the similarity between implicit semantic relations using web search engines
abstract
Measuring the similarity between implicit semantic relations is an important task in information retrieval and natural language processing. For example, consider the situation where you know an entity-pair (e.g. Google, YouTube), between which a particular relation holds (e.g. acquisition), and you are interested in retrieving other entity-pairs for which the same relation holds (e.g. Yahoo, Inktomi). Existing keyword-based search engines cannot be directly applied in this case because in keyword-based search, the goal is to retrieve documents that are relevant to the words used in the query -- not necessarily to the relations implied by a pair of words. Accurate measurement of relational similarity is an important step in numerous natural language processing tasks such as identification of word analogies, and classification of noun-modifier pairs. We propose a method that uses Web search engines to efficiently compute the relational similarity between two pairs of words. Our method consists of three components: representing the various semantic relations that exist between a pair of words using automatically extracted lexical patterns, clustering the extracted lexical patterns to identify the different semantic relations implied by them, and measuring the similarity between different semantic relations using an inter-cluster correlation matrix. We propose a pattern extraction algorithm to extract a large number of lexical patterns that express numerous semantic relations. We then present an efficient clustering algorithm to cluster the extracted lexical patterns. Finally, we measure the relational similarity between word-pairs using inter-cluster correlation. We evaluate the proposed method in a relation classification task. Experimental results on a dataset covering multiple relation types show a statistically significant improvement over the current state-of-the-art relational similarity measures.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WSDM1
2009 Measuring the similarity between implicit semantic relations from the web
abstract
Measuring the similarity between semantic relations that hold among entities is an important and necessary step in various Web related tasks such as relation extraction, information retrieval and analogy detection. For example, consider the case in which a person knows a pair of entities (e.g. Google, YouTube), between which a particular relation holds (e.g. acquisition). The person is interested in retrieving other such pairs with similar relations (e.g. Microsoft, Powerset). Existing keyword-based search engines cannot be applied directly in this case because, in keyword-based search, the goal is to retrieve documents that are relevant to the words used in a query -- not necessarily to the relations implied by a pair of words. We propose a relational similarity measure, using a Web search engine, to compute the similarity between semantic relations implied by two pairs of words. Our method has three components: representing the various semantic relations that exist between a pair of words using automatically extracted lexical patterns, clustering the extracted lexical patterns to identify the different patterns that express a particular semantic relation, and measuring the similarity between semantic relations using a metric learning approach. We evaluate the proposed method in two tasks: classifying semantic relations between named entities, and solving word-analogy questions. The proposed method outperforms all baselines in a relation classification task with a statistically significant average precision score of 0.74. Moreover, it reduces the time taken by Latent Relational Analysis to process 374 word-analogy questions from 9 days to less than 6 hours, with an SAT score of 51%.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WWW1
2008 WWW sits the SAT: Measuring Relational Similarity on the Web
abstract
Measuring relational similarity between words is important in numerous natural language processing tasks such as solving analogy questions and classifying noun-modifier relations. We propose a method to measure the similarity between semantic relations that hold between two pairs of words using a web search engine. First, each pair of words is represented by a vector of automatically extracted lexical patterns. Then a Support Vector Machine is trained to recognize word pairs with similar semantic relations. We evaluate the proposed method on SAT multiple-choice word-analogy questions. The proposed method achieves a score of 40% which is comparable with relational similarity measures which use manually created resources such as WordNet. The proposed method significantly reduces the time taken by previously proposed computationally intensive methods, such as latent relational analysis, to process 374 analogy questions from 8 days to less than 6 hours.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
ECAI1
2008 A Co-occurrence Graph-based Approach for Personal Name Alias Extraction from Anchor Texts
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IJCNLP1
2008 Mining for personal name aliases on the web
abstract
We propose a novel approach to find aliases of a given name from the web. We exploit a set of known names and their aliases as training data and extract lexical patterns that convey information related to aliases of names from text snippets returned by a web search engine. The patterns are then used to find candidate aliases of a given name. We use anchor texts and hyperlinks to design a word co-occurrence model and define numerous ranking scores to evaluate the association between a name and its candidate aliases. The proposed method outperforms numerous baselines and previous work on alias extraction on a dataset of personal names, achieving a statistically significant mean reciprocal rank of 0.6718. Moreover, the aliases extracted using the proposed method improve recall by 20% in a relation-detection task.
Danushka Bollegala, Taiki Honma, Yutaka Matsuo, Mitsuru Ishizuka
WWW1
2007 An Integrated Approach to Measuring Semantic Similarity between Words Using Information Available on the Web
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
HLT-NAACL1
2007 Measuring semantic similarity between words using web search engines
abstract
Article Share on Measuring semantic similarity between words using web search enginesWWW '07: Proceedings of the 16th international conference on World Wide WebMay 2007 Pages 757–766https://doi.org/10.1145/1242572.1242675Online:08 May 2007Publication History 123citation3,949DownloadsMetricsTotal Citations123Total Downloads3,949Last 12 Months94Last 6 weeks13 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WWW1
2006 Spinning Multiple Social Networks for Semantic Web
Yutaka Matsuo, Masahiro Hamasaki, Yoshiyuki Nakamura, Takuichi Nishimura, Kôiti Hasida, Hideaki Takeda 0001, Junichiro Mori, Danushka Bollegala, Mitsuru Ishizuka
AAAI8
2006 A Bottom-Up Approach to Sentence Ordering for Multi-Document Summarization
abstract
Ordering information is a difficult but important task for applications generating natural-language text. We present a bottom-up approach to arranging sentences extracted for multi-document summarization. To capture the association and order of two textual segments (eg, sentences), we define four criteria, chronology, topical-closeness, precedence, and succession. These criteria are integrated into a criterion by a supervised learning approach. We repeatedly concatenate two textual segments into one segment based on the criterion until we obtain the overall segment with all sentences arranged. Our experimental results show a significant improvement over existing sentence ordering strategies.
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka
ACL1
2006 Extracting Key Phrases to Disambiguate Personal Names on the Web
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
CICLing1
2006 Disambiguating Personal Names on the Web Using Automatically Extracted Key Phrases
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
ECAI1
2005 A Machine Learning Approach to Sentence Ordering for Multidocument Summarization and Its Evaluation
Danushka Bollegala, Naoaki Okazaki, Mitsuru Ishizuka
IJCNLP1