EDBT 2026 Demo / reviewers in the wild / expert
Nigel Collier
dblp:90/2619 · also Nigel H. Collier
· DBLP profile ↗
88ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0002-7230-4164ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 10 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Value of Information: A Framework for Human-Agent CommunicationabstractYijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulić, Andreea Bobu, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulic, Andreea Bobu, Nigel Collier |
ACL (1) | 7 |
| 2026 | Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement LearningabstractZheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi, Nigel Collier |
ACL (1) | 5 |
| 2026 | Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition BottleneckabstractDespite scaling to massive context windows, Large Language Models (LLMs) struggle with multi-hop reasoning due to inherent position bias, which causes them to overlook information at certain positions.Whether these failures stem from an inability to locate evidence (recognition failure) or integrate it (synthesis failure) is unclear.We introduce Multi-Focus Attention Instruction (MFAI), a semantic probe to disentangle these mechanisms by explicitly steering attention towards selected positions.Across 5 LLMs on two multi-hop QA tasks (MuSiQue and NeoQA), we identify the "Weakest Link Effect": in our 18document, 3-bucket setting, multi-hop reasoning performance collapses to the level of the least visible evidence, governed by absolute position rather than the linear distance between facts.While matched MFAI resolves recognition bottlenecks, improving accuracy by up to 11.49% in low-visibility positions, misleading MFAI yields divergent effects modulated by task topology: entity-centric tasks with vertical reasoning chains are vulnerable, whereas eventcentric tasks with horizontal evidence structures are more resilient.Finally, we demonstrate that "thinking" models utilizing System-2 reasoning effectively locate and integrate the required information, matching gold-only baselines even in noisy, long-context settings.Supplementary experiments on 2WikiMulti-HopQA, extended 3-4 hop counts, and a 32B model confirm these findings generalize across datasets, reasoning depths, and model scales. Meiru Zhang, Zaiqiao Meng, Nigel Collier |
ACL (1) | 3 |
| 2026 | LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form GenerationabstractHallucination remains a major challenge for the safe and trustworthy deployment of large language models (LLMs) in factual content generation.Prior work has explored confidence estimation as an effective approach to hallucination detection, but often relies on post-hoc self-consistency methods that require computationally expensive sampling.Verbalized confidence offers a more efficient alternative, but existing approaches are largely limited to shortform question answering (QA) tasks and do not generalize well to open-ended generation.In this paper, we propose LOVEC (Long-form Verbalized Confidence), a novel reinforcement learning (RL)-based method that trains LLMs to append an on-the-fly numerical confidence score to each generated statement during longform generation.The confidence score serves as a direct and interpretable signal of the factuality of generation.We introduce two evaluation settings, free-form tagging and iterative tagging, to assess different verbalized confidence estimation methods.Experiments on three long-form QA datasets show that our RLtrained models achieve better calibration and generalize robustly across domains.Also, our method is highly efficient, being 20× faster than traditional self-consistency methods while achieving better calibration. Caiqi Zhang, Chengzu Li, Nigel Collier, Andreas Vlachos 0001 |
ACL (1) | 4 |
| 2025 | iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to NewsabstractUnderstanding how individuals perceive and react to information is fundamental for advancing social and behavioral sciences and developing human-centered AI systems.Current approaches often lack the granular data needed to model these personalized responses, relying instead on aggregated labels that obscure the rich variability driven by individual differences.We introduce iNews, a novel large-scale dataset specifically designed to facilitate the modeling of personalized affective responses to news content.Our dataset comprises annotations from 291 demographically diverse UK participants across 2,899 multimodal Facebook news posts from major UK outlets, with an average of 5.18 annotators per sample.For each post, annotators provide multifaceted labels including valence, arousal, dominance, discrete emotions, content relevance judgments, sharing likelihood, and modality importance ratings.Crucially, we collect comprehensive annotator persona information covering demographics, personality, media trust, and consumption patterns, which explain 15.2% of annotation variance -substantially higher than existing NLP datasets.Incorporating this information yields a 7% accuracy gain in zero-shot prediction and remains beneficial even with 32-shot in-context learning. Tiancheng Hu, Nigel Collier |
ACL (1) | 2 |
| 2025 | 500xCompressor: Generalized Prompt Compression for Large Language ModelsabstractPrompt compression is important for large language models (LLMs) to increase inference speed, reduce costs, and improve user experience.However, current methods face challenges such as low compression ratios and potential training-test overlap during evaluation.To address these issues, we propose 500xCompressor, a method that compresses natural language contexts into a minimum of one special token and demonstrates strong generalization ability.The 500xCompressor introduces approximately 0.3% additional parameters and achieves compression ratios ranging from 6x to 500x, achieving 27-90% reduction in calculations and 55-83% memory savings when generating 100-400 tokens for new and reused prompts at 500x compression, while retaining 70-74% (F1) and 77-84% (Exact Match) of the LLM capabilities compared to using non-compressed prompts.It is designed to compress any text, answer various types of questions, and can be utilized by the original LLM without requiring fine-tuning.Initially, 500xCompressor was pretrained on the Arx-ivCorpus, followed by fine-tuning on the Arx-ivQA dataset, and subsequently evaluated on strictly unseen and cross-domain question answering (QA) datasets.This study shows that KV values outperform embeddings in preserving information at high compression ratios.The highly compressive nature of natural language prompts, even for detailed information, suggests potential for future applications and the development of a new LLM language.1 Zongqian Li, Yixuan Su, Nigel Collier |
ACL (1) | 3 |
| 2025 | LoGU: Long-form Generation with Uncertainty ExpressionsabstractRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Sen Yang, Nigel Collier, Dong Yu, Deqing Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Nigel Collier, Dong Yu 0001, Deqing Yang |
ACL (1) | 6 |
| 2025 | Conformity in Large Language ModelsabstractThe conformity effect describes the tendency of individuals to align their responses with the majority.Studying this bias in large language models (LLMs) is crucial, as LLMs are increasingly used in various information-seeking and decision-making tasks as conversation partners to improve productivity.Thus, conformity to incorrect responses can compromise their effectiveness.In this paper, we adapt psychological experiments to examine the extent of conformity in popular LLMs.Our findings reveal that all tested models exhibit varying levels of conformity toward the majority, regardless of their initial choice or correctness, across different knowledge domains.Notably, we are the first to show that LLMs are more likely to conform when they are more uncertain in their own prediction.We further explore factors that influence conformity, such as training paradigms and input characteristics, finding that instruction-tuned models are less susceptible to conformity, while increasing the naturalness of majority tones amplifies conformity.Finally, we propose two interventions, Devil's Advocate and Question Distillation, to mitigate conformity, providing insights into building more robust language models. What is the oldest college in Cambridge?It is Peterhouse College. What is the oldest college in Cambridge?King's. Caiqi Zhang, Tom Stafford 0002, Nigel Collier, Andreas Vlachos 0001 |
ACL (1) | 4 |
| 2025 | UNCLE: Benchmarking Uncertainty Expressions in Long-Form GenerationabstractLarge Language Models (LLMs) are prone to hallucination, particularly in long-form generations.A promising direction to mitigate hallucination is to teach LLMs to express uncertainty explicitly when they lack sufficient knowledge.However, existing work lacks direct and fair evaluation of LLMs' ability to express uncertainty effectively in long-form generation.To address this gap, we first introduce UNCLE, a benchmark designed to evaluate uncertainty expression in both long-and short-form question answering (QA).UNCLE covers five domains and includes more than 1,000 entities, each with paired short-and long-form QA items.Our dataset is the first to directly link short-and long-form QA through aligned questions and gold-standard answers.Along with UNCLE, we propose a suite of new metrics to assess the models' capabilities to selectively express uncertainty.We then demonstrate that current models fail to convey uncertainty appropriately in long-form generation.We further explore both prompt-based and training-based methods to improve models' performance, with the training-based methods yielding greater gains.Further analysis of alignment gaps between short-and long-form uncertainty expression highlights promising directions for future research using UNCLE. Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Dong Yu 0001, Nigel Collier, Deqing Yang |
EMNLP | 6 |
| 2025 | All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model ReasoningabstractConfidence estimation is essential for the reliable deployment of large language models (LLMs).Existing methods are primarily designed for factual QA tasks and often fail to generalize to reasoning tasks.To address this gap, we propose a set of training-free, graph-based confidence estimation methods tailored to reasoning tasks.Our approach models reasoning paths as directed graphs and estimates confidence by exploiting graph properties such as centrality, path convergence, and path weighting.Experiments with two LLMs on three reasoning datasets demonstrate improved confidence estimation and enhanced performance on two downstream tasks. Caiqi Zhang, Ehsan Shareghi, Nigel Collier |
EMNLP | 4 |
| 2025 | Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsabstractLarge Language Models (LLMs) are expected to be predictable and trustworthy to support reliable decision-making systems. Yet current LLMs often show inconsistencies in their judgments. In this work, we examine \textit{logical preference consistency} as a foundational requirement for building more dependable LLM systems, ensuring stable and coherent decision-making while minimizing erratic or contradictory outputs.
To quantify the logical preference consistency, we propose a universal evaluation framework based on three fundamental properties: *transitivity*, *commutativity* and *negation invariance*.
Through extensive experimentation across diverse LLMs, we demonstrate that these properties serve as strong indicators of judgment robustness.
Furthermore, we introduce a data refinement and augmentation technique, REPAIR, that enhances logical consistency while maintaining alignment with human preferences. Finally, we show that improving consistency leads to better performance in LLM-driven logic-based algorithms, reinforcing stability and coherence in decision-making systems. Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi, Ivan Vulic, Nigel Collier |
ICML | 6 |
| 2025 | Prompt Compression for Large Language Models: A SurveyabstractZongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier |
NAACL (Long Papers) | 4 |
| 2025 | PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt TuningabstractParameter-efficient fine-tuning (PEFT) methods have shown promise in adapting large language models, yet existing approaches exhibit counter-intuitive phenomena: integrating either matrix decomposition or mixture-of-experts (MoE) individually decreases performance across tasks, though decomposition improves results on specific domains despite reducing parameters, while MoE increases parameter count without corresponding decrease in training efficiency. Motivated by these observations and the modular nature of PT, we propose PT-MoE, a novel framework that integrates matrix decomposition with MoE routing for efficient PT. Evaluation results across 17 datasets demonstrate that PT-MoE achieves state-of-the-art performance in both question answering (QA) and mathematical problem solving tasks, improving F1 score by 1.49 points over PT and 2.13 points over LoRA in QA tasks, while improving mathematical accuracy by 10.75 points over PT and 0.44 points over LoRA, all while using 25% fewer parameters than LoRA. Our analysis reveals that while PT methods generally excel in QA tasks and LoRA-based methods in math datasets, the integration of matrix decomposition and MoE in PT-MoE yields complementary benefits: decomposition enables efficient parameter sharing across experts while MoE provides dynamic adaptation, collectively enabling PT-MoE to demonstrate cross-task consistency and generalization abilities. These findings, along with ablation studies on routing mechanisms and architectural components, provide insights for future PEFT methods. Zongqian Li, Yixuan Su, Nigel Collier |
NeurIPS | 3 |
| 2024 | BAND: Biomedical Alert News DatasetabstractInfectious disease outbreaks continue to pose a significant threat to human health and well-being. To improve disease surveillance and understanding of disease spread, several surveillance systems have been developed to monitor daily news alerts and social media. However, existing systems lack thorough epidemiological analysis in relation to corresponding alerts or news, largely due to the scarcity of well-annotated reports data. To address this gap, we introduce the Biomedical Alert News Dataset (BAND), which includes 1,508 samples from existing reported news articles, open emails, and alerts, as well as 30 epidemiology-related questions. These questions necessitate the model's expert reasoning abilities, thereby offering valuable insights into the outbreak of the disease. The BAND dataset brings new challenges to the NLP world, requiring better inference capability of the content and the ability to infer important information. We provide several benchmark tasks, including Named Entity Recognition (NER), Question Answering (QA), and Event Extraction (EE), to demonstrate existing models' capabilities and limitations in handling epidemiology-specific tasks. It is worth noting that some models may lack the human-like inference capability required to fully utilize the corpus. To the best of our knowledge, the BAND corpus is the largest corpus of well-annotated biomedical outbreak alert news with elaborately designed questions, making it a valuable resource for epidemiologists and NLP researchers alike. Meiru Zhang, Zaiqiao Meng, Yannan Shen, David L. Buckeridge, Nigel Collier |
AAAI | 6 |
| 2024 | Quantifying the Persona Effect in LLM SimulationsabstractLarge language models (LLMs) have shown remarkable promise in simulating human language and behavior.This study investigates how integrating persona variables-demographic, social, and behavioral factors-impacts LLMs' ability to simulate diverse perspectives.We find that persona variables account for <10% variance in annotations in existing subjective NLP datasets.Nonetheless, incorporating persona variables via prompting in LLMs provides modest but statistically significant improvements.Persona prompting is most effective in samples where many annotators disagree, but their disagreements are relatively minor.Notably, we find a linear relationship in our setting: the stronger the correlation between persona variables and human annotations, the more accurate the LLM predictions are using persona prompting.In a zero-shot setting, a powerful 70b model with persona prompting captures 81% of the annotation variance achievable by linear regression trained on ground truth annotations.However, for most subjective NLP datasets, where persona variables have limited explanatory power, the benefits of persona prompting are limited.1 Tiancheng Hu, Nigel Collier |
ACL (1) | 2 |
| 2024 | TopViewRS: Vision-Language Models as Top-View Spatial ReasonersabstractTop-view perspective denotes a typical way in which humans read and reason over different types of maps, and it is vital for localization and navigation of humans as well as of 'non-human' agents, such as the ones backed by large Vision-Language Models (VLMs).Nonetheless, spatial reasoning capabilities of modern VLMs in this setup remain unattested and underexplored.In this work, we study their capability to understand and reason over spatial relations from the top view.The focus on top view also enables controlled evaluations at different granularity of spatial reasoning; we clearly disentangle different abilities (e.g., recognizing particular objects versus understanding their relative positions).We introduce the TOPVIEWRS (Top-View Reasoning in Space) dataset, consisting of 11,384 multiple-choice questions with either realistic or semantic top-view map as visual input.We then use it to study and evaluate VLMs across 4 perception and reasoning tasks with different levels of complexity.Evaluation of 10 representative open-and closedsource VLMs reveals the gap of more than 50% compared to average human performance, and it is even lower than the random baseline in some cases.Although additional experiments show that Chain-of-Thought reasoning can boost model capabilities by 5.82% on average, the overall performance of VLMs remains limited.Our findings underscore the critical need for enhanced model capability in top-view spatial reasoning and set a foundation for further research towards human-level proficiency of VLMs in real-world multimodal tasks. Chengzu Li, Caiqi Zhang, Han Zhou 0010, Nigel Collier, Anna Korhonen, Ivan Vulic |
EMNLP | 4 |
| 2024 | LUQ: Long-text Uncertainty Quantification for LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable capability in a variety of NLP tasks.However, LLMs are also prone to generate nonfactual content.Uncertainty Quantification (UQ) is pivotal in enhancing our understanding of a model's confidence on its generation, thereby aiding in the mitigation of nonfactual outputs.Existing research on UQ predominantly targets short text generation, typically yielding brief, word-limited responses.However, real-world applications frequently necessitate much longer responses.Our study first highlights the limitations of current UQ methods in handling long text generation.We then introduce LUQ with its two variations: LUQ-ATOMIC and LUQ-PAIR, a series of novel sampling-based UQ approaches specifically designed for long text.Our findings reveal that LUQ outperforms existing baseline methods in correlating with the model's factuality scores (negative coefficient of -0.85 observed for Gemini Pro).To further improve the factuality of LLM responses, we propose LUQ-ENSEMBLE, a method that ensembles responses from multiple models and selects the response with the lowest uncertainty.The ensembling method greatly improves the response factuality upon the best standalone LLM. 1 * Now at Google DeepMind. Caiqi Zhang, Fangyu Liu 0001, Marco Basaldella, Nigel Collier |
EMNLP | 4 |
| 2024 | Fairer Preferences Elicit Improved Human-Aligned Large Language Model JudgmentsabstractLarge language models (LLMs) have shown promising abilities as cost-effective and reference-free evaluators for assessing language generation quality.In particular, pairwise LLM evaluators, which compare two generated texts and determine the preferred one, have been employed in a wide range of applications.However, LLMs exhibit preference biases and worrying sensitivity to prompt designs.In this work, we first reveal that the predictive preference of LLMs can be highly brittle and skewed, even with semantically equivalent instructions.We find that fairer predictive preferences from LLMs consistently lead to judgments that are better aligned with humans.Motivated by this phenomenon, we propose an automatic Zero-shot Evaluation-oriented Prompt Optimization framework, ZEPO, which aims to produce fairer preference decisions and improve the alignment of LLM evaluators with human judgments.To this end, we propose a zeroshot learning objective based on the preference decision fairness.ZEPO demonstrates substantial performance improvements over stateof-the-art LLM evaluators, without requiring labeled data, on representative meta-evaluation benchmarks.Our findings underscore the critical correlation between preference fairness and human alignment, positioning ZEPO as an efficient prompt optimizer for bridging the gap between LLM evaluators and human judgments.* Now at Google.Code is available at https://github. com/cambridgeltl/zepo.Generate new prompts Zero-shot Fairness ZEPO Biased Preference Fairer Preference Optimized Prompt Initial Prompt LLM Optimizer LLM Evaluator Which summary candidate has better coherence?If the candidate A is better, please return 'A'.If the candidate B is better, please return 'B'.Which one exhibits better coherence?Return 'A' for the rst summary or 'B' for the second.Only provide the letter of your choice. Han Zhou 0010, Xingchen Wan, Yinhong Liu, Nigel Collier, Ivan Vulic, Anna Korhonen |
EMNLP | 4 |
| 2023 | On the Effectiveness of Parameter-Efficient Fine-TuningabstractFine-tuning pre-trained models has been ubiquitously proven to be effective in a wide range of NLP tasks. However, fine-tuning the whole model is parameter inefficient as it always yields an entirely new model for each task. Currently, many research works propose to only fine-tune a small portion of the parameters while keeping most of the parameters shared across different tasks. These methods achieve surprisingly good performance and are shown to be more stable than their corresponding fully fine-tuned counterparts. However, such kind of methods is still not well understood. Some natural questions arise: How does the parameter sparsity lead to promising performance? Why is the model more stable than the fully fine-tuned models? How to choose the tunable parameters? In this paper, we first categorize the existing methods into random approaches, rule-based approaches, and projection-based approaches based on how they choose which parameters to tune. Then, we show that all of the methods are actually sparse fine-tuned models and conduct a novel theoretical analysis of them. We indicate that the sparsity is actually imposing a regularization on the original model by controlling the upper bound of the stability. Such stability leads to better generalization capability which has been empirically observed in a lot of recent research works. Despite the effectiveness of sparsity grounded by our theory, it still remains an open problem of how to choose the tunable parameters. Currently, the random and rule-based methods do not utilize task-specific data information while the projection-based approaches suffer from the projection discontinuity problem. To better choose the tunable parameters, we propose a novel Second-order Approximation Method (SAM) which approximates the original problem with an analytically solvable optimization function. The tunable parameters are determined by directly optimizing the approximation function. We conduct extensive experiments on several tasks. The experimental results show that our proposed SAM model outperforms many strong baseline models and it also verifies our theoretical analysis. The source code of this paper can be obtained from https://github.com/fuzihaofzh/AnalyzeParameterEff\/icientFinetune . Anthony Man-Cho So, Wai Lam, Lidong Bing, Nigel Collier |
AAAI | 6 |
| 2023 | MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart DerenderingabstractFangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Eisenschlos. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Fangyu Liu 0001, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Martin Eisenschlos |
ACL (1) | 8 |
| 2023 | On Reality and the Limits of Language Data: Aligning LLMs with Human Norms
Nigel Collier, Fangyu Liu 0001, Ehsan Shareghi |
CogSci | 1 |
| 2023 | Probing Cross-Lingual Lexical Knowledge from Multilingual Sentence EncodersabstractIvan Vulić, Goran Glavaš, Fangyu Liu, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ivan Vulic, Goran Glavas, Fangyu Liu 0001, Nigel Collier, Edoardo Maria Ponti, Anna Korhonen |
EACL | 4 |
| 2023 | Biomedical Named Entity Recognition via Dictionary-based Synonym GeneralizationabstractBiomedical named entity recognition is one of the core tasks in biomedical natural language processing (BioNLP).To tackle this task, numerous supervised/distantly supervised approaches have been proposed.Despite their remarkable success, these approaches inescapably demand laborious human effort.To alleviate the need of human effort, dictionarybased approaches have been proposed to extract named entities simply based on a given dictionary.However, one downside of existing dictionary-based approaches is that they are challenged to identify concept synonyms that are not listed in the given dictionary, which we refer as the synonym generalization problem.In this study, we propose a novel Synonym Generalization (SynGen) framework that recognizes the biomedical concepts contained in the input text using span-based predictions.In particular, SynGen introduces two regularization terms, namely, (1) a synonym distance regularizer; and (2) a noise perturbation regularizer, to minimize the synonym generalization error.To demonstrate the effectiveness of our approach, we provide a theoretical analysis of the bound of synonym generalization error.We extensively evaluate our approach on a wide range of benchmarks and the results verify that SynGen outperforms previous dictionary-based models by notable margins.Lastly, we provide a detailed analysis to further reveal the merits and inner-workings of our approach.1 Yixuan Su, Zaiqiao Meng, Nigel Collier |
EMNLP | 4 |
| 2023 | Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data PerspectiveabstractThere are a number of diverging hypotheses about the neural text degeneration problem, i.e., generating repetitive and dull loops, which makes this problem both interesting and confusing. In this work, we aim to advance our understanding by presenting a straightforward and fundamental explanation from the data perspective. Our preliminary investigation reveals a strong correlation between the degeneration issue and the presence of repetitions in training data. Subsequent experiments also demonstrate that by selectively dropping out the attention to repetitive words in training data, degeneration can be significantly minimized. Furthermore, our empirical analysis illustrates that prior works addressing the degeneration issue from various standpoints, such as the high-inflow words, the likelihood objective, and the self-reinforcement phenomenon, can be interpreted by one simple explanation. That is, penalizing the repetitions in training data is a common and fundamental factor for their effectiveness. Moreover, our experiments reveal that penalizing the repetitions in training data remains critical even when considering larger model sizes and instruction tuning. Tian Lan 0003, Deng Cai 0002, Lemao Liu, Nigel Collier, Taro Watanabe, Yixuan Su |
NeurIPS | 6 |
| 2023 | Visual Spatial ReasoningabstractAbstract Spatial relations are a basic part of human cognition. However, they are expressed in natural language in a variety of ways, and previous work has suggested that current vision-and-language models (VLMs) struggle to capture relational information. In this paper, we present Visual Spatial Reasoning (VSR), a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing). While using a seemingly simple annotation format, we show how the dataset includes challenging linguistic phenomena, such as varying reference frames. We demonstrate a large gap between human and model performance: The human ceiling is above 95%, while state-of-the-art models only achieve around 70%. We observe that VLMs’ by-relation performances have little correlation with the number of training examples and the tested models are in general incapable of recognising relations concerning the orientations of objects.1 Fangyu Liu 0001, Guy Emerson, Nigel Collier |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Incorporating Stock Market Signals for Twitter Stance DetectionabstractCostanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier |
ACL (1) | 6 |
| 2022 | Improving Word Translation via Two-Stage Contrastive LearningabstractWord translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages.In this work, we propose a robust and effective two-stage contrastive learning framework for the BLI task.At Stage C1, we propose to refine standard cross-lingual linear maps between static word embeddings (WEs) via a contrastive learning objective; we also show how to integrate it into the self-learning procedure for even more refined cross-lingual maps.In Stage C2, we conduct BLI-oriented contrastive fine-tuning of mBERT, unlocking its word translation capability.We also show that static WEs induced from the 'C2-tuned' mBERT complement static WEs from Stage C1.Comprehensive experiments on standard BLI datasets for diverse languages and different experimental setups demonstrate substantial gains achieved by our framework.While the BLI method from Stage C1 already yields substantial gains over all state-of-the-art BLI methods in our comparison, even stronger improvements are met with the full two-stage framework: e.g., we report gains for 112/112 BLI setups, spanning 28 language pairs. Yaoyiran Li, Fangyu Liu 0001, Nigel Collier, Anna Korhonen, Ivan Vulic |
ACL (1) | 3 |
| 2022 | Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language ModelsabstractZaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zaiqiao Meng, Fangyu Liu 0001, Ehsan Shareghi, Yixuan Su, Charlotte Collins, Nigel Collier |
ACL (1) | 6 |
| 2022 | Prix-LM: Pretraining for Multilingual Knowledge Base ConstructionabstractKnowledge bases (KBs) contain plenty of structured world and commonsense knowledge.As such, they often complement distributional text-based information and facilitate various downstream tasks.Since their manual construction is resource-and timeintensive, recent efforts have tried leveraging large pretrained language models (PLMs) to generate additional monolingual knowledge facts for KBs.However, such methods have not been attempted for building and enriching multilingual KBs.Besides wider application, such multilingual KBs can provide richer combined knowledge than monolingual (e.g., English) KBs.Knowledge expressed in different languages may be complementary and unequally distributed: this implies that the knowledge available in high-resource languages can be transferred to low-resource ones.To achieve this, it is crucial to represent multilingual knowledge in a shared/unified space.To this end, we propose a unified representation model, Prix-LM , for multilingual KB construction and completion.We leverage two types of knowledge, monolingual triples and cross-lingual links, extracted from existing multilingual KBs, and tune a multilingual language encoder XLM-R via a causal language modeling objective.Prix-LM integrates useful multilingual and KB-based factual knowledge into a single model.Experiments on standard entity-related tasks, such as link prediction in multiple languages, cross-lingual entity linking and bilingual lexicon induction, demonstrate its effectiveness, with gains reported over strong task-specialised baselines. Wenxuan Zhou 0002, Fangyu Liu 0001, Ivan Vulic, Nigel Collier, Muhao Chen 0001 |
ACL (1) | 4 |
| 2022 | A Contrastive Framework for Neural Text GenerationabstractText generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g., beam search) of neural language models often lead to degenerate solutions---the generated text is unnatural and contains undesirable repetitions. Existing approaches introduce stochasticity via sampling or modify training objectives to decrease the probabilities of certain tokens (e.g., unlikelihood training). However, they often lead to solutions that lack coherence. In this work, we show that an underlying reason for model degeneration is the anisotropic distribution of token representations. We present a contrastive solution: (i) SimCTG, a contrastive training objective to calibrate the model's representation space, and (ii) a decoding method---contrastive search---to encourage diversity while maintaining coherence in the generated text. Extensive experiments and analyses on three benchmarks from two languages demonstrate that our proposed approach outperforms state-of-the-art text generation methods as evaluated by both human and automatic metrics. Yixuan Su, Tian Lan 0003, Yan Wang 0060, Dani Yogatama, Lingpeng Kong, Nigel Collier |
NeurIPS | 6 |
| 2022 | BioCaster in 2021: automatic disease outbreaks detection from global news mediaabstractSUMMARY: BioCaster was launched in 2008 to provide an ontology-based text mining system for early disease detection from open news sources. Following a 6-year break, we have re-launched the system in 2021. Our goal is to systematically upgrade the methodology using state-of-the-art neural network language models, whilst retaining the original benefits that the system provided in terms of logical reasoning and automated early detection of infectious disease outbreaks. Here, we present recent extensions such as neural machine translation in 10 languages, neural classification of disease outbreak reports and a new cloud-based visualization dashboard. Furthermore, we discuss our vision for further improvements, including combining risk assessment with event semantics and assessing the risk of outbreaks with multi-granularity. We hope that these efforts will benefit the global public health community. AVAILABILITY AND IMPLEMENTATION: BioCaster web-portal is freely accessible at http://biocaster.org. Zaiqiao Meng, Anya Okhmatovskaia, Maxime Polleri, Yannan Shen, Guido Powell, Iris Ganser, Meiru Zhang, Nicholas B. King, David L. Buckeridge, Nigel Collier |
Bioinform. | 11 |
| 2022 | PheneBank: a literature-based database of phenotypesabstractMOTIVATION: Significant effort has been spent by curators to create coding systems for phenotypes such as the Human Phenotype Ontology, as well as disease-phenotype annotations. We aim to support the discovery of literature-based phenotypes and integrate them into the knowledge discovery process. RESULTS: PheneBank is a Web-portal for retrieving human phenotype-disease associations that have been text-mined from the whole of Medline. Our approach exploits state-of-the-art machine learning for concept identification by utilizing an expert annotated rare disease corpus from the PMC Text Mining subset. Evaluation of the system for entities is conducted on a gold-standard corpus of rare disease sentences and for associations against the Monarch initiative data. AVAILABILITY AND IMPLEMENTATION: The PheneBank Web-portal freely available at http://www.phenebank.org. Annotated Medline data is available from Zenodo at DOI: 10.5281/zenodo.1408800. Semantic annotation software is freely available for non-commercial use at GitHub: https://github.com/pilehvar/phenebank. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mohammad Taher Pilehvar, Adam Bernard, Damian Smedley, Nigel Collier |
Bioinform. | 4 |
| 2021 | Visual Pivoting for (Unsupervised) Entity AlignmentabstractThis work studies the use of visual semantic representations to align entities in heterogeneous knowledge graphs (KGs). Images are natural components of many existing KGs. By combining visual knowledge with other auxiliary information, we show that the proposed new approach, EVA, creates a holistic entity representation that provides strong signals for cross-graph entity alignment. Besides, previous entity alignment methods require human labelled seed alignment, restricting availability. EVA provides a completely unsupervised solution by leveraging the visual similarity of entities to create an initial seed dictionary (visual pivots). Experiments on benchmark data sets DBP15k and DWY15k show that EVA offers state-of-the-art performance on both monolingual and cross-lingual entity alignment tasks. Furthermore, we discover that images are particularly useful to align long-tail KG entities, which inherently lack the structural contexts necessary for capturing the correspondences. Code release: https://github.com/cambridgeltl/eva; project page: http://cogcomp.org/page/publication view/927. Fangyu Liu 0001, Muhao Chen 0001, Dan Roth 0001, Nigel Collier |
AAAI | 4 |
| 2021 | Dialogue Response Selection with Hierarchical Curriculum LearningabstractYixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060 |
ACL/IJCNLP (1) | 8 |
| 2021 | MirrorWiC: On Eliciting Word-in-Context Representations from Pretrained Language ModelsabstractRecent work indicated that pretrained language models (PLMs) such as BERT and RoBERTa can be transformed into effective sentence and word encoders even via simple self-supervised techniques.Inspired by this line of work, in this paper we propose a fully unsupervised approach to improving word-in-context (WiC) representations in PLMs, achieved via a simple and efficient WiC-targeted fine-tuning procedure: MIRROR-WIC.The proposed method leverages only raw texts sampled from Wikipedia, assuming no sense-annotated data, and learns contextaware word representations within a standard contrastive learning setup.We experiment with a series of standard and comprehensive WiC benchmarks across multiple languages.Our proposed fully unsupervised MIRROR-WIC models obtain substantial gains over offthe-shelf PLMs across all monolingual, multilingual and cross-lingual setups.Moreover, on some standard WiC benchmarks, MIRROR-WIC is even on-par with supervised models fine-tuned with in-task data and sense labels. Qianchu Liu, Fangyu Liu 0001, Nigel Collier, Anna Korhonen, Ivan Vulic |
CoNLL | 3 |
| 2021 | Non-Autoregressive Text Generation with Pre-trained Language ModelsabstractYixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, Nigel Collier. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yixuan Su, Deng Cai 0002, Yan Wang 0060, David Vandyke, Simon Baker, Piji Li, Nigel Collier |
EACL | 7 |
| 2021 | Visually Grounded Reasoning across Languages and CulturesabstractThe design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems. Fangyu Liu 0001, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, Desmond Elliott |
EMNLP (1) | 5 |
| 2021 | Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersabstractPrevious work has indicated that pretrained Masked Language Models (MLMs) are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further taskspecific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data.In this work, we demonstrate that it is possible to turn MLMs into effective lexical and sentence encoders even without any additional data, relying simply on self-supervision.We propose an extremely simple, fast, and effective contrastive learning technique, termed Mirror-BERT, which converts MLMs (e.g., BERT and RoBERTa) into such encoders in 20-30 seconds with no access to additional external knowledge.Mirror-BERT relies on identical and slightly modified string pairs as positive (i.e., synonymous) fine-tuning examples, and aims to maximise their similarity during "identity fine-tuning".We report huge gains over off-the-shelf MLMs with Mirror-BERT both in lexical-level and in sentencelevel tasks, across different domains and different languages.Notably, in sentence similarity (STS) and question-answer entailment (QNLI) tasks, our self-supervised Mirror-BERT model even matches the performance of the Sentence-BERT models from prior work which rely on annotated task data.Finally, we delve deeper into the inner workings of MLMs, and suggest some evidence on why this simple Mirror-BERT fine-tuning approach can yield effective universal lexical and sentence encoders. Fangyu Liu 0001, Ivan Vulic, Anna Korhonen, Nigel Collier |
EMNLP (1) | 4 |
| 2021 | Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERTabstractInfusing factual knowledge into pretrained models is fundamental for many knowledgeintensive tasks.In this paper, we propose Mixture-of-Partitions (MoP), an infusion approach that can handle a very large knowledge graph (KG) by partitioning it into smaller subgraphs and infusing their specific knowledge into various BERT models using lightweight adapters.To leverage the overall factual knowledge for a target task, these sub-graph adapters are further fine-tuned along with the underlying BERT through a mixture layer.We evaluate our MoP with three biomedical BERTs (SciBERT, BioBERT, PubmedBERT) on six downstream tasks (inc.NLI, QA, Classification), and the results show that our MoP consistently enhances the underlying BERTs in task performance, and achieves new SOTA performances on five evaluated datasets.1 Zaiqiao Meng, Fangyu Liu 0001, Thomas Hikaru Clark, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 5 |
| 2021 | Self-Alignment Pretraining for Biomedical Entity RepresentationsabstractFangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, Nigel Collier. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Fangyu Liu 0001, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, Nigel Collier |
NAACL-HLT | 5 |
| 2021 | PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval MemoryabstractThe ability of dialogue systems to express pre-specified style during conversations has a direct, positive impact on their usability and user satisfaction. While it has attracted much research interest, existing methods often generate stylistic responses at the cost of content quality. In this work, we introduce a prototype-to-style (PS) framework to tackle the challenge of stylistic dialogue generation. The proposed framework first exploits an Information Retrieval (IR) system and extracts a response prototype from the retrieved response. A stylistic response generator then takes the response prototype and the desired style as input to produce a high-quality and stylistic response. To effectively train the proposed model and imitate the real testing environment, we introduce a new style-aware learning objective and a denoising learning strategy. Results on three benchmark datasets (gender, emotion, and sentiment) from two languages demonstrate that the proposed approach significantly outperforms existing baselines both in terms of in-domain and cross-domain evaluations. Yixuan Su, Yan Wang 0060, Deng Cai 0002, Simon Baker, Anna Korhonen, Nigel Collier |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Will-They-Won't-They: A Very Large Dataset for Stance Detection on TwitterabstractCostanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Costanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou, Flavio Toxvaerd, Nigel Collier |
ACL | 6 |
| 2020 | COMETA: A Corpus for Medical Entity Linking in the Social MediaabstractWhilst there has been growing progress in Entity Linking (EL) for general language, existing datasets fail to address the complex nature of health terminology in layman's language.Meanwhile, there is a growing need for applications that can understand the public's voice in the health domain.To address this we introduce a new corpus called COMETA, consisting of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.Our corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by any of the existing resources in the field.Through benchmark experiments on 20 EL baselines from string-to neural-based models we shed light on the ability of these systems to perform complex inference on entities and concepts under 2 challenging evaluation scenarios.Our experimental results on COMETA illustrate that no golden bullet exists and even the best mainstream techniques still have a significant performance gap to fill, while the best solution relies on combining different views of data. Marco Basaldella, Fangyu Liu 0001, Ehsan Shareghi, Nigel Collier |
EMNLP (1) | 4 |
| 2019 | Unseen Word Representation by Aligning Heterogeneous Lexical Semantic SpacesabstractWord embedding techniques heavily rely on the abundance of training data for individual words. Given the Zipfian distribution of words in natural language texts, a large number of words do not usually appear frequently or at all in the training data. In this paper we put forward a technique that exploits the knowledge encoded in lexical resources, such as WordNet, to induce embeddings for unseen words. Our approach adapts graph embedding and cross-lingual vector space transformation techniques in order to merge lexical knowledge encoded in ontologies with that derived from corpus statistics. We show that the approach can provide consistent performance improvements across multiple evaluation benchmarks: in-vitro, on multiple rare word similarity datasets, and invivo, in two downstream text classification tasks. Victor Prokhorov, Mohammad Taher Pilehvar, Dimitri Kartsaklis, Pietro Liò, Nigel Collier |
AAAI | 5 |
| 2018 | Which Melbourne? Augmenting Geocoding with MapsabstractThe purpose of text geolocation is to associate geographic information contained in a document with a set (or sets) of coordinates, either implicitly by using linguistic features and/or explicitly by using geographic metadata combined with heuristics.We introduce a geocoder (location mention disambiguator) that achieves state-of-the-art (SOTA) results on three diverse datasets by exploiting the implicit lexical clues.Moreover, we propose a new method for systematic encoding of geographic metadata to generate two distinct views of the same text.To that end, we introduce the Map Vector (MapVec), a sparse representation obtained by plotting prior geographic probabilities, derived from population figures, on a World Map.We then integrate the implicit (language) and explicit (map) features to significantly improve a range of metrics.We also introduce an open-source dataset for geoparsing of news events covering global disease outbreaks and epidemics to help future evaluation in geoparsing. Milan Gritta, Mohammad Taher Pilehvar, Nigel Collier |
ACL (1) | 3 |
| 2018 | Mapping Text to Knowledge Graph Entities using Multi-Sense LSTMsabstractThis paper addresses the problem of mapping natural language text to knowledge base entities.The mapping process is approached as a composition of a phrase or a sentence into a point in a multi-dimensional entity space obtained from a knowledge graph.The compositional model is an LSTM equipped with a dynamic disambiguation mechanism on the input word embeddings (a Multi-Sense LSTM), addressing polysemy issues.Further, the knowledge base space is prepared by collecting random walks from a graph enhanced with textual features, which act as a set of semantic bridges between text and knowledge base entities.The ideas of this work are demonstrated on largescale text-to-entity mapping and entity classification tasks, with state of the art results.* This paper is dedicated to the memory of Euripides Kartsaklis, a man who loved technology. Dimitri Kartsaklis, Mohammad Taher Pilehvar, Nigel Collier |
EMNLP | 3 |
| 2018 | Large-scale Exploration of Neural Relation Classification ArchitecturesabstractExperimental performance on the task of relation classification has generally improved using deep neural network architectures.One major drawback of reported studies is that individual models have been evaluated on a very narrow range of datasets, raising questions about the adaptability of the architectures, while making comparisons between approaches difficult.In this work, we present a systematic large-scale analysis of neural relation classification architectures on six benchmark datasets with widely varying characteristics.We propose a novel multi-channel LSTM model combined with a CNN that takes advantage of all currently popular linguistic and architectural features.Our 'Man for All Seasons' approach achieves state-of-the-art performance on two datasets.More importantly, in our view, the model allowed us to obtain direct insights into the continued challenges faced by neural language models on this task. Hoang-Quynh Le, Duy-Cat Can, Sinh T. Vu, Mohammad Taher Pilehvar, Nigel Collier |
EMNLP | 6 |
| 2018 | Card-660: A Reliable Evaluation Framework for Rare Word Representation ModelsabstractRare word representation has recently enjoyed a surge of interest, owing to the crucial role that effective handling of infrequent words can play in accurate semantic understanding.However, there is a paucity of reliable benchmarks for evaluation and comparison of these techniques.We show in this paper that the only existing benchmark (the Stanford Rare Word dataset) suffers from low-confidence annotations and limited vocabulary; hence, it does not constitute a solid comparison framework.In order to fill this evaluation gap, we propose CAmbridge Rare word Dataset (CARD-660), an expert-annotated word similarity dataset which provides a highly reliable, yet challenging, benchmark for rare word representation techniques.Through a set of experiments we show that even the best mainstream word embeddings, with millions of words in their vocabularies, are unable to achieve performances higher than 0.43 (Pearson correlation) on the dataset, compared to a human-level upperbound of 0.90.We release the dataset and the annotation materials at https:// pilehvar.github.io/card-660/. Mohammad Taher Pilehvar, Dimitri Kartsaklis, Victor Prokhorov, Nigel Collier |
EMNLP | 4 |
| 2017 | Vancouver Welcomes You! Minimalist Location Metonymy ResolutionabstractNamed entities are frequently used in a metonymic manner.They serve as references to related entities such as people and organisations.Accurate identification and interpretation of metonymy can be directly beneficial to various NLP applications, such as Named Entity Recognition and Geographical Parsing.Until now, metonymy resolution (MR) methods mainly relied on parsers, taggers, dictionaries, external word lists and other handcrafted lexical resources.We show how a minimalist neural approach combined with a novel predicate window method can achieve competitive results on the Se-mEval 2007 task on Metonymy Resolution.Additionally, we contribute with a new Wikipedia-based MR dataset called RelocaR, which is tailored towards locations as well as improving previous deficiencies in annotation guidelines. Milan Gritta, Mohammad Taher Pilehvar, Nut Limsopatham, Nigel Collier |
ACL (1) | 4 |
| 2017 | Towards a Seamless Integration of Word Senses into Downstream NLP ApplicationsabstractLexical ambiguity can impede NLP systems from accurate understanding of semantics.Despite its potential benefits, the integration of sense-level information into NLP systems has remained understudied.By incorporating a novel disambiguation algorithm into a state-of-the-art classification model, we create a pipeline to integrate sense-level information into downstream NLP applications.We show that a simple disambiguation of the input text can lead to consistent performance improvement on multiple topic categorization and polarity detection datasets, particularly when the fine granularity of the underlying sense inventory is reduced and the document is sufficiently large.Our results also point to the need for sense representation research to focus more on in vivo evaluations which target the performance in downstream NLP applications rather than artificial benchmarks. Mohammad Taher Pilehvar, José Camacho-Collados, Roberto Navigli, Nigel Collier |
ACL (1) | 4 |
| 2017 | WSDM 2017 Workshop on Mining Online Health Reports: MOHRS 2017abstractThe workshop on Mining Online Health Reports (MOHRS) draws upon the rapidly developing field of Computational Health, focusing on textual content that has been generated through the various facets of Web activity. Online user-generated information mining, especially from social media platforms and search engines, has been in the forefront of many research efforts, especially in the fields of Information Retrieval and Natural Language Processing. The incorporation of such data and techniques in a number of health-oriented applications has provided strong evidence about the potential benefits, which include better population coverage, timeliness and the operational ability in places with less established health infrastructure. The workshop aims to create a platform where relevant state-of-the-art research is presented, but at the same time discussions among researchers with cross-disciplinary backgrounds can take place. It will focus on the characterisation of data sources, the essential methods for mining this textual information, as well as potential real-world applications and the arising ethical issues. MOHRS '17 will feature 3 keynote talks and 4 accepted paper presentations, together with a panel discussion session. Nigel Collier, Nut Limsopatham, Aron Culotta, Mike Conway, Ingemar J. Cox, Vasileios Lampos |
WSDM | 1 |
| 2016 | Normalising Medical Concepts in Social Media Texts by Learning Semantic RepresentationabstractAutomatically recognising medical concepts mentioned in social media messages (e.g.tweets) enables several applications for enhancing health quality of people in a community, e.g.real-time monitoring of infectious diseases in population.However, the discrepancy between the type of language used in social media and medical ontologies poses a major challenge.Existing studies deal with this challenge by employing techniques, such as lexical term matching and statistical machine translation.In this work, we handle the medical concept normalisation at the semantic level.We investigate the use of neural networks to learn the transition between layman's language used in social media messages and formal medical language used in the descriptions of medical concepts in a standard ontology.We evaluate our approaches using three different datasets, where social media texts are extracted from Twitter messages and blog posts.Our experimental results show that our proposed approaches significantly and consistently outperform existing effective baselines, which achieved state-of-the-art performance on several medical concept normalisation tasks, by up to 44%. Nut Limsopatham, Nigel Collier |
ACL (1) | 2 |
| 2016 | De-Conflated Semantic RepresentationsabstractOne major deficiency of most semantic representation techniques is that they usually model a word type as a single point in the semantic space, hence conflating all the meanings that the word can have.Addressing this issue by learning distinct representations for individual meanings of words has been the subject of several research studies in the past few years.However, the generated sense representations are either not linked to any sense inventory or are unreliable for infrequent word senses.We propose a technique that tackles these problems by de-conflating the representations of words based on the deep knowledge that can be derived from a semantic network.Our approach provides multiple advantages in comparison to the previous approaches, including its high coverage and the ability to generate accurate representations even for infrequent word senses.We carry out evaluations on six datasets across two semantic similarity tasks and report state-of-the-art results on most of them. Mohammad Taher Pilehvar, Nigel Collier |
EMNLP | 2 |
| 2016 | The digital revolution in phenotypingabstractPhenotypes have gained increased notoriety in the clinical and biological domain owing to their application in numerous areas such as the discovery of disease genes and drug targets, phylogenetics and pharmacogenomics. Phenotypes, defined as observable characteristics of organisms, can be seen as one of the bridges that lead to a translation of experimental findings into clinical applications and thereby support 'bench to bedside' efforts. However, to build this translational bridge, a common and universal understanding of phenotypes is required that goes beyond domain-specific definitions. To achieve this ambitious goal, a digital revolution is ongoing that enables the encoding of data in computer-readable formats and the data storage in specialized repositories, ready for integration, enabling translational research. While phenome research is an ongoing endeavor, the true potential hidden in the currently available data still needs to be unlocked, offering exciting opportunities for the forthcoming years. Here, we provide insights into the state-of-the-art in digital phenotyping, by means of representing, acquiring and analyzing phenotype data. In addition, we provide visions of this field for future research work that could enable better applications of phenotype data. Anika Oellrich, Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam H. Shah, Olivier Bodenreider, Mary Regina Boland, Ivo I. Georgiev, Kevin M. Livingston, Augustin Luna, Ann-Marie Mallon, Prashanti Manda, Peter N. Robinson, Gabriella Rustici, Michelle Simon, Rainer Winnenburg, Michel Dumontier |
Briefings Bioinform. | 2 |
| 2015 | Adapting Phrase-based Machine Translation to Normalise Medical Terms in Social Media MessagesabstractPrevious studies have shown that health reports in social media, such as Dai-lyStrength and Twitter, have potential for monitoring health conditions (e.g.adverse drug reactions, infectious diseases) in particular communities.However, in order for a machine to understand and make inferences on these health conditions, the ability to recognise when laymen's terms refer to a particular medical concept (i.e.text normalisation) is required.To achieve this, we propose to adapt an existing phrase-based machine translation (MT) technique and a vector representation of words to map between a social media phrase and a medical concept.We evaluate our proposed approach using a collection of phrases from tweets related to adverse drug reactions.Our experimental results show that the combination of a phrase-based MT technique and the similarity between word vector representations outperforms the baselines that apply only either of them by up to 55%. Nut Limsopatham, Nigel Collier |
EMNLP | 2 |
| 2015 | Crowdsourcing Twitter annotations to identify first-hand experiences of prescription drug useabstractSelf-reported patient data has been shown to be a valuable knowledge source for post-market pharmacovigilance. In this paper we propose using the popular micro-blogging service Twitter to gather evidence about adverse drug reactions (ADRs) after firstly having identified micro-blog messages (also know as "tweets") that report first-hand experience. In order to achieve this goal we explore machine learning with data crowdsourced from laymen annotators. With the help of lay annotators recruited from CrowdFlower we manually annotated 1548 tweets containing keywords related to two kinds of drugs: SSRIs (eg. Paroxetine), and cognitive enhancers (eg. Ritalin). Our results show that inter-annotator agreement (Fleiss' kappa) for crowdsourcing ranks in moderate agreement with a pair of experienced annotators (Spearman's Rho=0.471). We utilized the gold standard annotations from CrowdFlower for automatically training a range of supervised machine learning models to recognize first-hand experience. F-Score values are reported for 6 of these techniques with the Bayesian Generalized Linear Model being the best (F-Score=0.64 and Informedness=0.43) when combined with a selected set of features obtained by using information gain criteria. Nestor Alvaro, Mike Conway, Son Doan, Christoph Lofi, John P. Overington, Nigel Collier |
J. Biomed. Informatics | 6 |
| 2014 | Discriminating Rhetorical Analogies in Social MediaabstractAnalogies are considered to be one of the core concepts of human cognition and communication, and are very efficient at encoding complex information in a natural fashion.However, computational approaches towards largescale analysis of the semantics of analogies are hampered by the lack of suitable corpora with real-life example of analogies.In this paper we therefore propose a workflow for discriminating and extracting natural-language analogy statements from the Web, focusing on analogies between locations mined from travel reports, blogs, and the Social Web.For realizing this goal, we employ feature-rich supervised learning models which we extensively evaluate.We also showcase a crowd-supported workflow for building a suitable Gold dataset used for this purpose.The resulting system is able to successfully learn to identify analogies to a high degree of accuracy (F-Score 0.9) by using a high-dimensional subsequence feature space. Christoph Lofi, Christian Nieke, Nigel Collier |
EACL | 3 |
| 2013 | Using silver and semi-gold standard corpora to compare open named entity recognisersabstractOntologies have become a central resource for defining biomedical concepts but linkage to and from textual data is still an unresolved technology. In this paper we approach the task of concept recognition in text by comparing four extant systems (cTAKES, NCBO Annotator, BeCAS and Metamap) with default parameter settings. The systems are compared on benchmark data consisting of 2,163 scientific abstracts and 906 clinical trial reports using an automatically constructed “silver” standard and a random semi-gold standard evaluation methodology. Furthermore, evaluation is conducted on the basis of specific concept identifiers. Experimental results show: (i) Generally higher levels of concept recognition on clinical trial reports than on scientific abstracts; (ii) The best performing system we observed on the silver standard was cTAKES on both the abstract and clinical trial corpora, however NCBO Annotator performed stronger when considering only the selected broad semantic types; (iii) BeCAS and Metamap had a tendency to annotate coarser-grained annotations; (iv) the random semi-gold evaluation places an upper bound on the performance of systems. This shows broad agreement with the silver standard evaluation but highlights areas where the silver standard methodology might be improved. Tudor Groza, Anika Oellrich, Nigel Collier |
BIBM | 3 |
| 2013 | A partially supervised cross-collection topic model for cross-domain text classificationabstractCross-domain text classification aims to automatically train a precise text classifier for a target domain by using labelled text data from a related source domain. To this end, one of the most promising ideas is to induce a new feature representation so that the distributional difference between domains can be reduced and a more accurate classifier can be learned in this new feature space. However, most existing methods do not explore the duality of the marginal distribution of examples and the conditional distribution of class labels given labeled training examples in the source domain. Besides, few previous works attempt to explicitly distinguish the domain-independent and domain-specific latent features and align the domain-specific features to further improve the cross-domain learning. In this paper, we propose a model called Partially Supervised Cross-Collection LDA topic model (PSCCLDA) for cross-domain learning with the purpose of addressing these two issues in a unified way. Experimental results on nine datasets show that our model outperforms two standard classifiers and four state-of-the-art methods, which demonstrates the effectiveness of our proposed model. Yang Bao 0001, Nigel Collier, Anindya Datta |
CIKM | 2 |
| 2013 | Change-point detection in time-series data by relative density-ratio estimation
Song Liu 0002, Makoto Yamada, Nigel Collier, Masashi Sugiyama |
Neural Networks | 3 |
| 2012 | A Hybrid Approach to Finding Phenotype Candidates in Genetic Texts
Nigel Collier, Mai-Vu Tran, Hoang-Quynh Le, Anika Oellrich, Ai Kawazoe, Martin Hall-May, Dietrich Rebholz-Schuhmann |
COLING | 1 |
| 2012 | On-line Trend Analysis with Topic Models: \#twitter Trends Detection Topic Model Online
Jey Han Lau, Nigel Collier, Timothy Baldwin |
COLING | 2 |
| 2012 | GENI-DB: a database of global events for epidemic intelligenceabstractUNLABELLED: We present a novel public health database (GENI-DB) in which news events on the topic of over 176 infectious diseases and chemicals affecting human and animal health are compiled from surveillance of the global online news media in 10 languages. News event frequency data were gathered systematically through the BioCaster public health surveillance system from July 2009 to the present and is available to download by the research community for purposes of analyzing trends in the global burden of infectious diseases. Database search can be conducted by year, country, disease and language. AVAILABILITY: The GENI-DB is freely available via a web portal at http://born.nii.ac.jp/. Nigel Collier, Son Doan |
Bioinform. | 1 |
| 2010 | An ontology-driven system for detecting global health events
Nigel Collier, Reiko Matsuda Goodwin, John P. McCrae, Son Doan, Ai Kawazoe, Mike Conway, Asanee Kawtrakul, Koichi Takeuchi, Dinh Dien |
COLING | 1 |
| 2009 | Towards role-based filtering of disease outbreak reports
Son Doan, Ai Kawazoe, Mike Conway, Nigel Collier |
J. Biomed. Informatics | 4 |
| 2008 | The Choice of Features for Classification of Verbs in Biomedical Texts
Anna Korhonen, Yuval Krymolowski, Nigel Collier |
COLING | 3 |
| 2008 | Global Health Monitor - A Web-based System for Detecting and Mapping Infectious Diseases
Son Doan, Hung Quoc Ngo 0002, Ai Kawazoe, Nigel Collier |
IJCNLP | 4 |
| 2008 | BioCaster: detecting public health rumors with a Web-based text mining systemabstractSUMMARY: BioCaster is an ontology-based text mining system for detecting and tracking the distribution of infectious disease outbreaks from linguistic signals on the Web. The system continuously analyzes documents reported from over 1700 RSS feeds, classifies them for topical relevance and plots them onto a Google map using geocoded information. The background knowledge for bridging the gap between Layman's terms and formal-coding systems is contained in the freely available BioCaster ontology which includes information in eight languages focused on the epidemiological role of pathogens as well as geographical locations with their latitudes/longitudes. The system consists of four main stages: topic classification, named entity recognition (NER), disease/location detection and event recognition. Higher order event analysis is used to detect more precisely specified warning signals that can then be notified to registered users via email alerts. Evaluation of the system for topic recognition and entity identification is conducted on a gold standard corpus of annotated news articles. AVAILABILITY: The BioCaster map and ontology are freely available via a web portal at http://www.biocaster.org. Nigel Collier, Son Doan, Ai Kawazoe, Reiko Matsuda Goodwin, Mike Conway, Yoshio Tateno, Hung Quoc Ngo 0002, Dinh Dien, Asanee Kawtrakul, Koichi Takeuchi, Mika Shigematsu, Kiyosu Taniguchi |
Bioinform. | 1 |
| 2008 | Structuring an event ontology for disease outbreak detectionabstractBACKGROUND: This paper describes the design of an event ontology being developed for application in the machine understanding of infectious disease-related events reported in natural language text. This event ontology is designed to support timely detection of disease outbreaks and rapid judgment of their alerting status by 1) bridging a gap between layman's language used in disease outbreak reports and public health experts' deep knowledge, and 2) making multi-lingual information available. CONSTRUCTION AND CONTENT: This event ontology integrates a model of experts' knowledge for disease surveillance, and at the same time sets of linguistic expressions which denote disease-related events, and formal definitions of events. In this ontology, rather general event classes, which are suitable for application to language-oriented tasks such as recognition of event expressions, are placed on the upper-level, and more specific events of the experts' interest are in the lower level. Each class is related to other classes which represent participants of events, and linked with multi-lingual synonym sets and axioms. CONCLUSIONS: We consider that the design of the event ontology and the methodology introduced in this paper are applicable to other domains which require integration of natural language information and machine support for experts to assess them. The first version of the ontology, with about 40 concepts, will be available in March 2008. Ai Kawazoe, Hutchatai Chanlekha, Mika Shigematsu, Nigel Collier |
BMC Bioinform. | 4 |
| 2008 | Synonym set extraction from the biomedical literature by lexical pattern discoveryabstractBACKGROUND: Although there are a large number of thesauri for the biomedical domain many of them lack coverage in terms and their variant forms. Automatic thesaurus construction based on patterns was first suggested by Hearst 1, but it is still not clear how to automatically construct such patterns for different semantic relations and domains. In particular it is not certain which patterns are useful for capturing synonymy. The assumption of extant resources such as parsers is also a limiting factor for many languages, so it is desirable to find patterns that do not use syntactical analysis. Finally to give a more consistent and applicable result it is desirable to use these patterns to form synonym sets in a sound way. RESULTS: We present a method that automatically generates regular expression patterns by expanding seed patterns in a heuristic search and then develops a feature vector based on the occurrence of term pairs in each developed pattern. This allows for a binary classifications of term pairs as synonymous or non-synonymous. We then model this result as a probability graph to find synonym sets, which is equivalent to the well-studied problem of finding an optimal set cover. We achieved 73.2% precision and 29.7% recall by our method, out-performing hand-made resources such as MeSH and Wikipedia. CONCLUSION: We conclude that automatic methods can play a practical role in developing new thesauri or expanding on existing ones, and this can be done with only a small amount of training data and no need for resources such as parsers. We also concluded that the accuracy can be improved by grouping into synonym sets. John P. McCrae, Nigel Collier |
BMC Bioinform. | 2 |
| 2007 | Named entity recognition in Vietnamese using classifier votingabstractNamed entity recognition (NER) is one of the fundamental tasks in natural-language processing (NLP). Though the combination of different classifiers has been widely applied in several well-studied languages, this is the first time this method has been applied to Vietnamese. In this article, we describe how voting techniques can improve the performance of Vietnamese NER. By combining several state-of-the-art machine-learning algorithms using voting strategies, our final result outperforms individual algorithms and gained an F-measure of 89.12. A detailed discussion about the challenges of NER in Vietnamese is also presented. Pham Thi Xuan Thao, Tran Quoc Tri, Dinh Dien, Nigel Collier |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2006 | Automatic Classification of Verbs in Biomedical TextsabstractLexical classes, when tailored to the application and domain in question, can provide an effective means to deal with a number of natural language processing (NLP) tasks. While manual construction of such classes is difficult, recent research shows that it is possible to automatically induce verb classes from cross-domain corpora with promising accuracy. We report a novel experiment where similar technology is applied to the important, challenging domain of biomedicine. We show that the resulting classification, acquired from a corpus of biomedical journal articles, is highly accurate and strongly domain-specific. It can be used to aid BIO-NLP directly or as useful material for investigating the syntax and semantics of verbs in biomedical texts. Anna Korhonen, Yuval Krymolowski, Nigel Collier |
ACL | 3 |
| 2005 | Towards Semantic Role Labeling & IE in the Medical Literature
Yacov Kogan, Nigel Collier, Serguei V. S. Pakhomov, Michael Krauthammer |
AMIA | 2 |
| 2005 | Exploring Predicate-Argument Relations for Named Entity Recognition in the Molecular Biology Domain
Tuangthong Wattarujeekrit, Nigel Collier |
Discovery Science | 2 |
| 2005 | Bio-medical entity extraction using support vector machines
Koichi Takeuchi, Nigel Collier |
Artif. Intell. Medicine | 2 |
| 2004 | Sentiment Analysis using Support Vector Machines with Diverse Information Sources
Tony Mullen, Nigel Collier |
EMNLP | 2 |
| 2004 | Annotation of Coreference Relations Among Linguistic Expressions and Images in Biological Articles
Ai Kawazoe, Asanobu Kitamoto, Nigel Collier |
LREC | 3 |
| 2004 | An Annotation Scheme for a Rhetorical Analysis of Biology Articles
Yoko Mizuta, Nigel Collier |
LREC | 2 |
| 2004 | PASBio: predicate-argument structures for event extraction in molecular biologyabstractBACKGROUND: The exploitation of information extraction (IE), a technology aiming to provide instances of structured representations from free-form text, has been rapidly growing within the molecular biology (MB) research community to keep track of the latest results reported in literature. IE systems have traditionally used shallow syntactic patterns for matching facts in sentences but such approaches appear inadequate to achieve high accuracy in MB event extraction due to complex sentence structure. A consensus in the IE community is emerging on the necessity for exploiting deeper knowledge structures such as through the relations between a verb and its arguments shown by predicate-argument structure (PAS). PAS is of interest as structures typically correspond to events of interest and their participating entities. For this to be realized within IE a key knowledge component is the definition of PAS frames. PAS frames for non-technical domains such as newswire are already being constructed in several projects such as PropBank, VerbNet, and FrameNet. Knowledge from PAS should enable more accurate applications in several areas where sentence understanding is required like machine translation and text summarization. In this article, we explore the need to adapt PAS for the MB domain and specify PAS frames to support IE, as well as outlining the major issues that require consideration in their construction. RESULTS: We introduce PASBio by extending a model based on PropBank to the MB domain. The hypothesis we explore is that PAS holds the key for understanding relationships describing the roles of genes and gene products in mediating their biological functions. We chose predicates describing gene expression, molecular interactions and signal transduction events with the aim of covering a number of research areas in MB. Analysis was performed on sentences containing a set of verbal predicates from MEDLINE and full text journals. Results confirm the necessity to analyze PAS specifically for MB domain. CONCLUSIONS: At present PASBio contains the analyzed PAS of over 30 verbs, publicly available on the Internet for use in advanced applications. In the future we aim to expand the knowledge base to cover more verbs and the nominal form of each predicate. Tuangthong Wattarujeekrit, Parantu K. Shah, Nigel Collier |
BMC Bioinform. | 3 |
| 2004 | Comparison of character-level and part of speech features for name recognition in biomedical texts
Nigel Collier, Koichi Takeuchi |
J. Biomed. Informatics | 1 |
| 2003 | A Framework for Integrating Deep and Shallow Semantic Structures in Text Mining
Nigel Collier, Koichi Takeuchi, Ai Kawazoe, Tony Mullen, Tuangthong Wattarujeekrit |
KES | 1 |
| 2002 | Use of Support Vector Machines in Extended Named Entity Recognition
Koichi Takeuchi, Nigel Collier |
CoNLL | 2 |
| 2002 | PIA-Core: Semantic Annotation through Example-based Learning
Nigel Collier, Koichi Takeuchi |
LREC | 1 |
| 2002 | Progress on Multi-lingual Named Entity Annotation Guidelines using RDF (S)
Nigel Collier, Koichi Takeuchi, Chikashi Nobata, Jun-ichi Fukumoto, Norihiro Ogata |
LREC | 1 |
| 2000 | Extracting the Names of Genes and Gene Products with a Hidden Markov Model
Nigel Collier, Chikashi Nobata, Jun'ichi Tsujii |
COLING | 1 |
| 1999 | The GENIA project: corpus-based knowledge acquisition and information extraction from genome research papers
Nigel Collier, Hyun Seok Park, Norihiro Ogata, Yuka Tateisi, Chikashi Nobata, Tomoko Ohta, Tateshi Sekimizu, Hisao Imai, Katsutoshi Ibushi, Jun'ichi Tsujii |
EACL | 1 |
| 1999 | A Comparison of Query Translation Methods for English-Japanese Cross-Language Information Retrieval (poster abstract)
Gareth J. F. Jones, Tetsuya Sakai, Nigel Collier, Akira Kumano, Kazuo Sumita |
SIGIR | 3 |
| 1997 | Convergence Time Characteristics of an Associative Memory for Natural Language Processing
Nigel Collier |
IJCAI | 1 |