EDBT 2026 Demo / reviewers in the wild / expert
Trevor Cohn
dblp:66/4613 · also Trevor Anthony Cohn
· DBLP profile ↗
137ranked-venue papers
15as first author
42since 2021 · last 2025
0000-0003-4363-1673ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 126 · 13 first-author · 39 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Planning in the Dark: LLM-Symbolic Planning Pipeline Without ExpertsabstractLarge Language Models (LLMs) have shown promise in solving natural language-described planning tasks, but their direct use often leads to inconsistent reasoning and hallucination. While hybrid LLM-symbolic planning pipelines have emerged as a more robust alternative, they typically require extensive expert intervention to refine and validate generated action schemas. It not only limits scalability but also introduces a potential for biased interpretation, as a single expert's interpretation of ambiguous natural language descriptions might not align with the user's actual intent. To address this, we propose a novel approach that constructs an action schema library to generate multiple candidates, accounting for the diverse possible interpretations of natural language descriptions. We further introduce a semantic validation and ranking module that automatically filter and rank these candidates without expert-in-the-loop. The experiments showed our pipeline maintains superiority in planning over the direct LLM planning approach. These findings demonstrate the feasibility of a fully automated end-to-end LLM-symbolic planner that requires no expert intervention, opening up the possibility for a broader audience to engage with AI planning with less prerequisite of domain expertise. Sukai Huang, Nir Lipovetzky, Trevor Cohn |
AAAI | 3 |
| 2025 | LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian LanguagesabstractAs one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress.We introduce LO-RAXBENCH, a benchmark that focuses on lowresource languages of Indonesia and covers 6 diverse tasks: reading comprehension, opendomain QA, language inference, causal reasoning, translation, and cultural QA.Our dataset covers 20 languages, with the addition of two formality registers for three languages.We evaluate a diverse set of multilingual and regionfocused LLMs and found that this benchmark is challenging.We note a visible discrepancy between performance in Indonesian and other languages, especially the low-resource ones.There is no clear lead when using a regionspecific model as opposed to the general multilingual model.Lastly, we show that a change in register affects model performance, especially with registers not commonly found in social media, such as high-level politeness 'Krama' Javanese. Alham Fikri Aji, Trevor Cohn |
EMNLP | 2 |
| 2025 | Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan GenerationabstractThe capability of Large Language Models (LLMs) to plan remains a topic of debate. Some critics argue that strategies to boost LLMs' reasoning skills are ineffective in planning tasks, while others report strong outcomes merely from training models on a planning corpus. This paper revisits these claims by developing an end-to-end LLM-based planner and evaluating a range of reasoning-enhancement strategies --- including fine-tuning, Chain-of-Thought (CoT) prompting, and reinforcement learning (RL) --- across multiple dimensions of plan quality: validity, executability, goal satisfiability, and more. Our findings reveal fine-tuning alone is insufficient, especially on out-of-distribution tasks. Strategies like CoT prompting primarily enhance local coherence, yielding higher executability rates --- a necessary prerequisite for validity --- but provide only incremental gains and struggle to ensure global plan validity. Notably, RL guided by a novel Longest Contiguous Common Subsequence reward significantly enhances both executability and validity, particularly on longer-horizon problems. Overall, our research addresses key misconceptions in the LLM-planning literature and underscores reward-driven RL optimization as a promising direction for advancing robust LLM-based planning by jointly improving executability and validity. Sukai Huang, Trevor Cohn, Nir Lipovetzky |
ICAPS | 2 |
| 2025 | Improving Language Model Distillation through Hidden State MatchingabstractHidden State Matching is shown to improve knowledge distillation of language models by encouraging similarity between a student and its teacher's hidden states, as demonstrated by DistilBERT and its successors. This typically uses a cosine loss, which restricts the dimensionality of the student to the teacher's, severely limiting the compression ratio. We present an alternative technique using Centered Kernel Alignment (CKA) to match hidden states of different dimensionality, allowing for smaller students and higher compression ratios. We show the efficacy of our method using encoder--decoder (BART, mBART \& T5) and encoder-only (BERT) architectures across a range of tasks from classification to summarization and translation. Our technique is competitive with the current state-of-the-art distillation methods at comparable compression rates. It requires no pretrained student models, but rather can synthesize new student models from scratch through pretraining distillation. It can scale to students smaller than the current methods, is no slower in training and inference, and is considerably more flexible. Sayantan Dasgupta, Trevor Cohn |
ICLR | 2 |
| 2025 | Mufu: Multilingual Fused Learning for Low-Resource Translation with LLMabstractMultilingual large language models (LLMs) are great translators, but this is largely limited to high-resource languages. For many LLMs, translating in and out of low-resource languages remains a challenging task. To maximize data efficiency in this low-resource setting, we introduce Mufu, which includes a selection of automatically generated multilingual candidates and an instruction to correct inaccurate translations in the prompt. Mufu prompts turn a translation task into a postediting one, and seek to harness the LLM’s reasoning capability with auxiliary translation candidates, from which the model is required to assess the input quality, align the semantics cross-lingually, copy from relevant inputs and override instances that are incorrect. Our experiments on En-XX translations over the Flores-200 dataset show LLMs finetuned against Mufu-style prompts are robust to poor quality auxiliary translation candidates, achieving performance superior to NLLB 1.3B distilled model in 64% of low- and very-low-resource language pairs. We then distill these models to reduce inference cost, while maintaining on average 3.1 chrF improvement over finetune-only baseline in low-resource translations. Zheng Wei Lim, Nitish Gupta, Honglin Yu, Trevor Cohn |
ICLR | 4 |
| 2025 | MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model PretrainingabstractData quality is a critical driver of large language model performance, yet existing model-based selection methods focus almost exclusively on English, neglecting other languages that are essential in the training mix for multilingual LLMs. We introduce MuRating, a scalable framework that transfers high-quality English data-quality signals into a multilingual autorater, capable of handling 17 languages. MuRating aggregates multiple English autoraters via pairwise comparisons to learn unified document quality scores, then projects these judgments through translation to train a multilingual evaluator on monolingual, cross-lingual, and parallel text pairs. Applied to web data, MuRating selects balanced subsets of English and multilingual content to pretrain LLaMA-architecture models of 1.2B and 7B parameters. Compared to strong baselines, including QuRater, FineWeb2-HQ, AskLLM, DCLM, our approach increases average accuracy on both English benchmarks and multilingual evaluations. Extensive analyses further validate that pairwise training provides greater stability and robustness than pointwise scoring, underscoring the effectiveness of MuRating as a general multilingual data-selection framework. Zhixun Chen, Wenhan Han, Binbin Li 0001, Haobin Lin, Fengze Liu, Bingni Zhang, Taifeng Wang, Trevor Cohn |
NeurIPS | 12 |
| 2025 | Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationabstractSign Language Translation (SLT) aims to map sign language videos to spoken language text. A common approach relies on gloss annotations as an intermediate representation, decomposing SLT into two sub-tasks: video-to-gloss recognition and gloss-to-text translation. While effective, this paradigm depends on expert-annotated gloss labels, which are costly and rarely available in existing datasets, limiting its scalability. To address this challenge, we propose a gloss-free pseudo gloss generation framework that eliminates the need for human-annotated glosses while preserving the structured intermediate representation.
Specifically, we prompt a Large Language Model (LLM) with a few example text-gloss pairs using in-context learning to produce draft sign glosses from spoken language text.
To enhance the correspondence between LLM-generated pseudo glosses and the sign sequences in video, we correct the ordering in the pseudo glosses for better alignment via a weakly supervised learning process.
This reordering facilitates the incorporation of auxiliary alignment objectives, and allows for the use of efficient supervision via a Connectionist Temporal Classification (CTC) loss.
We train our SLT model—consisting of a vision encoder and a translator—through a three-stage pipeline, which progressively narrows the modality gap between sign language and spoken language.
Despite its simplicity, our approach outperforms previous state-of-the-art gloss-free frameworks on two SLT benchmarks and achieves competitive results compared to gloss-based methods. Jianyuan Guo, Peike Li, Trevor Cohn |
NeurIPS | 3 |
| 2025 | Zero-Shot Performance Prediction for Probabilistic Scaling LawsabstractThe prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task’s data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to $30$ LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes. Viktoria Schram, Markus Hiller, Daniel Beck, Trevor Cohn |
NeurIPS | 4 |
| 2025 | Few-Shot Multilingual Open-Domain QA from Five ExamplesabstractAbstract Recent approaches to multilingual open- domain question answering (MLODQA) have achieved promising results given abundant language-specific training data. However, the considerable annotation cost limits the application of these methods for underrepresented languages. We introduce a few-shot learning approach to synthesize large-scale multilingual data from large language models (LLMs). Our method begins with large-scale self-supervised pre-training using WikiData, followed by training on high-quality synthetic multilingual data generated by prompting LLMs with few-shot supervision. The final model, FsModQA, significantly outperforms existing few-shot and supervised baselines in MLODQA and cross-lingual and monolingual retrieval. We further show our method can be extended for effective zero-shot adaptation to new languages through a cross-lingual prompting strategy with only English-supervised data, making it a general and applicable solution for MLODQA tasks without costly large-scale annotation. Fan Jiang 0014, Tom Drummond, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | Pre-training Cross-lingual Open Domain Question Answering with Large-scale Synthetic SupervisionabstractCross-lingual open domain question answering (CLQA) is a complex problem, comprising cross-lingual retrieval from a multilingual knowledge base, followed by answer generation in the query language.Both steps are usually tackled by separate models, requiring substantial annotated datasets, and typically auxiliary resources, like machine translation systems to bridge between languages.In this paper, we show that CLQA can be addressed using a single encoder-decoder model.To effectively train this model, we propose a selfsupervised method based on exploiting the cross-lingual link structure within Wikipedia.We demonstrate how linked Wikipedia pages can be used to synthesise supervisory signals for cross-lingual retrieval, through a form of cloze query, and generate more natural questions to supervise answer generation.Together, we show our approach, CLASS, outperforms comparable methods on both supervised and zero-shot language adaptation settings, including those using machine translation.𝓒 𝐸𝑛 𝓒 𝐸𝑛 𝓒 𝐸𝑛 𝓒 𝑀𝑢𝑙𝑡𝑖 𝓒 𝑀𝑢𝑙𝑡𝑖 Parallel Sentence Mining Once in I n d i a , hippies went to many different destinations, on the beaches of Goa and Kovalam in Trivandrum (Kerala), or crossed the border into Nepal to spend months in Kathmandu.インドでは、ヒッピーは多くの異な る目的地へいったが、トリヴァンド ラム(ケーララ州)のゴアとコバラ ムのビーチに大量に集まったり、 国境を越えたネパールのカトマン ズで数ヶ月過ごしたりした。 q En : Once in [Mask], hippies went to many different destinations... Fan Jiang 0014, Tom Drummond, Trevor Cohn |
EMNLP | 3 |
| 2024 | Revisiting subword tokenization: A case study on affixal negation in large language modelsabstractThinh Hung Truong, Yulia Otmakhova, Karin Verspoor, Trevor Cohn, Timothy Baldwin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Thinh Truong, Yulia Otmakhova 0001, Karin Verspoor, Trevor Cohn, Timothy Baldwin |
NAACL-HLT | 4 |
| 2024 | Backdoor Attacks on Multilingual Machine TranslationabstractJun Wang, Qiongkai Xu, Xuanli He, Benjamin Rubinstein, Trevor Cohn. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jun Wang 0126, Qiongkai Xu, Xuanli He, Benjamin I. P. Rubinstein, Trevor Cohn |
NAACL-HLT | 5 |
| 2024 | SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning AttacksabstractAbstract Modern NLP models are often trained on public datasets drawn from diverse sources, rendering them vulnerable to data poisoning attacks. These attacks can manipulate the model’s behavior in ways engineered by the attacker. One such tactic involves the implantation of backdoors, achieved by poisoning specific training instances with a textual trigger and a target class label. Several strategies have been proposed to mitigate the risks associated with backdoor attacks by identifying and removing suspected poisoned examples. However, we observe that these strategies fail to offer effective protection against several advanced backdoor attacks. To remedy this deficiency, we propose a novel defensive mechanism that first exploits training dynamics to identify poisoned samples with high precision, followed by a label propagation step to improve recall and thus remove the majority of poisoned instances. Compared with recent advanced defense methods, our method considerably reduces the success rates of several backdoor attacks while maintaining high classification accuracy on clean test sets. Xuanli He, Qiongkai Xu, Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | Predicting Human Translation Difficulty with Neural Machine TranslationabstractAbstract Human translators linger on some words and phrases more than others, and predicting this variation is a step towards explaining the underlying cognitive processes. Using data from the CRITT Translation Process Research Database, we evaluate the extent to which surprisal and attentional features derived from a Neural Machine Translation (NMT) model account for reading and production times of human translators. We find that surprisal and attention are complementary predictors of translation difficulty, and that surprisal derived from a NMT model is the single most successful predictor of production duration. Our analyses draw on data from hundreds of translators operating across 13 language pairs, and represent the most comprehensive investigation of human translation difficulty to date. Zheng Wei Lim, Ekaterina Vylomova, Charles Kemp, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | A Survey for Efficient Open Domain Question AnsweringabstractQin Zhang, Shangsi Chen, Dongkuan Xu, Qingqing Cao, Xiaojun Chen, Trevor Cohn, Meng Fang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qin Zhang 0011, Shangsi Chen, Dongkuan Xu, Xiaojun Chen 0006, Trevor Cohn |
ACL (1) | 6 |
| 2023 | Fair Enough: Standardizing Evaluation and Model Selection for Fairness Research in NLPabstractModern NLP systems exhibit a range of biases, which a growing literature on model debiasing attempts to correct.However current progress is hampered by a plurality of definitions of bias, means of quantification, and oftentimes vague relation between debiasing algorithms and theoretical measures of bias.This paper seeks to clarify the current situation and plot a course for meaningful progress in fair learning, with two key contributions: (1) making clear inter-relations among the current gamut of methods, and their relation to fairness theory; and (2) addressing the practical problem of model selection, which involves a trade-off between fairness and accuracy and has led to systemic issues in fairness research.Putting them together, we make several recommendations to help shape future work. 1 Timothy Baldwin, Trevor Cohn |
EACL | 3 |
| 2023 | Don't Mess with Mister-in-Between: Improved Negative Search for Knowledge Graph CompletionabstractThe best methods for knowledge graph completion use a 'dual-encoding' framework, a form of neural model with a bottleneck that facilitates fast approximate search over a vast collection of candidates.These approaches are trained using contrastive learning to differentiate between known positive examples and sampled negative instances.The mechanism for sampling negatives to date has been very simple, driven by pragmatic engineering considerations (e.g., using mismatched instances from the same batch).We propose several novel means of finding more informative negatives, based on searching for candidates with high lexical overlaps, from the dual-encoder model and according to knowledge graph structures.Experimental results on four benchmarks show that our best single model improves consistently over previous methods and obtains new state-of-the-art performance, including the challenging large-scale Wikidata5M dataset.Combing different strategies through model ensembling results in a further performance boost. Fan Jiang 0014, Tom Drummond, Trevor Cohn |
EACL | 3 |
| 2023 | Probing Power by Prompting: Harnessing Pre-trained Language Models for Power Connotation FramingabstractSubtle changes in word choice in communication can evoke very different associations with the involved actors.For instance, a company 'employing workers' evokes a more positive connotation than the one 'exploiting' them.This concept is called connotation.This paper investigates whether pre-trained language models (PLMs) encode such subtle connotative information about power differentials between involved entities.We design a probing framework for power connotation, building on Sap et al. (2017)'s operationalization of connotation frames.We show that zero-shot prompting of PLMs leads to above chance prediction of power connotation, however fine-tuning PLMs using our framework drastically improves their accuracy.Using our fine-tuned models, we present a case study of power dynamics in US news reporting on immigration, showing the potential of our framework as a tool for understanding subtle bias in the media. 1 Shima Khanehzar, Trevor Cohn, Gosia Mikolajczak, Lea Frermann |
EACL | 2 |
| 2023 | Performance Prediction via Bayesian Matrix Factorisation for Multilingual Natural Language Processing TasksabstractPerformance prediction for Natural Language Processing (NLP) seeks to reduce the experimental burden resulting from the myriad of different evaluation scenarios, e.g., the combination of languages used in multilingual transfer.In this work, we explore the framework of Bayesian matrix factorisation for performance prediction, as many experimental settings in NLP can be naturally represented in matrix format.Our approach outperforms the stateof-the-art in several NLP benchmarks, including machine translation and cross-lingual entity linking.Furthermore, it also avoids hyperparameter tuning and is able to provide uncertainty estimates over predictions. Viktoria Schram, Daniel Beck, Trevor Cohn |
EACL | 3 |
| 2023 | Fingerprint Attack: Client De-Anonymization in Federated LearningabstractFederated Learning allows collaborative training without data sharing in settings where participants do not trust the central server and one another. Privacy can be further improved by ensuring that communication between the participants and the server is anonymized through a shuffle; decoupling the participant identity from their data. This paper seeks to examine whether such a defense is adequate to guarantee anonymity, by proposing a novel fingerprinting attack over gradients sent by the participants to the server. We show that clustering of gradients can easily break the anonymization in an empirical study of learning federated language models on two language corpora. We then show that training with differential privacy can provide a practical defense against our fingerprint attack. Qiongkai Xu, Trevor Cohn, Olga Ohrimenko |
ECAI | 2 |
| 2023 | Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationabstractModern NLP models are often trained over large untrusted datasets, raising the potential for a malicious adversary to compromise model behaviour.For instance, backdoors can be implanted through crafting training instances with a specific textual trigger and a target label.This paper posits that backdoor poisoning attacks exhibit spurious correlation between simple text features and classification labels, and accordingly, proposes methods for mitigating spurious correlation as means of defence.Our empirical study reveals that the malicious triggers are highly correlated to their target labels; therefore such correlations are extremely distinguishable compared to those scores of benign features, and can be used to filter out potentially problematic instances.Compared with several existing defences, our defence method significantly reduces attack success rates across backdoor attacks, and in the case of insertion-based attacks, our method provides a near-perfect defence. 1 Xuanli He, Qiongkai Xu, Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn |
EMNLP | 5 |
| 2023 | Everybody Needs Good Neighbours: An Unsupervised Locality-based Method for Bias Mitigation
Timothy Baldwin, Trevor Cohn |
ICLR | 3 |
| 2023 | The Next Chapter: A Study of Large Language Models in StorytellingabstractTo enhance the quality of generated stories, recent story generation models have been investigating the utilization of higher-level attributes like plots or commonsense knowledge.The application of prompt-based learning with large language models (LLMs), exemplified by GPT-3, has exhibited remarkable performance in diverse natural language processing (NLP) tasks.This paper conducts a comprehensive investigation, utilizing both automatic and human evaluation, to compare the story generation capacity of LLMs with recent models across three datasets with variations in style, register, and length of stories.The results demonstrate that LLMs generate stories of significantly higher quality compared to other story generation models.Moreover, they exhibit a level of performance that competes with human authors, albeit with the preliminary observation that they tend to replicate real stories in situations involving world knowledge, resembling a form of plagiarism. Zhuohan Xie, Trevor Cohn, Jey Han Lau |
INLG | 2 |
| 2022 | Incorporating Constituent Syntax for Coreference ResolutionabstractSyntax has been shown to benefit Coreference Resolution from incorporating long-range dependencies and structured information captured by syntax trees, either in traditional statistical machine learning based systems or recently proposed neural models. However, most leading systems use only dependency trees. We argue that constituent trees also encode important information, such as explicit span-boundary signals captured by nested multi-word phrases, extra linguistic labels and hierarchical structures useful for detecting anaphora. In this work, we propose a simple yet effective graph-based method to incorporate constituent syntactic structures. Moreover, we also explore to utilise higher-order neighbourhood information to encode rich structures in constituent trees. A novel message propagation mechanism is therefore proposed to enable information flow among elements in syntax trees. Experiments on the English and Chinese portions of OntoNotes 5.0 benchmark show that our proposed model either beats a strong baseline or achieves new state-of-the-art performance. Code is available at https://github.com/Fantabulous-J/Coref-Constituent-Graph. Fan Jiang 0014, Trevor Cohn |
AAAI | 2 |
| 2022 | Measuring and Mitigating Name Biases in Neural Machine TranslationabstractNeural Machine Translation (NMT) systems exhibit problematic biases, such as stereotypical gender bias in the translation of occupation terms into languages with grammatical gender.In this paper we describe a new source of bias prevalent in NMT systems, relating to translations of sentences containing person names.To correctly translate such sentences, a NMT system needs to estimate the gender of names.We show that leading systems are particularly poor at this task, especially for female given names.This bias is deeper than given name gender: we show that the translation of terms with ambiguous sentiment can also be affected by person names, and the same holds true for proper nouns denoting race.To mitigate these biases we propose a simple but effective data augmentation method based on randomly switching entities during translation, which effectively eliminates the problem without any effect on translation quality. Jun Wang 0126, Benjamin I. P. Rubinstein, Trevor Cohn |
ACL (1) | 3 |
| 2022 | The ChEMU 2022 Evaluation Campaign: Information Extraction in Chemical Patents
Yuan Li 0012, Biaoyan Fang, Estrid He, Hiyori Yoshikawa, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zenan Zhai, Zubair Afzal, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
ECIR (2) | 10 |
| 2022 | Balancing out Bias: Achieving Fairness Through Balanced TrainingabstractGroup bias in natural language processing tasks manifests as disparities in system error rates across texts authorized by different demographic groups, typically disadvantaging minority groups.Dataset balancing has been shown to be effective at mitigating bias, however existing approaches do not directly account for correlations between author demographics and linguistic variables, limiting their effectiveness.To achieve Equal Opportunity fairness, such as equal job opportunity without regard to demographics, this paper introduces a simple, but highly effective, objective for countering bias using balanced training.We extend the method in the form of a gated model, which incorporates protected attributes as input, and show that it is effective at reducing bias in predictions through demographic input perturbation, outperforming all other bias mitigation techniques when combined with balanced training.1 Timothy Baldwin, Trevor Cohn |
EMNLP | 3 |
| 2022 | Unsupervised Cross-Lingual Transfer of Structured Predictors without Source DataabstractKemal Kurniawan, Lea Frermann, Philip Schulz, Trevor Cohn. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kemal Kurniawan, Lea Frermann, Philip Schulz, Trevor Cohn |
NAACL-HLT | 4 |
| 2022 | Optimising Equal Opportunity Fairness in Model TrainingabstractAili Shen, Xudong Han, Trevor Cohn, Timothy Baldwin, Lea Frermann. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Aili Shen, Trevor Cohn, Timothy Baldwin, Lea Frermann |
NAACL-HLT | 3 |
| 2022 | Improving negation detection with negation-focused pre-trainingabstractThinh Truong, Timothy Baldwin, Trevor Cohn, Karin Verspoor. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Hung-Thinh Truong, Timothy Baldwin, Trevor Cohn, Karin Verspoor |
NAACL-HLT | 3 |
| 2021 | Commonsense Knowledge in Word Associations and ConceptNetabstractHumans use countless basic, shared facts about the world to efficiently navigate in their environment.This commonsense knowledge is rarely communicated explicitly, however, understanding how commonsense knowledge is represented in different paradigms is important for both deeper understanding of human cognition and for augmenting automatic reasoning systems.This paper presents an in-depth comparison of two large-scale resources of general knowledge: ConceptNet, an engineered relational database, and SWOW a knowledge graph derived from crowd-sourced word associations.We examine the structure, overlap and differences between the two graphs, as well as the extent to which they encode situational commonsense knowledge.We finally show empirically that both resources improve downstream task performance on commonsense reasoning benchmarks over text-only baselines, suggesting that large-scale word association data, which have been obtained for several languages through crowd-sourcing, can be a valuable complement to curated knowledge graphs. 1 Chunhua Liu, Trevor Cohn, Lea Frermann |
CoNLL | 2 |
| 2021 | Learning Coupled Policies for Simultaneous Machine Translation using Imitation LearningabstractWe present a novel approach to efficiently learn a simultaneous translation model with coupled programmer-interpreter policies. First, we present an algorithmic oracle to produce oracle READ/WRITE actions for training bilingual sentence-pairs using the notion of word alignments. This oracle actions are designed to capture enough information from the partial input before writing the output. Next, we perform a coupled scheduled sampling to effectively mitigate the exposure bias when learning both policies jointly with imitation learning. Experiments on six language-pairs show our method outperforms strong baselines in terms of translation quality while keeping the translation delay low. Philip Arthur, Trevor Cohn, Gholamreza Haffari |
EACL | 2 |
| 2021 | Diverse Adversaries for Mitigating Bias in TrainingabstractAdversarial learning can learn fairer and less biased models of language than standard methods.However, current adversarial techniques only partially mitigate model bias, added to which their training procedures are often unstable.In this paper, we propose a novel approach to adversarial learning based on the use of multiple diverse discriminators, whereby discriminators are encouraged to learn orthogonal hidden representations from one another.Experimental results show that our method substantially improves over standard adversarial removal methods, in terms of reducing bias and the stability of training. Timothy Baldwin, Trevor Cohn |
EACL | 3 |
| 2021 | PPT: Parsimonious Parser Transfer for Unsupervised Cross-Lingual AdaptationabstractCross-lingual transfer is a leading technique for parsing low-resource languages in the absence of explicit supervision.Simple 'direct transfer' of a learned model based on a multilingual input encoding has provided a strong benchmark.This paper presents a method for unsupervised cross-lingual transfer that improves over direct transfer systems by using their output as implicit supervision as part of self-training on unlabelled text in the target language.The method assumes minimal resources and provides maximal flexibility by (a) accepting any pre-trained arc-factored dependency parser; (b) assuming no access to source language data; (c) supporting both projective and non-projective parsing; and (d) supporting multi-source transfer.With English as the source language, we show significant improvements over state-of-the-art transfer models on both distant and nearby languages, despite our conceptually simpler approach.We provide analyses of the choice of source languages for multi-source transfer, and the advantage of non-projective parsing.Our code is available online. 1 Kemal Kurniawan, Lea Frermann, Philip Schulz, Trevor Cohn |
EACL | 4 |
| 2021 | ChEMU 2021: Reaction Reference Resolution and Anaphora Resolution in Chemical Patents
Estrid He, Biaoyan Fang, Hiyori Yoshikawa, Yuan Li 0012, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zubair Afzal, Zenan Zhai, Lawrence Cavedon, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
ECIR (2) | 11 |
| 2021 | Fairness-aware Class Imbalanced LearningabstractClass imbalance is a common challenge in many NLP tasks, and has clear connections to bias, in that bias in training data often leads to higher accuracy for majority groups at the expense of minority groups.However there has traditionally been a disconnect between research on class-imbalanced learning and mitigating bias, and only recently have the two been looked at through a common lens.In this work we evaluate long-tail learning methods for tweet sentiment and occupation classification, and extend a margin-loss based approach with methods to enforce fairness.We empirically show through controlled experiments that the proposed approaches help mitigate both class imbalance and demographic biases. 1 Shivashankar Subramanian, Afshin Rahimi 0001, Timothy Baldwin, Trevor Cohn, Lea Frermann |
EMNLP (1) | 4 |
| 2021 | Evaluating Debiasing Techniques for Intersectional BiasesabstractBias is pervasive in NLP models, motivating the development of automatic debiasing techniques.Evaluation of NLP debiasing methods has largely been limited to binary attributes in isolation, e.g., debiasing with respect to binary gender or race, however many corpora involve multiple such attributes, possibly with higher cardinality.In this paper we argue that a truly fair model must consider 'gerrymandering' groups which comprise not only single attributes, but also intersectional groups.We evaluate a form of bias-constrained model which is new to NLP, as well an extension of the iterative nullspace projection technique which can handle multiple protected attributes. Shivashankar Subramanian, Timothy Baldwin, Trevor Cohn, Lea Frermann |
EMNLP (1) | 4 |
| 2021 | It Is Not As Good As You Think! Evaluating Simultaneous Machine Translation on Interpretation DataabstractMost existing simultaneous machine translation (SiMT) systems are trained and evaluated on offline translation corpora.We argue that SiMT systems should be trained and tested on real interpretation data.To illustrate this argument, we propose an interpretation test set and conduct a realistic evaluation of SiMT trained on offline translations.Our results, on our test set along with 3 existing smaller scale language pairs, highlight the difference of up-to 13.83 BLEU score when SiMT models are evaluated on translation vs interpretation data.In the absence of interpretation training data, we propose a translationto-interpretation (T2I) style transfer method which allows converting existing offline translations into interpretation-style data, leading to up-to 2.8 BLEU improvement.However, the evaluation gap remains notable, calling for constructing large-scale interpretation corpora better suited for evaluating and developing SiMT systems. 1 Jinming Zhao, Philip Arthur, Gholamreza Haffari, Trevor Cohn, Ehsan Shareghi |
EMNLP (1) | 4 |
| 2021 | Generating Diverse Descriptions from Semantic GraphsabstractText generation from semantic graphs is traditionally performed with deterministic methods, which generate a unique description given an input graph. However, the generation problem admits a range of acceptable textual outputs, exhibiting lexical, syntactic and semantic variation. To address this disconnect, we present two main contributions. First, we propose a stochastic graph-to-text model, incorporating a latent variable in an encoder-decoder model, and its use in an ensemble. Second, to assess the diversity of the generated sentences, we propose a new automatic evaluation metric which jointly evaluates output diversity and quality in a multi-reference setting. We evaluate the models on WebNLG datasets in English and Russian, and show an ensemble of stochastic models produces diverse sets of generated sentences while, retaining similar quality to state-of-the-art models. Jiuzhou Han, Daniel Beck, Trevor Cohn |
INLG | 3 |
| 2021 | Incorporating Syntax and Semantics in Coreference Resolution with Heterogeneous Graph Attention NetworkabstractExternal syntactic and semantic information has been largely ignored by existing neural coreference resolution models.In this paper, we present a heterogeneous graph-based model to incorporate syntactic and semantic structures of sentences.The proposed graph contains a syntactic sub-graph where tokens are connected based on a dependency tree, and a semantic sub-graph that contains arguments and predicates as nodes and semantic role labels as edges.By applying a graph attention network, we can obtain syntactically and semantically augmented word representation, which can be integrated using an attentive integration layer and gating mechanism.Experiments on the OntoNotes 5.0 benchmark show the effectiveness of our proposed model. 1 Fan Jiang 0014, Trevor Cohn |
NAACL-HLT | 2 |
| 2021 | Framing Unpacked: A Semi-Supervised Interpretable Multi-View Model of Media FramesabstractShima Khanehzar, Trevor Cohn, Gosia Mikolajczak, Andrew Turpin, Lea Frermann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shima Khanehzar, Trevor Cohn, Gosia Mikolajczak, Andrew Turpin, Lea Frermann |
NAACL-HLT | 2 |
| 2021 | A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data PoisoningabstractAs modern neural machine translation (NMT) systems have been widely deployed, their security vulnerabilities require close scrutiny. Most recently, NMT systems have been found vulnerable to targeted attacks which cause them to produce specific, unsolicited, and even harmful translations. These attacks are usually exploited in a white-box setting, where adversarial inputs causing targeted translations are discovered for a known target system. However, this approach is less viable when the target system is black-box and unknown to the adversary (e.g., secured commercial systems). In this paper, we show that targeted attacks on black-box NMT systems are feasible, based on poisoning a small fraction of their parallel training data. We show that this attack can be realised practically via targeted corruption of web documents crawled to form the system’s training data. We then analyse the effectiveness of the targeted poisoning in two common NMT training scenarios: the from-scratch training and the pre-train & fine-tune paradigm. Our results are alarming: even on the state-of-the-art systems trained with massive parallel data (tens of millions), the attacks are still successful (over 50% success rate) under surprisingly low poisoning budgets (e.g., 0.006%). Lastly, we discuss potential defences to counter such attacks. Jun Wang 0126, Francisco Guzmán, Benjamin I. P. Rubinstein, Trevor Cohn |
WWW | 6 |
| 2020 | Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation MetricsabstractAutomatic metrics are fundamental for the development and evaluation of machine translation systems.Judging whether, and to what extent, automatic metrics concur with the gold standard of human evaluation is not a straightforward problem.We show that current methods for judging metrics are highly sensitive to the translations used for assessment, particularly the presence of outliers, which often leads to falsely confident conclusions about a metric's efficacy.Finally, we turn to pairwise system ranking, developing a method for thresholding performance improvement under an automatic metric against human judgements, which allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected.Together, these findings suggest improvements to the protocols for metric evaluation and system performance evaluation in machine translation. Nitika Mathur, Timothy Baldwin, Trevor Cohn |
ACL | 3 |
| 2020 | ChEMU: Named Entity Recognition and Event Extraction of Chemical Reactions from Patents
Dat Quoc Nguyen, Zenan Zhai, Hiyori Yoshikawa, Biaoyan Fang, Christian Druckenbrodt, Camilo Thorne, Ralph Hoessel, Saber A. Akhondi, Trevor Cohn, Timothy Baldwin, Karin Verspoor |
ECIR (2) | 9 |
| 2020 | Decoding As Dynamic Programming For Recurrent Autoregressive Models
Najam Zaidi, Trevor Cohn, Gholamreza Haffari |
ICLR | 2 |
| 2019 | Semi-supervised Stochastic Multi-Domain Learning using Variational InferenceabstractSupervised models of NLP rely on large collections of text which closely resemble the intended testing setting.Unfortunately matching text is often not available in sufficient quantity, and moreover, within any domain of text, data is often highly heterogenous.In this paper we propose a method to distill the important domain signal as part of a multi-domain learning system, using a latent variable model in which parts of a neural model are stochastically gated based on the inferred domain.We compare the use of discrete versus continuous latent variables, operating in a domain-supervised or a domain semi-supervised setting, where the domain is known only for a subset of training inputs.We show that our model leads to substantial performance improvements over competitive benchmark domain adaptation methods, including methods using adversarial learning. Yitong Li 0002, Timothy Baldwin, Trevor Cohn |
ACL (1) | 3 |
| 2019 | Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation EvaluationabstractAccurate, automatic evaluation of machine translation is critical for system tuning, and evaluating progress in the field.We proposed a simple unsupervised metric, and additional supervised metrics which rely on contextual word embeddings to encode the translation and reference sentences.We find that these models rival or surpass all existing metrics in the WMT 2017 sentence-level and systemlevel tracks, and our trained model has a substantially higher correlation with human judgements than all existing metrics on the WMT 2017 to-English sentence level dataset. Nitika Mathur, Timothy Baldwin, Trevor Cohn |
ACL (1) | 3 |
| 2019 | Massively Multilingual Transfer for NERabstractIn cross-lingual transfer, NLP models over one or more source languages are applied to a lowresource target language.While most prior work has used a single source model or a few carefully selected models, here we consider a "massive" setting with many such models.This setting raises the problem of poor transfer, particularly from distant languages.We propose two techniques for modulating the transfer, suitable for zero-shot or few-shot learning, respectively.Evaluating on named entity recognition, we show that our techniques are much more effective than strong baselines, including standard ensembling, and our unsupervised method rivals oracle selection of the single best individual model. 1 Afshin Rahimi 0001, Yuan Li 0012, Trevor Cohn |
ACL (1) | 3 |
| 2019 | Grounding learning of modifier dynamics: An application to color namingabstractXudong Han, Philip Schulz, Trevor Cohn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Philip Schulz, Trevor Cohn |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Deep Ordinal Regression for Pledge Specificity PredictionabstractShivashankar Subramanian, Trevor Cohn, Timothy Baldwin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shivashankar Subramanian, Trevor Cohn, Timothy Baldwin |
EMNLP/IJCNLP (1) | 2 |
| 2019 | A Unified Neural Architecture for Instrumental Audio TasksabstractWithin Music Information Retrieval (MIR), prominent tasks - including pitch-tracking, source-separation, super-resolution, and synthesis - typically call for specialised methods, despite their similarities. Conditional Generative Adversarial Networks (cGANs) have been shown to be highly versatile in learning general image-to-image translations, but have not yet been adapted across MIR. In this work, we present an end-to-end supervisable architecture to perform all aforementioned audio tasks, consisting of a WaveNet synthesiser conditioned on the output of a jointly-trained cGAN spectrogram translator. In doing so, we demonstrate the potential of such flexible techniques to unify MIR tasks, promote efficient transfer learning, and converge research to the improvement of powerful, general methods. Finally, to the best of our knowledge, we present the first application of GANs to guided instrument synthesis. Steven Spratley, Daniel Beck, Trevor Cohn |
ICASSP | 3 |
| 2019 | Exploiting Worker Correlation for Label Aggregation in CrowdsourcingabstractCrowdsourcing has emerged as a core component of data science pipelines. From collected noisy worker labels, aggregation models that incorporate worker reliability parameters aim to infer a latent true annotation. In this paper, we argue that existing crowdsourcing approaches do not sufficiently model worker correlations observed in practical settings; we propose in response an enhanced Bayesian classifier combination (EBCC) model, with inference based on a mean-field variational approach. An introduced mixture of intra-class reliabilities—connected to tensor decomposition and item clustering—induces inter-worker correlation. EBCC does not suffer the limitations of existing correlation models: intractable marginalisation of missing labels and poor scaling to large worker cohorts. Extensive empirical comparison on 17 real-world datasets sees EBCC achieving the highest mean accuracy across 10 benchmark crowdsourcing methods. Yuan Li 0012, Benjamin I. P. Rubinstein, Trevor Cohn |
ICML | 3 |
| 2019 | Truth Inference at Scale: A Bayesian Model for Adjudicating Highly Redundant Crowd AnnotationsabstractCrowd-sourcing is a cheap and popular means of creating training and evaluation datasets for machine learning, however it poses the problem of 'truth inference', as individual workers cannot be wholly trusted to provide reliable annotations. Research into models of annotation aggregation attempts to infer a latent 'true' annotation, which has been shown to improve the utility of crowd-sourced data. However, existing techniques beat simple baselines only in low redundancy settings, where the number of annotations per instance is low (= 3), or in situations where workers are unreliable and produce low quality annotations (e.g., through spamming, random, or adversarial behaviours.) As we show, datasets produced by crowd-sourcing are often not of this type: the data is highly redundantly annotated (= 5 annotations per instance), and the vast majority of workers produce high quality outputs. In these settings, the majority vote heuristic performs very well, and most truth inference models underperform this simple baseline. We propose a novel technique, based on a Bayesian graphical model with conjugate priors, and simple iterative expectation-maximisation inference. Our technique produces competitive performance to the state-of-the-art benchmark methods, and is the only method that significantly outperforms the majority vote heuristic at one-sided level 0.025, shown by significance tests. Moreover, our technique is simple, is implemented in only 50 lines of code, and trains in seconds. 1 Yuan Li 0012, Benjamin I. P. Rubinstein, Trevor Cohn |
WWW | 3 |
| 2019 | Gaussian Processes for Rumour Stance Classification in Social MediaabstractSocial media tend to be rife with rumours while new reports are released piecemeal during breaking news. Interestingly, one can mine multiple reactions expressed by social media users in those situations, exploring their stance towards rumours, ultimately enabling the flagging of highly disputed rumours as being potentially false. In this work, we set out to develop an automated, supervised classifier that uses multi-task learning to classify the stance expressed in each individual tweet in a conversation around a rumour as either supporting, denying or questioning the rumour. Using a Gaussian Process classifier, and exploring its effectiveness on two datasets with very different characteristics and varying distributions of stances, we show that our approach consistently outperforms competitive baseline classifiers. Our classifier is especially effective in estimating the distribution of different types of stance associated with a given rumour, which we set forth as a desired characteristic for a rumour-tracking system that will show both ordinary users of Twitter and professional news practitioners how others orient to the disputed veracity of a rumour, with the final aim of establishing its actual truth value. Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Arkaitz Zubiaga, Maria Liakata, Rob Procter |
ACM Trans. Inf. Syst. | 3 |
| 2018 | Deep-speare: A joint neural model of poetic language, meter and rhymeabstractIn this paper, we propose a joint architecture that captures language, rhyme and meter for sonnet modelling.We assess the quality of generated poems using crowd and expert judgements.The stress and rhyme models perform very well, as generated poems are largely indistinguishable from human-written poems.Expert evaluation, however, reveals that a vanilla language model captures meter implicitly, and that machine-generated poems still underperform in terms of readability and emotion.Our research shows the importance expert evaluation for poetry generation, and that future research should look beyond rhyme/meter and focus on poetic language. Jey Han Lau, Trevor Cohn, Timothy Baldwin, Julian Brooke, Adam Hammond |
ACL (1) | 2 |
| 2018 | Semi-supervised User Geolocation via Graph Convolutional NetworksabstractSocial media user geolocation is vital to many applications such as event detection.In this paper, we propose GCN, a multiview geolocation model based on Graph Convolutional Networks, that uses both text and network context.We compare GCN to the state-of-the-art, and to two baselines we propose, and show that our model achieves or is competitive with the stateof-the-art over three benchmark geolocation datasets when sufficient supervision is available.We also evaluate GCN under a minimal supervision scenario, and show it outperforms baselines.We find that highway network gates are essential for controlling the amount of useful neighbourhood expansion in GCN. Afshin Rahimi 0001, Trevor Cohn, Timothy Baldwin |
ACL (1) | 2 |
| 2018 | Graph-to-Sequence Learning using Gated Graph Neural NetworksabstractMany NLP applications can be framed as a graph-to-sequence learning problem. Previous work proposing neural architectures on graph-to-sequence obtained promising results compared to grammar-based approaches but still rely on linearisation heuristics and/or standard recurrent networks to achieve the best performance. In this work propose a new model that encodes the full structural information contained in the graph. Our architecture couples the recently proposed Gated Graph Neural Networks with an input transformation that allows nodes and edges to have their own hidden representations, while tackling the parameter explosion problem present in previous work. Experimental results shows that our model outperforms strong baselines in generation from AMR graphs and syntax-based neural machine translation. Daniel Beck, Gholamreza Haffari, Trevor Cohn |
ACL (1) | 3 |
| 2018 | A Stochastic Decoder for Neural Machine TranslationabstractThe process of translation is ambiguous, in that there are typically many valid translations for a given sentence.This gives rise to significant variation in parallel corpora, however, most current models of machine translation do not account for this variation, instead treating the problem as a deterministic process.To this end, we present a deep generative model of machine translation which incorporates a chain of latent variables, in order to account for local lexical and syntactic variation in parallel corpora.We provide an indepth analysis of the pitfalls encountered in variational inference for training deep generative models.Experiments on several different language pairs demonstrate that the model consistently improves over strong baselines.* Code and a workflow that reproduces the experiments are available at https://github.com/philschulz/ stochastic-decoder. Philip Schulz, Wilker Aziz, Trevor Cohn |
ACL (1) | 3 |
| 2018 | Evaluating the Utility of Hand-crafted Features in Sequence LabelingabstractConventional wisdom is that hand-crafted features are redundant for deep learning models, as they already learn adequate representations of text automatically from corpora.In this work, we test this claim by proposing a new method for exploiting handcrafted features as part of a novel hybrid learning approach, incorporating a feature auto-encoder loss component.We evaluate on the task of named entity recognition (NER), where we show that including manual features for partof-speech, word shapes and gazetteers can improve the performance of a neural CRF model.We obtain a F 1 of 91.89 for the CoNLL-2003 English shared task, which significantly outperforms a collection of highly competitive baseline models.We also present an ablation study showing the importance of autoencoding, over using features as either inputs or outputs alone, and moreover, show including the autoencoder components reduces training requirements to 60%, while retaining the same predictive accuracy. Minghao Wu, Fei Liu 0023, Trevor Cohn |
EMNLP | 3 |
| 2018 | Evaluation Phonemic Transcription of Low-Resource Tonal Languages for Language Documentation
Oliver Adams, Trevor Cohn, Graham Neubig, Hilaria Cruz, Steven Bird, Alexis Michaud |
LREC | 2 |
| 2018 | Hierarchical Structured Model for Fine-to-Coarse Manifesto Text AnalysisabstractShivashankar Subramanian, Trevor Cohn, Timothy Baldwin. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shivashankar Subramanian, Trevor Cohn, Timothy Baldwin |
NAACL-HLT | 2 |
| 2018 | Discourse-aware rumour stance classification in social media using sequential classifiers
Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Isabelle Augenstein |
Inf. Process. Manag. | 7 |
| 2017 | Topically Driven Neural Language ModelabstractLanguage models are typically applied at the sentence level, without access to the broader document context.We present a neural language model that incorporates document context in the form of a topic model-like architecture, thus providing a succinct representation of the broader document context outside of the current sentence.Experiments over a range of datasets demonstrate that our model outperforms a pure sentence-based model in terms of language model perplexity, and leads to topics that are potentially more coherent than those produced by a standard LDA topic model.Our model also has the ability to generate related sentences for a topic, providing another way to interpret topics. Jey Han Lau, Timothy Baldwin, Trevor Cohn |
ACL (1) | 3 |
| 2017 | Longitudinal Modeling of Social Media with Hawkes Process Based on Users and NetworksabstractOnline social media provide a platform for rapid network propagation of information at an unprecedented scale. In this paper, we study the evolution of information cascades in Twitter using a point process model of user activity. Twitter is rich with heterogenous information on users and network structure. We develop several Hawkes process models considering various properties of Twitter including conversational structure, users' connections and general features of users including the textual information, and show how they are helpful in modeling the social network activity. Evaluation on Twitter data sets shows that incorporating richer properties improves the performance in predicting future activity of users and memes. P. K. Srijith, Michal Lukasik, Kalina Bontcheva, Trevor Cohn |
ASONAM | 4 |
| 2017 | Multi-step prediction with missing smart sensor data using multi-task Gaussian processesabstractWith the proliferation of sensors and the increased connectivity of citizens, many global cities are increasingly adopting Smart City initiatives. Such initiatives provide real-time monitoring capabilities, and effective modelling techniques allow the prediction of future states in a city. For example, urban electricity smart meter data can be utilised to predict future demand in order to facilitate capacity planning. However, the accuracy of this foresight is often marred by low quality and missing sensor data in real-world systems. In this work, we focus on the problem of reliable forecasting by mitigating the effect of missing data on forecast accuracy. In order to mitigate the effects of missing data, we develop a multi-task learning scheme to jointly learn Gaussian Process Regression models between highly correlated sensors. We demonstrate that our methods are robust in a variety of error generation scenarios. We validate our methods based on publicly available and real-world datasets related to electricity smart meters in a university campus and pedestrian counts in a global city, where we achieve significant improvement over competitive baselines and other effective forecasting methods. Pasan Karunaratne, Masud Moshtaghi, Shanika Karunasekera, Aaron Harwood, Trevor Cohn |
IEEE BigData | 5 |
| 2017 | Cross-Lingual Word Embeddings for Low-Resource Language ModelingabstractOliver Adams, Adam Makarucha, Graham Neubig, Steven Bird, Trevor Cohn. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Oliver Adams, Adam J. Makarucha, Graham Neubig, Steven Bird, Trevor Cohn |
EACL (1) | 5 |
| 2017 | Multilingual Training of Crosslingual Word EmbeddingsabstractLong Duong, Hiroshi Kanayama, Tengfei Ma, Steven Bird, Trevor Cohn. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Long Duong, Hiroshi Kanayama, Tengfei Ma 0001, Steven Bird, Trevor Cohn |
EACL (1) | 5 |
| 2017 | Learning how to Active Learn: A Deep Reinforcement Learning ApproachabstractActive learning aims to select a small subset of data for annotation such that a classifier learned on the data is highly accurate.This is usually done using heuristic selection methods, however the effectiveness of such methods is limited and moreover, the performance of heuristics varies between datasets.To address these shortcomings, we introduce a novel formulation by reframing the active learning as a reinforcement learning problem and explicitly learning a data selection policy, where the policy takes the role of the active learning heuristic.Importantly, our method allows the selection policy learned using simulation on one language to be transferred to other languages.We demonstrate our method using cross-lingual named entity recognition, observing uniform improvements over traditional active learning. Yuan Li 0012, Trevor Cohn |
EMNLP | 3 |
| 2017 | Towards Decoding as Continuous Optimisation in Neural Machine TranslationabstractWe propose a novel decoding approach for neural machine translation (NMT) based on continuous optimisation.We reformulate decoding, a discrete optimization problem, into a continuous problem, such that optimization can make use of efficient gradient-based techniques.Our powerful decoding framework allows for more accurate decoding for standard neural machine translation models, as well as enabling decoding in intractable models such as intersection of several different NMT models.Our empirical results show that our decoding framework is effective, and can leads to substantial improvements in translations, especially in situations where greedy search and beam search are not feasible.Finally, we show how the technique is highly competitive with, and complementary to, reranking. Cong Duy Vu Hoang, Gholamreza Haffari, Trevor Cohn |
EMNLP | 3 |
| 2017 | Sequence Effects in Crowdsourced AnnotationsabstractManual data annotation is a vital component of NLP research.When designing annotation tasks, properties of the annotation interface can lead to unintentional artefacts in the resulting dataset, biasing the evaluation.In this paper, we explore sequence effects where annotations of an item are affected by the preceding items.Having assigned one label to an instance, the annotator may be less (or more) likely to assign the same label to the next.During rating tasks, seeing a low quality item may affect the score given to the next item either positively or negatively.We see clear evidence of both types of effects using auto-correlation studies over three different crowdsourced datasets.We then recommend a simple way to minimise sequence effects. Nitika Mathur, Timothy Baldwin, Trevor Cohn |
EMNLP | 3 |
| 2017 | Continuous Representation of Location for Geolocation and Lexical Dialectology using Mixture Density NetworksabstractWe propose a method for embedding twodimensional locations in a continuous vector space using a neural network-based model incorporating mixtures of Gaussian distributions, presenting two model variants for text-based geolocation and lexical dialectology.Evaluated over Twitter data, the proposed model outperforms conventional regression-based geolocation and provides a better estimate of uncertainty.We also show the effectiveness of the representation for predicting words from location in lexical dialectology, and evaluate it using the DARE dataset. Afshin Rahimi 0001, Timothy Baldwin, Trevor Cohn |
EMNLP | 3 |
| 2017 | Modelling the Working Week for Multi-Step Forecasting using Gaussian Process RegressionabstractIn time-series forecasting, regression is a popular method, with Gaussian Process Regression widely held to be the state of the art. The versatility of Gaussian Processes has led to them being used in many varied application domains. However, though many real-world applications involve data which follows a working-week structure, where weekends exhibit substantially different behavior to weekdays, methods for explicit modelling of working-week effects in Gaussian Process Regression models have not been proposed. Not explicitly modelling the working week fails to incorporate a significant source of information which can be invaluable in forecasting scenarios. In this work we provide novel kernel-combination methods to explicitly model working-week effects in time-series data for more accurate predictions using Gaussian Process Regression. Further, we demonstrate that prediction accuracy can be improved by constraining the non-convex optimization process of finding optimal hyperparameter values. We validate the effectiveness of our methods by performing multi-step prediction on two real-world publicly available time-series datasets - one relating to electricity Smart Meter data of the University of Melbourne, and the other relating to the counts of pedestrians in the City of Melbourne. Pasan Karunaratne, Masud Moshtaghi, Shanika Karunasekera, Aaron Harwood, Trevor Cohn |
IJCAI | 5 |
| 2017 | Compressed Nonparametric Language ModellingabstractHierarchical Pitman-Yor Process priors are compelling for learning language models, outperforming point-estimate based methods. However, these models remain unpopular due to computational and statistical inference issues, such as memory and time usage, as well as poor mixing of sampler. In this work we propose a novel framework which represents the HPYP model compactly using compressed suffix trees. Then, we develop an efficient approximate inference scheme in this framework that has a much lower memory footprint compared to full HPYP and is fast in the inference time. The experimental results illustrate that our model can be built on significantly larger datasets compared to previous HPYP models, while being several orders of magnitudes smaller, fast for training and inference, and outperforming the perplexity of the state-of-the-art Modified Kneser-Ney count-based LM smoothing by up to 15%. Ehsan Shareghi, Gholamreza Haffari, Trevor Cohn |
IJCAI | 3 |
| 2017 | End-to-end Network for Twitter Geolocation Prediction and HashingabstractWe propose an end-to-end neural network to predict the geolocation of a tweet. The network takes as input a number of raw Twitter metadata such as the tweet message and associated user account information. Our model is language independent, and despite minimal feature engineering, it is interpretable and capable of learning location indicative words and timing patterns. Compared to state-of-the-art systems, our model outperforms them by 2%-6%. Additionally, we propose extensions to the model to compress representation learnt by the network into binary codes. Experiments show that it produces compact codes compared to benchmark hashing algorithms. An implementation of the model is released publicly. Jey Han Lau, Lianhua Chi, Khoi-Nguyen Tran, Trevor Cohn |
IJCNLP(1) | 4 |
| 2017 | Capturing Long-range Contextual Dependencies with Memory-enhanced Conditional Random FieldsabstractDespite successful applications across a broad range of NLP tasks, conditional random fields (“CRFs”), in particular the linear-chain variant, are only able to model local features. While this has important benefits in terms of inference tractability, it limits the ability of the model to capture long-range dependencies between items. Attempts to extend CRFs to capture long-range dependencies have largely come at the cost of computational complexity and approximate inference. In this work, we propose an extension to CRFs by integrating external memory, taking inspiration from memory networks, thereby allowing CRFs to incorporate information far beyond neighbouring steps. Experiments across two tasks show substantial improvements over strong CRF and LSTM baselines. Fei Liu 0023, Timothy Baldwin, Trevor Cohn |
IJCNLP(1) | 3 |
| 2016 | Convolution Kernels for Discriminative Learning from Streaming TextabstractTime series modeling is an important problem with many applications in different domains. Here we consider discriminative learning from time series, where we seek to predict an output response variable based on time series input. We develop a method based on convolution kernels to model discriminative learning over streams of text. Our method outperforms competitive baselines in three synthetic and two real datasets, rumour frequency modeling and popularity prediction tasks. Michal Lukasik, Trevor Cohn |
AAAI | 2 |
| 2016 | Take and Took, Gaggle and Goose, Book and Read: Evaluating the Utility of Vector Differences for Lexical Relation LearningabstractRecent work has shown that simple vector subtraction over word embeddings is surprisingly effective at capturing different lexical relations, despite lacking explicit supervision.Prior work has evaluated this intriguing result using a word analogy prediction formulation and hand-selected relations, but the generality of the finding over a broader range of lexical relation types and different learning settings has not been evaluated.In this paper, we carry out such an evaluation in two learning settings:(1) spectral clustering to induce word relations, and ( 2) supervised learning to classify vector differences into relation types.We find that word embeddings capture a surprising amount of information, and that, under suitable supervised training, vector subtraction generalises well to a broad range of relations, including over unseen lexical items. Ekaterina Vylomova, Laura Rimell, Trevor Cohn, Timothy Baldwin |
ACL (1) | 3 |
| 2016 | Exploring Prediction Uncertainty in Machine Translation Quality EstimationabstractMachine Translation Quality Estimation is a notoriously difficult task, which lessens its usefulness in real-world translation environments.Such scenarios can be improved if quality predictions are accompanied by a measure of uncertainty.However, models in this task are traditionally evaluated only in terms of point estimate metrics, which do not take prediction uncertainty into account.We investigate probabilistic methods for Quality Estimation that can provide well-calibrated uncertainty estimates and evaluate them in terms of their full posterior predictive distributions.We also show how this posterior information can be useful in an asymmetric risk scenario, which aims to capture typical situations in translation workflows. Daniel Beck, Lucia Specia, Trevor Cohn |
CoNLL | 3 |
| 2016 | Learning when to trust distant supervision: An application to low-resource POS tagging using cross-lingual projectionabstractCross lingual projection of linguistic annotation suffers from many sources of bias and noise, leading to unreliable annotations that cannot be used directly.In this paper, we introduce a novel approach to sequence tagging that learns to correct the errors from cross-lingual projection using an explicit debiasing layer.This is framed as joint learning over two corpora, one tagged with gold standard and the other with projected tags.We evaluated with only 1,000 tokens tagged with gold standard tags, along with more plentiful parallel data.Our system equals or exceeds the state-of-the-art on eight simulated lowresource settings, as well as two real lowresource languages, Malagasy and Kinyarwanda. Trevor Cohn |
CoNLL | 2 |
| 2016 | Learning a Lexicon and Translation Model from Phoneme LatticesabstractLanguage documentation begins by gathering speech.Manual or automatic transcription at the word level is typically not possible because of the absence of an orthography or prior lexicon, and though manual phonemic transcription is possible, it is prohibitively slow.On the other hand, translations of the minority language into a major language are more easily acquired.We propose a method to harness such translations to improve automatic phoneme recognition.The method assumes no prior lexicon or translation model, instead learning them from phoneme lattices and translations of the speech being transcribed.Experiments demonstrate phoneme error rate improvements against two baselines and the model's ability to learn useful bilingual lexical entries. Oliver Adams, Graham Neubig, Trevor Cohn, Steven Bird, Quoc Truong Do, Satoshi Nakamura 0001 |
EMNLP | 3 |
| 2016 | Learning Crosslingual Word Embeddings without Bilingual CorporaabstractCrosslingual word embeddings represent lexical items from different languages in the same vector space, enabling transfer of NLP tools. However, previous attempts had expensive resource requirements, difficulty incorporating monolingual data or were unable to handle polysemy. We address these drawbacks in our method which takes advantage of a high coverage dictionary in an EM style training algorithm over monolingual corpora in two languages. Our model achieves state-of-the-art performance on bilingual lexicon induction task exceeding models using large bilingual corpora, and competitive results on the monolingual word similarity and cross-lingual document classification task. Long Duong, Hiroshi Kanayama, Tengfei Ma 0001, Steven Bird, Trevor Cohn |
EMNLP | 5 |
| 2016 | Learning Robust Representations of TextabstractDeep neural networks have achieved remarkable results across many language processing tasks, however these methods are highly sensitive to noise and adversarial attacks.We present a regularization based method for limiting network sensitivity to its inputs, inspired by ideas from computer vision, thus learning models that are more robust.Empirical evaluation over a range of sentiment datasets with a convolutional neural network shows that, compared to a baseline model and the dropout method, our method achieves superior performance over noisy inputs and out-of-domain data. 1 Yitong Li 0002, Trevor Cohn, Timothy Baldwin |
EMNLP | 2 |
| 2016 | Richer Interpolative Smoothing Based on Modified Kneser-Ney Language ModelingabstractIn this work we present a generalisation of the Modified Kneser-Ney interpolative smoothing for richer smoothing via additional discount parameters.We provide mathematical underpinning for the estimator of the new discount parameters, and showcase the utility of our rich MKN language models on several European languages.We further explore the interdependency among the training data size, language model order, and number of discount parameters.Our empirical results illustrate that larger number of discount parameters, i) allows for better allocation of mass in the smoothing process, particularly on small data regime where statistical sparsity is severe, and ii) leads to significant reduction in perplexity, particularly for out-of-domain test sets which introduce higher ratio of out-ofvocabulary words. 1 Ehsan Shareghi, Trevor Cohn, Gholamreza Haffari |
EMNLP | 2 |
| 2016 | Learning a Translation Model from Word LatticesabstractTranslation models have been used to improve automatic speech recognition when speech input is paired with a written translation, primarily for the task of computer-aided translation. Existing approaches require large amounts of parallel text for training the translation models, but for many language pairs this data is not available. We propose a model for learning lexical translation parameters directly from the word lattices for which a transcription is sought. The model is expressed through composition of each lattice with a weighted finite-state transducer representing the translation model, where inference is performed by sampling paths through the composed finitestate transducer. We show consistent word error rate reductions in two datasets, using between just 20 minutes and 4 hours of speech input, additionally outperforming a translation model trained on the 1-best path. Oliver Adams, Graham Neubig, Trevor Cohn, Steven Bird |
INTERSPEECH | 3 |
| 2016 | Studying the Temporal Dynamics of Word Co-occurrences: An Application to Event Detection
Daniel Preotiuc-Pietro, P. K. Srijith, Mark Hepple, Trevor Cohn |
LREC | 4 |
| 2016 | Incorporating Structural Alignment Biases into an Attentional Neural Translation ModelabstractNeural encoder-decoder models of machine translation have achieved impressive results, rivalling traditional translation models. However their modelling formulation is overly simplistic, and omits several key inductive biases built into traditional models. In this paper we extend the attentional neural translation model to include structural biases from word based alignment models, including positional bias, Markov conditioning, fertility and agreement over translation directions. We show improvements over a baseline attentional model and standard phrase-based model over several language pairs, evaluating on difficult languages in a low resource setting. Trevor Cohn, Cong Duy Vu Hoang, Ekaterina Vymolova, Kaisheng Yao, Chris Dyer, Gholamreza Haffari |
HLT-NAACL | 1 |
| 2016 | An Attentional Model for Speech Translation Without TranscriptionabstractLong Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, Trevor Cohn. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Long Duong, Antonios Anastasopoulos, David Chiang 0001, Steven Bird, Trevor Cohn |
HLT-NAACL | 5 |
| 2016 | Incorporating Side Information into Recurrent Neural Network Language ModelsabstractCong Duy Vu Hoang, Trevor Cohn, Gholamreza Haffari. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Cong Duy Vu Hoang, Trevor Cohn, Gholamreza Haffari |
HLT-NAACL | 2 |
| 2016 | Fast, Small and Exact: Infinite-order Language Modelling with Compressed Suffix TreesabstractEfficient methods for storing and querying are critical for scaling high-order m-gram language models to large corpora. We propose a language model based on compressed suffix trees, a representation that is highly compact and can be easily held in memory, while supporting queries needed in computing language model probabilities on-the-fly. We present several optimisations which improve query runtimes up to 2500×, despite only incurring a modest increase in construction time and memory usage. For large corpora and high Markov orders, our method is highly competitive with the state-of-the-art KenLM package. It imposes much lower memory requirements, often by orders of magnitude, and has runtimes that are either similar (for training) or comparable (for querying). Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, Trevor Cohn |
Trans. Assoc. Comput. Linguistics | 4 |
| 2015 | Predicting Peer-to-Peer Loan Rates Using Bayesian Non-Linear RegressionabstractPeer-to-peer lending is a new highly liquid market for debt, which is rapidly growing in popularity. Here we consider modelling market rates, developing a non-linear Gaussian Process regression method which incorporates both structured data and unstructured text from the loan application. We show that the peer-to-peer market is predictable, and identify a small set of key factors with high predictive power. Our approach outperforms baseline methods for predicting market rates, and generates substantial profit in a trading simulation. Zsolt Bitvai, Trevor Cohn |
AAAI | 2 |
| 2015 | Cross-lingual Transfer for Unsupervised Dependency Parsing Without Parallel DataabstractCross-lingual transfer has been shown to produce good results for dependency parsing of resource-poor languages.Although this avoids the need for a target language treebank, most approaches have still used large parallel corpora.However, parallel data is scarce for low-resource languages, and we report a new method that does not need parallel data.Our method learns syntactic word embeddings that generalise over the syntactic contexts of a bilingual vocabulary, and incorporates these into a neural network parser.We show empirical improvements over a baseline delexicalised parser on both the CoNLL and Universal Dependency Treebank datasets.We analyse the importance of the source languages, and show that combining multiple source-languages leads to a substantial improvement. Long Duong, Trevor Cohn, Steven Bird, Paul Cook |
CoNLL | 2 |
| 2015 | A Neural Network Model for Low-Resource Universal Dependency ParsingabstractAccurate dependency parsing requires large treebanks, which are only available for a few languages. We propose a method that takes advantage of shared structure across languages to build a mature parser using less training data. We propose a model for learning a shared “univer-sal ” parser that operates over an inter-lingual continuous representation of lan-guage, along with language-specific map-ping components. Compared with super-vised learning, our methods give a con-sistent 8-10 % improvement across several treebanks in low-resource simulations. 1 Long Duong, Trevor Cohn, Steven Bird, Paul Cook |
EMNLP | 2 |
| 2015 | Classifying Tweet Level Judgements of Rumours in Social MediaabstractSocial media is a rich source of rumours and corresponding community reactions.Rumours reflect different characteristics, some shared and some individual.We formulate the problem of classifying tweet level judgements of rumours as a supervised learning task.Both supervised and unsupervised domain adaptation are considered, in which tweets from a rumour are classified on the basis of other annotated rumours.We demonstrate how multi-task learning helps achieve good results on rumours from the 2011 England riots. Michal Lukasik, Trevor Cohn, Kalina Bontcheva |
EMNLP | 2 |
| 2015 | Modeling Tweet Arrival Times using Log-Gaussian Cox ProcessesabstractResearch on modeling time series text corpora has typically focused on predicting what text will come next, but less well studied is predicting when the next text event will occur.In this paper we address the latter case, framed as modeling continuous inter-arrival times under a log-Gaussian Cox process, a form of inhomogeneous Poisson process which captures the varying rate at which the tweets arrive over time.In an application to rumour modeling of tweets surrounding the 2014 Ferguson riots, we show how interarrival times between tweets can be accurately predicted, and that incorporating textual features further improves predictions. Michal Lukasik, P. K. Srijith, Trevor Cohn, Kalina Bontcheva |
EMNLP | 3 |
| 2015 | Compact, Efficient and Unlimited Capacity: Language Modeling with Compressed Suffix TreesabstractEfficient methods for storing and querying language models are critical for scaling to large corpora and high Markov orders.In this paper we propose methods for modeling extremely large corpora without imposing a Markov condition.At its core, our approach uses a succinct index -a compressed suffix tree -which provides near optimal compression while supporting efficient search.We present algorithms for on-the-fly computation of probabilities under a Kneser-Ney language model.Our technique is exact and although slower than leading LM toolkits, it shows promising scaling properties, which we demonstrate through ∞-order modeling over the full Wikipedia collection. Ehsan Shareghi, Matthias Petri, Gholamreza Haffari, Trevor Cohn |
EMNLP | 4 |
| 2015 | Exploiting Text and Network Context for Geolocation of Social Media UsersabstractAfshin Rahimi, Duy Vu, Trevor Cohn, Timothy Baldwin. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Afshin Rahimi 0001, Duy Vu, Trevor Cohn, Timothy Baldwin |
HLT-NAACL | 3 |
| 2015 | Structured Prediction of Sequences and Trees Using Infinite Contexts
Ehsan Shareghi, Gholamreza Haffari, Trevor Cohn, Ann E. Nicholson |
ECML/PKDD (2) | 3 |
| 2015 | Day trading profit maximization with multi-task learning and technical analysis
Zsolt Bitvai, Trevor Cohn |
Mach. Learn. | 2 |
| 2015 | A Bayesian non-linear method for feature selection in machine translation quality estimation
Kashif Shah, Trevor Cohn, Lucia Specia |
Mach. Transl. | 2 |
| 2015 | Learning Structural Kernels for Natural Language ProcessingabstractStructural kernels are a flexible learning paradigm that has been widely used in Natural Language Processing. However, the problem of model selection in kernel-based methods is usually overlooked. Previous approaches mostly rely on setting default values for kernel hyperparameters or using grid search, which is slow and coarse-grained. In contrast, Bayesian methods allow efficient model selection by maximizing the evidence on the training data through gradient-based methods. In this paper we show how to perform this in the context of structural kernels by using Gaussian Processes. Experimental results on tree kernels show that this procedure results in better prediction performance compared to hyperparameter optimization via grid search. The framework proposed in this paper can be adapted to other structures besides trees, e.g., strings and graphs, thereby extending the utility of kernel-based methods. Daniel Beck, Trevor Cohn, Christian Hardmeier, Lucia Specia |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Factored Markov Translation with Robust ModelingabstractPhrase-based translation models usually memorize local translation literally and make independent assumption between phrases which makes it neither generalize well on unseen data nor model sentence-level effects between phrases. In this pa-per we present a new method to model correlations between phrases as a Markov model and meanwhile employ a robust smoothing strategy to provide better gen-eralization. This method defines a re-cursive estimation process and backs off in parallel paths to infer richer structures. Our evaluation shows an 1.1–3.2 % BLEU improvement over competitive baselines for Chinese-English and Arabic-English translation. 1 Yang Feng 0004, Trevor Cohn, Xinkai Du |
CoNLL | 2 |
| 2014 | Predicting and Characterising User Impact on TwitterabstractThe open structure of online social networks and their uncurated nature give rise to problems of user credibility and influence.In this paper, we address the task of predicting the impact of Twitter users based only on features under their direct control, such as usage statistics and the text posted in their tweets.We approach the problem as regression and apply linear as well as nonlinear learning methods to predict a user impact score, estimated by combining the numbers of the user's followers, followees and listings.The experimental results point out that a strong prediction performance is achieved, especially for models based on the Gaussian Processes framework.Hence, we can interpret various modelling components, transforming them into indirect 'suggestions' for impact boosting. Vasileios Lampos, Nikolaos Aletras, Daniel Preotiuc-Pietro, Trevor Cohn |
EACL | 4 |
| 2014 | Data selection for discriminative training in statistical machine translation
Xingyi Song, Lucia Specia, Trevor Cohn |
EAMT | 3 |
| 2014 | Joint Emotion Analysis via Multi-task Gaussian ProcessesabstractWe propose a model for jointly predicting multiple emotions in natural language sentences.Our model is based on a low-rank coregionalisation approach, which combines a vector-valued Gaussian Process with a rich parameterisation scheme.We show that our approach is able to learn correlations and anti-correlations between emotions on a news headlines dataset.The proposed model outperforms both singletask baselines and other multi-task approaches. Daniel Beck, Trevor Cohn, Lucia Specia |
EMNLP | 2 |
| 2014 | What Can We Get From 1000 Tokens? A Case Study of Multilingual POS Tagging For Resource-Poor Languagesabstractby stating that they required an external tag dictionary.We have corrected these inaccuracies to reflect their modest data requirements. Long Duong, Trevor Cohn, Karin Verspoor, Steven Bird, Paul Cook |
EMNLP | 2 |
| 2013 | An Infinite Hierarchical Bayesian Model of Phrasal Translation
Trevor Cohn, Gholamreza Haffari |
ACL (1) | 1 |
| 2013 | Modelling Annotator Bias with Multi-task Gaussian Processes: An Application to Machine Translation Quality Estimation
Trevor Cohn, Lucia Specia |
ACL (1) | 1 |
| 2013 | A Markov Model of Machine Translation using Non-parametric Bayesian Inference
Yang Feng 0004, Trevor Cohn |
ACL (1) | 2 |
| 2013 | A user-centric model of voting intention from Social Media
Vasileios Lampos, Daniel Preotiuc-Pietro, Trevor Cohn |
ACL (1) | 3 |
| 2013 | Topic-Oriented Words as Features for Named Entity Recognition
Ziqi Zhang 0001, Trevor Cohn, Fabio Ciravegna |
CICLing (1) | 2 |
| 2013 | A temporal model of text periodicities using Gaussian ProcessesabstractTemporal variations of text are usually ignored in NLP applications.However, text use changes with time, which can affect many applications.In this paper we model periodic distributions of words over time.Focusing on hashtag frequency in Twitter, we first automatically identify the periodic patterns.We use this for regression in order to forecast the volume of a hashtag based on past data.We use Gaussian Processes, a state-ofthe-art bayesian non-parametric model, with a novel periodic kernel.We demonstrate this in a text classification setting, assigning the tweet hashtag based on the rest of its text.This method shows significant improvements over competitive baselines. Daniel Preotiuc-Pietro, Trevor Cohn |
EMNLP | 2 |
| 2013 | Adaptation of lecture speech recognition system with machine translation outputabstractIn spoken language translation, integration of the ASR and MT components is critical for good performance. In this paper, we consider the recognition setting where a text translation of each utterance is also available. We present experiments with different ASR system adaptation techniques to exploit MT system outputs. In particular, N-best MT outputs are represented as an utterance-specific language model, which are then used to rescore ASR lattices. We show that this method improves significantly over ASR alone, resulting in an absolute WER reduction of more than 6% for both indomain and out-of-domain acoustic models. Raymond W. M. Ng, Thomas Hain, Trevor Cohn |
ICASSP | 3 |
| 2013 | An abstractive approach to sentence compressionabstractIn this article we generalize the sentence compression task. Rather than simply shorten a sentence by deleting words or constituents, as in previous work, we rewrite it using additional operations such as substitution, reordering, and insertion. We present an experimental study showing that humans can naturally create abstractive sentences using a variety of rewrite operations, not just deletion. We next create a new corpus that is suited to the abstractive compression task and formulate a discriminative tree-to-tree transduction model that can account for structural and lexical mismatches. The model incorporates a grammar extraction method, uses a language model for coherent output, and can be easily tuned to a wide range of compression-specific loss functions. Trevor Cohn, Mirella Lapata |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | Left-to-Right Tree-to-String Decoding with Prediction
Yang Feng 0004, Yang Liu 0005, Qun Liu 0001, Trevor Cohn |
EMNLP-CoNLL | 4 |
| 2012 | Evaluating a Morphological Analyser of Inuktitut
Jeremy Nicholson, Trevor Cohn, Timothy Baldwin |
HLT-NAACL | 2 |
| 2011 | A Hierarchical Pitman-Yor Process HMM for Unsupervised Part of Speech Induction
Phil Blunsom, Trevor Cohn |
ACL | 2 |
| 2010 | Multi-Document Summarization Using A* Search and Discriminative Learning
Ahmet Aker, Trevor Cohn, Robert J. Gaizauskas |
EMNLP | 2 |
| 2010 | Unsupervised Induction of Tree Substitution Grammars for Dependency Parsing
Phil Blunsom, Trevor Cohn |
EMNLP | 2 |
| 2010 | Inducing Synchronous Grammars with Slice Sampling
Phil Blunsom, Trevor Cohn |
HLT-NAACL | 2 |
| 2010 | Inducing Tree-Substitution Grammars
Trevor Cohn, Phil Blunsom, Sharon Goldwater |
J. Mach. Learn. Res. | 1 |
| 2009 | A Gibbs Sampler for Phrasal Synchronous Grammar Induction
Phil Blunsom, Trevor Cohn, Chris Dyer, Miles Osborne |
ACL/IJCNLP | 2 |
| 2009 | Word Lattices for Multi-Source Translation
Josh Schroeder, Trevor Cohn, Philipp Koehn |
EACL | 2 |
| 2009 | A Bayesian Model of Syntax-Directed Tree to String Grammar Induction
Trevor Cohn, Phil Blunsom |
EMNLP | 1 |
| 2009 | Inducing Compact but Accurate Tree-Substitution Grammars
Trevor Cohn, Sharon Goldwater, Phil Blunsom |
HLT-NAACL | 1 |
| 2009 | Sentence Compression as Tree TransductionabstractThis paper presents a tree-to-tree transduction method for sentence compression. Our model is based on synchronous tree substitution grammar, a formalism that allows local distortion of the tree topology and can thus naturally capture structural mismatches. We describe an algorithm for decoding in this framework and show how the model can be trained discriminatively within a large margin framework. Experimental results on sentence compression bring significant improvements over a state-of-the-art model. Trevor Cohn, Mirella Lapata |
J. Artif. Intell. Res. | 1 |
| 2008 | A Discriminative Latent Variable Model for Statistical Machine Translation
Phil Blunsom, Trevor Cohn, Miles Osborne |
ACL | 2 |
| 2008 | ParaMetric: An Automatic Evaluation Metric for Paraphrasing
Chris Callison-Burch, Trevor Cohn, Mirella Lapata |
COLING | 2 |
| 2008 | Sentence Compression Beyond Word Deletion
Trevor Cohn, Mirella Lapata |
COLING | 1 |
| 2008 | Bayesian Synchronous Grammar InductionabstractWe present a novel method for inducing synchronous context free grammars (SCFGs) from a corpus of parallel string pairs. SCFGs can model equivalence between strings in terms of substitutions, insertions and deletions, and the reordering of sub-strings. We develop a non-parametric Bayesian model and apply it to a machine translation task, using priors to replace the various heuristics commonly used in this field. Using a variational Bayes training procedure, we learn the latent structure of translation equivalence through the induction of synchronous grammar categories for phrasal translations, showing improvements in translation performance over previously proposed maximum likelihood models. Phil Blunsom, Trevor Cohn, Miles Osborne |
NIPS | 2 |
| 2008 | Constructing Corpora for the Development and Evaluation of Paraphrase SystemsabstractAutomatic paraphrasing is an important component in many natural language processing tasks. In this article we present a new parallel corpus with paraphrase annotations. We adopt a definition of paraphrase based on word alignments and show that it yields high inter-annotator agreement. As Kappa is suited to nominal data, we employ an alternative agreement statistic which is appropriate for structured alignment tasks. We discuss how the corpus can be usefully employed in evaluating paraphrase systems automatically (e.g., by measuring precision, recall, and F1) and also in developing linguistically rich paraphrase models based on syntactic structure. Trevor Cohn, Chris Callison-Burch, Mirella Lapata |
Comput. Linguistics | 1 |
| 2007 | Machine Translation by Triangulation: Making Effective Use of Multi-Parallel Corpora
Trevor Cohn, Mirella Lapata |
ACL | 1 |
| 2007 | Large Margin Synchronous Generation and its Application to Sentence Compression
Trevor Cohn, Mirella Lapata |
EMNLP-CoNLL | 1 |
| 2006 | Discriminative Word Alignment with Conditional Random FieldsabstractIn this paper we present a novel approach for inducing word alignments from sentence aligned data. We use a Conditional Random Field (CRF), a discriminative model, which is estimated on a small supervised training set. The CRF is conditioned on both the source and target texts, and thus allows for the use of arbitrary and overlapping features over these data. Moreover, the CRF has efficient training and decoding processes which both find globally optimal solutions.We apply this alignment model to both French-English and Romanian-English language pairs. We show how a large number of highly predictive features can be easily incorporated into the CRF, and demonstrate that even with only a few hundred word-aligned training sentences, our model improves over the current state-of-the-art with alignment error rates of 5.29 and 25.8 for the two tasks respectively. Phil Blunsom, Trevor Cohn |
ACL | 2 |
| 2006 | Efficient Inference in Large Conditional Random Fields
Trevor Cohn |
ECML | 1 |
| 2005 | Scaling Conditional Random Fields Using Error-Correcting CodesabstractConditional Random Fields (CRFs) have been applied with considerable success to a number of natural language processing tasks. However, these tasks have mostly involved very small label sets. When deployed on tasks with larger label sets, the requirements for computational resources mean that training becomes intractable.This paper describes a method for training CRFs on such tasks, using error correcting output codes (ECOC). A number of CRFs are independently trained on the separate binary labelling tasks of distinguishing between a subset of the labels and its complement. During decoding, these models are combined to produce a predicted label sequence which is resilient to errors by individual models.Error-correcting CRF training is much less resource intensive and has a much faster training time than a standardly formulated CRF, while decoding performance remains quite comparable. This allows us to scale CRFs to previously impossible tasks, as demonstrated by our experiments with large label sets. Trevor Cohn, Andrew Smith 0002, Miles Osborne |
ACL | 1 |
| 2005 | Logarithmic Opinion Pools for Conditional Random FieldsabstractRecent work on Conditional Random Fields (CRFs) has demonstrated the need for regularisation to counter the tendency of these models to overfit. The standard approach to regularising CRFs involves a prior distribution over the model parameters, typically requiring search over a hyperparameter space. In this paper we address the overfitting problem from a different perspective, by factoring the CRF distribution into a weighted product of individual "expert" CRF distributions. We call this model a logarithmic opinion pool (LOP) of CRFs (LOP-CRFs). We apply the LOP-CRF to two sequencing tasks. Our results show that unregularised expert CRFs with an unregularised CRF under a LOP can outperform the unregularised CRF, and attain a performance level close to the regularised CRF. LOP-CRFs therefore provide a viable alternative to CRF regularisation without the need for hyperparameter search. Andrew Smith 0002, Trevor Cohn, Miles Osborne |
ACL | 2 |
| 2005 | Semantic Role Labelling with Tree Conditional Random Fields
Trevor Cohn, Phil Blunsom |
CoNLL | 1 |