EDBT 2026 Demo / reviewers in the wild / expert
Phil Blunsom
dblp:96/4705
· DBLP profile ↗
89ranked-venue papers
11as first author
19since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 85 · 10 first-author · 19 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rope to Nope and Back Again: A New Hybrid Attention StrategyabstractLong-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu et al., 2024c; Peng et al., 2023). By adjusting RoPE parameters and incorporating training data with extended contexts, we can train performant models with considerably longer input sequences. However, existing RoPE-based methods exhibit performance limitations when applied to extended context lengths. This paper presents a comprehensive analysis of various attention mechanisms, including RoPE, No Positional Embedding (NoPE), and Query-Key Normalization (QK-Norm), identifying their strengths and shortcomings in long-context modeling. Our investigation identifies distinctive attention patterns in these methods and highlights their impact on long-context performance, providing valuable insights for architectural design. on long context performance, providing valuable insights for architectural design. Building on these findings, we propose a novel architecture featuring a hybrid attention mechanism that integrates global and local attention spans. This design not only surpasses conventional RoPE-based transformer models with full attention in both long and short context tasks but also delivers substantial efficiency gains during training and inference. Bharat Venkitesh, Dwaraknath Gnaneshwar, David Cairuz, Phil Blunsom, Acyr Locatelli |
NeurIPS | 6 |
| 2024 | Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelabstractAhmet Üstün, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, Sara Hooker. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ahmet Üstün, Viraat Aryabumi, Wei-Yin Ko, Daniel D'souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, Sara Hooker |
ACL (1) | 12 |
| 2024 | Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete FunctionsabstractIn order to understand the in-context learning phenomenon, recent works have adopted a stylized experimental framework and demonstrated that Transformers can match the performance of gradient-based learning algorithms for various classes of real-valued functions. However, the limitations of Transformers in implementing learning algorithms, and their ability to learn other forms of algorithms are not well understood. Additionally, the degree to which these capabilities are confined to attention-based models is unclear. Furthermore, it remains to be seen whether the insights derived from these stylized settings can be extrapolated to pretrained Large Language Models (LLMs). In this work, we take a step towards answering these questions by demonstrating the following: (a) On a test-bed with a variety of Boolean function classes, we find that Transformers can nearly match the optimal learning algorithm for 'simpler' tasks, while their performance deteriorates on more 'complex' tasks. Additionally, we find that certain attention-free models perform (almost) identically to Transformers on a range of tasks. (b) When provided a *teaching sequence*, i.e. a set of examples that uniquely identifies a function in a class, we show that Transformers learn more sample-efficiently. Interestingly, our results show that Transformers can learn to implement *two distinct* algorithms to solve a *single* task, and can adaptively select the more sample-efficient algorithm depending on the sequence of in-context examples. (c) Lastly, we show that extant LLMs, e.g. LLaMA-2, GPT-4, can compete with nearest-neighbor baselines on prediction tasks that are guaranteed to not be in their training set. Satwik Bhattamishra, Arkil Patel, Phil Blunsom, Varun Kanade |
ICLR | 3 |
| 2024 | Human Feedback is not Gold StandardabstractHuman feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single `preference' score captures. We hypothesise that preference scores are subjective and open to undesirable biases. We critically analyse the use of human feedback for both training and evaluation, to verify whether it fully captures a range of crucial error criteria. We find that while preference scores have fairly good coverage, they under-represent important aspects like factuality. We further hypothesise that both preference scores and error annotation may be affected by confounders, and leverage instruction-tuned models to generate outputs that vary along two possible confounding dimensions: assertiveness and complexity. We find that the assertiveness of an output skews the perceived rate of factuality errors, indicating that human annotations are not a fully reliable evaluation metric or training objective. Finally, we offer preliminary evidence that using human feedback as a training objective disproportionately increases the assertiveness of model outputs. We encourage future work to carefully consider whether preference scores are well aligned with the desired objective. Tom Hosking, Phil Blunsom, Max Bartolo |
ICLR | 2 |
| 2024 | Separations in the Representational Capabilities of Transformers and Recurrent ArchitecturesabstractTransformer architectures have been widely adopted in foundation models. Due to their high inference costs, there is renewed interest in exploring the potential of efficient recurrent architectures (RNNs). In this paper, we analyze the differences in the representational capabilities of Transformers and RNNs across several tasks of practical relevance, including index lookup, nearest neighbor, recognizing bounded Dyck languages, and string equality. For the tasks considered, our results show separations based on the size of the model required for different architectures. For example, we show that a one-layer Transformer of logarithmic width can perform index lookup, whereas an RNN requires a hidden state of linear size. Conversely, while constant-size RNNs can recognize bounded Dyck languages, we show that one-layer Transformers require a linear size for this task. Furthermore, we show that two-layer Transformers of logarithmic size can perform decision tasks such as string equality or disjointness, whereas both one-layer Transformers and recurrent models require linear size for these tasks. We also show that a log-size two-layer Transformer can implement the nearest neighbor algorithm in its forward pass; on the other hand recurrent models require linear size. Our constructions are based on the existence of $N$ nearly orthogonal vectors in $O(\log N)$ dimensional space and our lower bounds are based on reductions from communication complexity problems. We supplement our theoretical results with experiments that highlight the differences in the performance of these architectures on practical-size sequences. Satwik Bhattamishra, Michael Hahn 0001, Phil Blunsom, Varun Kanade |
NeurIPS | 3 |
| 2024 | BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsabstractMixture of Experts (MoE) framework has become a popular architecture for large language models due to its superior performance compared to dense models. However, training MoEs from scratch in a large-scale regime is prohibitively expensive. Previous work addresses this challenge by independently training multiple dense expert models and using them to initialize an MoE. In particular, state-of-the-art approaches initialize MoE layers using experts' feed-forward parameters while merging all other parameters, limiting the advantages of the specialized dense models when upcycling them as MoEs. We propose BAM (Branch-Attend-Mix), a simple yet effective improvement to MoE training. BAM makes full use of specialized dense models by not only using their feed-forward network (FFN) to initialize the MoE layers but also leveraging experts' attention weights fully by leveraging them as mixture-of-attention (MoA) layers. We explore two methods for upcycling MoA layers: 1) initializing separate attention experts from dense models including key, value, and query matrices; and 2) initializing only Q projections while sharing key-value pairs across all experts to facilitate efficient inference. Our experiments using seed models ranging from 590 million to 2 billion parameters show that our approach outperforms state-of-the-art approaches under the same data and compute budget in both perplexity and downstream tasks evaluations, confirming the effectiveness of BAM. Qizhen Zhang 0002, Nikolas Gritsch, Dwaraknath Gnaneshwar, Simon Guo 0003, David Cairuz, Bharat Venkitesh, Jakob N. Foerster, Phil Blunsom, Sebastian Ruder, Ahmet Üstün, Acyr Locatelli |
NeurIPS | 8 |
| 2023 | Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean FunctionsabstractDespite the widespread success of Transformers on NLP tasks, recent works have found that they struggle to model several formal languages when compared to recurrent models.This raises the question of why Transformers perform well in practice and whether they have any properties that enable them to generalize better than recurrent models.In this work, we conduct an extensive empirical study on Boolean functions to demonstrate the following: (i) Random Transformers are relatively more biased towards functions of low sensitivity.(ii) When trained on Boolean functions, both Transformers and LSTMs prioritize learning functions of low sensitivity, with Transformers ultimately converging to functions of lower sensitivity.(iii) On sparse Boolean functions which have low sensitivity, we find that Transformers generalize near perfectly even in the presence of noisy labels whereas LSTMs overfit and achieve poor generalization accuracy.Overall, our results provide strong quantifiable evidence that suggests differences in the inductive biases of Transformers and recurrent models which may help explain Transformer's effective generalization performance despite relatively limited expressiveness. Satwik Bhattamishra, Arkil Patel, Varun Kanade, Phil Blunsom |
ACL (1) | 4 |
| 2023 | On "Scientific Debt" in NLP: A Case for More Rigour in Language Model Pre-Training ResearchabstractMade Nindyatama Nityasya, Haryo Wibowo, Alham Fikri Aji, Genta Winata, Radityo Eko Prasojo, Phil Blunsom, Adhiguna Kuncoro. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Made Nindyatama Nityasya, Haryo Akbarianto Wibowo, Alham Fikri Aji, Genta Indra Winata, Radityo Eko Prasojo, Phil Blunsom, Adhiguna Kuncoro |
ACL (1) | 6 |
| 2023 | Intriguing Properties of Quantization at ScaleabstractEmergent properties have been widely adopted as a term to describe behavior not present in smaller models but observed in larger models (Wei et al., 2022a). Recent work suggests that the trade-off incurred by quantization is also an emergent property, with sharp drops in performance in models over 6B parameters. In this work, we ask _are quantization cliffs in performance solely a factor of scale?_ Against a backdrop of increased research focus on why certain emergent properties surface at scale, this work provides a useful counter-example. We posit that it is possible to optimize for a quantization friendly training recipe that suppresses large activation magnitude outliers. Here, we find that outlier dimensions are not an inherent product of scale, but rather sensitive to the optimization conditions present during pre-training. This both opens up directions for more efficient quantization, and poses the question of whether other emergent properties are inherent or can be altered and conditioned by optimization and architecture design choices. We successfully quantize models ranging in size from 410M to 52B with minimal degradation in performance. Arash Ahmadian, Saurabh Dash, Hongyu Chen 0008, Bharat Venkitesh, Stephen Zhen Gou, Phil Blunsom, Ahmet Üstün, Sara Hooker |
NeurIPS | 6 |
| 2022 | A Systematic Investigation of Commonsense Knowledge in Large Language ModelsabstractXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, Aida Nematzadeh. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xiang Li 0069, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume, Phil Blunsom, Aida Nematzadeh |
EMNLP | 5 |
| 2022 | StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsabstractKnowledge and language understanding of models evaluated through question answering (QA) has been usually studied on static snapshots of knowledge, like Wikipedia. However, our world is dynamic, evolves over time, and our models’ knowledge becomes outdated. To study how semi-parametric QA models and their underlying parametric language models (LMs) adapt to evolving knowledge, we construct a new large-scale dataset, StreamingQA, with human written and generated questions asked on a given date, to be answered from 14 years of time-stamped news articles. We evaluate our models quarterly as they read new articles not seen in pre-training. We show that parametric models can be updated without full retraining, while avoiding catastrophic forgetting. For semi-parametric models, adding new articles into the search space allows for rapid adaptation, however, models with an outdated underlying LM under-perform those with a retrained LM. For questions about higher-frequency named entities, parametric updates are particularly beneficial. In our dynamic world, the StreamingQA dataset enables a more realistic evaluation of QA models, and our experiments highlight several promising directions for future research. Adam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi, Eren Sezener, Devang Agrawal, Cyprien de Masson d'Autume, Tim Scholtes, Manzil Zaheer, Susannah Young, Ellen Gilsenan-McMahon, Sophia Austin, Phil Blunsom, Angeliki Lazaridou |
ICML | 13 |
| 2022 | Mutual Information Constraints for Monte-Carlo Objectives to Prevent Posterior Collapse Especially in Language ModellingabstractPosterior collapse is a common failure mode of density models trained as variational autoencoders, wherein they model the data without relying on their latent variables, rendering these variables useless. We focus on two factors contributing to posterior collapse, that have been studied separately in the literature. First, the underspecification of the model, which in an extreme but common case allows posterior collapse to be the theoretical optimium. Second, the looseness of the variational lower bound and the related underestimation of the utility of the latents. We weave these two strands of research together, specifically the tighter bounds of multi-sample Monte-Carlo objectives and constraints on the mutual information between the observable and the latent variables. The main obstacle is that the usual method of estimating the mutual information as the average Kullback-Leibler divergence between the easily available variational posterior q(z|x) and the prior does not work with Monte-Carlo objectives because their q(z|x) is not a direct approximation to the model's true posterior p(z|x). Hence, we construct estimators of the Kullback-Leibler divergence of the true posterior from the prior by recycling samples used in the objective, with which we train models of continuous and discrete latents at much improved rate-distortion and no posterior collapse. While alleviated, the tradeoff between modelling the data and using the latents still remains, and we urge for evaluating inference methods across a range of mutual information values. Gábor Melis, András György 0001, Phil Blunsom |
J. Mach. Learn. Res. | 3 |
| 2022 | Relational Memory-Augmented Language ModelsabstractAbstract We present a memory-augmented approach to condition an autoregressive language model on a knowledge graph. We represent the graph as a collection of relation triples and retrieve relevant relations for a given context to improve text generation. Experiments on WikiText-103, WMT19, and enwik8 English datasets demonstrate that our approach produces a better language model in terms of perplexity and bits per character. We also show that relational memory improves coherence, is complementary to token-based memory, and enables causal interventions. Our model provides a simple yet effective way to combine an autoregressive language model and a knowledge graph for more coherent and logical generation. Qi Liu 0049, Dani Yogatama, Phil Blunsom |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Transformer Grammars: Augmenting Transformer Language Models with Syntactic Inductive Biases at ScaleabstractAbstract We introduce Transformer Grammars (TGs), a novel class of Transformer language models that combine (i) the expressive power, scalability, and strong performance of Transformers and (ii) recursive syntactic compositions, which here are implemented through a special attention mask and deterministic transformation of the linearized tree. We find that TGs outperform various strong baselines on sentence-level language modeling perplexity, as well as on multiple syntax-sensitive language modeling evaluation metrics. Additionally, we find that the recursive syntactic composition bottleneck which represents each sentence as a single vector harms perplexity on document-level language modeling, providing evidence that a different kind of memory mechanism—one that is independent of composed syntactic representations—plays an important role in current successful models of long text. Laurent Sartran, Samuel Barrett, Adhiguna Kuncoro, Milos Stanojevic, Phil Blunsom, Chris Dyer |
Trans. Assoc. Comput. Linguistics | 5 |
| 2021 | A Generative Framework for Simultaneous Machine TranslationabstractWe propose a generative framework for simultaneous machine translation.Conventional approaches use a fixed number of source words to translate or learn dynamic policies for the number of source words by reinforcement learning.Here we formulate simultaneous translation as a structural sequence-tosequence learning problem.A latent variable is introduced to model read or translate actions at every time step, which is then integrated out to consider all the possible translation policies.A re-parameterised Poisson prior is used to regularise the policies which allows the model to explicitly balance translation quality and latency.The experiments demonstrate the effectiveness and robustness of the generative framework, which achieves the best BLEU scores given different average translation latencies on benchmark datasets. Yishu Miao, Phil Blunsom, Lucia Specia |
EMNLP (1) | 2 |
| 2021 | Counterfactual Data Augmentation for Neural Machine TranslationabstractWe propose a data augmentation method for neural machine translation.It works by interpreting language models and phrasal alignment causally.Specifically, it creates augmented parallel translation corpora by generating (path-specific) counterfactual aligned phrases.We generate these by sampling new source phrases from a masked language model, then sampling an aligned counterfactual target phrase by noting that a translation language model can be interpreted as a Gumbel-Max Structural Causal Model (Oberst and Sontag, 2019).Compared to previous work, our method takes both context and alignment into account to maintain the symmetry between source and target sequences.Experiments on IWSLT'15 English → Vietnamese, WMT'17 English → German, WMT'18 English → Turkish, and WMT'19 robust English → French show that the method can improve the performance of translation, backtranslation and translation robustness. Qi Liu 0049, Matt J. Kusner, Phil Blunsom |
NAACL-HLT | 3 |
| 2021 | Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsabstractOur world is open-ended, non-stationary, and constantly evolving; thus what we talk about and how we talk about it change over time. This inherent dynamic nature of language contrasts with the current static language modelling paradigm, which trains and evaluates models on utterances from overlapping time periods. Despite impressive recent progress, we demonstrate that Transformer-XL language models perform worse in the realistic setup of predicting future utterances from beyond their training period, and that model performance becomes increasingly worse with time. We find that, while increasing model size alone—a key driver behind recent progress—does not solve this problem, having models that continually update their knowledge with new information can indeed mitigate this performance degradation over time. Hence, given the compilation of ever-larger language modelling datasets, combined with the growing list of language-model-based NLP applications that require up-to-date factual knowledge about the world, we argue that now is the right time to rethink the static way in which we currently train and evaluate our language models, and develop adaptive language models that can remain up-to-date with respect to our ever-changing and non-stationary world. We publicly release our dynamic, streaming language modelling benchmarks for WMT and arXiv to facilitate language model evaluation that takes temporal dynamics into account. Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d'Autume, Tomás Kociský, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, Phil Blunsom |
NeurIPS | 14 |
| 2021 | Pretraining the Noisy Channel Model for Task-Oriented DialogueabstractAbstract Direct decoding for task-oriented dialogue is known to suffer from the explaining-away effect, manifested in models that prefer short and generic responses. Here we argue for the use of Bayes’ theorem to factorize the dialogue task into two models, the distribution of the context given the response, and the prior for the response itself. This approach, an instantiation of the noisy channel model, both mitigates the explaining-away effect and allows the principled incorporation of large pretrained models for the response prior. We present extensive experiments showing that a noisy channel model decodes better responses compared to direct decoding and that a two-stage pretraining strategy, employing both open-domain and task-oriented dialogue data, improves over randomly initialized models. Qi Liu 0049, Lei Yu 0008, Laura Rimell, Phil Blunsom |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Learning With Stochastic Guidance for Robot NavigationabstractDue to the sparse rewards and high degree of environmental variation, reinforcement learning approaches, such as deep deterministic policy gradient (DDPG), are plagued by issues of high variance when applied in complex real-world environments. We present a new framework for overcoming these issues by incorporating a stochastic switch, allowing an agent to choose between high- and low-variance policies. The stochastic switch can be jointly trained with the original DDPG in the same framework. In this article, we demonstrate the power of the framework in a navigation task, where the robot can dynamically choose to learn through exploration or to use the output of a heuristic controller as guidance. Instead of starting from completely random actions, the navigation capability of a robot can be quickly bootstrapped by several simple independent controllers. The experimental results show that with the aid of stochastic guidance, we are able to effectively and efficiently train DDPG navigation policies and achieve significantly better performance than state-of-the-art baseline models. Linhai Xie, Yishu Miao, Sen Wang 0002, Phil Blunsom, Zhihua Wang 0005, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Make Up Your Mind! Adversarial Generation of Inconsistent Natural Language ExplanationsabstractTo increase trust in artificial intelligence systems, a promising research direction consists of designing neural models capable of generating natural language explanations for their predictions.In this work, we show that such models are nonetheless prone to generating mutually inconsistent explanations, such as "Because there is a dog in the image."and "Because there is no dog in the [same] image.",exposing flaws in either the decision-making process of the model or in the generation of the explanations.We introduce a simple yet effective adversarial framework for sanity checking models against the generation of inconsistent natural language explanations.Moreover, as part of the framework, we address the problem of adversarial attacks with full target sequences, a scenario that was not previously addressed in sequence-to-sequence attacks.Finally, we apply our framework on a state-of-the-art neural natural language inference model that provides natural language explanations for its predictions.Our framework shows that this model is capable of generating a significant number of inconsistent explanations.PREMISE: A guy in a red jacket is snowboarding in midair. Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, Phil Blunsom |
ACL | 5 |
| 2020 | Learning to Segment Actions from Observation and NarrationabstractWe apply a generative segmental model of task structure, guided by narration, to action segmentation in video.We focus on unsupervised and weakly-supervised settings where no action labels are known during training.Despite its simplicity, our model performs competitively with previous work on a dataset of naturalistic instructional videos.Our model allows us to vary the sources of supervision used in training, and we find that both task structure and narrative language provide large benefits in segmentation quality. Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, Aida Nematzadeh |
ACL | 3 |
| 2020 | Simulating Early Word Learning in Situated Connectionist Agents
Felix Hill, Stephen Clark, Phil Blunsom, Karl Moritz Hermann |
CogSci | 3 |
| 2020 | Visual Grounding in Video for Unsupervised Word TranslationabstractThere are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to improve unsupervised word mapping between languages. The key idea is to establish a common visual representation between two languages by learning embeddings from unpaired instructional videos narrated in the native language. Given this shared embedding we demonstrate that (i) we can map words between the languages, particularly the 'visual' words; (ii) that the shared embedding provides a good initialization for existing unsupervised text-based word translation techniques, forming the basis for our proposed hybrid visual-text mapping algorithm, MUVE; and (iii) our approach achieves superior performance by addressing the shortcomings of text-based methods -- it is more robust, handles datasets with less commonality, and is applicable to low-resource languages. We apply these methods to translate words from English to French, Korean, and Japanese -- all without any parallel corpora and simply by watching many videos of people speaking while doing things. Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira 0001, Phil Blunsom, Andrew Zisserman |
CVPR | 7 |
| 2020 | Mogrifier LSTM
Gábor Melis, Tomás Kociský, Phil Blunsom |
ICLR | 3 |
| 2020 | Syntactic Structure Distillation Pretraining for Bidirectional EncodersabstractTextual representation learners trained on large amounts of data have achieved notable success on downstream tasks; intriguingly, they have also performed well on challenging tests of syntactic competence. Hence, it remains an open question whether scalable learners like BERT can become fully proficient in the syntax of natural language by virtue of data scale alone, or whether they still benefit from more explicit syntactic biases. To answer this question, we introduce a knowledge distillation strategy for injecting syntactic biases into BERT pretraining, by distilling the syntactically informative predictions of a hierarchical—albeit harder to scale—syntactic language model. Since BERT models masked words in bidirectional context, we propose to distill the approximate marginal distribution over words in context from the syntactic LM. Our approach reduces relative error by 2–21% on a diverse set of structured prediction tasks, although we obtain mixed results on the GLUE benchmark. Our findings demonstrate the benefits of syntactic biases, even for representation learners that exploit large amounts of data, and contribute to a better understanding of where syntactic biases are helpful in benchmarks of natural language understanding. Adhiguna Kuncoro, Lingpeng Kong, Daniel Fried, Dani Yogatama, Laura Rimell, Chris Dyer, Phil Blunsom |
Trans. Assoc. Comput. Linguistics | 7 |
| 2020 | Better Document-Level Machine Translation with Bayes' RuleabstractWe show that Bayes’ rule provides an effective mechanism for creating document translation models that can be learned from only parallel sentences and monolingual documents a compelling benefit because parallel documents are not always available. In our formulation, the posterior probability of a candidate translation is the product of the unconditional (prior) probability of the candidate output document and the “reverse translation probability” of translating the candidate output back into the source language. Our proposed model uses a powerful autoregressive language model as the prior on target language documents, but it assumes that each sentence is translated independently from the target to the source language. Crucially, at test time, when a source document is observed, the document language model prior induces dependencies between the translations of the source sentences in the posterior. The model’s independence assumption not only enables efficient use of available data, but it additionally admits a practical left-to-right beam-search algorithm for carrying out inference. Experiments show that our model benefits from using cross-sentence context in the language model, and it outperforms existing document translation approaches. Lei Yu 0008, Laurent Sartran, Wojciech Stokowiec, Wang Ling, Lingpeng Kong, Phil Blunsom, Chris Dyer |
Trans. Assoc. Comput. Linguistics | 6 |
| 2019 | MotionTransformer: Transferring Neural Inertial Tracking between DomainsabstractInertial information processing plays a pivotal role in egomotion awareness for mobile agents, as inertial measurements are entirely egocentric and not environment dependent. However, they are affected greatly by changes in sensor placement/orientation or motion dynamics, and it is infeasible to collect labelled data from every domain. To overcome the challenges of domain adaptation on long sensory sequences, we propose MotionTransformer - a novel framework that extracts domain-invariant features of raw sequences from arbitrary domains, and transforms to new domains without any paired data. Through the experiments, we demonstrate that it is able to efficiently and effectively convert the raw sequence from a new unlabelled target domain into an accurate inertial trajectory, benefiting from the motion knowledge transferred from the labelled source domain. We also conduct real-world experiments to show our framework can reconstruct physically meaningful trajectories from raw IMU measurements obtained with a standard mobile phone in various attachments. Changhao Chen, Yishu Miao, Xiaoxuan Lu 0001, Linhai Xie, Phil Blunsom, Andrew Markham, Agathoniki Trigoni |
AAAI | 5 |
| 2019 | Learning to Discover, Ground and Use Words with Segmental Neural Language ModelsabstractWe propose a segmental neural language model that combines the generalization power of neural networks with the ability to discover word-like units that are latent in unsegmented character sequences.In contrast to previous segmentation models that treat word segmentation as an isolated task, our model unifies word discovery, learning how words fit together to form sentences, and, by conditioning the model on visual context, how words' meanings ground in representations of nonlinguistic modalities.Experiments show that the unconditional model learns predictive distributions better than character LSTM models, discovers words competitively with nonparametric Bayesian word segmentation models, and that modeling language conditional on visual context improves performance on both. Kazuya Kawakami, Chris Dyer, Phil Blunsom |
ACL (1) | 3 |
| 2019 | Scalable Syntax-Aware Language Models Using Knowledge DistillationabstractPrior work has shown that, on small amounts of training data, syntactic neural language models learn structurally sensitive generalisations more successfully than sequential language models. However, their computational complexity renders scaling difficult, and it remains an open question whether structural biases are still necessary when sequential models have access to ever larger amounts of training data. To answer this question, we introduce an efficient knowledge distillation (KD) technique that transfers knowledge from a syntactic language model trained on a small corpus to an LSTM language model, hence enabling the LSTM to develop a more structurally sensitive representation of the larger training data it learns from. On targeted syntactic evaluations, we find that, while sequential LSTMs perform much better than previously reported, our proposed technique substantially improves on this baseline, yielding a new state of the art. Our findings and analysis affirm the importance of structural biases, even in models that learn from large amounts of data. Adhiguna Kuncoro, Chris Dyer, Laura Rimell, Stephen Clark, Phil Blunsom |
ACL (1) | 5 |
| 2019 | WikiCREM: A Large Unsupervised Corpus for Coreference ResolutionabstractVid Kocijan, Oana-Maria Camburu, Ana-Maria Cretu, Yordan Yordanov, Phil Blunsom, Thomas Lukasiewicz. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Vid Kocijan, Oana-Maria Camburu, Ana-Maria Cretu 0002, Yordan Yordanov, Phil Blunsom, Thomas Lukasiewicz |
EMNLP/IJCNLP (1) | 5 |
| 2018 | LSTMs Can Learn Syntax-Sensitive Dependencies Well, But Modeling Structure Makes Them BetterabstractAdhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, Phil Blunsom. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, Phil Blunsom |
ACL (1) | 6 |
| 2018 | On the State of the Art of Evaluation in Neural Language Models
Gábor Melis, Chris Dyer, Phil Blunsom |
ICLR (Poster) | 3 |
| 2018 | Memory Architectures in Recurrent Neural Network Language Models
Dani Yogatama, Yishu Miao, Gábor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, Phil Blunsom |
ICLR (Poster) | 7 |
| 2018 | Neural Syntactic Generative Models with Exact MarginalizationabstractWe present neural syntactic generative models with exact marginalization that support both dependency parsing and language modeling.Exact marginalization is made tractable through dynamic programming over shiftreduce parsing and minimal RNN-based feature sets.Our algorithms complement previous approaches by supporting batched training and enabling online computation of next word probabilities.For supervised dependency parsing, our model achieves a stateof-the-art result among generative approaches.We also report empirical results on unsupervised syntactic models and their role in language modeling.We find that our model formulation of latent dependencies with exact marginalization do not lead to better intrinsic language modeling performance than vanilla RNNs, and that parsing accuracy is not correlated with language modeling perplexity in stack-based models. Jan Buys, Phil Blunsom |
NAACL-HLT | 2 |
| 2018 | e-SNLI: Natural Language Inference with Natural Language ExplanationsabstractIn order for machine learning to garner widespread public adoption, models must be able to provide interpretable and robust explanations for their decisions, as well as learn from human-provided explanations at train time. In this work, we extend the Stanford Natural Language Inference dataset with an additional layer of human-annotated natural language explanations of the entailment relations. We further implement models that incorporate these explanations into their training process and output them at test time. We show how our corpus of explanations, which we call e-SNLI, can be used for various goals, such as obtaining full sentence justifications of a model’s decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets. Our dataset thus opens up a range of research directions for using natural language explanations, both for improving models and for asserting their trust Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, Phil Blunsom |
NeurIPS | 4 |
| 2018 | Neural Arithmetic Logic UnitsabstractNeural networks can learn to represent and manipulate numerical information, but they seldom generalize well outside of the range of numerical values encountered during training. To encourage more systematic numerical extrapolation, we propose an architecture that represents numerical quantities as linear activations which are manipulated using primitive arithmetic operators, controlled by learned gates. We call this module a neural arithmetic logic unit (NALU), by analogy to the arithmetic logic unit in traditional processors. Experiments show that NALU-enhanced neural networks can learn to track time, perform arithmetic over images of numbers, translate numerical language into real-valued scalars, execute computer code, and count objects in images. In contrast to conventional architectures, we obtain substantially better generalization both inside and outside of the range of numerical values encountered during training, often extrapolating orders of magnitude beyond trained numerical ranges. Andrew Trask, Felix Hill, Scott E. Reed, Jack W. Rae, Chris Dyer, Phil Blunsom |
NeurIPS | 6 |
| 2018 | The NarrativeQA Reading Comprehension ChallengeabstractReading comprehension (RC)—in contrast to information retrieval—requires integrating information and reasoning about events, entities, and their relations across a full document. Question answering is conventionally used to assess RC ability, in both artificial agents and children learning to read. However, existing RC datasets and tasks are dominated by questions that can be solved by selecting answers using superficial information (e.g., local context similarity or global term frequency); they thus fail to test for the essential integrative aspect of RC. To encourage progress on deeper comprehension of language, we present a new dataset and set of tasks in which the reader must answer questions about stories by reading entire books or movie scripts. These tasks are designed so that successfully answering their questions requires understanding the underlying narrative rather than relying on shallow pattern matching or salience. We show that although humans solve the tasks easily, standard RC models struggle on the tasks presented here. We provide an analysis of the dataset and the challenges it presents. Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, Edward Grefenstette |
Trans. Assoc. Comput. Linguistics | 3 |
| 2017 | Robust Incremental Neural Semantic Graph ParsingabstractParsing sentences to linguisticallyexpressive semantic representations is a key goal of Natural Language Processing.Yet statistical parsing has focussed almost exclusively on bilexical dependencies or domain-specific logical forms.We propose a neural encoder-decoder transition-based parser which is the first full-coverage semantic graph parser for Minimal Recursion Semantics (MRS).The model architecture uses stack-based embedding features, predicting graphs jointly with unlexicalized predicates and their token alignments.Our parser is more accurate than attention-based baselines on MRS, and on an additional Abstract Meaning Representation (AMR) benchmark, and GPU batch processing makes it an order of magnitude faster than a high-precision grammar-based parser.Further, the 86.69%Smatch score of our MRS parser is higher than the upper-bound on AMR parsing, making MRS an attractive choice as a semantic representation. Jan Buys, Phil Blunsom |
ACL (1) | 2 |
| 2017 | Learning to Create and Reuse Words in Open-Vocabulary Neural Language ModelingabstractFixed-vocabulary language models fail to account for one of the most characteristic statistical facts of natural language: the frequent creation and reuse of new word types.Although character-level language models offer a partial solution in that they can create word types not attested in the training corpus, they do not capture the "bursty" distribution of such words.In this paper, we augment a hierarchical LSTM language model that generates sequences of word tokens character by character with a caching mechanism that learns to reuse previously generated words.To validate our model we construct a new open-vocabulary language modeling corpus (the Multilingual Wikipedia Corpus; MWC) from comparable Wikipedia articles in 7 typologically diverse languages and demonstrate the effectiveness of our model across this range of languages. Kazuya Kawakami, Chris Dyer, Phil Blunsom |
ACL (1) | 3 |
| 2017 | Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word ProblemsabstractSolving algebraic word problems requires executing a series of arithmetic operations-a program-to obtain a final answer.However, since programs can be arbitrarily complicated, inducing them directly from question-answer pairs is a formidable challenge.To make this task more feasible, we solve these problems by generating answer rationales, sequences of natural language and human-readable mathematical expressions that derive the final answer through a series of small steps.Although rationales do not explicitly specify programs, they provide a scaffolding for their structure via intermediate milestones.To evaluate our approach, we have created a new 100,000-sample dataset of questions, answers and rationales.Experimental results show that indirect supervision of program learning via answer rationales is a promising strategy for inducing arithmetic programs. Wang Ling, Dani Yogatama, Chris Dyer, Phil Blunsom |
ACL (1) | 4 |
| 2017 | Reference-Aware Language ModelsabstractWe propose a general class of language models that treat reference as discrete stochastic latent variables.This decision allows for the creation of entity mentions by accessing external databases of referents (required by, e.g., dialogue generation) or past internal state (required to explicitly model coreferentiality).Beyond simple copying, our coreference model can additionally refer to a referent using varied mention forms (e.g., a reference to "Jane" can be realized as "she"), a characteristic feature of reference in natural languages.Experiments on three representative applications show our model variants outperform models based on deterministic attention and standard language modeling baselines. Phil Blunsom, Chris Dyer, Wang Ling |
EMNLP | 2 |
| 2017 | Learning to Compose Words into Sentences with Reinforcement Learning
Dani Yogatama, Phil Blunsom, Chris Dyer, Edward Grefenstette, Wang Ling |
ICLR (Poster) | 2 |
| 2017 | The Neural Noisy Channel
Lei Yu 0008, Phil Blunsom, Chris Dyer, Edward Grefenstette, Tomás Kociský |
ICLR (Poster) | 2 |
| 2017 | Discovering Discrete Latent Topics with Neural Variational InferenceabstractTopic models have been widely explored as probabilistic generative models of documents. Traditional inference methods have sought closed-form derivations for updating the models, however as the expressiveness of these models grows, so does the difficulty of performing fast and accurate inference over their parameters. This paper presents alternative neural approaches to topic modelling by providing parameterisable distributions over topics which permit training by backpropagation in the framework of neural variational inference. In addition, with the help of a stick-breaking construction, we propose a recurrent network that is able to discover a notionally unbounded number of topics, analogous to Bayesian non-parametric topic models. Experimental results on the MXM Song Lyrics, 20NewsGroups and Reuters News datasets demonstrate the effectiveness and efficiency of these neural topic models. Yishu Miao, Edward Grefenstette, Phil Blunsom |
ICML | 3 |
| 2017 | Latent Intention Dialogue ModelsabstractDeveloping a dialogue agent that is capable of making autonomous decisions and communicating by natural language is one of the long-term goals of machine learning research. The traditional approaches either rely on hand-crafting a small state-action set for applying reinforcement learning that is not scalable or constructing deterministic models for learning dialogue sentences that fail to capture the conversational stochasticity. In this paper, however, we propose a Latent Intention Dialogue Model that employs a discrete latent variable to learn underlying dialogue intentions in the framework of neural variational inference. Additionally, in a goal-oriented dialogue scenario, the latent intentions can be interpreted as actions guiding the generation of machine responses, which can be further refined autonomously by reinforcement learning. The experiments demonstrate the effectiveness of discrete latent variable models on learning goal-oriented dialogues, and the results outperform the published benchmarks on both corpus-based evaluation and human evaluation. Tsung-Hsien Wen, Yishu Miao, Phil Blunsom, Steve J. Young |
ICML | 3 |
| 2016 | Latent Predictor Networks for Code GenerationabstractWang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomáš Kočiský, Fumin Wang, Andrew Senior. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, Andrew W. Senior |
ACL (1) | 2 |
| 2016 | Semantic Parsing with Semi-Supervised Sequential AutoencodersabstractWe present a novel semi-supervised approach for sequence transduction and apply it to semantic parsing.The unsupervised component is based on a generative model in which latent sentences generate the unpaired logical forms.We apply this method to a number of semantic parsing tasks focusing on domains with limited access to labelled training data and extend those datasets with synthetically generated logical forms. Tomás Kociský, Gábor Melis, Edward Grefenstette, Chris Dyer, Wang Ling, Phil Blunsom, Karl Moritz Hermann |
EMNLP | 6 |
| 2016 | Language as a Latent Variable: Discrete Generative Models for Sentence CompressionabstractIn this work we explore deep generative models of text in which the latent representation of a document is itself drawn from a discrete language model distribution.We formulate a variational auto-encoder for inference in this model and apply it to the task of compressing sentences.In this application the generative model first draws a latent summary sentence from a background language model, and then subsequently draws the observed sentence conditioned on this latent summary.In our empirical evaluation we show that generative formulations of both abstractive and extractive compression yield state-of-the-art results when trained on a large amount of supervised data.Further, we explore semi-supervised compression scenarios where we show that it is possible to achieve performance competitive with previously proposed supervised models while training on a fraction of the supervised data. Yishu Miao, Phil Blunsom |
EMNLP | 2 |
| 2016 | Online Segment to Segment Neural TransductionabstractWe introduce an online neural sequence to sequence model that learns to alternate between encoding and decoding segments of the input as it is read.By independently tracking the encoding and decoding representations our algorithm permits exact polynomial marginalization of the latent segmentation during training, and during decoding beam search is employed to find the best alignment path together with the predicted output sequence.Our model tackles the bottleneck of vanilla encoder-decoders that have to read and memorize the entire input sequence in their fixedlength hidden states before producing any output.It is different from previous attentive models in that, instead of treating the attention weights as output of a deterministic function, our model assigns attention weights to a sequential latent variable which can be marginalized out and permits online generation.Experiments on abstractive sentence summarization and morphological inflection show significant performance gains over the baseline encoder-decoders. Lei Yu 0008, Jan Buys, Phil Blunsom |
EMNLP | 3 |
| 2016 | Neural Variational Inference for Text ProcessingabstractRecent advances in neural variational inference have spawned a renaissance in deep latent variable models. In this paper we introduce a generic variational inference framework for generative and conditional models of text. While traditional variational methods derive an analytic approximation for the intractable distributions over latent variables, here we construct an inference network conditioned on the discrete text input to provide the variational distribution. We validate this framework on two very different text modelling applications, generative document modelling and supervised question answering. Our neural variational document model combines a continuous stochastic document representation with a bag-of-words generative model and achieves the lowest reported perplexities on two standard test corpora. The neural answer selection model employs a stochastic representation layer within an attention mechanism to extract the semantics between a question and answer pair. On two question answering benchmarks this model exceeds all previous published benchmarks. Yishu Miao, Lei Yu 0008, Phil Blunsom |
ICML | 3 |
| 2015 | Detection of Steganographic Techniques on TwitterabstractWe propose a method to detect hidden data in English text. We target a system pre-viously thought secure, which hides mes-sages in tweets. The method brings ideas from image steganalysis into the linguis-tic domain, including the training of a feature-rich model for detection. To iden-tify Twitter users guilty of steganography, we aggregate evidence; a first, in any do-main. We test our system on a set of 1M steganographic tweets, and show it to be effective. 1 Alex Wilson 0001, Phil Blunsom, Andrew D. Ker |
EMNLP | 2 |
| 2015 | Pragmatic Neural Language Modelling in Machine TranslationabstractThis paper presents an in-depth investigation on integrating neural language models in translation systems. Scaling neural language models is a difficult task, but crucial for real-world applications. This paper evaluates the impact on end-to-end MT quality of both new and existing scaling techniques. We show when explicitly normalising neural models is necessary and what optimisation tricks one should use in such scenarios. We also focus on scalable training algorithms and investigate noise contrastive estimation and diagonal contexts as sources for further speed improvements. We explore the trade-offs between neural models and back-off n-gram models and find that neural models make strong candidates for natural language applications in memory constrained environments, yet still lag behind traditional models in raw translation quality. We conclude with a set of recommendations one should follow to build a scalable neural language model for MT. Paul Baltescu, Phil Blunsom |
HLT-NAACL | 2 |
| 2015 | Learning to Transduce with Unbounded MemoryabstractRecently, strong results have been demonstrated by Deep Recurrent Neural Networks on natural language transduction problems. In this paper we explore the representational power of these models using synthetic grammars designed to exhibit phenomena similar to those found in real transduction problems such as machine translation. These experiments lead us to propose new memory-based recurrent networks that implement continuously differentiable analogues of traditional data structures such as Stacks, Queues, and DeQues. We show that these architectures exhibit superior generalisation performance to Deep RNNs and are often able to learn the underlying generating algorithms in our transduction experiments. Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, Phil Blunsom |
NIPS | 4 |
| 2015 | Teaching Machines to Read and ComprehendabstractTeaching machines to read natural language documents remains an elusive challenge. Machine reading systems can be tested on their ability to answer questions posed on the contents of documents that they have seen, but until now large scale training and test datasets have been missing for this type of evaluation. In this work we define a new methodology that resolves this bottleneck and provides large scale supervised reading comprehension data. This allows us to develop a class of attention based deep neural networks that learn to read real documents and answer complex questions with minimal prior knowledge of language structure. Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, Phil Blunsom |
NIPS | 7 |
| 2015 | Non-Line-of-Sight Identification and Mitigation Using Received Signal StrengthabstractIndoor wireless systems often operate under non-line-of-sight (NLOS) conditions that can cause ranging errors for location-based applications. As such, these applications could benefit greatly from NLOS identification and mitigation techniques. These techniques have been primarily investigated for ultra-wide band (UWB) systems, but little attention has been paid to WiFi systems, which are far more prevalent in practice. In this study, we address the NLOS identification and mitigation problems using multiple received signal strength (RSS) measurements from WiFi signals. Key to our approach is exploiting several statistical features of the RSS time series, which are shown to be particularly effective. We develop and compare two algorithms based on machine learning and a third based on hypothesis testing to separate LOS/NLOS measurements. Extensive experiments in various indoor environments show that our techniques can distinguish between LOS/NLOS conditions with an accuracy of around 95%. Furthermore, the presented techniques improve distance estimation accuracy by 60% as compared to state-of-the-art NLOS mitigation techniques. Finally, improvements in distance estimation accuracy of 50% are achieved even without environment-specific training data, demonstrating the practicality of our approach to real world implementations. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, Phil Blunsom, Jeff Frolik |
IEEE Trans. Wirel. Commun. | 5 |
| 2014 | Multilingual Models for Compositional Distributed SemanticsabstractWe present a novel technique for learning semantic representations, which extends the distributional hypothesis to multilingual data and joint-space embeddings. Our models leverage parallel data and learn to strongly align the embeddings of semantically equivalent sentences, while maintaining sufficient distance between those of dissimilar sentences. The models do not rely on word alignments or any syntactic information and are successfully applied to a number of diverse languages. We extend our approach to learn semantic representations at the document level, too. We evaluate these models on two cross-lingual document classification tasks, outperforming the prior state of the art. Through qualitative analysis and the study of pivoting effects we demonstrate that our representations are semantically plausible and can capture semantic relationships across languages without parallel data. Karl Moritz Hermann, Phil Blunsom |
ACL (1) | 2 |
| 2014 | A Convolutional Neural Network for Modelling SentencesabstractThe ability to accurately represent sentences is central to language understanding. We describe a convolutional architecture dubbed the Dynamic Convolutional Neural Network (DCNN) that we adopt for the semantic modelling of sentences. The network uses Dynamic k-Max Pooling, a global pooling operation over linear sequences. The network handles input sentences of varying length and induces a feature graph over the sentence that is capable of explicitly capturing short and long-range relations. The network does not rely on a parse tree and is easily applicable to any language. We test the DCNN in four experiments: small scale binary and multi-class sentiment prediction, six-way question classification and Twitter sentiment prediction by distant supervision. The network achieves excellent performance in the first three tasks and a greater than 25% error reduction in the last task with respect to the strongest baseline. Nal Kalchbrenner, Edward Grefenstette, Phil Blunsom |
ACL (1) | 3 |
| 2014 | Modelling the Lexicon in Unsupervised Part of Speech InductionabstractAutomatically inducing the syntactic partof-speech categories for words in text is a fundamental task in Computational Linguistics.While the performance of unsupervised tagging models has been slowly improving, current state-of-the-art systems make the obviously incorrect assumption that all tokens of a given word type must share a single part-of-speech tag.This one-tag-per-type heuristic counters the tendency of Hidden Markov Model based taggers to over generate tags for a given word type.However, it is clearly incompatible with basic syntactic theory.In this paper we extend a state-ofthe-art Pitman-Yor Hidden Markov Model tagger with an explicit model of the lexicon.In doing so we are able to incorporate a soft bias towards inducing few tags per type.We develop a particle filter for drawing samples from the posterior of our model and present empirical results that show that our model is competitive with and faster than the state-of-the-art without making any unrealistic restrictions. Gregory Dubbin, Phil Blunsom |
EACL | 2 |
| 2014 | Dynamic Topic Adaptation for Phrase-based MTabstractTranslating text from diverse sources poses a challenge to current machine translation systems which are rarely adapted to structure beyond corpus level. We explore topic adaptation on a diverse data set and present a new bilingual vari-ant of Latent Dirichlet Allocation to com-pute topic-adapted, probabilistic phrase translation features. We dynamically in-fer document-specific translation proba-bilities for test sets of unknown origin, thereby capturing the effects of document context on phrase translations. We show gains of up to 1.26 BLEU over the base-line and 1.04 over a domain adaptation benchmark. We further provide an anal-ysis of the domain-specific data and show additive gains of our model in combination with other types of topic-adapted features. 1 Eva Hasler, Phil Blunsom, Philipp Koehn, Barry Haddow |
EACL | 2 |
| 2014 | Compositional Morphology for Word Representations and Language ModellingabstractThis paper presents a scalable method for integrating compositional morphological representations into a vector-based probabilistic language model. Our approach is evaluated in the context of log-bilinear language models, rendered suitably efficient for implementation inside a machine translation decoder by factoring the vocabulary. We perform both intrinsic and extrinsic evaluations, presenting results on a range of languages which demonstrate that our model learns morphological representations that both perform well on word similarity tasks and lead to substantial reductions in perplexity. When used for translation into morphologically rich languages with large vocabularies, our models obtain improvements of up to 1.2 BLEU points relative to a baseline system using back-off n-gram models. Jan A. Botha, Phil Blunsom |
ICML | 2 |
| 2013 | The Role of Syntax in Vector Space Models of Compositional Semantics
Karl Moritz Hermann, Phil Blunsom |
ACL (1) | 2 |
| 2013 | Collapsed Variational Bayesian Inference for Hidden Markov ModelsabstractApproximate inference for Bayesian models is dominated by two approaches, variational Bayesian inference and Markov Chain Monte Carlo. Both approaches have their own advantages and disadvantages, and they can complement each other. Recently researchers have proposed collapsed variational Bayesian inference to combine the advantages of both. Such inference methods have been successful in several models whose hidden variables are conditionally independent given the parameters. In this paper we propose two collapsed variational Bayesian inference algorithms for hidden Markov models, a popular framework for representing time series data. We validate our algorithms on the natural language processing task of unsupervised part-of-speech induction, showing that they are both more computationally efficient than sampling, and more accurate than standard variational Bayesian inference for HMMs. Phil Blunsom |
AISTATS | 2 |
| 2013 | Collapsed Variational Bayesian Inference for PCFGs
Phil Blunsom |
CoNLL | 2 |
| 2013 | Adaptor Grammars for Learning Non-Concatenative MorphologyabstractThis paper contributes an approach for expressing non-concatenative morphological phenomena, such as stem derivation in Semitic languages, in terms of a mildly context-sensitive grammar formalism.This offers a convenient level of modelling abstraction while remaining computationally tractable.The nonparametric Bayesian framework of adaptor grammars is extended to this richer grammar formalism to propose a probabilistic model that can learn word segmentation and morpheme lexicons, including ones with discontiguous strings as elements, from unannotated data.Our experiments on Hebrew and three variants of Arabic data find that the additional expressiveness to capture roots and templates as atomic units improves the quality of concatenative segmentation and stem identification.We obtain 74% accuracy in identifying triliteral Hebrew roots, while performing morphological segmentation with an F1-score of 78.1. Jan A. Botha, Phil Blunsom |
EMNLP | 2 |
| 2013 | Recurrent Continuous Translation ModelsabstractWe introduce a class of probabilistic continuous translation models called Recurrent Continuous Translation Models that are purely based on continuous representations for words, phrases and sentences and do not rely on alignments or phrasal translation units.The models have a generation and a conditioning aspect.The generation of the translation is modelled with a target Recurrent Language Model, whereas the conditioning on the source sentence is modelled with a Convolutional Sentence Model.Through various experiments, we show first that our models obtain a perplexity with respect to gold translations that is > 43% lower than that of stateof-the-art alignment-based translation models.Secondly, we show that they are remarkably sensitive to the word order, syntax, and meaning of the source sentence despite lacking alignments.Finally we show that they match a state-of-the-art system when rescoring n-best lists of translations. Nal Kalchbrenner, Phil Blunsom |
EMNLP | 2 |
| 2013 | On Assessing the Accuracy of Positioning Systems in Indoor Environments
Hongkai Wen 0001, Zhuoling Xiao, Agathoniki Trigoni, Phil Blunsom |
EWSN | 4 |
| 2013 | A Systematic Bayesian Treatment of the IBM Alignment Models
Yarin Gal, Phil Blunsom |
HLT-NAACL | 2 |
| 2013 | Identification and mitigation of non-line-of-sight conditions using received signal strengthabstractVarious applications, such as localisation of persons and objects could benefit greatly from non-line-of-sight (NLOS) identification and mitigation techniques. However, such techniques have been primarily investigated for ultra-wide band (UWB) signals, leaving the area of WiFi signals untouched. In this study, we propose two accurate approaches using only received signal strength (RSS) measurements from WiFi signals to identify NLOS conditions and mitigate the effects. We first explore several features from the RSS which are later demonstrated as very effective in identifying and mitigating NLOS conditions. After that, we develop and compare two major optimization problems based on a machine learning technique and hypothesis testing according to different user requirements and information available. Extensive experiments in various indoor environments have shown that our techniques can not only accurately distinguish between LOS/NLOS conditions, but also mitigate the impact of NLOS conditions as well. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, Phil Blunsom, Jeff Frolik |
WiMob | 5 |
| 2012 | Bayesian Language Modelling of German Compounds
Jan A. Botha, Chris Dyer, Phil Blunsom |
COLING | 3 |
| 2012 | A Bayesian Model for Learning SCFGs with Discontiguous Rules
Abby D. Levenberg, Chris Dyer, Phil Blunsom |
EMNLP-CoNLL | 3 |
| 2012 | Unsupervised Bayesian Part of Speech Inference with Particle Gibbs
Gregory Dubbin, Phil Blunsom |
ECML/PKDD (1) | 2 |
| 2011 | A Hierarchical Pitman-Yor Process HMM for Unsupervised Part of Speech Induction
Phil Blunsom, Trevor Cohn |
ACL | 1 |
| 2010 | Unsupervised Induction of Tree Substitution Grammars for Dependency Parsing
Phil Blunsom, Trevor Cohn |
EMNLP | 1 |
| 2010 | Inducing Synchronous Grammars with Slice Sampling
Phil Blunsom, Trevor Cohn |
HLT-NAACL | 1 |
| 2010 | Inducing Tree-Substitution Grammars
Trevor Cohn, Phil Blunsom, Sharon Goldwater |
J. Mach. Learn. Res. | 2 |
| 2010 | Monte Carlo techniques for phrase-based translation
Abhishek Arun, Barry Haddow, Philipp Koehn, Adam Lopez, Chris Dyer, Phil Blunsom |
Mach. Transl. | 6 |
| 2010 | Metrics for MT evaluation: evaluating reordering
Alexandra Birch, Miles Osborne, Phil Blunsom |
Mach. Transl. | 3 |
| 2009 | A Gibbs Sampler for Phrasal Synchronous Grammar Induction
Phil Blunsom, Trevor Cohn, Chris Dyer, Miles Osborne |
ACL/IJCNLP | 1 |
| 2009 | Monte Carlo inference and maximization for phrase-based translation
Abhishek Arun, Chris Dyer, Barry Haddow, Phil Blunsom, Adam Lopez, Philipp Koehn |
CoNLL | 4 |
| 2009 | A Bayesian Model of Syntax-Directed Tree to String Grammar Induction
Trevor Cohn, Phil Blunsom |
EMNLP | 2 |
| 2009 | Inducing Compact but Accurate Tree-Substitution Grammars
Trevor Cohn, Sharon Goldwater, Phil Blunsom |
HLT-NAACL | 3 |
| 2009 | Learning Machine Translation - Cyril Goutte, Nicola Cancedda, Marc Dymetman, and George Foster (editors) The MIT Press, 2009, xii+316 pp; ISBN 978-0-262-07297-7
Phil Blunsom |
Comput. Linguistics | 1 |
| 2008 | A Discriminative Latent Variable Model for Statistical Machine Translation
Phil Blunsom, Trevor Cohn, Miles Osborne |
ACL | 1 |
| 2008 | Probabilistic Inference for Machine Translation
Phil Blunsom, Miles Osborne |
EMNLP | 1 |
| 2008 | Bayesian Synchronous Grammar InductionabstractWe present a novel method for inducing synchronous context free grammars (SCFGs) from a corpus of parallel string pairs. SCFGs can model equivalence between strings in terms of substitutions, insertions and deletions, and the reordering of sub-strings. We develop a non-parametric Bayesian model and apply it to a machine translation task, using priors to replace the various heuristics commonly used in this field. Using a variational Bayes training procedure, we learn the latent structure of translation equivalence through the induction of synchronous grammar categories for phrasal translations, showing improvements in translation performance over previously proposed maximum likelihood models. Phil Blunsom, Trevor Cohn, Miles Osborne |
NIPS | 1 |
| 2006 | Discriminative Word Alignment with Conditional Random FieldsabstractIn this paper we present a novel approach for inducing word alignments from sentence aligned data. We use a Conditional Random Field (CRF), a discriminative model, which is estimated on a small supervised training set. The CRF is conditioned on both the source and target texts, and thus allows for the use of arbitrary and overlapping features over these data. Moreover, the CRF has efficient training and decoding processes which both find globally optimal solutions.We apply this alignment model to both French-English and Romanian-English language pairs. We show how a large number of highly predictive features can be easily incorporated into the CRF, and demonstrate that even with only a few hundred word-aligned training sentences, our model improves over the current state-of-the-art with alignment error rates of 5.29 and 25.8 for the two tasks respectively. Phil Blunsom, Trevor Cohn |
ACL | 1 |
| 2006 | Multilingual Deep Lexical Acquisition for HPSGs via Supertagging
Phil Blunsom, Timothy Baldwin |
EMNLP | 1 |
| 2006 | Question classification with log-linear modelsabstractQuestion classification has become a crucial step in modern question answering systems. Previous work has demonstrated the effectiveness of statistical machine learning approaches to this problem. This paper presents a new approach to building a question classifier using log-linear models. Evidence from a rich and diverse set of syntactic and semantic features is evaluated, as well as approaches which exploit the hierarchical structure of the question classes. Phil Blunsom, Krystle Kocik, James R. Curran |
SIGIR | 1 |
| 2005 | Semantic Role Labelling with Tree Conditional Random Fields
Trevor Cohn, Phil Blunsom |
CoNLL | 2 |