EDBT 2026 Demo / reviewers in the wild / expert
James Henderson 0001
dblp:h/JamesHenderson · also James Brinton Henderson
· DBLP profile ↗
60ranked-venue papers
17as first author
14since 2021 · last 2025
0000-0003-3714-4799ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 17 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reduction of supervision for biomedical knowledge discoveryabstractBACKGROUND: Knowledge discovery in scientific literature is hindered by the increasing volume of publications and the scarcity of extensive annotated data. To tackle the challenge of information overload, it is essential to employ automated methods for knowledge extraction and processing. Finding the right balance between the level of supervision and the effectiveness of models poses a significant challenge. While supervised techniques generally result in better performance, they have the major drawback of demanding labeled data. This requirement is labor-intensive, time-consuming, and hinders scalability when exploring new domains. METHODS AND RESULTS: In this context, our study addresses the challenge of identifying semantic relationships between biomedical entities (e.g., diseases, proteins, medications) in unstructured text while minimizing dependency on supervision. We introduce a suite of unsupervised algorithms based on dependency trees and attention mechanisms and employ a range of pointwise binary classification methods. Transitioning from weakly supervised to fully unsupervised settings, we assess the methods' ability to learn from data with noisy labels. The evaluation on four biomedical benchmark datasets explores the effectiveness of the methods, demonstrating their potential to enable scalable knowledge discovery systems less reliant on annotated datasets. CONCLUSION: Our approach tackles a central issue in knowledge discovery: balancing performance with minimal supervision which is crucial to adapting models to varied and changing domains. This study also investigates the use of pointwise binary classification techniques within a weakly supervised framework for knowledge discovery. By gradually decreasing supervision, we assess the robustness of these techniques in handling noisy labels, revealing their capability to shift from weakly supervised to entirely unsupervised scenarios. Comprehensive benchmarking offers insights into the effectiveness of these techniques, examining how unsupervised methods can reliably capture complex relationships in biomedical texts. These results suggest an encouraging direction toward scalable, adaptable knowledge discovery systems, representing progress in creating data-efficient methodologies for extracting useful insights when annotated data is limited. Christos Theodoropoulos 0001, Andrei Catalin, James Henderson 0001, Marie-Francine Moens |
BMC Bioinform. | 3 |
| 2024 | TESS: Text-to-Text Self-Conditioned Simplex DiffusionabstractRabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson, Iz Beltagy, Matthew Peters, Arman Cohan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Rabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson 0001, Iz Beltagy, Matthew E. Peters, Arman Cohan |
EACL (1) | 4 |
| 2023 | HyperMixer: An MLP-based Low Cost Alternative to TransformersabstractFlorian Mai, Arnaud Pannatier, Fabio Fehr, Haolin Chen, Francois Marelli, Francois Fleuret, James Henderson. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Florian Mai, Arnaud Pannatier, Fabio Fehr, François Marelli, François Fleuret, James Henderson 0001 |
ACL (1) | 7 |
| 2023 | A VAE for Transformers with Nonparametric Variational Information Bottleneck
James Henderson 0001, Fabio Fehr |
ICLR | 1 |
| 2022 | Prompt-free and Efficient Few-shot Learning with Language ModelsabstractRabeeh Karimi Mahabadi, Luke Zettlemoyer, James Henderson, Lambert Mathias, Marzieh Saeidi, Veselin Stoyanov, Majid Yazdani. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Rabeeh Karimi Mahabadi, Luke Zettlemoyer, James Henderson 0001, Lambert Mathias, Marzieh Saeidi, Veselin Stoyanov, Majid Yazdani |
ACL (1) | 3 |
| 2022 | SMaLL-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource LanguagesabstractAlireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson, Laurent Besacier. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Alireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson 0001, Laurent Besacier |
EMNLP | 5 |
| 2021 | Parameter-efficient Multi-task Fine-tuning for Transformers via Shared HypernetworksabstractRabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, James Henderson. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani 0001, James Henderson 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Imposing Relation Structure in Language-Model Embeddings Using Contrastive LearningabstractThough language model text embeddings have revolutionized NLP research, their ability to capture high-level semantic information, such as relations between entities in text, is limited.In this paper, we propose a novel contrastive learning framework that trains sentence embeddings to encode the relations in a graph structure.Given a sentence (unstructured text) and its graph, we use contrastive learning to impose relation-related structure on the tokenlevel representations of the sentence obtained with a CharacterBERT (El Boukkouri et al., 2020) model.The resulting relation-aware sentence embeddings achieve state-of-the-art results on the relation extraction task using only a simple KNN classifier, thereby demonstrating the success of the proposed method.Additional visualization by a tSNE analysis shows the effectiveness of the learned representation space compared to baselines.Furthermore, we show that we can learn a different space for named entity recognition, again using a contrastive learning objective, and demonstrate how to successfully combine both representation spaces in an entity-relation task. Christos Theodoropoulos 0001, James Henderson 0001, Andrei Catalin, Marie-Francine Moens |
CoNLL | 2 |
| 2021 | Variational Information Bottleneck for Effective Low-Resource Fine-Tuning
Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson 0001 |
ICLR | 3 |
| 2021 | Measuring Societal Biases from Text Corpora with Smoothed First-Order Co-occurrence
Navid Rekabsaz, Robert West 0001, James Henderson 0001, Allan Hanbury |
ICWSM | 3 |
| 2021 | Multi-Adversarial Learning for Cross-Lingual Word EmbeddingsabstractGenerative adversarial networks (GANs) have succeeded in inducing cross-lingual word embeddings -maps of matching words across languages-without supervision.Despite these successes, GANs' performance for the difficult case of distant languages is still not satisfactory.These limitations have been explained by GANs' incorrect assumption that source and target embedding spaces are related by a single linear mapping and are approximately isomorphic.We assume instead that, especially across distant languages, the mapping is only piece-wise linear, and propose a multi-adversarial learning method.This novel method induces the seed cross-lingual dictionary through multiple mappings, each induced to fit the mapping for one subspace.Our experiments on unsupervised bilingual lexicon induction and cross-lingual document classification show that this method improves performance over previous single-mapping methods, especially for distant languages. Haozhou Wang, James Henderson 0001, Paola Merlo |
NAACL-HLT | 2 |
| 2021 | Compacter: Efficient Low-Rank Hypercomplex Adapter LayersabstractAdapting large-scale pretrained language models to downstream tasks via fine-tuning is the standard method for achieving state-of-the-art performance on NLP benchmarks. However, fine-tuning all weights of models with millions or billions of parameters is sample-inefficient, unstable in low-resource settings, and wasteful as it requires storing a separate copy of the model for each task. Recent work has developed parameter-efficient fine-tuning methods, but these approaches either still require a relatively large number of parameters or underperform standard fine-tuning. In this work, we propose Compacter, a method for fine-tuning large-scale language models with a better trade-off between task performance and the number of trainable parameters than prior work. Compacter accomplishes this by building on top of ideas from adapters, low-rank optimization, and parameterized hypercomplex multiplication layers.Specifically, Compacter inserts task-specific weight matrices into a pretrained model's weights, which are computed efficiently as a sum of Kronecker products between shared slow'' weights andfast'' rank-one matrices defined per Compacter layer. By only training 0.047% of a pretrained model's parameters, Compacter performs on par with standard fine-tuning on GLUE and outperforms standard fine-tuning on SuperGLUE and low-resource settings. Our code is publicly available at https://github.com/rabeehk/compacter. Rabeeh Karimi Mahabadi, James Henderson 0001, Sebastian Ruder |
NeurIPS | 2 |
| 2021 | Towards syntax-aware token embeddingsabstractAbstract Distributional semantic word representations are at the basis of most modern NLP systems. Their usefulness has been proven across various tasks, particularly as inputs to deep learning models. Beyond that, much work investigated fine-tuning the generic word embeddings to leverage linguistic knowledge from large lexical resources. Some work investigated context-dependent word token embeddings motivated by word sense disambiguation, using sequential context and large lexical resources. More recently, acknowledging the need for an in-context representation of words, some work leveraged information derived from language modelling and large amounts of data to induce contextualised representations. In this paper, we investigate Syntax-Aware word Token Embeddings (SATokE) as a way to explicitly encode specific information derived from the linguistic analysis of a sentence in vectors which are input to a deep learning model. We propose an efficient unsupervised learning algorithm based on tensor factorisation for computing these token embeddings given an arbitrary graph of linguistic structure. Applying this method to syntactic dependency structures, we investigate the usefulness of such token representations as part of deep learning models of text understanding. We encode a sentence either by learning embeddings for its tokens and the relations between them from scratch or by leveraging pre-trained relation embeddings to infer token representations. Given sufficient data, the former is slightly more accurate than the latter, yet both provide more informative token embeddings than standard word representations, even when the word representations have been learned on the same type of context from larger corpora (namely pre-trained dependency-based word embeddings). We use a large set of supervised tasks and two major deep learning families of models for sentence understanding to evaluate our proposal. We empirically demonstrate the superiority of the token representations compared to popular distributional representations of words for various sentence and sentence pair classification tasks. Diana Nicoleta Popa, Julien Perez, James Henderson 0001, Éric Gaussier |
Nat. Lang. Eng. | 3 |
| 2021 | Recursive Non-Autoregressive Graph-to-Graph Transformer for Dependency Parsing with Iterative RefinementabstractWe propose the Recursive Non-autoregressive Graph-to-Graph Transformer architecture (RNGTr) for the iterative refinement of arbitrary graphs through the recursive application of a non-autoregressive Graph-to-Graph Transformer and apply it to syntactic dependency parsing. We demonstrate the power and effectiveness of RNGTr on several dependency corpora, using a refinement model pre-trained with BERT. We also introduce Syntactic Transformer (SynTr), a non-recursive parser similar to our refinement model. RNGTr can improve the accuracy of a variety of initial parsers on 13 languages from the Universal Dependencies Treebanks, English and Chinese Penn Treebanks, and the German CoNLL2009 corpus, even improving over the new state-of-the-art results achieved by SynTr, significantly improving the state-of-the-art for all corpora tested. Alireza Mohammadshahi, James Henderson 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | The Unstoppable Rise of Computational Linguistics in Deep LearningabstractIn this paper, we trace the history of neural networks applied to natural language understanding tasks, and identify key contributions which the nature of language has made to the development of neural network architectures.We focus on the importance of variable binding and its instantiation in attention-based models, and argue that Transformer is not a sequence model but an induced-structure model.This perspective leads to predictions of the challenges facing research in deep learning architectures for natural language understanding. James Henderson 0001 |
ACL | 1 |
| 2020 | End-to-End Bias Mitigation by Modelling Biases in CorporaabstractSeveral recent studies have shown that strong natural language understanding (NLU) models are prone to relying on unwanted dataset biases without learning the underlying task, resulting in models that fail to generalize to out-of-domain datasets and are likely to perform poorly in real-world scenarios.We propose two learning strategies to train neural models, which are more robust to such biases and transfer better to out-of-domain datasets.The biases are specified in terms of one or more bias-only models, which learn to leverage the dataset biases.During training, the bias-only models' predictions are used to adjust the loss of the base model to reduce its reliance on biases by down-weighting the biased examples and focusing training on the hard examples.We experiment on large-scale natural language inference and fact verification benchmarks, evaluating on out-of-domain datasets that are specifically designed to assess the robustness of models against known biases in the training data.Results show that our debiasing methods greatly improve robustness in all settings and better transfer to other textual entailment datasets. Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson 0001 |
ACL | 3 |
| 2020 | Plug and Play Autoencoders for Conditional Text GenerationabstractText autoencoders are commonly used for conditional generation tasks such as style transfer.We propose methods which are plug and play, where any pretrained autoencoder can be used, and only require learning a mapping within the autoencoder's embedding space, training embedding-to-embedding (Emb2Emb).This reduces the need for labeled training data for the task and makes the training procedure more efficient.Crucial to the success of this method is a loss term for keeping the mapped embedding on the manifold of the autoencoder and a mapping which is trained to navigate the manifold by learning offset vectors.Evaluations on style transfer tasks both with and without sequence-to-sequence supervision show that our method performs better than or comparable to strong baselines while being up to four times faster. Florian Mai, Nikolaos Pappas 0002, Ivan Montero, Noah A. Smith, James Henderson 0001 |
EMNLP (1) | 5 |
| 2019 | Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word EmbeddingsabstractHaozhou Wang, James Henderson, Paola Merlo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Haozhou Wang, James Henderson 0001, Paola Merlo |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Deep Residual Output Layers for Neural Language GenerationabstractMany tasks, including language generation, benefit from learning the structure of the output space, particularly when the space of output labels is large and the data is sparse. State-of-the-art neural language models indirectly capture the output space structure in their classifier weights since they lack parameter sharing across output labels. Learning shared output label mappings helps, but existing methods have limited expressivity and are prone to overfitting. In this paper, we investigate the usefulness of more powerful shared mappings for output labels, and propose a deep residual output mapping with dropout between layers to better capture the structure of the output space and avoid overfitting. Evaluations on three language generation tasks show that our output label mapping can match or improve state-of-the-art recurrent and self-attention architectures, and suggest that the classifier does not necessarily need to be high-rank to better model natural language if it is better at capturing the structure of the output space. Nikolaos Pappas 0002, James Henderson 0001 |
ICML | 2 |
| 2019 | GILE: A Generalized Input-Label Embedding for Text ClassificationabstractNeural text classification models typically treat output labels as categorical variables that lack description and semantics. This forces their parametrization to be dependent on the label set size, and, hence, they are unable to scale to large label sets and generalize to unseen ones. Existing joint input-label text models overcome these issues by exploiting label descriptions, but they are unable to capture complex label relationships, have rigid parametrization, and their gains on unseen labels happen often at the expense of weak performance on the labels seen during training. In this paper, we propose a new input-label model that generalizes over previous such models, addresses their limitations, and does not compromise performance on seen labels. The model consists of a joint nonlinear input-label embedding with controllable capacity and a joint-space-dependent classification unit that is trained with cross-entropy loss to optimize classification performance. We evaluate models on full-resource and low- or zero-resource text classification of multilingual news and biomedical text with a large label set. Our model outperforms monolingual and multilingual models that do not leverage label semantics and previous joint input-label space models in both scenarios. Nikolaos Pappas 0002, James Henderson 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | Document-Level Neural Machine Translation with Hierarchical Attention NetworksabstractNeural Machine Translation (NMT) can be improved by including document-level contextual information.For this purpose, we propose a hierarchical attention model to capture the context in a structured and dynamic manner.The model is integrated in the original NMT architecture as another level of abstraction, conditioning on the NMT model's own previous hidden states.Experiments show that hierarchical attention significantly improves the BLEU score over a strong NMT baseline with the state-of-the-art in context-aware methods, and that both the encoder and decoder benefit from context in complementary ways. Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas 0002, James Henderson 0001 |
EMNLP | 4 |
| 2018 | Integrating Weakly Supervised Word Sense Disambiguation into Neural Machine TranslationabstractThis paper demonstrates that word sense disambiguation (WSD) can improve neural machine translation (NMT) by widening the source context considered when modeling the senses of potentially ambiguous words. We first introduce three adaptive clustering algorithms for WSD, based on k-means, Chinese restaurant processes, and random walks, which are then applied to large word contexts represented in a low-rank space and evaluated on SemEval shared-task data. We then learn word vectors jointly with sense vectors defined by our best WSD method, within a state-of-the-art NMT system. We show that the concatenation of these vectors, and the use of a sense selection mechanism based on the weighted average of sense vectors, outperforms several baselines including sense-aware ones. This is demonstrated by translation on five language pairs. The improvements are more than 1 BLEU point over strong NMT baselines, +4% accuracy over all ambiguous nouns and verbs, or +20% when scored manually over several challenging words. Xiao Pu 0001, Nikolaos Pappas 0002, James Henderson 0001, Andrei Popescu-Belis |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | A Vector Space for Distributional Semantics for EntailmentabstractDistributional semantics creates vectorspace representations that capture many forms of semantic similarity, but their relation to semantic entailment has been less clear.We propose a vector-space model which provides a formal foundation for a distributional semantics of entailment.Using a mean-field approximation, we develop approximate inference procedures and entailment operators over vectors of probabilities of features being known (versus unknown).We use this framework to reinterpret an existing distributionalsemantic model (Word2Vec) as approximating an entailment-based model of the distributions of words in contexts, thereby predicting lexical entailment relations.In both unsupervised and semi-supervised experiments on hyponymy detection, we get substantial improvements over previous results. James Henderson 0001, Diana Nicoleta Popa |
ACL (1) | 1 |
| 2015 | Incremental Recurrent Neural Network Dependency Parser with Search-based Discriminative TrainingabstractWe propose a discriminatively trained recurrent neural network (RNN) that predicts the actions for a fast and accurate shift-reduce dependency parser.The RNN uses its output-dependent model structure to compute hidden vectors that encode the preceding partial parse, and uses them to estimate probabilities of parser actions.Unlike a similar previous generative model ( Henderson and Titov, 2010), the RNN is trained discriminatively to optimize a fast beam search.This beam search prunes after each shift action, so we add a correctness probability to each shift action and train this score to discriminate between correct and incorrect sequences of parser actions.We also speed up parsing time by caching computations for frequent feature combinations, including during training, giving us both faster training and a form of backoff smoothing.The resulting parser is over 35 times faster than its generative counterpart with nearly the same accuracy, producing state-of-art dependency parsing results while requiring minimal feature engineering. Majid Yazdani, James Henderson 0001 |
CoNLL | 2 |
| 2015 | Named entity recognition with document-specific KB tag gazetteersabstractWe consider a novel setting for Named Entity Recognition (NER) where we have access to document-specific knowledge base tags.These tags consist of a canonical name from a knowledge base (KB) and entity type, but are not aligned to the text.We explore how to use KB tags to create document-specific gazetteers at inference time to improve NER.We find that this kind of supervision helps recognise organisations more than standard widecoverage gazetteers.Moreover, augmenting document-specific gazetteers with KB information lets users specify fewer tags for the same performance, reducing cost. Will Radford, Xavier Carreras, James Henderson 0001 |
EMNLP | 3 |
| 2015 | Learning Semantic Composition to Detect Non-compositionality of Multiword ExpressionsabstractNon-compositionality of multiword expressions is an intriguing problem that can be the source of error in a variety of NLP tasks such as language generation, machine translation and word sense disambiguation. We present methods of non-compositionality detection for English noun compounds using the unsupervised learning of a semantic composition function. Compounds which are not well modeled by the learned semantic composition function are considered noncompositional. We explore a range of distributional vector-space models for semantic composition, empirically evaluate these models, and propose additional methods which improve results further. We show that a complex function such as polynomial projection can learn semantic composition and identify non-compositionality in an unsupervised way, beating all other baselines ranging from simple to complex. We show that enforcing sparsity is a useful regularizer in learning complex composition functions. We show further improvements by training a decomposition function in addition to the composition function. Finally, we propose an EM algorithm over latent compositionality annotations that also improves the performance. Majid Yazdani, Meghdad Farahmand, James Henderson 0001 |
EMNLP | 3 |
| 2015 | A Model of Zero-Shot Learning of Spoken Language UnderstandingabstractWhen building spoken dialogue systems for a new domain, a major bottleneck is developing a spoken language understanding (SLU) module that handles the new domain's terminology and semantic concepts.We propose a statistical SLU model that generalises to both previously unseen input words and previously unseen output classes by leveraging unlabelled data.After mapping the utterance into a vector space, the model exploits the structure of the output labels by mapping each label to a hyperplane that separates utterances with and without that label.Both these mappings are initialised with unsupervised word embeddings, so they can be computed even for words or concepts which were not in the SLU training data. Majid Yazdani, James Henderson 0001 |
EMNLP | 2 |
| 2014 | Undirected Machine Translation with Discriminative Reinforcement LearningabstractWe present a novel Undirected Machine Translation model of Hierarchical MT that is not constrained to the standard bottomup inference order.Removing the ordering constraint makes it possible to condition on top-down structure and surrounding context.This allows the introduction of a new class of contextual features that are not constrained to condition only on the bottom-up context.The model builds translation-derivations efficiently in a greedy fashion.It is trained to learn to choose jointly the best action and the best inference order.Experiments show that the decoding time is halved and forestrescoring is 6 times faster, while reaching accuracy not significantly different from state of the art. Andrea Gesmundo, James Henderson 0001 |
EACL | 2 |
| 2014 | The PARLANCE mobile application for interactive search in English and MandarinabstractHelen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gašić, James Henderson, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazon-Terrazas, Majid Yazdani, Steve Young, Yanchao Yu. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazón-Terrazas, Majid Yazdani, Steve J. Young, Yanchao Yu |
SIGDIAL Conference | 9 |
| 2013 | Graph-Based Seed Set Expansion for Relation Extraction Using Random Walk Hitting Times
Joel Lang, James Henderson 0001 |
HLT-NAACL | 2 |
| 2013 | Demonstration of the PARLANCE system: a data-driven incremental, spoken dialogue system for interactive search
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay |
SIGDIAL Conference | 7 |
| 2013 | Multilingual Joint Parsing of Syntactic and Semantic Dependencies with a Latent Variable ModelabstractCurrent investigations in data-driven models of parsing have shifted from purely syntactic analysis to richer semantic representations, showing that the successful recovery of the meaning of text requires structured analyses of both its grammar and its semantics. In this article, we report on a joint generative history-based model to predict the most likely derivation of a dependency parser for both syntactic and semantic dependencies, in multiple languages. Because these two dependency structures are not isomorphic, we propose a weak synchronization at the level of meaningful subsequences of the two derivations. These synchronized subsequences encompass decisions about the left side of each individual word. We also propose novel derivations for semantic dependency structures, which are appropriate for the relatively unconstrained nature of these graphs. To train a joint model of these synchronized derivations, we make use of a latent variable model of parsing, the Incremental Sigmoid Belief Network (ISBN) architecture. This architecture induces latent feature representations of the derivations, which are used to discover correlations both within and between the two derivations, providing the first application of ISBNs to a multi-task learning problem. This joint model achieves competitive performance on both syntactic and semantic dependency parsing for several languages. Because of the general nature of the approach, this extension of the ISBN architecture to weakly synchronized syntactic-semantic derivations is also an exemplification of its applicability to other problems where two independent, but related, representations are being learned. James Henderson 0001, Paola Merlo, Ivan Titov 0001, Gabriele Musillo |
Comput. Linguistics | 1 |
| 2011 | Heuristic Search for Non-Bottom-Up Tree Structure Prediction
Andrea Gesmundo, James Henderson 0001 |
EMNLP | 2 |
| 2010 | Incremental Sigmoid Belief Networks for Grammar Learning
James Henderson 0001, Ivan Titov 0001 |
J. Mach. Learn. Res. | 1 |
| 2009 | Online Graph Planarisation for Synchronous Parsing of Semantic and Syntactic Dependencies
Ivan Titov 0001, James Henderson 0001, Paola Merlo, Gabriele Musillo |
IJCAI | 2 |
| 2009 | Automatic annotation of context and speech acts for dialogue corporaabstractAbstract Richly annotated dialogue corpora are essential for new research directions in statistical learning approaches to dialogue management, context-sensitive interpretation, and context-sensitive speech recognition. In particular, large dialogue corpora annotated with contextual information and speech acts are urgently required. We explore how existing dialogue corpora (usually consisting of utterance transcriptions) can be automatically processed to yield new corpora where dialogue context and speech acts are accurately represented. We present a conceptual and computational framework for generating such corpora. As an example, we present and evaluate an automatic annotation system which builds ‘Information State Update’ (ISU) representations of dialogue context for the Communicator (2000 and 2001) corpora of human–machine dialogues (2,331 dialogues). The purposes of this annotation are to generate corpora for reinforcement learning of dialogue policies, for building user simulations, for evaluating different dialogue strategies against a baseline, and for training models for context-dependent interpretation and speech recognition. The automatic annotation system parses system and user utterances into speech acts and builds up sequences of dialogue context representations using an ISU dialogue manager. We present the architecture of the automatic annotation system and a detailed example to illustrate how the system components interact to produce the annotations. We also evaluate the annotations, with respect to the task completion metrics of the original corpus and in comparison to hand-annotated data and annotations produced by a baseline automatic system. The automatic annotations perform well and largely outperform the baseline automatic annotations in all measures. The resulting annotated corpus has been used to train high-quality user simulations and to learn successful dialogue strategies. The final corpus will be made publicly available. Kallirroi Georgila, Oliver Lemon, James Henderson 0001, Johanna D. Moore |
Nat. Lang. Eng. | 3 |
| 2008 | A Latent Variable Model of Synchronous Parsing for Syntactic and Semantic Dependencies
James Henderson 0001, Paola Merlo, Gabriele Musillo, Ivan Titov 0001 |
CoNLL | 1 |
| 2008 | Hybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data SetsabstractWe propose a method for learning dialogue management policies from a fixed data set. The method addresses the challenges posed by Information State Update (ISU)-based dialogue systems, which represent the state of a dialogue as a large set of features, resulting in a very large state space and a huge policy space. To address the problem that any fixed data set will only provide information about small portions of these state and policy spaces, we propose a hybrid model that combines reinforcement learning with supervised learning. The reinforcement learning is used to optimize a measure of dialogue reward, while the supervised learning is used to restrict the learned policy to the portions of these spaces for which we have data. We also use linear function approximation to address the need to generalize from a fixed amount of data to large state spaces. To demonstrate the effectiveness of this method on this challenging task, we trained this model on the COMMUNICATOR corpus, to which we have added annotations for user actions and Information States. When tested with a user simulation trained on a different part of the same data set, our hybrid model outperforms a pure supervised learning model and a pure reinforcement learning model. It also outperforms the hand-crafted systems on the COMMUNICATOR data, according to automatic evaluation measures, improving over the average COMMUNICATOR system policy by 10%. The proposed method will improve techniques for bootstrapping and automatic optimization of dialogue management policies from limited initial data sets. James Henderson 0001, Oliver Lemon, Kallirroi Georgila |
Comput. Linguistics | 1 |
| 2007 | Constituent Parsing with Incremental Sigmoid Belief Networks
Ivan Titov 0001, James Henderson 0001 |
ACL | 2 |
| 2007 | Fast and Robust Multilingual Dependency Parsing with a Generative Latent Variable Model
Ivan Titov 0001, James Henderson 0001 |
EMNLP-CoNLL | 2 |
| 2007 | Incremental Bayesian networks for structure predictionabstractWe propose a class of graphical models appropriate for structure prediction problems where the model structure is a function of the output structure. Incremental Sigmoid Belief Networks (ISBNs) avoid the need to sum over the possible model structures by using directed arcs and incrementally specifying the model structure. Exact inference in such directed models is not tractable, but we derive two efficient approximations based on mean field methods, which prove effective in artificial experiments. We then demonstrate their effectiveness on a benchmark natural language parsing task, where they achieve state-of-the-art accuracy. Also, the model which is a closer approximation to an ISBN has better parsing accuracy, suggesting that ISBNs are an appropriate abstract model of structure prediction tasks. Ivan Titov 0001, James Henderson 0001 |
ICML | 2 |
| 2006 | Porting Statistical Parsers with Data-Defined Kernels
Ivan Titov 0001, James Henderson 0001 |
CoNLL | 2 |
| 2006 | An ISU Dialogue System Exhibiting Reinforcement Learning of Dialogue Policies: Generic Slot-Filling in the TALK In-car System
Oliver Lemon, Kallirroi Georgila, James Henderson 0001, Matthew N. Stuttle |
EACL | 3 |
| 2006 | Loss Minimization in Parse Reranking
Ivan Titov 0001, James Henderson 0001 |
EMNLP | 2 |
| 2006 | User simulation for spoken dialogue systems: learning and evaluationabstractWe propose the “advanced ” n-grams as a new technique for simulating user behaviour in spoken dialogue systems, and we compare it with two methods used in our prior work, i.e. linear feature combination and “normal ” n-grams. All methods operate on the intention level and can incorporate speech recognition and understanding errors. In the linear feature combination model user actions (lists of 〈 speech act, task 〉 pairs) are selected, based on features of the current dialogue state which encodes the whole history of the dialogue. The user simulation based on “normal ” n-grams treats a dialogue as a sequence of lists of 〈 speech act, task 〉 pairs. Here the length of the history considered is restricted by the order of the n-gram. The “advanced ” n-grams are a variation of the normal ngrams, where user actions are conditioned not only on speech acts and tasks but also on the current status of the tasks, i.e. whether Kallirroi Georgila, James Henderson 0001, Oliver Lemon |
INTERSPEECH | 2 |
| 2006 | Evaluating Effectiveness and Portability of Reinforcement Learned Dialogue Strategies with Real Users: the Talk Towninfo EvaluationabstractWe report evaluation results for real users of a learnt dialogue management policy versus a hand-coded policy in the TALK project's "Townlnfo" tourist information system. The learnt policy, for filling and confirming information slots, was derived from COMMUNICATOR (flight-booking) data using reinforcement learning (RL) as described in [2], ported to the tourist information domain (using a general method that we propose here), and tested using 18 human users in 180 dialogues, who also used a state-of-the-art hand- coded dialogue policy embedded in an otherwise identical system. We found that users of the (ported) learned policy had an average gain in perceived task completion of 14.2% (from 67.6% to 81.8% at p < .03), that the hand-coded policy dialogues had on average 3.3 more system turns (p < .01), and that the user satisfaction results were comparable, even though the policy was learned for a different domain. Combining these in a dialogue reward score, we found a 14.4% increase for the learnt policy (a 23.8% relative increase, p < .03). These results are important because they show a) that results for real users are consistent with results for automatic evaluation [2] of learned policies using simulated users [3, 4], b) that a policy learned using linear function approximation over a very large policy space [2] is effective for real users, and c) that policies learned using data for one domain can be used successfully in other domains. We also present a qualitative discussion of the learnt policy. Oliver Lemon, Kallirroi Georgila, James Henderson 0001 |
SLT | 3 |
| 2005 | Data-Defined Kernels for Parse Reranking Derived from Probabilistic ModelsabstractPrevious research applying kernel methods to natural language parsing have focussed on proposing kernels over parse trees, which are hand-crafted based on domain knowledge and computational considerations. In this paper we propose a method for defining kernels in terms of a probabilistic model of parsing. This model is then trained, so that the parameters of the probabilistic model reflect the generalizations in the training data. The method we propose then uses these trained parameters to define a kernel for reranking parse trees. In experiments, we use a neural network based statistical parser as the probabilistic model, and use the resulting kernel with the Voted Perceptron algorithm to rerank the top 20 parses from the probabilistic model. This method achieves a significant improvement over the accuracy of the probabilistic model. James Henderson 0001, Ivan Titov 0001 |
ACL | 1 |
| 2005 | Deriving kernels from MLP probability estimators for large categorization problemsabstractIn multi-class categorization problems with a very large or unbounded number of classes, it is often not computationally feasible to train and/or test a kernel-based classifier. One solution is to use a fast computation to pre-select a subset of the classes for reranking with a kernel method, but even then tractability can be a problem. We investigate using trained multilayer perceptron probability estimators to derive appropriate kernels for such problems. We propose a kernel derivation method which is specifically designed for reranking problems, and a more efficient variant of this method which is specifically designed for neural networks with large numbers of output units. When applied to a neural network model of natural language parsing, these new methods achieve state-of-the-art performance which improves over the original model. Ivan Titov 0001, James Henderson 0001 |
IJCNN | 2 |
| 2005 | Learning user simulations for information state update dialogue systemsabstractThis paper describes and compares two methods for simulating user behaviour in spoken dialogue systems. User simulations are important for automatic dialogue strategy learning and the evaluation of competing strategies. Our methods are designed for use with "Information State Update" (ISU)-based dialogue systems. The first method is based on supervised learning using linear feature combination and a normalised exponential output function. The user is modelled as a stochastic process which selects user actions ( pairs) based on features of the current dialogue state, which encodes the whole history of the dialogue. The second method uses n-grams of speech act, task pairs, restricting the length of the history considered by the order of the n-gram. Both models were trained and evaluated on a subset of the COMMUNICATOR corpus, to which we added annotations for user actions and Information States. The model based on linear feature combination has a perplexity of 2.08 whereas the best n-gram (4-gram) has a perplexity of 3.58. Each one of the user models ran against a system policy trained on the same corpus with a method similar to the one used for our linear feature combination model. The quality of the simulated dialogues produced was then measured as a function of the filled slots, confirmed slots, and number of actions performed by the system in each dialogue. In this experiment both the linear feature combination model and the best n-grams (5-gram and 4-gram) produced similar quality simulated dialogues. Kallirroi Georgila, James Henderson 0001, Oliver Lemon |
INTERSPEECH | 2 |
| 2004 | Discriminative Training of a Neural Network Statistical ParserabstractDiscriminative methods have shown significant improvements over traditional generative methods in many machine learning applications, but there has been difficulty in extending them to natural language parsing.One problem is that much of the work on discriminative methods conflates changes to the learning method with changes to the parameterization of the problem.We show how a parser can be trained with a discriminative learning method while still parameterizing the problem according to a generative probability model.We present three methods for training a neural network to estimate the probabilities for a statistical parser, one generative, one discriminative, and one where the probability model is generative but the training criteria is discriminative.The latter model outperforms the previous two, achieving state-ofthe-art levels of performance (90.1% F-measure on constituents). James Henderson 0001 |
ACL | 1 |
| 2004 | Estimating probabilities for unbounded categorization problems
James Henderson 0001 |
Neurocomputing | 1 |
| 2003 | Neural Network Probability Estimation for Broad Coverage Parsing
James Henderson 0001 |
EACL | 1 |
| 2003 | Structural Bias in Inducing Representations for Probabilistic Natural Language Parsing
James Henderson 0001 |
ICANN | 1 |
| 2003 | Inducing History Representations for Broad Coverage Statistical Parsing
James Henderson 0001 |
HLT-NAACL | 1 |
| 2003 | Towards Effective Parsing with Neural Networks: Inherent Generalisations and Bounded Resource Effects
Peter C. R. Lane, James Henderson 0001 |
Appl. Intell. | 2 |
| 2002 | Using Syntactic Analysis to Increase Efficiency in Visualizing Text Collections
James Henderson 0001, Paola Merlo, Ivan Petroff, Gerold Schneider |
COLING | 1 |
| 2002 | Estimating probabilities for unbounded categorization problems
James Henderson 0001 |
ESANN | 1 |
| 2001 | Incremental Syntactic Parsing of Natural Language Corpora with Simple Synchrony NetworksabstractThe article explores the use of Simple Synchrony Networks (SSNs) for learning to parse English sentences drawn from a corpus of naturally occurring text. Parsing natural language sentences requires taking a sequence of words and outputting a hierarchical structure representing how those words fit together to form constituents. Feedforward and simple recurrent networks have had great difficulty with this task, in part because the number of relationships required to specify a structure is too large for the number of unit outputs they have available. SSNs have the representational power to output the necessary O(n/sup 2/) possible structural relationships because SSNs extend the O(n) incremental outputs of simple recurrent networks with the O(n) entity outputs provided by temporal synchrony variable binding. The article presents an incremental representation of constituent structures which allows SSNs to make effective use of both these dimensions. Experiments on learning to parse naturally occurring text show that this output format supports both effective representation and effective generalization in SSNs. To emphasize the importance of this generalization ability, the article also proposes a short-term memory mechanism for retaining a bounded number of constituents during parsing. This mechanism improves the O(n/sup 2/) speed of the basic SSN architecture to linear time, but experiments confirm that the generalization ability of SSN networks is maintained. Peter C. R. Lane, James Henderson 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1992 | A Connectionist Parser for Structure Unification GrammarabstractThis paper presents a connectionist syntactic parser which uses Structure Unification Grammar as its grammatical framework. The parser is implemented in a connectionist architecture which stores and dynamically manipulates symbolic representations, but which can't represent arbitrary disjunction and has bounded memory. These problems can be overcome with Structure Unification Grammar's extensive use of partial descriptions. James Henderson 0001 |
ACL | 1 |
| 1991 | An Incremental Connectionist Phrase Structure ParserabstractNo abstract available. James Henderson 0001 |
ACL | 1 |