Ivan Titov 0001

dblp:08/5391-1 · DBLP profile ↗
← Back
103ranked-venue papers
16as first author
35since 2021 · last 2026
0000-0002-2583-1893ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 101 · 15 first-author · 35 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
abstract
Abstract Human users increasingly communicate with large language models (LLMs), but LLMs suffer from frequent overconfidence in their output, even when its accuracy is questionable, which undermines their trustworthiness and perceived legitimacy. Therefore, there is a need for language models to signal their confidence in order to reap the benefits of human-machine collaboration and mitigate potential harms. Verbalized uncertainty is the expression of confidence with linguistic means, an approach that integrates perfectly into language-based interfaces. Most recent research in natural language processing (NLP) overlooks the nuances surrounding human uncertainty communication and the biases that influence the communication of and with machines. We argue for anthropomimetic uncertainty, the principle that intuitive and trustworthy uncertainty communication requires a degree of imitation of human linguistic behaviors. We present a thorough overview of the research in human uncertainty communication, survey ongoing research in NLP, and perform additional analyses to demonstrate so-far underexplored biases in verbalized uncertainty. We conclude by pointing out unique factors in human-machine uncertainty and outlining future research directions towards implementing anthropomimetic uncertainty. kaleidophon/anthropomimetic-uncertainty
Dennis Ulmer, Alexandra Lorson, Ivan Titov 0001, Christian Hardmeier
Trans. Assoc. Comput. Linguistics3
2025 Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
abstract
Zihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Ivan Titov 0001, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin
ACL (1)7
2025 Explanation Regularisation through the Lens of Attributions
abstract
Explanation regularisation (ER) has been introduced as a way to guide text classifiers to form their predictions relying on input tokens that humans consider plausible. This is achieved by introducing an auxiliary explanation loss that measures how well the output of an input attribution technique for the model agrees with human-annotated rationales. The guidance appears to benefit performance in out-of-domain (OOD) settings, presumably due to an increased reliance on plausible tokens. However, previous work has under-explored the impact of guidance on that reliance, particularly when reliance is measured using attribution techniques different from those used to guide the model. In this work, we seek to close this gap, and also explore the relationship between reliance on plausible features and OOD performance. We find that the connection between ER and the ability of a classifier to rely on plausible features has been overstated and that a stronger reliance on plausible tokens does not seem to be the cause for OOD improvements.
Ivan Titov 0001, Wilker Aziz
COLING2
2025 M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
abstract
Multilingual LLM performance is often critically dependent on model size.With an eye on efficiency, this has led to a surge in interest in one-shot pruning methods that retain the benefits of large-scale pretraining while shrinking the model size.However, as pruning tends to come with performance loss, it is important to understand the trade-offs between multilinguality and sparsification.In this work, we study multilingual performance under different sparsity constraints and show that moderate ratios already substantially harm performance.To help bridge this gap, we propose M-Wanda, a pruning method that models cross-lingual variation by incorporating language-aware activation statistics into its pruning criterion and dynamically adjusts layerwise sparsity based on cross-lingual importance.We show that M-Wanda consistently improves performance at minimal additional costs.We are the first to explicitly optimize pruning to retain multilingual performance, and hope to inspire future advances in multilingual pruning. 1
Rochelle Choenni, Ivan Titov 0001
EMNLP2
2025 Enhancing RLHF with Human Gaze Modeling
abstract
Reinforcement Learning from Human Feedback (RLHF) aligns language models with human preferences but is computationally expensive.We explore two approaches that leverage human gaze modeling to enhance RLHF:(1) gaze-aware reward models and (2) gazebased distribution of sparse rewards at token level.Our experiments demonstate that gazeinformed RLHF achieves faster convergence while maintaining or slightly improving performance, thus, reducing computational costs during policy optimization.These results show that human gaze provides a valuable and underused signal for policy optimization, pointing to a promising direction for improving RLHF efficiency.
Karim Galliamov, Ivan Titov 0001, Ilya Pershin
EMNLP2
2025 Language Agents Meet Causality - Bridging LLMs and Causal World Models
abstract
Large Language Models (LLMs) have recently shown great promise in planning and reasoning applications. These tasks demand robust systems, which arguably require a causal understanding of the environment. While LLMs can acquire and reflect common sense causal knowledge from their pretraining data, this information is often incomplete, incorrect, or inapplicable to a specific environment. In contrast, causal representation learning (CRL) focuses on identifying the underlying causal structure within a given environment. We propose a framework that integrates CRLs with LLMs to enable causally-aware reasoning and planning. This framework learns a causal world model, with causal variables linked to natural language expressions. This mapping provides LLMs with a flexible interface to process and generate descriptions of actions and states in text form. Effectively, the causal world model acts as a simulator that the LLM can query and interact with. We evaluate the framework on causal inference and planning tasks across temporal scales and environmental complexities. Our experiments demonstrate the effectiveness of the approach, with the causally-aware method outperforming LLM-based reasoners, especially for longer planning horizons.
John Gkountouras, Matthias Lindemann, Phillip Lippe, Efstratios Gavves, Ivan Titov 0001
ICLR5
2025 Post-hoc Reward Calibration: A Case Study on Length Bias
abstract
Reinforcement Learning from Human Feedback aligns the outputs of Large Language Models with human values and preferences. Central to this process is the reward model (RM), which translates human feedback into training signals for optimising LLM behaviour. However, RMs can develop biases by exploiting spurious correlations in their training data, such as favouring outputs based on length or style rather than true quality. These biases can lead to incorrect output rankings, sub-optimal model evaluations, and the amplification of undesirable behaviours in LLMs alignment. This paper addresses the challenge of correcting such biases without additional data and training, introducing the concept of Post-hoc Reward Calibration. We first propose to use local average reward to estimate the bias term and, thus, remove it to approximate the underlying true reward. We then extend the approach to a more general and robust form with the Locally Weighted Regression. Focusing on the prevalent length bias, we validate our proposed approaches across three experimental settings, demonstrating consistent improvements: (1) a 3.11 average performance gain across 33 reward models on the RewardBench dataset; (2) improved agreement of RM produced rankings with GPT-4 evaluations and human preferences based on the AlpacaEval benchmark; and (3) improved Length-Controlled win rate (Dubois et al., 2024) of the RLHF process in multiple LLM–RM combinations. According to our experiments, our method is computationally efficient and generalisable to other types of bias and RMs, offering a scalable and robust solution for mitigating biases in LLM alignment and evaluation.
Zihan Qiu, Edoardo Maria Ponti, Ivan Titov 0001
ICLR5
2025 What's New in My Data? Novelty Exploration via Contrastive Generation
abstract
Fine-tuning is widely used to adapt language models for specific goals, often leveraging real-world data such as patient records, customer-service interactions, or web content in languages not covered in pre-training. These datasets are typically massive, noisy, and often confidential, making their direct inspection challenging. However, understanding them is essential for guiding model deployment and informing decisions about data cleaning or suppressing any harmful behaviors learned during fine-tuning. In this study, we introduce the task of novelty discovery through generation, which aims to identify novel domains of a fine-tuning dataset by generating examples that illustrate these properties. Our approach - Contrastive Generative Exploration (CGE) - assumes no direct access to the data but instead relies on a pre-trained model and the same model after fine-tuning. By contrasting the predictions of these two models, CGE can generate examples that highlight novel domains of the fine-tuning data. However, this simple approach may produce examples that are too similar to one another, failing to capture the full range of novel domains present in the dataset. We address this by introducing an iterative version of CGE, where the previously generated examples are used to update the pre-trained model, and this updated model is then contrasted with the fully fine-tuned model to generate the next example, promoting diversity in the generated outputs. Our experiments demonstrate the effectiveness of CGE in detecting novel domains, such as toxic language, as well as new natural and programming languages. Furthermore, we show that CGE remains effective even when models are fine-tuned using differential privacy techniques.
Masaru Isonuma, Ivan Titov 0001
ICLR2
2025 Layerwise Recurrent Router for Mixture-of-Experts
abstract
The scaling of large language models (LLMs) has revolutionized their capabilities in various tasks, yet this growth must be matched with efficient computational strategies. The Mixture-of-Experts (MoE) architecture stands out for its ability to scale model size without significantly increasing training costs. Despite their advantages, current MoE models often display parameter inefficiency. For instance, a pre-trained MoE-based LLM with 52 billion parameters might perform comparably to a standard model with 6.7 billion. Being a crucial part of MoE, current routers in different layers independently assign tokens without leveraging historical routing information, potentially leading to suboptimal token-expert combinations and the parameter inefficiency problem. To alleviate this issue, we introduce the Layerwise Recurrent Router for Mixture-of-Experts (RMoE). RMoE leverages a Gated Recurrent Unit (GRU) to establish dependencies between routing decisions across consecutive layers. Such layerwise recurrence can be efficiently parallelly computed for input tokens and introduces negotiable costs. Our extensive empirical evaluations demonstrate that RMoE-based language models consistently outperform a spectrum of baseline models. Furthermore, RMoE integrates a novel computation stage orthogonal to existing methods, allowing seamless compatibility with other MoE architectures. Our analyses attribute RMoE's gains to its effective cross-layer information sharing, which also improves expert selection and diversity.
Zihan Qiu, Shuang Cheng, Yizhi Zhou, Ivan Titov 0001, Jie Fu 0001
ICLR6
2025 Joint Localization and Activation Editing for Low-Resource Fine-Tuning
abstract
Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, are commonly used to adapt LLMs. However, the effectiveness of standard PEFT methods is limited in low-resource scenarios with only a few hundred examples. Recent advances in interpretability research have inspired the emergence of activation editing (or steering) techniques, which modify the activations of specific model components. Due to their extremely small parameter counts, these methods show promise for small datasets. However, their performance is highly dependent on identifying the correct modules to edit and often lacks stability across different datasets. In this paper, we propose Joint Localization and Activation Editing (JoLA), a method that jointly learns (1) which heads in the Transformer to edit (2) whether the intervention should be additive, multiplicative, or both and (3) the intervention parameters themselves - the vectors applied as additive offsets or multiplicative scalings to the head output. Through evaluations on three benchmarks spanning commonsense reasoning, natural language understanding, and natural language generation, we demonstrate that JoLA consistently outperforms existing methods.
Wen Lai, Alexander Fraser 0001, Ivan Titov 0001
ICML3
2025 A Controllable Examination for Long-Context Language Models
abstract
Existing frameworks for evaluating long-context language models (LCLM) can be broadly categorized into real-world applications (e.g, document summarization) and synthetic tasks (e.g, needle-in-a-haystack). Despite their utility, both approaches are accompanied by certain intrinsic limitations. Real-world tasks often involve complexity that makes interpretation challenging and suffer from data contamination, whereas synthetic tasks frequently lack meaningful coherence between the target information ("needle") and its surrounding context ("haystack"), undermining their validity as proxies for realistic applications. In response to these challenges, we posit that an ideal long-context evaluation framework should be characterized by three essential features: 1) seamless context: coherent contextual integration between target information and its surrounding context; 2) controllable setting: an extensible task setup that enables controlled studies—for example, incorporating additional required abilities such as numerical reasoning; and 3) sound evaluation: avoiding LLM-as-Judge and conduct exact-match to ensure deterministic and reproducible evaluation results.This study introduces $\textbf{LongBioBench}$, a benchmark that utilizes artificially generated biographies as a controlled environment for assessing LCLMs across dimensions of $\textit{understanding}$, $\textit{reasoning}$, and $\textit{trustworthiness}$. Our experimental evaluation, which includes $\textbf{18}$ LCLMs in total, demonstrates that most models still exhibit deficiencies in semantic understanding and elementary reasoning over retrieved results and are less trustworthy as context length increases.Our further analysis indicates some design choices employed by existing synthetic benchmarks, such as contextual non-coherence, numerical needles, and the absence of distractors, rendering them vulnerable to test the model's long-context capabilities.Moreover, we also reveal that long-context continual pretraining primarily adjusts RoPE embedding to accommodate extended context lengths, which in turn yields only marginal improvements in the model’s true capabilities. To sum up, compared to previous synthetic benchmarks, LongBioBench achieves a better trade-off between mirroring authentic language tasks and maintaining controllability, and is highly interpretable and configurable.
Zihan Qiu, Fei Yuan 0006, Jeff Z. Pan, Ivan Titov 0001
NeurIPS7
2024 Unlearning Traces the Influential Training Data of Language Models
abstract
Identifying the training datasets that influence a language model's outputs is essential for minimizing the generation of harmful content and enhancing its performance.Ideally, we can measure the influence of each dataset by removing it from training; however, it is prohibitively expensive to retrain a model multiple times.This paper presents UnTrac: Unlearning Traces the influence of a training dataset on the model's performance.UnTrac is extremely simple; each training dataset is unlearned by gradient ascent, and we evaluate how much the model's predictions change after unlearning.Furthermore, we propose a more scalable approach, UnTrac-Inv, which unlearns a test dataset and evaluates the unlearned model on training datasets.UnTrac-Inv resembles Un-Trac, while being efficient for massive training datasets.In the experiments, we examine if our methods can assess the influence of pretraining datasets on generating toxic, biased, and untruthful content.Our methods estimate their influence much more accurately than existing methods while requiring neither excessive memory space nor multiple checkpoints.
Masaru Isonuma, Ivan Titov 0001
ACL (1)2
2024 SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
abstract
Strong inductive biases enable learning from little data and help generalization outside of the training distribution.Popular neural architectures such as Transformers lack strong structural inductive biases for seq2seq NLP tasks on their own.Consequently, they struggle with systematic generalization beyond the training distribution, e.g. with extrapolating to longer inputs, even when pre-trained on large amounts of text.We show how a structural inductive bias can be efficiently injected into a seq2seq model by pre-training it to simulate structural transformations on synthetic data.Specifically, we inject an inductive bias towards Finite State Transducers (FSTs) into a Transformer by pretraining it to simulate FSTs given their descriptions.Our experiments show that our method imparts the desired inductive bias, resulting in improved systematic generalization and better few-shot learning for FST-like tasks.Our analysis shows that fine-tuned models accurately capture the state dynamics of the unseen underlying FSTs, suggesting that the simulation process is internalized by the fine-tuned model. 1
Matthias Lindemann, Alexander Koller, Ivan Titov 0001
ACL (1)3
2024 Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
abstract
Models need appropriate inductive biases to effectively learn from small amounts of data and generalize systematically outside of the training distribution.While Transformers are highly versatile and powerful, they can still benefit from enhanced structural inductive biases for seq2seq tasks, especially those involving syntactic transformations, such as converting active to passive voice or semantic parsing.In this paper, we propose to strengthen the structural inductive bias of a Transformer by intermediate pre-training to perform synthetically generated syntactic transformations of dependency trees given a description of the transformation.Our experiments confirm that this helps with fewshot learning of syntactic tasks such as chunking, and also improves structural generalization for semantic parsing.Our analysis shows that the intermediate pre-training leads to attention heads that keep track of which syntactic transformation needs to be applied to which token, and that the model can leverage these attention heads on downstream tasks. 1
Matthias Lindemann, Alexander Koller, Ivan Titov 0001
EMNLP3
2024 Autoencoding Conditional Neural Processes for Representation Learning
abstract
Conditional neural processes (CNPs) are a flexible and efficient family of models that learn to learn a stochastic process from data. They have seen particular application in contextual image completion - observing pixel values at some locations to predict a distribution over values at other unobserved locations. However, the choice of pixels in learning CNPs is typically either random or derived from a simple statistical measure (e.g. pixel variance). Here, we turn the problem on its head and ask: which pixels would a CNP like to observe - do they facilitate fitting better CNPs, and do such pixels tell us something meaningful about the underlying image? To this end we develop the Partial Pixel Space Variational Autoencoder (PPS-VAE), an amortised variational framework that casts CNP context as latent variables learnt simultaneously with the CNP. We evaluate PPS-VAE over a number of tasks across different visual data, and find that not only can it facilitate better-fit CNPs, but also that the spatial arrangement and values meaningfully characterise image information - evaluated through the lens of classification on both within and out-of-data distributions. Our model additionally allows for dynamic adaption of context-set size and the ability to scale-up to larger images, providing a promising avenue to explore learning meaningful and effective visual representations.
Victor Prokhorov, Ivan Titov 0001, N. Siddharth 0001
ICML2
2023 Compositional Generalization without Trees using Multiset Tagging and Latent Permutations
abstract
Seq2seq models have been shown to struggle with compositional generalization in semantic parsing, i.e. generalizing to unseen compositions of phenomena that the model handles correctly in isolation.We phrase semantic parsing as a two-step process: we first tag each input token with a multiset of output tokens.Then we arrange the tokens into an output sequence using a new way of parameterizing and predicting permutations.We formulate predicting a permutation as solving a regularized linear program and we backpropagate through the solver.In contrast to prior work, our approach does not place a priori restrictions on possible permutations, making it very expressive.Our model outperforms pretrained seq2seq models and prior work on realistic semantic parsing tasks that require generalization to longer examples.We also outperform non-tree-based models on structural generalization on the COGS benchmark.For the first time, we show that a model without an inductive bias provided by trees achieves high accuracy on generalization to deeper recursion depth. 1
Matthias Lindemann, Alexander Koller, Ivan Titov 0001
ACL (1)3
2023 Compositional Generalisation with Structured Reordering and Fertility Layers
abstract
Seq2seq models have been shown to struggle with compositional generalisation, i.e. generalising to new and potentially more complex structures than seen during training.Taking inspiration from grammar-based models that excel at compositional generalisation, we present a flexible end-to-end differentiable neural model that composes two structural operations: a fertility step, which we introduce in this work, and a reordering step based on previous work (Wang et al., 2021).To ensure differentiability, we use the expected value of each step, which we compute using dynamic programming.Our model outperforms seq2seq models by a wide margin on challenging compositional splits of realistic semantic parsing tasks that require generalisation to longer examples.It also compares favourably to other models targeting compositional generalisation. 1
Matthias Lindemann, Alexander Koller, Ivan Titov 0001
EACL3
2023 Cross-Modal Conceptualization in Bottleneck Models
abstract
Concept Bottleneck Models (CBMs) (Koh et al., 2020) assume that training examples (e.g., x-ray images) are annotated with high-level concepts (e.g., types of abnormalities), and perform classification by first predicting the concepts, followed by predicting the label relying on these concepts.The main difficulty in using CBMs comes from having to choose concepts that are predictive of the label and then having to label training examples with these concepts.In our approach, we adopt a more moderate assumption and instead use text descriptions (e.g., radiology reports), accompanying the images in training, to guide the induction of concepts.Our cross-modal approach treats concepts as discrete latent variables and promotes concepts that (1) are predictive of the label, and (2) can be predicted reliably from both the image and text.Through experiments conducted on datasets ranging from synthetic datasets (e.g., synthetic images with generated descriptions) to realistic medical imaging datasets, we demonstrate that cross-modal learning encourages the induction of interpretable concepts while also facilitating disentanglement.Our results also suggest that this guidance leads to increased robustness by suppressing the reliance on shortcut features.
Danis Alukaev, Semen Kiselev, Ilya Pershin, Bulat Ibragimov, Alexey Kornaev, Ivan Titov 0001
EMNLP7
2023 Memorisation Cartography: Mapping out the Memorisation-Generalisation Continuum in Neural Machine Translation
abstract
When training a neural network, it will quickly memorise some source-target mappings from your dataset but never learn some others.Yet, memorisation is not easily expressed as a binary feature that is good or bad: individual datapoints lie on a memorisation-generalisation continuum.What determines a datapoint's position on that spectrum, and how does that spectrum influence neural models' performance?We address these two questions for neural machine translation (NMT) models.We use the counterfactual memorisation metric to (1) build a resource that places 5M NMT datapoints on a memorisation-generalisation map, (2) illustrate how the datapoints' surface-level characteristics and a models' per-datum training signals are predictive of memorisation in NMT, (3) and describe the influence that subsets of that map have on NMT systems' performance.1 * Work partially conducted during an internship at FAIR. 1 Click here to interactively explore the NMT memorisation maps in our demo.
Verna Dankers, Ivan Titov 0001, Dieuwke Hupkes
EMNLP2
2023 Theoretical and Practical Perspectives on what Influence Functions Do
abstract
Influence functions (IF) have been seen as a technique for explaining model predictions through the lens of the training data. Their utility is assumed to be in identifying training examples "responsible" for a prediction so that, for example, correcting a prediction is possible by intervening on those examples (removing or editing them) and retraining the model. However, recent empirical studies have shown that the existing methods of estimating IF predict the leave-one-out-and-retrain effect poorly. In order to understand the mismatch between the theoretical promise and the practical results, we analyse five assumptions made by IF methods which are problematic for modern-scale deep neural networks and which concern convexity, numeric stability, training trajectory and parameter divergence. This allows us to clarify what can be expected theoretically from IF. We show that while most assumptions can be addressed successfully, the parameter divergence poses a clear limitation on the predictive power of IF: influence fades over training time even with deterministic training. We illustrate this theoretical result with BERT and ResNet models. Another conclusion from the theoretical analysis is that IF are still useful for model debugging and correcting even though some of the assumptions made in prior work do not hold: using natural language processing and computer vision tasks, we verify that mis-predictions can be successfully corrected by taking only a few fine-tuning steps on influential examples.
Andrea Schioppa, Katja Filippova, Ivan Titov 0001, Polina Zablotskaia
NeurIPS3
2022 Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine Translation
abstract
Unlike literal expressions, idioms' meanings do not directly follow from their parts, posing a challenge for neural machine translation (NMT). NMT models are often unable to translate idioms accurately and over-generate compositional, literal translations. In this work, we investigate whether the non-compositionality of idioms is reflected in the mechanics of the dominant NMT model, Transformer, by analysing the hidden states and attention patterns for models with English as source language and one of seven European languages as target language. When Transformer emits a non-literal translation - i.e. identifies the expression as idiomatic - the encoder processes idioms more strongly as single lexical units compared to literal expressions. This manifests in idioms' parts being grouped through attention and in reduced interaction between idioms and their context. In the decoder's cross-attention, figurative inputs result in reduced attention on source-side tokens. These results suggest that Transformer's tendency to process idioms as compositional expressions contributes to literal translations of idioms.
Verna Dankers, Christopher G. Lucas, Ivan Titov 0001
ACL (1)3
2022 Hierarchical Phrase-Based Sequence-to-Sequence Learning
abstract
We describe a neural transducer that maintains the flexibility of standard sequence-to-sequence (seq2seq) models while incorporating hierarchical phrases as a source of inductive bias during training and as explicit constraints during inference.Our approach trains two models: a discriminative parser based on a bracketing transduction grammar whose derivation tree hierarchically aligns source and target phrases, and a neural seq2seq model that learns to translate the aligned phrases one-by-one.We use the same seq2seq model to translate at all phrase scales, which results in two inference modes: one mode in which the parser is discarded and only the seq2seq component is used at the sequence-level, and another in which the parser is combined with the seq2seq model.Decoding in the latter mode is done with the cube-pruned CKY algorithm, which is more involved but can make use of new translation rules during inference.We formalize our model as a sourceconditioned synchronous grammar and develop an efficient variational inference algorithm for training.When applied on top of both randomly initialized and pretrained seq2seq models, we find that both inference modes performs well compared to baselines on small scale machine translation benchmarks.
Bailin Wang, Ivan Titov 0001, Jacob Andreas
EMNLP2
2021 Meta-Learning to Compositionally Generalize
abstract
Henry Conklin, Bailin Wang, Kenny Smith, Ivan Titov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Henry Conklin, Bailin Wang, Kenny Smith, Ivan Titov 0001
ACL/IJCNLP (1)4
2021 Analyzing the Source and Target Contributions to Predictions in Neural Machine Translation
abstract
Elena Voita, Rico Sennrich, Ivan Titov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Elena Voita, Rico Sennrich, Ivan Titov 0001
ACL/IJCNLP (1)3
2021 Beyond Sentence-Level End-to-End Speech Translation: Context Helps
abstract
Biao Zhang, Ivan Titov, Barry Haddow, Rico Sennrich. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Biao Zhang 0006, Ivan Titov 0001, Barry Haddow, Rico Sennrich
ACL/IJCNLP (1)2
2021 Learning Opinion Summarizers by Selecting Informative Reviews
abstract
Opinion summarization has been traditionally approached with unsupervised, weaklysupervised and few-shot learning techniques.In this work, we collect a large dataset of summaries paired with user reviews for over 31,000 products, enabling supervised training.However, the number of reviews per product is large (320 on average), making summarization -and especially training a summarizerimpractical.Moreover, the content of many reviews is not reflected in the human-written summaries, and, thus, the summarizer trained on random review subsets hallucinates.In order to deal with both of these challenges, we formulate the task as jointly learning to select informative subsets of reviews and summarizing the opinions expressed in these subsets.The choice of the review subset is treated as a latent variable, predicted by a small and simple selector.The subset is then fed into a more powerful summarizer.For joint training, we use amortized variational inference and policy gradient methods.Our experiments demonstrate the importance of selecting informative reviews resulting in improved quality of summaries and reduced hallucinations.
Arthur Brazinskas, Mirella Lapata, Ivan Titov 0001
EMNLP (1)3
2021 Editing Factual Knowledge in Language Models
abstract
The factual knowledge acquired during pretraining and stored in the parameters of Language Models (LMs) can be useful in downstream tasks (e.g., question answering or textual inference).However, some facts can be incorrectly induced or become obsolete over time.We present KNOWLEDGEEDITOR, a method which can be used to edit this knowledge and, thus, fix 'bugs' or unexpected predictions without the need for expensive retraining or fine-tuning.Besides being computationally efficient, KNOWLEDGEEDITOR does not require any modifications in LM pretraining (e.g., the use of meta-learning).In our approach, we train a hyper-network with constrained optimization to modify a fact without affecting the rest of the knowledge; the trained hyper-network is then used to predict the weight update at test time.We show KNOWL-EDGEEDITOR's efficacy with two popular architectures and knowledge-intensive tasks: i) a BERT model fine-tuned for fact-checking, and ii) a sequence-to-sequence BART model for question answering.With our method, changing a prediction on the specific wording of a query tends to result in a consistent change in predictions also for its paraphrases.We show that this can be further encouraged by exploiting (e.g., automatically-generated) paraphrases during training.Interestingly, our hyper-network can be regarded as a 'probe' revealing which components need to be changed to manipulate factual knowledge; our analysis shows that the updates tend to be concentrated on a small subset of components. 1 How is Namibia's capital city called?Semantically equivalent Answers Scores Namibia Nigeria Nibia Namibia Tasman -0.43 -0.69 -0.89 -1.08 -1.19What is the capital of Namibia?Answers Scores Namibia Nigeria Nibia Tasman Namibia -0.32 -0.79 -0.87 -1.14 -1.16What is the capital of Russia?Answers Scores
Nicola De Cao, Wilker Aziz, Ivan Titov 0001
EMNLP (1)3
2021 Highly Parallel Autoregressive Entity Linking with Discriminative Correction
abstract
Generative approaches have been recently shown to be effective for both Entity Disambiguation and Entity Linking (i.e., joint mention detection and disambiguation).However, the previously proposed autoregressive formulation for EL suffers from i) high computational cost due to a complex (deep) decoder, ii) non-parallelizable decoding that scales with the source sequence length, and iii) the need for training on a large amount of data.In this work, we propose a very efficient approach that parallelizes autoregressive linking across all potential mentions and relies on a shallow and efficient decoder.Moreover, we augment the generative objective with an extra discriminative component, i.e. a correction term which lets us directly optimize the generator's ranking.When taken together, these techniques tackle all the above issues: our model is >70 times faster and more accurate than the previous generative method, outperforming stateof-the-art approaches on the standard English dataset AIDA-CoNLL. 1
Nicola De Cao, Wilker Aziz, Ivan Titov 0001
EMNLP (1)3
2021 A Differentiable Relaxation of Graph Segmentation and Alignment for AMR Parsing
abstract
Meaning Representations (AMR) are a broad-coverage semantic formalism which represents sentence meaning as a directed acyclic graph.To train most AMR parsers, one needs to segment the graph into subgraphs and align each such subgraph to a word in a sentence; this is normally done at preprocessing, relying on hand-crafted rules.In contrast, we treat both alignment and segmentation as latent variables in our model and induce them as part of end-to-end training.As marginalizing over the structured latent variables is infeasible, we use the variational autoencoding framework.To ensure end-to-end differentiable optimization, we introduce a differentiable relaxation of the segmentation and alignment problems.We observe that inducing segmentation yields substantial gains over using a 'greedy' segmentation heuristic.The performance of our method also approaches that of a model that relies on the segmentation rules of Lyu and Titov (2018), which were hand-crafted to handle individual AMR constructions.
Chunchuan Lyu, Shay B. Cohen, Ivan Titov 0001
EMNLP (1)3
2021 Language Modeling, Lexical Translation, Reordering: The Training Process of NMT through the Lens of Classical SMT
abstract
Differently from the traditional statistical MT that decomposes the translation task into distinct separately learned components, neural machine translation uses a single neural network to model the entire translation process.Despite neural machine translation being defacto standard, it is still not clear how NMT models acquire different competences over the course of training, and how this mirrors the different models in traditional SMT.In this work, we look at the competences related to three core SMT components and find that during training, NMT first focuses on learning targetside language modeling, then improves translation quality approaching word-by-word translation, and finally learns more complicated reordering patterns.We show that this behavior holds for several models and language pairs.Additionally, we explain how such an understanding of the training process can be useful in practice and, as an example, show how it can be used to improve vanilla nonautoregressive neural machine translation by guiding teacher model selection.
Elena Voita, Rico Sennrich, Ivan Titov 0001
EMNLP (1)3
2021 Sparse Attention with Linear Units
abstract
Recently, it has been argued that encoderdecoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants.In this work, we introduce a novel, simple method for achieving sparsity in attention: we replace the softmax activation with a ReLU, and show that sparsity naturally emerges from such a formulation.Training stability is achieved with layer normalization with either a specialized initialization or an additional gating function.Our model, which we call Rectified Linear Attention (ReLA), is easy to implement and more efficient than previously proposed sparse attention mechanisms.We apply ReLA to the Transformer and conduct experiments on five machine translation tasks.ReLA achieves translation performance comparable to several strong baselines, with training and decoding speed similar to that of the vanilla attention.Our analysis shows that ReLA delivers high sparsity rate and head diversity, and the induced cross attention achieves better accuracy with respect to source-target word alignment than recent sparsified softmax-based models.Intriguingly, ReLA heads also learn to attend to nothing (i.e.'switch off') for some queries, which is not possible with sparsified softmax alternatives.1
Biao Zhang 0006, Ivan Titov 0001, Rico Sennrich
EMNLP (1)2
2021 Interpreting Graph Neural Networks for NLP With Differentiable Edge Masking
Michael Sejr Schlichtkrull, Nicola De Cao, Ivan Titov 0001
ICLR3
2021 Meta-Learning for Domain Generalization in Semantic Parsing
abstract
The importance of building semantic parsers which can be applied to new domains and generate programs unseen at training has long been acknowledged, and datasets testing outof-domain performance are becoming increasingly available.However, little or no attention has been devoted to learning algorithms or objectives which promote domain generalization, with virtually all existing approaches relying on standard supervised learning.In this work, we use a meta-learning framework which targets zero-shot domain generalization for semantic parsing.We apply a modelagnostic training algorithm that simulates zeroshot parsing by constructing virtual train and test sets from disjoint domains.The learning objective capitalizes on the intuition that gradient steps that improve source-domain performance should also improve target-domain performance, thus encouraging a parser to generalize to unseen target domains.Experimental results on the (English) Spider and Chinese Spider datasets show that the meta-learning objective significantly boosts the performance of a baseline parser.
Bailin Wang, Mirella Lapata, Ivan Titov 0001
NAACL-HLT3
2021 Learning from Executions for Semantic Parsing
abstract
Semantic parsing aims at translating natural language (NL) utterances onto machineinterpretable programs, which can be executed against a real-world environment.The expensive annotation of utterance-program pairs has long been acknowledged as a major bottleneck for the deployment of contemporary neural models to real-life applications.In this work, we focus on the task of semi-supervised learning where a limited amount of annotated data is available together with many unlabeled NL utterances.Based on the observation that programs which correspond to NL utterances must be always executable, we propose to encourage a parser to generate executable programs for unlabeled utterances.Due to the large search space of executable programs, conventional methods that use approximations based on beam-search such as self-training and top-k marginal likelihood training, do not perform as well.Instead, we view the problem of learning from executions from the perspective of posterior regularization and propose a set of new training objectives.Experimental results on OVERNIGHT and GEOQUERY show that our new objectives outperform conventional methods, bridging the gap between semi-supervised and supervised learning.
Bailin Wang, Mirella Lapata, Ivan Titov 0001
NAACL-HLT3
2021 Structured Reordering for Modeling Latent Alignments in Sequence Transduction
abstract
Despite success in many domains, neural models struggle in settings where train and test examples are drawn from different distributions. In particular, in contrast to humans, conventional sequence-to-sequence (seq2seq) models fail to generalize systematically, i.e., interpret sentences representing novel combinations of concepts (e.g., text segments) seen in training. Traditional grammar formalisms excel in such settings by implicitly encoding alignments between input and output segments, but are hard to scale and maintain. Instead of engineering a grammar, we directly model segment-to-segment alignments as discrete structured latent variables within a neural seq2seq model. To efficiently explore the large space of alignments, we introduce a reorder-first align-later framework whose central component is a neural reordering module producing separable permutations. We present an efficient dynamic programming algorithm performing exact marginal inference of separable permutations, and, thus, enabling end-to-end differentiable training of our model. The resulting seq2seq model exhibits better systematic generalization than standard models on synthetic problems and NLP tasks (i.e., semantic parsing and machine translation).
Bailin Wang, Mirella Lapata, Ivan Titov 0001
NeurIPS3
2020 Unsupervised Opinion Summarization as Copycat-Review Generation
abstract
Opinion summarization is the task of automatically creating summaries that reflect subjective information expressed in multiple documents, such as product reviews.While the majority of previous work has focused on the extractive setting, i.e., selecting fragments from input reviews to produce a summary, we let the model generate novel sentences and hence produce abstractive summaries.Recent progress in summarization has seen the development of supervised models which rely on large quantities of document-summary pairs.Since such training data is expensive to acquire, we instead consider the unsupervised setting, in other words, we do not use any summaries in training.We define a generative model for a review collection which capitalizes on the intuition that when generating a new review given a set of other reviews of a product, we should be able to control the "amount of novelty" going into the new review or, equivalently, vary the extent to which it deviates from the input.At test time, when generating summaries, we force the novelty to be minimal, and produce a text reflecting consensus opinions.We capture this intuition by defining a hierarchical variational autoencoder model.Both individual reviews and the products they correspond to are associated with stochastic latent codes, and the review generator ("decoder") has direct access to the text of input reviews through the pointergenerator mechanism.Experiments on Amazon and Yelp datasets, show that setting at test time the review's latent code to its mean, allows the model to produce fluent and coherent summaries reflecting common opinions.
Arthur Brazinskas, Mirella Lapata, Ivan Titov 0001
ACL3
2020 Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
abstract
Massively multilingual models for neural machine translation (NMT) are theoretically attractive, but often underperform bilingual models and deliver poor zero-shot translations.In this paper, we explore ways to improve them.We argue that multilingual NMT requires stronger modeling capacity to support language pairs with varying typological characteristics, and overcome this bottleneck via language-specific components and deepening NMT architectures.We identify the off-target translation issue (i.e.translating into a wrong target language) as the major source of the inferior zero-shot performance, and propose random online backtranslation to enforce the translation of unseen training language pairs.Experiments on OPUS-100 (a novel multilingual dataset with 100 languages) show that our approach substantially narrows the performance gap with bilingual models in both oneto-many and many-to-many settings, and improves zero-shot performance by ∼10 BLEU, approaching conventional pivot-based methods. 1
Biao Zhang 0006, Philip Williams, Ivan Titov 0001, Rico Sennrich
ACL3
2020 Few-Shot Learning for Opinion Summarization
abstract
Opinion summarization is the automatic creation of text reflecting subjective information expressed in multiple documents, such as user reviews of a product.The task is practically important and has attracted a lot of attention.However, due to the high cost of summary production, datasets large enough for training supervised models are lacking.Instead, the task has been traditionally approached with extractive methods that learn to select text fragments in an unsupervised or weakly-supervised way.Recently, it has been shown that abstractive summaries, potentially more fluent and better at reflecting conflicting information, can also be produced in an unsupervised fashion.However, these models, not being exposed to actual summaries, fail to capture their essential properties.In this work, we show that even a handful of summaries is sufficient to bootstrap generation of the summary text with all expected properties, such as writing style, informativeness, fluency, and sentiment preservation.We start by training a conditional Transformer language model to generate a new product review given other available reviews of the product.The model is also conditioned on review properties that are directly related to summaries; the properties are derived from reviews with no manual effort.In the second stage, we fine-tune a plug-in module that learns to predict property values on a handful of summaries.This lets us switch the generator to the summarization mode.We show on Amazon and Yelp datasets that our approach substantially outperforms previous extractive and abstractive methods in automatic and human evaluation.
Arthur Brazinskas, Mirella Lapata, Ivan Titov 0001
EMNLP (1)3
2020 How do Decisions Emerge across Layers in Neural Models? Interpretation with Differentiable Masking
abstract
Attribution methods assess the contribution of inputs to the model prediction.One way to do so is erasure: a subset of inputs is considered irrelevant if it can be removed without affecting the prediction.Though conceptually simple, erasure's objective is intractable and approximate search remains expensive with modern deep NLP models.Erasure is also susceptible to the hindsight bias: the fact that an input can be dropped does not mean that the model 'knows' it can be dropped.The resulting pruning is over-aggressive and does not reflect how the model arrives at the prediction.To deal with these challenges, we introduce Differentiable Masking.DIFFMASK learns to maskout subsets of the input while maintaining differentiability.The decision to include or disregard an input token is made with a simple model based on intermediate hidden layers of the analyzed model.First, this makes the approach efficient because we predict rather than search.Second, as with probing classifiers, this reveals what the network 'knows' at the corresponding layers.This lets us not only plot attribution heatmaps but also analyze how decisions are formed across network layers.We use DIFFMASK to study BERT models on sentiment classification and question answering.1Question: Where did the Broncos practice for the Super Bowl ?
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, Ivan Titov 0001
EMNLP (1)4
2020 Detecting Word Sense Disambiguation Biases in Machine Translation for Model-Agnostic Adversarial Attacks
abstract
Word sense disambiguation is a well-known source of translation errors in NMT.We posit that some of the incorrect disambiguation choices are due to models' over-reliance on dataset artifacts found in training data, specifically superficial word co-occurrences, rather than a deeper understanding of the source text.We introduce a method for the prediction of disambiguation errors based on statistical data properties, demonstrating its effectiveness across several domains and model types.Moreover, we develop a simple adversarial attack strategy that minimally perturbs sentences in order to elicit disambiguation errors to further probe the robustness of translation models.Our findings indicate that disambiguation robustness varies substantially between domains and that different models trained on the same data are vulnerable to different attacks. 1
Denis Emelin, Ivan Titov 0001, Rico Sennrich
EMNLP (1)2
2020 Graph Convolutions over Constituent Trees for Syntax-Aware Semantic Role Labeling
abstract
Semantic role labeling (SRL) is the task of identifying predicates and labeling argument spans with semantic roles.Even though most semantic-role formalisms are built upon constituent syntax, and only syntactic constituents can be labeled as arguments (e.g., FrameNet and PropBank), all the recent work on syntaxaware SRL relies on dependency representations of syntax.In contrast, we show how graph convolutional networks (GCNs) can be used to encode constituent structures and inform an SRL system.Nodes in our SpanGCN correspond to constituents.The computation is done in 3 stages.First, initial node representations are produced by 'composing' word representations of the first and last words in the constituent.Second, graph convolutions relying on the constituent tree are performed, yielding syntacticallyinformed constituent representations.Finally, the constituent representations are 'decomposed' back into word representations, which are used as input to the SRL classifier.We evaluate SpanGCN against alternatives, including a model using GCNs over dependency trees, and show its effectiveness on standard English SRL benchmarks CoNLL-2005, CoNLL-2012, and FrameNet.
Diego Marcheggiani, Ivan Titov 0001
EMNLP (1)2
2020 Information-Theoretic Probing with Minimum Description Length
abstract
To measure how well pretrained representations encode some linguistic property, it is common to use accuracy of a probe, i.e. a classifier trained to predict the property from the representations.Despite widespread adoption of probes, differences in their accuracy fail to adequately reflect differences in representations.For example, they do not substantially favour pretrained representations over randomly initialized ones.Analogously, their accuracy can be similar when probing for genuine linguistic labels and probing for random synthetic tasks.To see reasonable differences in accuracy with respect to these random baselines, previous work had to constrain either the amount of probe training data or its model size.Instead, we propose an alternative to the standard probes, information-theoretic probing with minimum description length (MDL).With MDL probing, training a probe to predict labels is recast as teaching it to effectively transmit the data.Therefore, the measure of interest changes from probe accuracy to the description length of labels given representations.In addition to probe quality, the description length evaluates 'the amount of effort' needed to achieve the quality.This amount of effort characterizes either (i) size of a probing model, or (ii) the amount of data needed to achieve the high quality.We consider two methods for estimating MDL which can be easily implemented on top of the standard probing pipelines: variational coding and online coding.We show that these methods agree in results and are more informative and stable than the standard probes.1
Elena Voita, Ivan Titov 0001
EMNLP (1)2
2020 Visually Grounded Compound PCFGs
abstract
Exploiting visual groundings for language understanding has recently been drawing much attention.In this work, we study visually grounded grammar induction and learn a constituency parser from both unlabeled text and its visual groundings.Existing work on this task (Shi et al., 2019) optimizes a parser via REINFORCE and derives the learning signal only from the alignment of images and sentences.While their model is relatively accurate overall, its error distribution is very uneven, with low performance on certain constituents types (e.g., 26.2% recall on verb phrases, VPs) and high on others (e.g., 79.6% recall on noun phrases, NPs).This is not surprising as the learning signal is likely insufficient for deriving all aspects of phrasestructure syntax and gradient estimates are noisy.We show that using an extension of probabilistic context-free grammar model we can do fully-differentiable end-to-end visually grounded learning.Additionally, this enables us to complement the image-text alignment loss with a language modeling objective.On the MSCOCO test captions, our model establishes a new state of the art, outperforming its non-grounded version and, thus, confirming the effectiveness of visual groundings in constituency grammar induction.It also substantially outperforms the previous grounded model, with largest improvements on more 'abstract' categories (e.g., +55.1% recall on VPs). 1
Yanpeng Zhao, Ivan Titov 0001
EMNLP (1)2
2019 Interpretable Neural Predictions with Differentiable Binary Variables
abstract
The success of neural networks comes hand in hand with a desire for more interpretability.We focus on text classifiers and make them more interpretable by having them provide a justification-a rationale-for their predictions.We approach this problem by jointly training two neural network models: a latent model that selects a rationale (i.e. a short and informative part of the input text), and a classifier that learns from the words in the rationale alone.Previous work proposed to assign binary latent masks to input positions and to promote short selections via sparsityinducing penalties such as L 0 regularisation.We propose a latent model that mixes discrete and continuous behaviour allowing at the same time for binary selections and gradient-based training without REINFORCE.In our formulation, we can tractably compute the expected value of penalties such as L 0 , which allows us to directly optimise the model towards a prespecified text selection rate.We show that our approach is competitive with previous work on rationale extraction, and explore further uses in attention mechanisms.
Jasmijn Bastings, Wilker Aziz, Ivan Titov 0001
ACL (1)3
2019 Learning Latent Trees with Stochastic Perturbations and Differentiable Dynamic Programming
abstract
We treat projective dependency trees as latent variables in our probabilistic model and induce them in such a way as to be beneficial for a downstream task, without relying on any direct tree supervision.Our approach relies on Gumbel perturbations and differentiable dynamic programming.Unlike previous approaches to latent tree learning, we stochastically sample global structures and our parser is fully differentiable.We illustrate its effectiveness on sentiment analysis and natural language inference tasks.We also study its properties on a synthetic structure induction task.Ablation studies emphasize the importance of both stochasticity and constraining latent structures to be projective trees.
Caio F. Corro, Ivan Titov 0001
ACL (1)2
2019 Boosting Entity Linking Performance by Leveraging Unlabeled Documents
abstract
Modern entity linking systems rely on large collections of documents specifically annotated for the task (e.g., AIDA CoNLL).In contrast, we propose an approach which exploits only naturally occurring information: unlabeled documents and Wikipedia.Our approach consists of two stages.First, we construct a high recall list of candidate entities for each mention in an unlabeled document.Second, we use the candidate lists as weak supervision to constrain our document-level entity linking model.The model treats entities as latent variables and, when estimated on a collection of unlabelled texts, learns to choose entities relying both on local context of each mention and on coherence with other entities in the document.The resulting approach rivals fully-supervised state-of-the-art systems on standard test sets.It also approaches their performance in the very challenging setting: when tested on a test set sampled from the data used to estimate the supervised systems.By comparing to Wikipedia-only training of our model, we demonstrate that modeling unlabeled documents is beneficial.
Phong Le, Ivan Titov 0001
ACL (1)2
2019 Distant Learning for Entity Linking with Automatic Noise Detection
abstract
Accurate entity linkers have been produced for domains and languages where annotated data (i.e., texts linked to a knowledge base) is available.However, little progress has been made for the settings where no or very limited amounts of labeled data are present (e.g., legal or most scientific domains).In this work, we show how we can learn to link mentions without having any labeled examples, only a knowledge base and a collection of unannotated texts from the corresponding domain.In order to achieve this, we frame the task as a multi-instance learning problem and rely on surface matching to create initial noisy labels.As the learning signal is weak and our surrogate labels are noisy, we introduce a noise detection component in our model: it lets the model detect and disregard examples which are likely to be noisy.Our method, jointly learning to detect noise and link entities, greatly outperforms the surface matching baseline.For a subset of entity categories, it even approaches the performance of supervised learning.
Phong Le, Ivan Titov 0001
ACL (1)2
2019 When a Good Translation is Wrong in Context: Context-Aware Machine Translation Improves on Deixis, Ellipsis, and Lexical Cohesion
abstract
Though machine translation errors caused by the lack of context beyond one sentence have long been acknowledged, the development of context-aware NMT systems is hampered by several problems.Firstly, standard metrics are not sensitive to improvements in consistency in document-level translations.Secondly, previous work on context-aware NMT assumed that the sentence-aligned parallel data consisted of complete documents while in most practical scenarios such document-level data constitutes only a fraction of the available parallel data.To address the first issue, we perform a human study on an English-Russian subtitles dataset and identify deixis, ellipsis and lexical cohesion as three main sources of inconsistency.We then create test sets targeting these phenomena.To address the second shortcoming, we consider a set-up in which a much larger amount of sentence-level data is available compared to that aligned at the document level.We introduce a model that is suitable for this scenario and demonstrate major gains over a context-agnostic baseline on our new benchmarks without sacrificing performance as measured with BLEU. 1
Elena Voita, Rico Sennrich, Ivan Titov 0001
ACL (1)3
2019 Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
abstract
Multi-head self-attention is a key component of the Transformer, a state-of-the-art architecture for neural machine translation.In this work we evaluate the contribution made by individual attention heads in the encoder to the overall performance of the model and analyze the roles played by them.We find that the most important and confident heads play consistent and often linguistically-interpretable roles.When pruning heads using a method based on stochastic gates and a differentiable relaxation of the L 0 penalty, we observe that specialized heads are last to be pruned.Our novel pruning method removes the vast majority of heads without seriously affecting performance.For example, on the English-Russian WMT dataset, pruning 38 out of 48 encoder heads results in a drop of only 0.15 BLEU. 1
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, Ivan Titov 0001
ACL (1)5
2019 Capturing Argument Interaction in Semantic Role Labeling with Capsule Networks
abstract
Xinchi Chen, Chunchuan Lyu, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xinchi Chen, Chunchuan Lyu, Ivan Titov 0001
EMNLP/IJCNLP (1)3
2019 Semantic Role Labeling with Iterative Structure Refinement
abstract
Chunchuan Lyu, Shay B. Cohen, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chunchuan Lyu, Shay B. Cohen, Ivan Titov 0001
EMNLP/IJCNLP (1)3
2019 Context-Aware Monolingual Repair for Neural Machine Translation
abstract
Elena Voita, Rico Sennrich, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Elena Voita, Rico Sennrich, Ivan Titov 0001
EMNLP/IJCNLP (1)3
2019 The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives
abstract
Elena Voita, Rico Sennrich, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Elena Voita, Rico Sennrich, Ivan Titov 0001
EMNLP/IJCNLP (1)3
2019 Learning Semantic Parsers from Denotations with Latent Structured Alignments and Abstract Programs
abstract
Bailin Wang, Ivan Titov, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Bailin Wang, Ivan Titov 0001, Mirella Lapata
EMNLP/IJCNLP (1)2
2019 Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention
abstract
Biao Zhang, Ivan Titov, Rico Sennrich. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Biao Zhang 0006, Ivan Titov 0001, Rico Sennrich
EMNLP/IJCNLP (1)2
2019 Differentiable Perturb-and-Parse: Semi-Supervised Parsing with a Structured Variational Autoencoder
Caio F. Corro, Ivan Titov 0001
ICLR (Poster)2
2019 Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft
MTSummit (2)3
2019 Block Neural Autoregressive Flow
Nicola De Cao, Wilker Aziz, Ivan Titov 0001
UAI3
2018 AMR Parsing as Graph Prediction with Latent Alignment
abstract
meaning representations (AMRs) are broad-coverage sentence-level semantic representations.AMRs represent sentences as rooted labeled directed acyclic graphs.AMR parsing is challenging partly due to the lack of annotated alignments between nodes in the graphs and words in the corresponding sentences.We introduce a neural parser which treats alignments as latent variables within a joint probabilistic model of concepts, relations and alignments.As exact inference requires marginalizing over alignments and is infeasible, we use the variational autoencoding framework and a continuous relaxation of the discrete alignments.We show that joint modeling is preferable to using a pipeline of align and parse.The parser achieves the best reported results on the standard benchmark (74.4% on LDC2016E25).
Chunchuan Lyu, Ivan Titov 0001
ACL (1)2
2018 Improving Entity Linking by Modeling Latent Relations between Mentions
abstract
Entity linking involves aligning textual mentions of named entities to their corresponding entries in a knowledge base.Entity linking systems often exploit relations between textual mentions in a document (e.g., coreference) to decide if the linking decisions are compatible.Unlike previous approaches, which relied on supervised systems or heuristics to predict these relations, we treat relations as latent variables in our neural entity-linking model.We induce the relations without any supervision while optimizing the entity-linking system in an end-to-end fashion.Our multirelational model achieves the best reported scores on the standard benchmark (AIDA-CoNLL) and substantially outperforms its relation-agnostic version.Its training also converges much faster, suggesting that the injected structural bias helps to explain regularities in the training data.
Phong Le, Ivan Titov 0001
ACL (1)2
2018 Context-Aware Neural Machine Translation Learns Anaphora Resolution
abstract
Standard machine translation systems process sentences in isolation and hence ignore extra-sentential information, even though extended context can both prevent mistakes in ambiguous cases and improve translation coherence.We introduce a context-aware neural machine translation model designed in such way that the flow of information from the extended context to the translation model can be controlled and analyzed.We experiment with an English-Russian subtitles dataset, and observe that much of what is captured by our model deals with improving pronoun translation.We measure correspondences between induced attention distributions and coreference relations and observe that the model implicitly captures anaphora.It is consistent with gains for sentences where pronouns need to be gendered in translation.Beside improvements in anaphoric cases, the model also improves in overall BLEU, both over its context-agnostic version (+0.7) and over simple concatenation of the context and source sentences (+0.6).
Elena Voita, Pavel Serdyukov, Rico Sennrich, Ivan Titov 0001
ACL (1)4
2018 Embedding Words as Distributions with a Bayesian Skip-gram Model
abstract
We introduce a method for embedding words as probability densities in a low-dimensional space. Rather than assuming that a word embedding is fixed across the entire text collection, as in standard word embedding methods, in our Bayesian model we generate it from a word-specific prior density for each occurrence of a given word. Intuitively, for each word, the prior density encodes the distribution of its potential ‘meanings’. These prior densities are conceptually similar to Gaussian embeddings of ėwcitevilnis2014word. Interestingly, unlike the Gaussian embeddings, we can also obtain context-specific densities: they encode uncertainty about the sense of a word given its context and correspond to the approximate posterior distributions within our model. The context-dependent densities have many potential applications: for example, we show that they can be directly used in the lexical substitution task. We describe an effective estimation method based on the variational autoencoding framework. We demonstrate the effectiveness of our embedding technique on a range of standard benchmarks.
Arthur Brazinskas, Serhii Havrylov, Ivan Titov 0001
COLING3
2018 Modeling Relational Data with Graph Convolutional Networks
Michael Sejr Schlichtkrull, Thomas Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov 0001, Max Welling
ESWC5
2017 Optimizing Differentiable Relaxations of Coreference Evaluation Metrics
abstract
Coreference evaluation metrics are hard to optimize directly as they are nondifferentiable functions, not easily decomposable into elementary decisions.Consequently, most approaches optimize objectives only indirectly related to the end goal, resulting in suboptimal performance.Instead, we propose a differentiable relaxation that lends itself to gradient-based optimisation, thus bypassing the need for reinforcement learning or heuristic modification of cross-entropy.We show that by modifying the training objective of a competitive neural coreference system, we obtain a substantial gain in performance.This suggests that our approach can be regarded as a viable alternative to using reinforcement learning or more computationally expensive imitation learning.
Phong Le, Ivan Titov 0001
CoNLL2
2017 A Simple and Accurate Syntax-Agnostic Neural Model for Dependency-based Semantic Role Labeling
abstract
We introduce a simple and accurate neural model for dependency-based semantic role labeling.Our model predicts predicate-argument dependencies relying on states of a bidirectional LSTM encoder.The semantic role labeler achieves competitive performance on English, even without any kind of syntactic information and only using local inference.However, when automatically predicted partof-speech tags are provided as input, it substantially outperforms all previous local models and approaches the best reported results on the English CoNLL-2009 dataset.We also consider Chinese, Czech and Spanish where our approach also achieves competitive results.Syntactic parsers are unreliable on out-of-domain data, so standard (i.e., syntactically-informed) SRL models are hindered when tested in this setting.Our syntax-agnostic model appears more robust, resulting in the best reported results on standard out-of-domain test sets.
Diego Marcheggiani, Anton Frolov 0001, Ivan Titov 0001
CoNLL3
2017 Graph Convolutional Encoders for Syntax-aware Neural Machine Translation
abstract
We present a simple and effective approach to incorporating syntactic structure into neural attention-based encoderdecoder models for machine translation.We rely on graph-convolutional networks (GCNs), a recent class of neural networks developed for modeling graph-structured data.Our GCNs use predicted syntactic dependency trees of source sentences to produce representations of words (i.e.hidden states of the encoder) that are sensitive to their syntactic neighborhoods.GCNs take word representations as input and produce word representations as output, so they can easily be incorporated as layers into standard encoders (e.g., on top of bidirectional RNNs or convolutional neural networks).We evaluate their effectiveness with English-German and English-Czech translation experiments for different types of encoders and observe substantial improvements over their syntax-agnostic versions in all the considered setups.
Jasmijn Bastings, Ivan Titov 0001, Wilker Aziz, Diego Marcheggiani, Khalil Sima'an
EMNLP2
2017 Encoding Sentences with Graph Convolutional Networks for Semantic Role Labeling
abstract
Semantic role labeling (SRL) is the task of identifying the predicate-argument structure of a sentence.It is typically regarded as an important step in the standard NLP pipeline.As the semantic representations are closely related to syntactic ones, we exploit syntactic information in our model.We propose a version of graph convolutional networks (GCNs), a recent class of neural networks operating on graphs, suited to model syntactic dependency graphs.GCNs over syntactic dependency trees are used as sentence encoders, producing latent feature representations of words in a sentence.We observe that GCN layers are complementary to LSTM ones: when we stack both GCN and LSTM layers, we obtain a substantial improvement over an already state-of-theart LSTM SRL model, resulting in the best reported scores on the standard benchmark (CoNLL-2009) both for Chinese and English.
Diego Marcheggiani, Ivan Titov 0001
EMNLP2
2017 Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols
abstract
Learning to communicate through interaction, rather than relying on explicit supervision, is often considered a prerequisite for developing a general AI. We study a setting where two agents engage in playing a referential game and, from scratch, develop a communication protocol necessary to succeed in this game. Unlike previous work, we require that messages they exchange, both at train and test time, are in the form of a language (i.e. sequences of discrete symbols). We compare a reinforcement learning approach and one using a differentiable relaxation (straight-through Gumbel-softmax estimator) and observe that the latter is much faster to converge and it results in more effective protocols. Interestingly, we also observe that the protocol we induce by optimizing the communication success exhibits a degree of compositionality and variability (i.e. the same information can be phrased in different ways), both properties characteristic of natural languages. As the ultimate goal is to ensure that communication is accomplished in natural language, we also perform experiments where we inject prior information about natural language into our model and study properties of the resulting protocol.
Serhii Havrylov, Ivan Titov 0001
NIPS2
2017 Modelling Semantic Expectation: Using Script Knowledge for Referent Prediction
abstract
Recent research in psycholinguistics has provided increasing evidence that humans predict upcoming content. Prediction also affects perception and might be a key to robustness in human language processing. In this paper, we investigate the factors that affect human prediction by building a computational model that can predict upcoming discourse referents based on linguistic knowledge alone vs. linguistic knowledge jointly with common-sense knowledge in the form of scripts. We find that script knowledge significantly improves model estimates of human predictions. In a second study, we test the highly controversial hypothesis that predictability influences referring expression type but do not find evidence for such an effect.
Ashutosh Modi, Ivan Titov 0001, Vera Demberg, Asad B. Sayeed, Manfred Pinkal
Trans. Assoc. Comput. Linguistics2
2016 Bilingual Learning of Multi-sense Embeddings with Discrete Autoencoders
abstract
We present an approach to learning multi-sense word embeddings relying both on monolingual and bilingual information.Our model consists of an encoder, which uses monolingual and bilingual context (i.e. a parallel sentence) to choose a sense for a given word, and a decoder which predicts context words based on the chosen sense.The two components are estimated jointly.We observe that the word representations induced from bilingual data outperform the monolingual counterparts across a range of evaluation tasks, even though crosslingual information is not available at test time.
Simon Suster, Ivan Titov 0001, Gertjan van Noord
HLT-NAACL2
2016 Adapting to All Domains at Once: Rewarding Domain Invariance in SMT
abstract
Existing work on domain adaptation for statistical machine translation has consistently assumed access to a small sample from the test distribution (target domain) at training time. In practice, however, the target domain may not be known at training time or it may change to match user needs. In such situations, it is natural to push the system to make safer choices, giving higher preference to domain-invariant translations, which work well across domains, over risky domain-specific alternatives. We encode this intuition by (1) inducing latent subdomains from the training data only; (2) introducing features which measure how specialized phrases are to individual induced sub-domains; (3) estimating feature weights on out-of-domain data (rather than on the target domain). We conduct experiments on three language pairs and a number of different domains. We observe consistent improvements over a baseline which does not explicitly reward domain invariance.
Cuong Hoang, Khalil Sima'an, Ivan Titov 0001
Trans. Assoc. Comput. Linguistics3
2016 Discrete-State Variational Autoencoders for Joint Discovery and Factorization of Relations
abstract
We present a method for unsupervised open-domain relation discovery. In contrast to previous (mostly generative and agglomerative clustering) approaches, our model relies on rich contextual features and makes minimal independence assumptions. The model is composed of two parts: a feature-rich relation extractor, which predicts a semantic relation between two entities, and a factorization model, which reconstructs arguments (i.e., the entities) relying on the predicted relation. The two components are estimated jointly so as to minimize errors in recovering arguments. We study factorization models inspired by previous work in relation factorization and selectional preference modeling. Our models substantially outperform the generative and agglomerative-clustering counterparts and achieve state-of-the-art performance.
Diego Marcheggiani, Ivan Titov 0001
Trans. Assoc. Comput. Linguistics2
2015 Unsupervised Induction of Semantic Roles within a Reconstruction-Error Minimization Framework
abstract
We introduce a new approach to unsupervised estimation of feature-rich semantic role labeling models.Our model consists of two components: (1) an encoding component: a semantic role labeling model which predicts roles given a rich set of syntactic and lexical features; (2) a reconstruction component: a tensor factorization model which relies on roles to predict argument fillers.When the components are estimated jointly to minimize errors in argument reconstruction, the induced roles largely correspond to roles defined in annotated resources.Our method performs on par with most accurate role induction methods on English and German, even though, unlike these previous approaches, we do not incorporate any prior linguistic knowledge about the languages.
Ivan Titov 0001, Ehsan Khoddam
HLT-NAACL1
2014 Inducing Neural Models of Script Knowledge
abstract
Induction of common sense knowledge about prototypical sequence of events has recently received much attention (e.g., Chambers and Jurafsky (2008); Regneri et al. (2010)). Instead of inducing this knowledge in the form of graphs, as in much of the previous work, in our method, distributed representations of event real-izations are computed based on distributed representations of predicates and their ar-guments, and then these representations are used to predict prototypical event or-derings. The parameters of the composi-tional process for computing the event rep-resentations and the ranking component of the model are jointly estimated. We show that this approach results in a sub-stantial boost in performance on the event ordering task with respect to the previous approaches, both on natural and crowd-sourced texts. 1
Ashutosh Modi, Ivan Titov 0001
CoNLL2
2014 A Hierarchical Bayesian Model for Unsupervised Induction of Script Knowledge
abstract
Scripts representing common sense knowledge about stereotyped sequences of events have been shown to be a valuable resource for NLP applications.We present a hierarchical Bayesian model for unsupervised learning of script knowledge from crowdsourced descriptions of human activities.Events and constraints on event ordering are induced jointly in one unified framework.We use a statistical model over permutations which captures event ordering constraints in a more flexible way than previous approaches.In order to alleviate the sparsity problem caused by using relatively small datasets, we incorporate in our hierarchical model an informed prior on word distributions.The resulting model substantially outperforms a state-of-the-art method on the event ordering task.
Lea Frermann, Ivan Titov 0001, Manfred Pinkal
EACL2
2014 Improved Estimation of Entropy for Evaluation of Word Sense Induction
abstract
Information-theoretic measures are among the most standard techniques for evaluation of clustering methods including word sense induction (WSI) systems. Such measures rely on sample-based estimates of the entropy. However, the standard maximum likelihood estimates of the entropy are heavily biased with the bias dependent on, among other things, the number of clusters and the sample size. This makes the measures unreliable and unfair when the number of clusters produced by different systems vary and the sample size is not exceedingly large. This corresponds exactly to the setting of WSI evaluation where a ground-truth cluster sense number arguably does not exist and the standard evaluation scenarios use a small number of instances of each word to compute the score. We describe more accurate entropy estimators and analyze their performance both in simulations and on evaluation of WSI systems.
Linlin Li 0001, Ivan Titov 0001, Caroline Sporleder
Comput. Linguistics2
2013 Cross-lingual Transfer of Semantic Role Labeling Models
Mikhail Kozhevnikov, Ivan Titov 0001
ACL (1)2
2013 A Bayesian Model for Joint Unsupervised Induction of Sentiment, Aspect and Discourse Representations
Angeliki Lazaridou, Ivan Titov 0001, Caroline Sporleder
ACL (1)2
2013 Predicting the Resolution of Referring Expressions from User Behavior
abstract
We present a statistical model for predicting how the user of an interactive, situated NLP system resolved a referring expression.The model makes an initial prediction based on the meaning of the utterance, and revises it continuously based on the user's behavior.The combined model outperforms its components in predicting reference resolution and when to give feedback.
Nikolaos Engonopoulos, Martin Villalba, Ivan Titov 0001, Alexander Koller
EMNLP3
2013 Translating Video Content to Natural Language Descriptions
abstract
Humans use rich natural language to describe and communicate visual perceptions. In order to provide natural language descriptions for visual content, this paper combines two important ingredients. First, we generate a rich semantic representation of the visual content including e.g. object and activity labels. To predict the semantic representation we learn a CRF to model the relationships between different components of the visual input. And second, we propose to formulate the generation of natural language as a machine translation problem using the semantic representation as source language and the generated sentences as target language. For this we exploit the power of a parallel corpus of videos and textual descriptions and adapt statistical machine translation to translate between our two languages. We evaluate our video descriptions on the TACoS dataset, which contains video snippets aligned with sentence descriptions. Using automatic evaluation and human judgments we show significant improvements over several baseline approaches, motivated by prior work. Our translation approach also shows improvements over related work on an image description task.
Marcus Rohrbach, Ivan Titov 0001, Stefan Thater, Manfred Pinkal, Bernt Schiele
ICCV3
2013 Semantic Role Labeling
Martha Palmer, Ivan Titov 0001, Shumin Wu
HLT-NAACL2
2013 Multilingual Joint Parsing of Syntactic and Semantic Dependencies with a Latent Variable Model
abstract
Current investigations in data-driven models of parsing have shifted from purely syntactic analysis to richer semantic representations, showing that the successful recovery of the meaning of text requires structured analyses of both its grammar and its semantics. In this article, we report on a joint generative history-based model to predict the most likely derivation of a dependency parser for both syntactic and semantic dependencies, in multiple languages. Because these two dependency structures are not isomorphic, we propose a weak synchronization at the level of meaningful subsequences of the two derivations. These synchronized subsequences encompass decisions about the left side of each individual word. We also propose novel derivations for semantic dependency structures, which are appropriate for the relatively unconstrained nature of these graphs. To train a joint model of these synchronized derivations, we make use of a latent variable model of parsing, the Incremental Sigmoid Belief Network (ISBN) architecture. This architecture induces latent feature representations of the derivations, which are used to discover correlations both within and between the two derivations, providing the first application of ISBNs to a multi-task learning problem. This joint model achieves competitive performance on both syntactic and semantic dependency parsing for several languages. Because of the general nature of the approach, this extension of the ISBN architecture to weakly synchronized syntactic-semantic derivations is also an exemplification of its applicability to other problems where two independent, but related, representations are being learned.
James Henderson 0001, Paola Merlo, Ivan Titov 0001, Gabriele Musillo
Comput. Linguistics3
2012 Crosslingual Induction of Semantic Roles
Ivan Titov 0001, Alexandre Klementiev
ACL (1)1
2012 Inducing Crosslingual Distributed Representations of Words
Alexandre Klementiev, Ivan Titov 0001, Binod Bhattarai
COLING2
2012 Semi-Supervised Semantic Role Labeling: Approaching from an Unsupervised Perspective
Ivan Titov 0001, Alexandre Klementiev
COLING1
2012 A Bayesian Approach to Unsupervised Semantic Role Induction
Ivan Titov 0001, Alexandre Klementiev
EACL1
2011 Domain Adaptation by Constraining Inter-Domain Variability of Latent Feature Representation
Ivan Titov 0001
ACL1
2011 A Bayesian Model for Unsupervised Semantic Parsing
Ivan Titov 0001, Alexandre Klementiev
ACL1
2010 Bootstrapping Semantic Analyzers from Non-Contradictory Texts
Ivan Titov 0001, Mikhail Kozhevnikov
ACL1
2010 Multi-document topic segmentation
abstract
Multiple documents describing the same or closely related sets of events are common and often easy to obtain: for example, consider document clusters on a news aggregator site or multiple reviews of the same product or service. Even though each such document discusses a similar set of topics, they provide alternative views or complimentary information on each of these topics. We argue that revealing hidden relations by jointly segmenting the documents, or, equivalently, predicting links between topically related segments in different documents would help to visualize documents of interest and construct friendlier user interfaces. In this paper, we refer to this problem as multi-document topic segmentation. We propose an unsupervised Bayesian model for the considered problem that models both shared and document-specific topics, and utilizes Dirichlet process priors to determine the effective number of topics. We show that topic segmentation can be inferred efficiently using a simple split-merge sampling algorithm. The resulting method outperforms baseline models on four datasets for multi-document topic segmentation.
Minwoo Jeong, Ivan Titov 0001
CIKM2
2010 Incremental Sigmoid Belief Networks for Grammar Learning
James Henderson 0001, Ivan Titov 0001
J. Mach. Learn. Res.2
2009 Unsupervised Rank Aggregation with Domain-Specific Expertise
Alexandre Klementiev, Dan Roth 0001, Kevin Small, Ivan Titov 0001
IJCAI4
2009 Online Graph Planarisation for Synchronous Parsing of Semantic and Syntactic Dependencies
Ivan Titov 0001, James Henderson 0001, Paola Merlo, Gabriele Musillo
IJCAI1
2008 A Joint Model of Text and Aspect Ratings for Sentiment Summarization
Ivan Titov 0001, Ryan T. McDonald
ACL1
2008 A Latent Variable Model of Synchronous Parsing for Syntactic and Semantic Dependencies
James Henderson 0001, Paola Merlo, Gabriele Musillo, Ivan Titov 0001
CoNLL4
2008 Modeling online reviews with multi-grain topic models
abstract
In this paper we present a novel framework for extracting the ratable aspects of objects from online user reviews. Extracting such aspects is an important challenge in automatically mining product opinions from the web and in generating opinion-based summaries of user reviews [18, 19, 7, 12, 27, 36, 21]. Our models are based on extensions to standard topic modeling methods such as LDA and PLSA to induce multi-grain topics. We argue that multi-grain models are more appropriate for our task since standard models tend to produce topics that correspond to global properties of objects (e.g., the brand of a product type) rather than the aspects of an object that tend to be rated by a user. The models we present not only extract ratable aspects, but also cluster them into coherent topics, e.g., 'waitress' and 'bartender' are part of the same topic 'staff' for restaurants. This differentiates it from much of the previous work which extracts aspects through term frequency analysis with minimal clustering. We evaluate the multi-grain models both qualitatively and quantitatively to show that they improve significantly upon standard topic models.
Ivan Titov 0001, Ryan T. McDonald
WWW1
2007 Constituent Parsing with Incremental Sigmoid Belief Networks
Ivan Titov 0001, James Henderson 0001
ACL1
2007 Fast and Robust Multilingual Dependency Parsing with a Generative Latent Variable Model
Ivan Titov 0001, James Henderson 0001
EMNLP-CoNLL1
2007 Incremental Bayesian networks for structure prediction
abstract
We propose a class of graphical models appropriate for structure prediction problems where the model structure is a function of the output structure. Incremental Sigmoid Belief Networks (ISBNs) avoid the need to sum over the possible model structures by using directed arcs and incrementally specifying the model structure. Exact inference in such directed models is not tractable, but we derive two efficient approximations based on mean field methods, which prove effective in artificial experiments. We then demonstrate their effectiveness on a benchmark natural language parsing task, where they achieve state-of-the-art accuracy. Also, the model which is a closer approximation to an ISBN has better parsing accuracy, suggesting that ISBNs are an appropriate abstract model of structure prediction tasks.
Ivan Titov 0001, James Henderson 0001
ICML1
2006 Porting Statistical Parsers with Data-Defined Kernels
Ivan Titov 0001, James Henderson 0001
CoNLL1
2006 Loss Minimization in Parse Reranking
Ivan Titov 0001, James Henderson 0001
EMNLP1
2005 Data-Defined Kernels for Parse Reranking Derived from Probabilistic Models
abstract
Previous research applying kernel methods to natural language parsing have focussed on proposing kernels over parse trees, which are hand-crafted based on domain knowledge and computational considerations. In this paper we propose a method for defining kernels in terms of a probabilistic model of parsing. This model is then trained, so that the parameters of the probabilistic model reflect the generalizations in the training data. The method we propose then uses these trained parameters to define a kernel for reranking parse trees. In experiments, we use a neural network based statistical parser as the probabilistic model, and use the resulting kernel with the Voted Perceptron algorithm to rerank the top 20 parses from the probabilistic model. This method achieves a significant improvement over the accuracy of the probabilistic model.
James Henderson 0001, Ivan Titov 0001
ACL2
2005 Deriving kernels from MLP probability estimators for large categorization problems
abstract
In multi-class categorization problems with a very large or unbounded number of classes, it is often not computationally feasible to train and/or test a kernel-based classifier. One solution is to use a fast computation to pre-select a subset of the classes for reranking with a kernel method, but even then tractability can be a problem. We investigate using trained multilayer perceptron probability estimators to derive appropriate kernels for such problems. We propose a kernel derivation method which is specifically designed for reranking problems, and a more efficient variant of this method which is specifically designed for neural networks with large numbers of output units. When applied to a neural network model of natural language parsing, these new methods achieve state-of-the-art performance which improves over the original model.
Ivan Titov 0001, James Henderson 0001
IJCNN1