Wilker Aziz

dblp:51/10489 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-2093-3866ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Explanation Regularisation through the Lens of Attributions
abstract
Explanation regularisation (ER) has been introduced as a way to guide text classifiers to form their predictions relying on input tokens that humans consider plausible. This is achieved by introducing an auxiliary explanation loss that measures how well the output of an input attribution technique for the model agrees with human-annotated rationales. The guidance appears to benefit performance in out-of-domain (OOD) settings, presumably due to an increased reliance on plausible tokens. However, previous work has under-explored the impact of guidance on that reliance, particularly when reliance is measured using attribution techniques different from those used to guide the model. In this work, we seek to close this gap, and also explore the relationship between reliance on plausible features and OOD performance. We find that the connection between ER and the ability of a classifier to rely on plausible features has been overstated and that a stronger reliance on plausible tokens does not seem to be the cause for OOD improvements.
Ivan Titov 0001, Wilker Aziz
COLING3
2023 What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production Variability
abstract
In Natural Language Generation (NLG) tasks, for any input, multiple communicative goals are plausible, and any goal can be put into words, or produced, in multiple ways.We characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, connecting human production variability to aleatoric or data uncertainty.We then inspect the space of output strings shaped by a generation system's predicted probability distribution and decoding algorithm to probe its uncertainty.For each test input, we measure the generator's calibration to human production variability.Following this instance-level approach, we analyse NLG models and decoding strategies, demonstrating that probing a generator with multiple samples and, when possible, multiple references, provides the level of detail necessary to gain understanding of a model's representation of uncertainty. 1 * Equal contribution. 1 https://github.com/dmg-illc/nlg-uncertainty-probes
Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, Barbara Plank
EMNLP3
2022 GoURMET - Machine Translation for Low-Resourced Languages
abstract
The GoURMET project, funded by the European Commission’s H2020 program (under grant agreement 825299), develops models for machine translation, in particular for low-resourced languages. Data, models and software releases as well as the GoURMET Translate Tool are made available as open source.
Peggy van der Kreeft, Alexandra Birch, Sevi Sariisik, Felipe Sánchez-Martínez, Wilker Aziz
EAMT5
2022 Stop Measuring Calibration When Humans Disagree
abstract
Calibration is a popular framework to evaluate whether a classifier knows when it does not know-i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct.Correctness is commonly estimated against the human majority class.Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies.We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instancelevel measures of calibration that capture key statistical properties of human judgementsclass frequency, ranking and entropy. 1
Joris Baan, Wilker Aziz, Barbara Plank, Raquel Fernández
EMNLP2
2022 Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine Translation
abstract
In NMT we search for the mode of the model distribution to form predictions.The mode and other high-probability translations found by beam search have been shown to often be inadequate in a number of ways.This prevents improving translation quality through better search, as these idiosyncratic translations end up selected by the decoding algorithm, a problem known as the beam search curse.Recently, an approximation to minimum Bayes risk (MBR) decoding has been proposed as an alternative decision rule that would likely not suffer from the same problems.We analyse this approximation and establish that it has no equivalent to the beam search curse.We then design approximations that decouple the cost of exploration from the cost of robust estimation of expected utility.This allows for much larger hypothesis spaces, which we show to be beneficial.We also show that mode-seeking strategies can aid in constructing compact sets of promising hypotheses and that MBR is effective in identifying good translations in them.We conduct experiments on three language pairs varying in amounts of resources available: English into and from German, Romanian, and Nepali. 1
Bryan Eikema, Wilker Aziz
EMNLP2
2022 Sparse Communication via Mixed Distributions
António Farinhas, Wilker Aziz, Vlad Niculae, André F. T. Martins
ICLR2
2021 Editing Factual Knowledge in Language Models
abstract
The factual knowledge acquired during pretraining and stored in the parameters of Language Models (LMs) can be useful in downstream tasks (e.g., question answering or textual inference).However, some facts can be incorrectly induced or become obsolete over time.We present KNOWLEDGEEDITOR, a method which can be used to edit this knowledge and, thus, fix 'bugs' or unexpected predictions without the need for expensive retraining or fine-tuning.Besides being computationally efficient, KNOWLEDGEEDITOR does not require any modifications in LM pretraining (e.g., the use of meta-learning).In our approach, we train a hyper-network with constrained optimization to modify a fact without affecting the rest of the knowledge; the trained hyper-network is then used to predict the weight update at test time.We show KNOWL-EDGEEDITOR's efficacy with two popular architectures and knowledge-intensive tasks: i) a BERT model fine-tuned for fact-checking, and ii) a sequence-to-sequence BART model for question answering.With our method, changing a prediction on the specific wording of a query tends to result in a consistent change in predictions also for its paraphrases.We show that this can be further encouraged by exploiting (e.g., automatically-generated) paraphrases during training.Interestingly, our hyper-network can be regarded as a 'probe' revealing which components need to be changed to manipulate factual knowledge; our analysis shows that the updates tend to be concentrated on a small subset of components. 1 How is Namibia's capital city called?Semantically equivalent Answers Scores Namibia Nigeria Nibia Namibia Tasman -0.43 -0.69 -0.89 -1.08 -1.19What is the capital of Namibia?Answers Scores Namibia Nigeria Nibia Tasman Namibia -0.32 -0.79 -0.87 -1.14 -1.16What is the capital of Russia?Answers Scores
Nicola De Cao, Wilker Aziz, Ivan Titov 0001
EMNLP (1)2
2021 Highly Parallel Autoregressive Entity Linking with Discriminative Correction
abstract
Generative approaches have been recently shown to be effective for both Entity Disambiguation and Entity Linking (i.e., joint mention detection and disambiguation).However, the previously proposed autoregressive formulation for EL suffers from i) high computational cost due to a complex (deep) decoder, ii) non-parallelizable decoding that scales with the source sequence length, and iii) the need for training on a large amount of data.In this work, we propose a very efficient approach that parallelizes autoregressive linking across all potential mentions and relies on a shallow and efficient decoder.Moreover, we augment the generative objective with an extra discriminative component, i.e. a correction term which lets us directly optimize the generator's ranking.When taken together, these techniques tackle all the above issues: our model is >70 times faster and more accurate than the previous generative method, outperforming stateof-the-art approaches on the standard English dataset AIDA-CoNLL. 1
Nicola De Cao, Wilker Aziz, Ivan Titov 0001
EMNLP (1)2
2021 Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months
abstract
In the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English.
Alexandra Birch, Barry Haddow, Antonio Valerio Miceli Barone, Jindrich Helcl, Jonas Waldendorf, Felipe Sánchez-Martínez, Mikel L. Forcada, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Miquel Esplà-Gomis, Wilker Aziz, Lina Murady, Sevi Sariisik, Peggy van der Kreeft, Kay Macquarrie
MTSummit (1)11
2020 Effective Estimation of Deep Generative Language Models
abstract
Advances in variational inference enable parameterisation of probabilistic models by deep neural networks.This combines the statistical transparency of the probabilistic modelling framework with the representational power of deep learning.Yet, due to a problem known as posterior collapse, it is difficult to estimate such models in the context of language modelling effectively.We concentrate on one such model, the variational auto-encoder, which we argue is an important building block in hierarchical probabilistic models of language.This paper contributes a sober view of the problem, a survey of techniques to address it, novel techniques, and extensions to the model.To establish a ranking of techniques, we perform a systematic comparison using Bayesian optimisation and find that many techniques perform reasonably similar, given enough resources.Still, a favourite can be named based on convenience.We also make several empirical observations and recommendations of best practices that should help researchers interested in this exciting field.
Tom Pelsmaeker, Wilker Aziz
ACL2
2020 Is MAP Decoding All You Need? The Inadequacy of the Mode in Neural Machine Translation
abstract
Recent studies have revealed a number of pathologies of neural machine translation (NMT) systems.Hypotheses explaining these mostly suggest there is something fundamentally wrong with NMT as a model or its training algorithm, maximum likelihood estimation (MLE).Most of this evidence was gathered using maximum a posteriori (MAP) decoding, a decision rule aimed at identifying the highest-scoring translation, i.e. the mode.We argue that the evidence corroborates the inadequacy of MAP decoding more than casts doubt on the model and its training algorithm.In this work, we show that translation distributions do reproduce various statistics of the data well, but that beam search strays from such statistics.We show that some of the known pathologies and biases of NMT are due to MAP decoding and not to NMT's statistical assumptions nor MLE.In particular, we show that the most likely translations under the model accumulate so little probability mass that the mode can be considered essentially arbitrary.We therefore advocate for the use of decision rules that take into account the translation distribution holistically.We show that an approximation to minimum Bayes risk decoding gives competitive results confirming that NMT models do capture important aspects of translation well in expectation.
Bryan Eikema, Wilker Aziz
COLING2
2020 How do Decisions Emerge across Layers in Neural Models? Interpretation with Differentiable Masking
abstract
Attribution methods assess the contribution of inputs to the model prediction.One way to do so is erasure: a subset of inputs is considered irrelevant if it can be removed without affecting the prediction.Though conceptually simple, erasure's objective is intractable and approximate search remains expensive with modern deep NLP models.Erasure is also susceptible to the hindsight bias: the fact that an input can be dropped does not mean that the model 'knows' it can be dropped.The resulting pruning is over-aggressive and does not reflect how the model arrives at the prediction.To deal with these challenges, we introduce Differentiable Masking.DIFFMASK learns to maskout subsets of the input while maintaining differentiability.The decision to include or disregard an input token is made with a simple model based on intermediate hidden layers of the analyzed model.First, this makes the approach efficient because we predict rather than search.Second, as with probing classifiers, this reveals what the network 'knows' at the corresponding layers.This lets us not only plot attribution heatmaps but also analyze how decisions are formed across network layers.We use DIFFMASK to study BERT models on sentiment classification and question answering.1Question: Where did the Broncos practice for the Super Bowl ?
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, Ivan Titov 0001
EMNLP (1)3
2020 A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
Duygu Ataman, Wilker Aziz, Alexandra Birch
ICLR2
2020 Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
abstract
Training neural network models with discrete (categorical or structured) latent variables can be computationally challenging, due to the need for marginalization over large or combinatorial sets. To circumvent this issue, one typically resorts to sampling-based approximations of the true marginal, requiring noisy gradient estimators (e.g., score function estimator) or continuous relaxations with lower-variance reparameterized gradients (e.g., Gumbel-Softmax). In this paper, we propose a new training strategy which replaces these estimators by an exact yet efficient marginalization. To achieve this, we parameterize discrete distributions over latent assignments using differentiable sparse mappings: sparsemax and its structured counterparts. In effect, the support of these distributions is greatly reduced, which enables efficient marginalization. We report successful results in three tasks covering a range of latent variable modeling applications: a semisupervised deep generative model, a latent communication game, and a generative model with a bit-vector latent representation. In all cases, we obtain good performance while still achieving the practicality of sampling-based approximations.
Gonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. Martins
NeurIPS3
2019 Interpretable Neural Predictions with Differentiable Binary Variables
abstract
The success of neural networks comes hand in hand with a desire for more interpretability.We focus on text classifiers and make them more interpretable by having them provide a justification-a rationale-for their predictions.We approach this problem by jointly training two neural network models: a latent model that selects a rationale (i.e. a short and informative part of the input text), and a classifier that learns from the words in the rationale alone.Previous work proposed to assign binary latent masks to input positions and to promote short selections via sparsityinducing penalties such as L 0 regularisation.We propose a latent model that mixes discrete and continuous behaviour allowing at the same time for binary selections and gradient-based training without REINFORCE.In our formulation, we can tractably compute the expected value of penalties such as L 0 , which allows us to directly optimise the model towards a prespecified text selection rate.We show that our approach is competitive with previous work on rationale extraction, and explore further uses in attention mechanisms.
Jasmijn Bastings, Wilker Aziz, Ivan Titov 0001
ACL (1)2
2019 Latent Variable Model for Multi-modal Translation
abstract
In this work, we propose to model the interaction between visual and textual features for multi-modal neural machine translation (MMT) through a latent variable model.This latent variable can be seen as a multi-modal stochastic embedding of an image and its description in a foreign language.It is used in a target-language decoder and also to predict image features.Importantly, our model formulation utilises visual and textual inputs during training but does not require that images be available at test time.We show that our latent variable MMT formulation improves considerably over strong baselines, including a multi-task learning approach (Elliott and Kádár, 2017) and a conditional variational auto-encoder approach (Toyama et al., 2016).Finally, we show improvements due to (i) predicting image features in addition to only conditioning on them, (ii) imposing a constraint on the KL term to promote models with nonnegligible mutual information between inputs and latent variable, and (iii) by training on additional target-language image descriptions (i.e.synthetic data).
Iacer Calixto, Miguel Ángel Ríos-Gaona, Wilker Aziz
ACL (1)3
2019 Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft
MTSummit (2)11
2019 Block Neural Autoregressive Flow
Nicola De Cao, Wilker Aziz, Ivan Titov 0001
UAI2
2018 A Stochastic Decoder for Neural Machine Translation
abstract
The process of translation is ambiguous, in that there are typically many valid translations for a given sentence.This gives rise to significant variation in parallel corpora, however, most current models of machine translation do not account for this variation, instead treating the problem as a deterministic process.To this end, we present a deep generative model of machine translation which incorporates a chain of latent variables, in order to account for local lexical and syntactic variation in parallel corpora.We provide an indepth analysis of the pitfalls encountered in variational inference for training deep generative models.Experiments on several different language pairs demonstrate that the model consistently improves over strong baselines.* Code and a workflow that reproduces the experiments are available at https://github.com/philschulz/ stochastic-decoder.
Philip Schulz, Wilker Aziz, Trevor Cohn
ACL (1)2
2018 Deep Generative Model for Joint Alignment and Word Representation
abstract
Miguel Rios, Wilker Aziz, Khalil Sima’an. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Miguel Ángel Ríos-Gaona, Wilker Aziz, Khalil Sima'an
NAACL-HLT2
2017 Graph Convolutional Encoders for Syntax-aware Neural Machine Translation
abstract
We present a simple and effective approach to incorporating syntactic structure into neural attention-based encoderdecoder models for machine translation.We rely on graph-convolutional networks (GCNs), a recent class of neural networks developed for modeling graph-structured data.Our GCNs use predicted syntactic dependency trees of source sentences to produce representations of words (i.e.hidden states of the encoder) that are sensitive to their syntactic neighborhoods.GCNs take word representations as input and produce word representations as output, so they can easily be incorporated as layers into standard encoders (e.g., on top of bidirectional RNNs or convolutional neural networks).We evaluate their effectiveness with English-German and English-Czech translation experiments for different types of encoders and observe substantial improvements over their syntax-agnostic versions in all the considered setups.
Jasmijn Bastings, Ivan Titov 0001, Wilker Aziz, Diego Marcheggiani, Khalil Sima'an
EMNLP3
2016 Fast Collocation-Based Bayesian HMM Word Alignment
abstract
We present a new Bayesian HMM word alignment model for statistical machine translation. The model is a mixture of an alignment model and a language model. The alignment component is a Bayesian extension of the standard HMM. The language model component is responsible for the generation of words needed for source fluency reasons from source language context. This allows for untranslatable source words to remain unaligned and at the same time avoids the introduction of artificial NULL words which introduces unusually long alignment jumps. Existing Bayesian word alignment models are unpractically slow because they consider each target position when resampling a given alignment link. The sampling complexity therefore grows linearly in the target sentence length. In order to make our model useful in practice, we devise an auxiliary variable Gibbs sampler that allows us to resample alignment links in constant time independently of the target sentence length. This leads to considerable speed improvements. Experimental results show that our model performs as well as existing word alignment toolkits in terms of resulting BLEU score.
Philip Schulz, Wilker Aziz
COLING2
2016 The Trouble with Machine Translation Coherence
Karin Sim Smith, Wilker Aziz, Lucia Specia
EAMT2
2016 Cohere: A Toolkit for Local Coherence
Karin Sim Smith, Wilker Aziz, Lucia Specia
LREC2
2015 Quality estimation for asr k-best list rescoring in spoken language translation
abstract
Spoken language translation (SLT) combines automatic speech recognition (ASR) and machine translation (MT). During the decoding stage, the best hypothesis produced by the ASR system may not be the best input candidate to the MT system, but making use of multiple sub-optimal ASR results in SLT has been shown to be too complex computationally. This paper presents a method to rescore the k-best ASR output such as to improve translation quality. A translation quality estimation model is trained on a large number of features which aim to capture complementary information from both ASR and MT on translation difficulty and adequacy, as well as syntactic properties of the SLT inputs and outputs. Based on the predicted quality score, the ASR hypotheses are rescored before they are fed to the MT system. ASR confidence is found to be crucial in guiding the rescoring step. In an English-to-French speech-to-text translation task, the coupling of ASR and MT systems led to an increase of 0.5 BLEU points in translation quality.
Raymond W. M. Ng, Kashif Shah, Wilker Aziz, Lucia Specia, Thomas Hain
ICASSP3
2014 Exact Decoding for Phrase-Based Statistical Machine Translation
abstract
The combinatorial space of translation derivations in phrase-based statistical ma-chine translation is given by the intersec-tion between a translation lattice and a tar-get language model. We replace this in-tractable intersection by a tractable relax-ation which incorporates a low-order up-perbound on the language model. Exact optimisation is achieved through a coarse-to-fine strategy with connections to adap-tive rejection sampling. We perform ex-act optimisation with unpruned language models of order 3 to 5 and show search-error curves for beam search and cube pruning on standard test sets. This is the first work to tractably tackle exact opti-misation with language models of orders higher than 3. 1
Wilker Aziz, Marc Dymetman, Lucia Specia
EMNLP1
2013 Multilingual WSD-like Constraints for Paraphrase Extraction
Wilker Aziz, Lucia Specia
CoNLL1
2012 Cross-lingual Sentence Compression for Subtitles
Wilker Aziz, Sheila C. M. de Sousa, Lucia Specia
EAMT1
2012 PET: a Tool for Post-editing and Assessing Machine Translation
Wilker Aziz, Sheila Castilho, Lucia Specia
LREC1
2011 Predicting Machine Translation Adequacy
Lucia Specia, Najeh Hajlaoui, Catalina Hallett, Wilker Aziz
MTSummit4
2010 Learning an Expert from Human Annotations in Statistical Machine Translation: the Case of Out-of-Vocabulary Words
Wilker Aziz, Marc Dymetman, Lucia Specia, Shachar Mirkin
EAMT1