André F. T. Martins

dblp:m/AndreFTMartins · also André Filipe Torres Martins · DBLP profile ↗
← Back
104ranked-venue papers
23as first author
55since 2021 · last 2026
0000-0001-8282-625XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 99 · 21 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 TOWER+: Bridging Generality and Translation Specialization in Multilingual LLMs
abstract
Ricardo Rei, Nuno M Guerreiro, José Pombal, João Alves, Amin Farajian, Pedro Teixeirinha, Andre Martins. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ricardo Rei, Nuno Miguel Guerreiro, José Pombal, João Alves 0003, M. Amin Farajian, Pedro Teixeirinha, André F. T. Martins
ACL (1)7
2026 Does Speech Translation Meet Users' Needs? An English to Portuguese Study Across Demographics
abstract
This paper introduces Ouvia, a research project to assess user-perceived usability and reliability of modern speech translation tools in En\rightarrowPt scenarios. The project centers on a user study in which we simulate real-life daily interactions by recruiting crowdworkers online from different sociodemographic groups. We collect their spoken requests and self-assessments about quality, satisfaction, and reliability. Here, we describe the project’s motivation and objectives, the study design, and the expected outcomes we will provide to speech translation practitioners.
Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky, Matteo Negri, Marine Carpuat, André F. T. Martins
EAMT (2)6
2026 A Context-aware Framework for Translation-mediated Conversations
abstract
Abstract Automatic translation systems offer a powerful solution to bridge language barriers in scenarios where participants do not share a common language. However, these systems can introduce errors leading to misunderstandings and conversation breakdown. A key issue is that current systems fail to incorporate the rich contextual information necessary to resolve ambiguities and omitted details, resulting in literal, inappropriate, or misaligned translations. In this work, we present a framework to improve large language model-based translation systems by incorporating contextual information in bilingual conversational settings during training and inference. We validate our proposed framework on two task-oriented domains: customer chat and user-assistant interaction. Across both settings, the system produced by our framework—TowerChat—consistently results in better translations than state-of-the-art systems like GPT-4o and TowerInstruct, as measured by multiple automatic translation quality metrics on several language pairs. We also show that the resulting model leverages context in an intended and interpretable way, improving consistency between the conveyed message and the generated translations.1
José Pombal, Sweta Agrawal, Emmanouil Zaranis, Patrick Fernandes, André F. T. Martins
Trans. Assoc. Comput. Linguistics5
2026 Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
abstract
Abstract Reinforcement learning (RL) has been proven to be an effective and robust method for training neural machine translation systems, especially when paired with powerful reward models that accurately assess translation quality. However, most research has focused on RL methods that use sentence-level feedback, leading to inefficient learning signals due to the reward sparsity problem—the model receives a single score for the entire sentence. To address this, we propose a novel approach that leverages fine-grained, token-level quality assessments along with error severity levels using RL methods. Specifically, we use xCOMET, a state-of-the-art quality estimation system, as our token-level reward model. We conduct experiments on small and large translation datasets with standard encoder-decoder and large language models-based machine translation systems, comparing the impact of sentence-level versus fine-grained reward signals on translation quality. Our results show that training with token-level rewards improves translation quality across language pairs over baselines according to both automatic and human evaluation. Furthermore, token-level reward optimization improves training stability, evidenced by a steady increase in mean rewards over training epochs.
Miguel Moura Ramos, Tomás Almeida, Daniel Vareta, Filipe Parrado de Azevedo, Sweta Agrawal, Patrick Fernandes, André F. T. Martins
Trans. Assoc. Comput. Linguistics7
2025 Did Translation Models Get More Robust Without Anyone Even Noticing?
abstract
Neural machine translation (MT) models achieve strong results across a variety of settings, but it is widely believed that they are highly sensitive to "noisy" inputs, such as spelling errors, abbreviations, and other formatting issues.In this paper, we revisit this insight in light of recent multilingual MT models and large language models (LLMs) applied to machine translation.Somewhat surprisingly, we show through controlled experiments that these models are far more robust to many kinds of noise than previous models, even when they perform similarly on clean data.This is notable because, even though LLMs have more parameters and more complex training processes than past models, none of the open ones we consider use any techniques specifically designed to encourage robustness.Next, we show that similar trends hold for social media translation experiments -LLMs are more robust to social media text.We include an analysis of the circumstances in which source correction techniques can be used to mitigate the effects of noise.Altogether, we show that robustness to many types of noise has increased.
Ben Peters, André F. T. Martins
ACL (1)2
2025 Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
abstract
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker
ACL (1)17
2025 Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
abstract
Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding.While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked.Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability.This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT).Experiments with state-ofthe-art QE metrics across multiple domains, datasets, and languages reveal significant bias.When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized.Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones.Moreover, a biased QE metric affects data filtering and quality-aware decoding.Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender. 1
Emmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. Martins
ACL (1)4
2025 Sparse Activations as Conformal Predictors
abstract
Conformal prediction is a distribution-free framework for uncertainty quantification that replaces point predictions with sets, offering marginal coverage guarantees (i.e., ensuring that the sets contain the true label with a specified probability, in expectation). In this paper, we uncover a novel connection between conformal prediction and sparse "softmax-like" transformations, such as sparsemax and $\gamma$-entmax (with $\gamma> 1$), which assign nonzero probability only to some labels. We introduce new non-conformity scores for classification that make the calibration process correspond to the widely used temperature scaling method. At test time, applying these sparse transformations with the calibrated temperature leads to a support set (i.e., the set of labels with nonzero probability) that automatically inherits the coverage guarantees of conformal prediction. Through experiments on computer vision and text classification benchmarks, we demonstrate that the proposed method achieves competitive results in terms of coverage, efficiency, and adaptiveness compared to standard non-conformity scores based on softmax.
Margarida M. Campos, João Cálem, Sophia Sklaviadis, Mário A. T. Figueiredo, André F. T. Martins
AISTATS5
2025 An Interdisciplinary Approach to Human-Centered Machine Translation
abstract
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon
EMNLP15
2025 Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
abstract
Larger models often outperform smaller ones but come with high computational costs.Cascading offers a potential solution.By default, it uses smaller models and defers only some instances to larger, more powerful models.However, designing effective deferral rules remains a challenge.In this paper, we propose a simple yet effective approach for machine translation, using existing quality estimation (QE) metrics as deferral rules.We show that QE-based deferral allows a cascaded system to match the performance of a larger model while invoking it for a small fraction (30% to 50%) of the examples, significantly reducing computational costs.We validate this approach through both automatic and human evaluation.
António Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei, André F. T. Martins
EMNLP5
2025 AdaSplash: Adaptive Sparse Flash Attention
abstract
The computational cost of softmax-based attention in transformers limits their applicability to long-context tasks. Adaptive sparsity, of which $\alpha$-entmax attention is an example, offers a flexible data-dependent alternative, but existing implementations are inefficient and do not leverage the sparsity to obtain runtime and memory gains. In this work, we propose AdaSplash, which combines the efficiency of GPU-optimized algorithms with the sparsity benefits of $\alpha$-entmax. We first introduce a hybrid Halley-bisection algorithm, resulting in a 7-fold reduction in the number of iterations needed to compute the $\alpha$-entmax transformation. Then, we implement custom Triton kernels to efficiently handle adaptive sparsity. Experiments with RoBERTa and ModernBERT for text classification and single-vector retrieval, along with GPT-2 for language modeling, show that our method achieves substantial improvements in runtime and memory efficiency compared to existing $\alpha$-entmax implementations. It approaches—and in some cases surpasses—the efficiency of highly optimized softmax implementations like FlashAttention-2, enabling long-context training while maintaining strong task performance.
Marcos V. Treviso, André F. T. Martins
ICML3
2025 ∞-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation
abstract
Current video-language models struggle with long-video understanding due to limited context lengths and reliance on sparse frame subsampling, which often leads to information loss. In this paper, we introduce $\infty$-Video, which is able to process arbitrarily long videos through a continuous-time long-term memory (LTM) consolidation mechanism. Our framework augments video Q-formers by making them able to process unbounded video contexts efficiently and without requiring additional training. Through continuous attention, our approach dynamically allocates higher granularity to the most relevant video segments, forming "sticky" memories which evolve over time. Experiments with Video-LLaMA and VideoChat2 demonstrate improved performance in video question-answering tasks, showcasing the potential of continuous-time LTM mechanisms to enable scalable and training-free comprehension of long videos.
Saul José Rodrigues dos Santos, António Farinhas, Daniel C. McNamee, André F. T. Martins
ICML4
2025 Hopfield-Fenchel-Young Networks: A Unified Framework for Associative Memory Retrieval
abstract
Associative memory models, such as Hopfield networks and their modern variants, have garnered renewed interest due to advancements in memory capacity and connections with self-attention in transformers. In this work, we introduce a unified framework-Hopfield-Fenchel-Young networks-which generalizes these models to a broader family of energy functions. Our energies are formulated as the difference between two Fenchel-Young losses: one, parameterized by a generalized entropy, defines the Hopfield scoring mechanism, while the other applies a post-transformation to the Hopfield output. By utilizing Tsallis and norm entropies, we derive end-to-end differentiable update rules that enable sparse transformations, uncovering new connections between loss margins, sparsity, and exact retrieval of single memory patterns. We further extend this framework to structured Hopfield networks using the SparseMAP transformation, allowing the retrieval of pattern associations rather than a single pattern. Our framework unifies and extends traditional and modern Hopfield networks and provides an energy minimization perspective for widely used post-transformations like $\ell_2$-normalization and layer normalization-all through suitable choices of Fenchel-Young losses and by using convex analysis as a building block. Finally, we validate our Hopfield-Fenchel-Young networks on diverse memory recall tasks, including free and sequential recall. Experiments on simulated data, image retrieval, multiple instance learning, and text rationalization demonstrate the effectiveness of our approach.
Saul José Rodrigues dos Santos, Vlad Niculae, Daniel C. McNamee, André F. T. Martins
J. Mach. Learn. Res.4
2025 Adding Chocolate to Mint : Mitigating Metric Interference in Machine Translation
abstract
Abstract As automatic metrics become increasingly stronger and widely adopted, the risk of unintentionally “gaming the metric” during model development rises. This issue is caused by metric interference (Mint), i.e., the use of the same or related metrics for both model tuning and evaluation. Mint can misguide practitioners into being overoptimistic about the performance of their systems: As system outputs become a function of the interfering metric, their estimated quality loses correlation with human judgments. In this work, we analyze two common cases of Mint in machine translation-related tasks: Filtering of training data, and decoding with quality signals. Importantly, we find that Mint strongly distorts instance-level metric scores, even when metrics are not directly optimized for—questioning the common strategy of leveraging a different, yet related metric for evaluation that is not used for tuning. To address this problem, we propose MintAdjust, a method for more reliable evaluation under Mint. On the WMT24 MT shared task test set, MintAdjust ranks translations and systems more accurately than state-of-the-art-metrics across a majority of language pairs, especially for high-quality systems. Furthermore, MintAdjust outperforms AutoRank, the ensembling method used by the organizers.1 We will release a codebase for replicating the results in this work upon publication.
José Pombal, Nuno Miguel Guerreiro, Ricardo Rei, André F. T. Martins
Trans. Assoc. Comput. Linguistics4
2024 The Center for Responsible AI Project
abstract
This paper describes the project “NextGenAI: Center for Responsible AI”, a 39-month Mobilizing and Green Agenda for Business Innovation funded by the Portuguese Recovery and Resilience Plan, under the Recovery and Resilience Facility (RRF). The project aims to create a new Center for Responsible AI in Portugal, capable of delivering more than 20 AI products in crucial areas like “Life Sciences”, many of which use generative AI, particularly NLP models such as those for Machine Translation, contributing to translating into legislation the European Law included in the EU AI Act, and creating a critical mass in the development of responsible AI technologies. To accomplish this mission, the Center for Responsible AI is formed by an ecosystem of startups and research institutions driving research in a virtuous way by addressing real market needs and opportunities in Responsible AI.
Maria Ana Henriques, Ana C. Farinha, Nuno André, António Novais, Sara Guerreiro de Sousa, Bruno Prezado Silva, Helena Moniz, André F. T. Martins, Paulo Dimas
EAMT (2)9
2024 Aligning Neural Machine Translation Models: Human Feedback in Training and Inference
abstract
Reinforcement learning from human feedback (RLHF) is a recent technique to improve the quality of the text generated by a language model, making it closer to what humans would generate.A core ingredient in RLHF’s success in aligning and improving large language models (LLMs) is its \textit{reward model}, trained using human feedback on model outputs. In machine translation (MT), where metrics trained from human annotations can readily be used as reward models, recent methods using \textit{minimum Bayes risk} decoding and reranking have succeeded in improving the final quality of translation.In this study, we comprehensively explore and compare techniques for integrating quality metrics as reward models into the MT pipeline. This includes using the reward model for data filtering, during the training phase through RL, and at inference time by employing reranking techniques, and we assess the effects of combining these in a unified approach.Our experimental results, conducted across multiple translation tasks, underscore the crucial role of effective data filtering, based on estimated quality, in harnessing the full potential of RL in enhancing MT quality.Furthermore, our findings demonstrate the effectiveness of combining RL training with reranking techniques, showcasing substantial improvements in translation quality.
Miguel Moura Ramos, Patrick Fernandes, António Farinhas, André F. T. Martins
EAMT (1)4
2024 Can Automatic Metrics Assess High-Quality Translations?
abstract
Automatic metrics for evaluating translation quality are typically validated by measuring how well they correlate with human assessments.However, correlation methods tend to capture only the ability of metrics to differentiate between good and bad source-translation pairs, overlooking their reliability in distinguishing alternative translations for the same source.In this paper, we confirm that this is indeed the case by showing that current metrics are insensitive to nuanced differences in translation quality.This effect is most pronounced when the quality is high and the variance among alternatives is low.Given this finding, we shift towards detecting high-quality correct translations, an important problem in practical decision-making scenarios where a binary check of correctness is prioritized over a nuanced evaluation of quality.Using the MQM framework as the gold standard, we systematically stress-test the ability of current metrics to identify translations with no errors as marked by humans.Our findings reveal that current metrics often over or underestimate translation quality, indicating significant room for improvement in machine translation evaluation.
Sweta Agrawal, António Farinhas, Ricardo Rei, André F. T. Martins
EMNLP4
2024 Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
abstract
Sweta Agrawal, José G. C. De Souza, Ricardo Rei, António Farinhas, Gonçalo Faria, Patrick Fernandes, Nuno M Guerreiro, Andre Martins. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Sweta Agrawal, José Guilherme Camargo de Souza, Ricardo Rei, António Farinhas, Gonçalo Rui Alves Faria, Patrick Fernandes, Nuno Miguel Guerreiro, André F. T. Martins
EMNLP8
2024 Non-Exchangeable Conformal Risk Control
abstract
Split conformal prediction has recently sparked great interest due to its ability to provide formally guaranteed uncertainty sets or intervals for predictions made by black-box neural models, ensuring a predefined probability of containing the actual ground truth. While the original formulation assumes data exchangeability, some extensions handle non-exchangeable data, which is often the case in many real-world scenarios. In parallel, some progress has been made in conformal methods that provide statistical guarantees for a broader range of objectives, such as bounding the best $F_1$-score or minimizing the false negative rate in expectation. In this paper, we leverage and extend these two lines of work by proposing non-exchangeable conformal risk control, which allows controlling the expected value of any monotone loss function when the data is not exchangeable. Our framework is flexible, makes very few assumptions, and allows weighting the data based on its relevance for a given test example; a careful choice of weights may result in tighter bounds, making our framework useful in the presence of change points, time series, or other forms of distribution drift. Experiments with both synthetic and real world data show the usefulness of our method.
António Farinhas, Chrysoula Zerva, Dennis Ulmer, André F. T. Martins
ICLR4
2024 Sparse and Structured Hopfield Networks
abstract
Modern Hopfield networks have enjoyed recent interest due to their connection to attention in transformers. Our paper provides a unified framework for sparse Hopfield networks by establishing a link with Fenchel-Young losses. The result is a new family of Hopfield-Fenchel-Young energies whose update rules are end-to-end differentiable sparse transformations. We reveal a connection between loss margins, sparsity, and exact memory retrieval. We further extend this framework to structured Hopfield networks via the SparseMAP transformation, which can retrieve pattern associations instead of a single pattern. Experiments on multiple instance learning and text rationalization demonstrate the usefulness of our approach.
Saul José Rodrigues dos Santos, Vlad Niculae, Daniel C. McNamee, André F. T. Martins
ICML4
2024 QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
abstract
An important challenge in machine translation (MT) is to generate high-quality and diverse translations. Prior work has shown that the estimated likelihood from the MT model correlates poorly with translation quality. In contrast, quality evaluation metrics (such as COMET or BLEURT) exhibit high correlations with human judgments, which has motivated their use as rerankers (such as quality-aware and minimum Bayes risk decoding). However, relying on a single translation with high estimated quality increases the chances of "gaming the metric''. In this paper, we address the problem of sampling a set of high-quality and diverse translations. We provide a simple and effective way to avoid over-reliance on noisy quality estimates by using them as the energy function of a Gibbs distribution. Instead of looking for a mode in the distribution, we generate multiple samples from high-density areas through the Metropolis-Hastings algorithm, a simple Markov chain Monte Carlo approach. The results show that our proposed method leads to high-quality and diverse outputs across multiple language pairs (English$\leftrightarrow$\{German, Russian\}) with two strong decoder-only LLMs (Alma-7b, Tower-7b).
Gonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, André F. T. Martins
NeurIPS6
2024 Reranking Laws for Language Generation: A Communication-Theoretic Perspective
abstract
To ensure large language models (LLMs) are used safely, one must reduce their propensity to hallucinate or to generate unacceptable answers. A simple and often used strategy is to first let the LLM generate multiple hypotheses and then employ a reranker to choose the best one. In this paper, we draw a parallel between this strategy and the use of redundancy to decrease the error rate in noisy communication channels. We conceptualize the generator as a sender transmitting multiple descriptions of a message through parallel noisy channels. The receiver decodes the message by ranking the (potentially corrupted) descriptions and selecting the one found to be most reliable. We provide conditions under which this protocol is asymptotically error-free (i.e., yields an acceptable answer almost surely) even in scenarios where the reranker is imperfect (governed by Mallows or Zipf-Mandelbrot models) and the channel distributions are statistically dependent. We use our framework to obtain reranking laws which we validate empirically on two real-world tasks using LLMs: text-to-code generation with DeepSeek-Coder 7B and machine translation of medical data with TowerInstruct 13B.
António Farinhas, Haau-Sing Li, André F. T. Martins
NeurIPS3
2024 Assessing the Role of Context in Chat Translation Evaluation: Is Context Helpful and Under What Conditions?
abstract
Abstract Despite the recent success of automatic metrics for assessing translation quality, their application in evaluating the quality of machine-translated chats has been limited. Unlike more structured texts like news, chat conversations are often unstructured, short, and heavily reliant on contextual information. This poses questions about the reliability of existing sentence-level metrics in this domain as well as the role of context in assessing the translation quality. Motivated by this, we conduct a meta-evaluation of existing automatic metrics, primarily designed for structured domains such as news, to assess the quality of machine-translated chats. We find that reference-free metrics lag behind reference-based ones, especially when evaluating translation quality in out-of-English settings. We then investigate how incorporating conversational contextual information in these metrics for sentence-level evaluation affects their performance. Our findings show that augmenting neural learned metrics with contextual information helps improve correlation with human judgments in the reference-free scenario and when evaluating translations in out-of-English settings. Finally, we propose a new evaluation metric, Context-MQM, that utilizes bilingual context with a large language model (LLM) and further validate that adding context helps even for LLM-based evaluation metrics.
Sweta Agrawal, M. Amin Farajian, Patrick Fernandes, Ricardo Rei, André F. T. Martins
Trans. Assoc. Comput. Linguistics5
2024 Conformal Prediction for Natural Language Processing: A Survey
abstract
Abstract The rapid proliferation of large language models and natural language processing (NLP) applications creates a crucial need for uncertainty quantification to mitigate risks such as Hallucinations and to enhance decision-making reliability in critical applications. Conformal prediction is emerging as a theoretically sound and practically useful framework, combining flexibility with strong statistical guarantees. Its model-agnostic and distribution-free nature makes it particularly promising to address the current shortcomings of NLP systems that stem from the absence of uncertainty quantification. This paper provides a comprehensive survey of conformal prediction techniques, their guarantees, and existing applications in NLP, pointing to directions for future research and open challenges.
Margarida M. Campos, António Farinhas, Chrysoula Zerva, Mário A. T. Figueiredo, André F. T. Martins
Trans. Assoc. Comput. Linguistics5
2024 xcomet : Transparent Machine Translation Evaluation through Fine-grained Error Detection
abstract
Abstract Widely used learned metrics for machine translation evaluation, such as Comet and Bleurt, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xcomet, an open-source learned metric designed to bridge the gap between these approaches. xcomet integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xcomet is largely capable of identifying localized critical errors and hallucinations.
Nuno Miguel Guerreiro, Ricardo Rei, Daan van Stigt, Luísa Coheur, Pierre Colombo, André F. T. Martins
Trans. Assoc. Comput. Linguistics6
2024 Conformalizing Machine Translation Evaluation
abstract
Abstract Several uncertainty estimation methods have been recently proposed for machine translation evaluation. While these methods can provide a useful indication of when not to trust model predictions, we show in this paper that the majority of them tend to underestimate model uncertainty, and as a result, they often produce misleading confidence intervals that do not cover the ground truth. We propose as an alternative the use of conformal prediction, a distribution-free method to obtain confidence intervals with a theoretically established guarantee on coverage. First, we demonstrate that split conformal prediction can “correct” the confidence intervals of previous methods to yield a desired coverage level, and we demonstrate these findings across multiple machine translation evaluation metrics and uncertainty quantification methods. Further, we highlight biases in estimated confidence intervals, reflected in imbalanced coverage for different attributes, such as the language and the quality of translations. We address this by applying conditional conformal prediction techniques to obtain calibration subsets for each data subgroup, leading to equalized coverage. Overall, we show that, provided access to a calibration set, conformal prediction can help identify the most suitable uncertainty quantification methods and adapt the predicted confidence intervals to ensure fairness with respect to different attributes.1
Chrysoula Zerva, André F. T. Martins
Trans. Assoc. Comput. Linguistics2
2023 When Does Translation Require Context? A Data-driven, Multilingual Exploration
abstract
Although proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics.Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way.In this paper, we develop the Multilingual Discourse-Aware (MUDA) benchmark, a series of taggers that identify and evaluate model performance on discourse phenomena in any given dataset.The choice of phenomena is inspired by a novel methodology to systematically identify translations requiring context.We confirm the difficulty of previously studied phenomena while uncovering others that were previously unaddressed.We find that common context-aware MT models make only marginal improvements over context-agnostic models, which suggests these models do not handle these ambiguities effectively.We release code and data for 14 language pairs to encourage the MT community to focus on accurately capturing discourse phenomena.1
Patrick Fernandes, Kayo Yin, Emmy Liu, André F. T. Martins, Graham Neubig
ACL (1)4
2023 Optimal Transport for Unsupervised Hallucination Detection in Neural Machine Translation
abstract
Neural machine translation (NMT) has become the de-facto standard in real-world machine translation applications.However, NMT models can unpredictably produce severely pathological translations, known as hallucinations, that seriously undermine user trust.It becomes thus crucial to implement effective preventive strategies to guarantee their proper functioning.In this paper, we address the problem of hallucination detection in NMT by following a simple intuition: as hallucinations are detached from the source content, they exhibit cross-attention patterns that are statistically different from those of good quality translations.We frame this problem with an optimal transport formulation and propose a fully unsupervised, plug-in detector that can be used with any attention-based NMT model.Experimental results show that our detector not only outperforms all previous model-based detectors, but is also competitive with detectors that employ external models trained on millions of samples for related tasks such as quality estimation and cross-lingual sentence similarity.
Nuno Miguel Guerreiro, Pierre Colombo, Pablo Piantanida, André F. T. Martins
ACL (1)4
2023 Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
abstract
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André Martins, François Yvon, Hinrich Schütze. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André F. T. Martins, François Yvon, Hinrich Schütze
ACL (1)9
2023 Python Code Generation by Asking Clarification Questions
abstract
Code generation from text requires understanding the user's intent from a natural language description and generating an executable code snippet that satisfies this intent.While recent pretrained language models demonstrate remarkable performance for this task, these models fail when the given natural language description is under-specified.In this work, we introduce a novel and more realistic setup for this task.We hypothesize that the underspecification of a natural language description can be resolved by asking clarification questions.Therefore, we collect and introduce a new dataset named CodeClarQA containing pairs of natural language descriptions and code with created synthetic clarification questions and answers.The empirical results of our evaluation of pretrained language model performance on code generation show that clarifications result in more precisely generated code, as shown by the substantial improvement of model performance in all evaluation metrics.Alongside this, our task and dataset introduce new challenges to the community, including when and what clarification questions should be asked.Our code and dataset are available on GitHub.1
Haau-Sing Li, Mohsen Mesgar, André F. T. Martins, Iryna Gurevych
ACL (1)3
2023 CREST: A Joint Framework for Rationalization and Counterfactual Text Generation
abstract
Selective rationales and counterfactual examples have emerged as two effective, complementary classes of interpretability methods for analyzing and training NLP models.However, prior work has not explored how these methods can be integrated to combine their complementary advantages.We overcome this limitation by introducing CREST (ContRastive Edits with Sparse raTionalization), a joint framework for selective rationalization and counterfactual text generation, and show that this framework leads to improvements in counterfactual quality, model robustness, and interpretability.First, CREST generates valid counterfactuals that are more natural than those produced by previous methods, and subsequently can be used for data augmentation at scale, reducing the need for human-generated examples.Second, we introduce a new loss function that leverages CREST counterfactuals to regularize selective rationales and show that this regularization improves both model robustness and rationale quality, compared to methods that do not leverage CREST counterfactuals.Our results demonstrate that CREST successfully bridges the gap between selective rationales and counterfactual examples, addressing the limitations of existing methods and providing a more comprehensive view of a model's predictions.
Marcos V. Treviso, Alexis Ross, Nuno Miguel Guerreiro, André F. T. Martins
ACL (1)4
2023 Looking for a Needle in a Haystack: A Comprehensive Study of Hallucinations in Neural Machine Translation
abstract
Although the problem of hallucinations in neural machine translation (NMT) has received some attention, research on this highly pathological phenomenon lacks solid ground.Previous work has been limited in several ways: it often resorts to artificial settings where the problem is amplified, it disregards some (common) types of hallucinations, and it does not validate adequacy of detection heuristics.In this paper, we set foundations for the study of NMT hallucinations.First, we work in a natural setting, i.e., in-domain data without artificial noise neither in training nor in inference.Next, we annotate a dataset of over 3.4k sentences indicating different kinds of critical errors and hallucinations.Then, we turn to detection methods and both revisit methods used previously and propose using glass-box uncertainty-based detectors.Overall, we show that for preventive settings, (i) previously used methods are largely inadequate, (ii) sequence log-probability works best and performs on par with reference-based methods.Finally, we propose DEHALLUCINATOR, a simple method for alleviating hallucinations at test time which significantly reduces the hallucinatory rate.
Nuno Miguel Guerreiro, Elena Voita, André F. T. Martins
EACL3
2023 BLEU Meets COMET: Combining Lexical and Neural Metrics Towards Robust Machine Translation Evaluation
abstract
Although neural-based machine translation evaluation metrics, such as COMET or BLEURT, have achieved strong correlations with human judgements, they are sometimes unreliable in detecting certain phenomena that can be considered as critical errors, such as deviations in entities and numbers. In contrast, traditional evaluation metrics such as BLEU or chrF, which measure lexical or character overlap between translation hypotheses and human references, have lower correlations with human judgements but are sensitive to such deviations. In this paper, we investigate several ways of combining the two approaches in order to increase robustness of state-of-the-art evaluation methods to translations with critical errors. We show that by using additional information during training, such as sentence-level features and word-level tags, the trained metrics improve their capability to penalize translations with specific troublesome phenomena, which leads to gains in correlations with humans and on the recent DEMETR benchmark on several language pairs.
Taisiya Glushkova, Chrysoula Zerva, André F. T. Martins
EAMT3
2023 Empirical Assessment of kNN-MT for Real-World Translation Scenarios
abstract
This paper aims to investigate the effectiveness of the k-Nearest Neighbor Machine Translation model (kNN-MT) in real-world scenarios. kNN-MT is a retrieval-augmented framework that combines the advantages of parametric models with non-parametric datastores built using a set of parallel sentences. Previous studies have primarily focused on evaluating the model using only the BLEU metric and have not tested kNN-MT in real world scenarios. Our study aims to fill this gap by conducting a comprehensive analysis on various datasets comprising different language pairs and different domains, using multiple automatic metrics and expert evaluated Multidimensional Quality Metrics (MQM). We compare kNN-MT with two alternate strategies: fine-tuning all the model parameters and adapter-based finetuning. Finally, we analyze the effect of the datastore size on translation quality, and we examine the number of entries necessary to bootstrap and configure the index.
Pedro Henrique Martins, João Alves 0003, Tânia Vaz, Madalena Gonçalves, Beatriz Silva, Marianna Buchicchio, José Guilherme Camargo de Souza, André F. T. Martins
EAMT8
2023 An Empirical Study of Translation Hypothesis Ensembling with Large Language Models
abstract
Large language models (LLMs) are becoming a one-fits-many solution, but they sometimes hallucinate or produce unreliable output.In this paper, we investigate how hypothesis ensembling can improve the quality of the generated text for the specific problem of LLM-based machine translation.We experiment with several techniques for ensembling hypotheses produced by LLMs such as ChatGPT, LLaMA, and Alpaca.We provide a comprehensive study along multiple dimensions, including the method to generate hypotheses (multiple prompts, temperaturebased sampling, and beam search) and the strategy to produce the final translation (instructionbased, quality-based reranking, and minimum Bayes risk (MBR) decoding).Our results show that MBR decoding is a very effective method, that translation quality can be improved using a small number of samples, and that instruction tuning has a strong impact on the relation between the diversity of the hypotheses and the sampling temperature.Our code is available at https://github.com/deep-spin/ translation-hypothesis-ensembling.
António Farinhas, José Guilherme Camargo de Souza, André F. T. Martins
EMNLP3
2023 Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
abstract
Abstract Natural language generation has witnessed significant advancements due to the training of large language models on vast internet-scale datasets. Despite these advancements, there exists a critical challenge: These models can inadvertently generate content that is toxic, inaccurate, and unhelpful, and existing automatic evaluation metrics often fall short of identifying these shortcomings. As models become more capable, human feedback is an invaluable signal for evaluating and improving models. This survey aims to provide an overview of recent research that has leveraged human feedback to improve natural language generation. First, we introduce a taxonomy distilled from existing research to categorize and organize the varied forms of feedback. Next, we discuss how feedback can be described by its format and objective, and cover the two approaches proposed to use feedback (either for training or decoding): directly using feedback or training feedback models. We also discuss existing datasets for human-feedback data collection, and concerns surrounding feedback collection. Finally, we provide an overview of the nascent field of AI feedback, which uses large language models to make judgments based on a set of principles and minimize the need for human intervention. We also release a website of this survey at feedback-gap-survey.info.
Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José Guilherme Camargo de Souza, Shuyan Zhou, Sherry Tongshuang Wu, Graham Neubig, André F. T. Martins
Trans. Assoc. Comput. Linguistics11
2023 Hallucinations in Large Multilingual Translation Models
abstract
Abstract Hallucinated translations can severely undermine and raise safety issues when machine translation systems are deployed in the wild. Previous research on the topic focused on small bilingual models trained on high-resource languages, leaving a gap in our understanding of hallucinations in multilingual models across diverse translation scenarios. In this work, we fill this gap by conducting a comprehensive analysis—over 100 language pairs across various resource levels and going beyond English-centric directions—on both the M2M neural machine translation (NMT) models and GPT large language models (LLMs). Among several insights, we highlight that models struggle with hallucinations primarily in low-resource directions and when translating out of English, where, critically, they may reveal toxic patterns that can be traced back to the training data. We also find that LLMs produce qualitatively different hallucinations to those of NMT models. Finally, we show that hallucinations are hard to reverse by merely scaling models trained with the same data. However, employing more diverse models, trained on different data or with different procedures, as fallback systems can improve translation quality and virtually eliminate certain pathologies.
Nuno Miguel Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, André F. T. Martins
Trans. Assoc. Comput. Linguistics7
2023 Efficient Methods for Natural Language Processing: A Survey
abstract
Abstract Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources include data, time, storage, or energy, all of which are naturally limited and unevenly distributed. This motivates research into efficient methods that require fewer resources to achieve similar results. This survey synthesizes and relates current methods and findings in efficient NLP. We aim to provide both guidance for conducting NLP under limited resources, and point towards promising research directions for developing more efficient methods.
Marcos V. Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro Henrique Martins, André F. T. Martins, Jessica Zosa Forde, Peter A. Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz 0001
Trans. Assoc. Comput. Linguistics12
2022 ∞-former: Infinite Memory Transformer
abstract
Transformers are unable to model long-term memories effectively, since the amount of computation they need to perform grows with the context length. While variations of efficient transformers have been proposed, they all have a finite memory capacity and are forced to drop old information. In this paper, we propose the ∞-former, which extends the vanilla transformer with an unbounded long-term memory. By making use of a continuous-space attention mechanism to attend over the long-term memory, the ∞-former's attention complexity becomes independent of the context length, trading off memory length with precision.In order to control where precision is more important, ∞-former maintains "sticky memories," being able to model arbitrarily long contexts while keeping the computation budget fixed.Experiments on a synthetic sorting task, language modeling, and document grounded dialogue generation demonstrate the ∞-former's ability to retain information from long sequences.
Pedro Henrique Martins, Zita Marinho, André F. T. Martins
ACL (1)3
2022 DeepSPIN: Deep Structured Prediction for Natural Language Processing
abstract
DeepSPIN is a research project funded by the European Research Council (ERC) whose goal is to develop new neural structured prediction methods, models, and algorithms for improving the quality, interpretability, and data-efficiency of natural language processing (NLP) systems, with special emphasis on machine translation and quality estimation. We describe in this paper the latest findings from this project.
André F. T. Martins, Ben Peters, Chrysoula Zerva, Chunchuan Lyu, Gonçalo M. Correia, Marcos V. Treviso, Pedro Henrique Martins, Tsvetomila Mihaylova
EAMT1
2022 Searching for COMETINHO: The Little Metric That Could
abstract
In recent years, several neural fine-tuned machine translation evaluation metrics such as COMET and BLEURT have been proposed. These metrics achieve much higher correlations with human judgments than lexical overlap metrics at the cost of computational efficiency and simplicity, limiting their applications to scenarios in which one has to score thousands of translation hypothesis (e.g. scoring multiple systems or Minimum Bayes Risk decoding). In this paper, we explore optimization techniques, pruning, and knowledge distillation to create more compact and faster COMET versions. Our results show that just by optimizing the code through the use of caching and length batching we can reduce inference time between 39% and 65% when scoring multiple systems. Also, we show that pruning COMET can lead to a 21% model reduction without affecting the model’s accuracy beyond 0.01 Kendall tau correlation. Furthermore, we present DISTIL-COMET a lightweight distilled version that is 80% smaller and 2.128x faster while attaining a performance close to the original model and above strong baselines such as BERTSCORE and PRISM.
Ricardo Rei, Ana C. Farinha, José Guilherme Camargo de Souza, Pedro G. Ramos, André F. T. Martins, Luísa Coheur, Alon Lavie
EAMT5
2022 QUARTZ: Quality-Aware Machine Translation
abstract
This paper presents QUARTZ, QUality-AwaRe machine Translation, a project led by Unbabel which aims at developing machine translation systems that are more robust and produce fewer critical errors. With QUARTZ we want to enable machine translation for user-generated conversational content types that do not tolerate critical errors in automatic translations.
José Guilherme Camargo de Souza, Ricardo Rei, Ana C. Farinha, Helena Moniz, André F. T. Martins
EAMT5
2022 Chunk-based Nearest Neighbor Machine Translation
abstract
Semi-parametric models, which augment generation with retrieval, have led to impressive results in language modeling and machine translation, due to their ability to retrieve fine-grained information from a datastore of examples.One of the most prominent approaches, kNN-MT, exhibits strong domain adaptation capabilities by retrieving tokens from domain-specific datastores (Khandelwal et al., 2021).However, kNN-MT requires an expensive retrieval operation for every single generated token, leading to a very low decoding speed (around 8 times slower than a parametric model).In this paper, we introduce a chunk-based kNN-MT model which retrieves chunks of tokens from the datastore, instead of a single token.We propose several strategies for incorporating the retrieved chunks into the generation process, and for selecting the steps at which the model needs to search for neighbors in the datastore.Experiments on machine translation in two settings, static and "on-the-fly" domain adaptation, show that the chunk-based kNN-MT model leads to significant speed-ups (up to 4 times) with only a small drop in translation quality. 1
Pedro Henrique Martins, Zita Marinho, André F. T. Martins
EMNLP3
2022 Disentangling Uncertainty in Machine Translation Evaluation
abstract
Trainable evaluation metrics for machine translation (MT) exhibit strong correlation with human judgements, but they are often hard to interpret and might produce unreliable scores under noisy or out-of-domain data.Recent work has attempted to mitigate this with simple uncertainty quantification techniques (Monte Carlo dropout and deep ensembles), however these techniques (as we show) are limited in several ways -for example, they are unable to distinguish between different kinds of uncertainty, and they are time and memory consuming.In this paper, we propose more powerful and efficient uncertainty predictors for MT evaluation, and we assess their ability to target different sources of aleatoric and epistemic uncertainty.To this end, we develop and compare training objectives for the COMET metric to enhance it with an uncertainty prediction output, including heteroscedastic regression, divergence minimization, and direct uncertainty prediction.Our experiments show improved results on uncertainty prediction for the WMT metrics task datasets, with a substantial reduction in computational costs.Moreover, they demonstrate the ability of these predictors to address specific uncertainty causes in MT evaluation, such as low quality references and outof-domain data. 1
Chrysoula Zerva, Taisiya Glushkova, Ricardo Rei, André F. T. Martins
EMNLP4
2022 Sparse Communication via Mixed Distributions
António Farinhas, Wilker Aziz, Vlad Niculae, André F. T. Martins
ICLR4
2022 Modeling Structure with Undirected Neural Networks
abstract
Neural networks are powerful function estimators, leading to their status as a paradigm of choice for modeling structured data. However, unlike other structured representations that emphasize the modularity of the problem {–} e.g., factor graphs {–} neural networks are usually monolithic mappings from inputs to outputs, with a fixed computation order. This limitation prevents them from capturing different directions of computation and interaction between the modeled variables. In this paper, we combine the representational strengths of factor graphs and of neural networks, proposing undirected neural networks (UNNs): a flexible framework for specifying computations that can be performed in any order. For particular choices, our proposed models subsume and extend many existing architectures: feed-forward, recurrent, self-attention networks, auto-encoders, and networks with implicit layers. We demonstrate the effectiveness of undirected neural architectures, both unstructured and structured, on a range of tasks: tree-constrained dependency parsing, convolutional image classification, and sequence completion with attention. By varying the computation order, we show how a single UNN can be used both as a classifier and a prototype generator, and how it can fill in missing parts of an input sequence, making them a promising field for further research.
Tsvetomila Mihaylova, Vlad Niculae, André F. T. Martins
ICML3
2022 MLQE-PE: A Multilingual Quality Estimation and Post-Editing Dataset
abstract
We present MLQE-PE, a new dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE). The dataset contains annotations for eleven language pairs, including both high- and low-resource languages. Specifically, it is annotated for translation quality with human labels for up to 10,000 translations per language pair in the following formats: sentence-level direct assessments and post-editing effort, and word-level binary good/bad labels. Apart from the quality-related scores, each source-translation sentence pair is accompanied by the corresponding post-edited sentence, as well as titles of the articles where the sentences were extracted from, and information on the neural MT models used to translate the text. We provide a thorough description of the data collection and annotation process as well as an analysis of the annotation distribution for each language pair. We also report the performance of baseline systems trained on the MLQE-PE dataset. The dataset is freely available and has already been used for several WMT shared tasks.
Marina Fomicheva, Erick Rocha Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, André F. T. Martins
LREC10
2022 Quality-Aware Decoding for Neural Machine Translation
abstract
Patrick Fernandes, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, Andre Martins. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Patrick Fernandes, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, Perez Ogayo, Graham Neubig, André F. T. Martins
NAACL-HLT7
2022 Learning to Scaffold: Optimizing Model Explanations for Teaching
abstract
Modern machine learning models are opaque, and as a result there is a burgeoning academic subfield on methods that explain these models' behavior. However, what is the precise goal of providing such explanations, and how can we demonstrate that explanations achieve this goal? Some research argues that explanations should help teach a student (either human or machine) to simulate the model being explained, and that the quality of explanations can be measured by the simulation accuracy of students on unexplained examples. In this work, leveraging meta-learning techniques, we extend this idea to improve the quality of the explanations themselves, specifically by optimizing explanations such that student models more effectively learn to simulate the original model. We train models on three natural language processing and computer vision tasks, and find that students trained with explanations extracted with our framework are able to simulate the teacher significantly more effectively than ones produced with previous methods. Through human annotations and a user study, we further find that these learned explanations more closely align with how humans would explain the required decisions in these tasks. Our code is available at https://github.com/coderpat/learning-scaffold.
Patrick Fernandes, Marcos V. Treviso, Danish Pruthi, André F. T. Martins, Graham Neubig
NeurIPS4
2022 Sparse Continuous Distributions and Fenchel-Young Losses
abstract
Exponential families are widely used in machine learning, including many distributions in continuous and discrete domains (e.g., Gaussian, Dirichlet, Poisson, and categorical distributions via the softmax transformation). Distributions in each of these families have fixed support. In contrast, for finite domains, recent work on sparse alternatives to softmax (e.g., sparsemax, $\alpha$-entmax, and fusedmax), has led to distributions with varying support. This paper develops sparse alternatives to continuous distributions, based on several technical contributions: First, we define $\Omega$-regularized prediction maps and Fenchel-Young losses for arbitrary domains (possibly countably infinite or continuous). For linearly parametrized families, we show that minimization of Fenchel-Young losses is equivalent to moment matching of the statistics, generalizing a fundamental property of exponential families. When $\Omega$ is a Tsallis negentropy with parameter $\alpha$, we obtain “deformed exponential families,” which include $\alpha$-entmax and sparsemax ($\alpha=2$) as particular cases. For quadratic energy functions, the resulting densities are $\beta$-Gaussians, an instance of elliptical distributions that contain as particular cases the Gaussian, biweight, triweight, and Epanechnikov densities, and for which we derive closed-form expressions for the variance, Tsallis entropy, and Fenchel-Young loss. When $\Omega$ is a total variation or Sobolev regularizer, we obtain a continuous version of the fusedmax. Finally, we introduce continuous-domain attention mechanisms, deriving efficient gradient backpropagation algorithms for $\alpha \in \{1,\frac{4}{3}, \frac{3}{2}, 2\}$. Using these algorithms, we demonstrate our sparse continuous distributions for attention-based audio classification and visual question answering, showing that they allow attending to time intervals and compact regions.
André F. T. Martins, Marcos V. Treviso, António Farinhas, Pedro M. Q. Aguiar, Mário A. T. Figueiredo, Mathieu Blondel, Vlad Niculae
J. Mach. Learn. Res.1
2021 Measuring and Increasing Context Usage in Context-Aware Machine Translation
abstract
Patrick Fernandes, Kayo Yin, Graham Neubig, André F. T. Martins. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Patrick Fernandes, Kayo Yin, Graham Neubig, André F. T. Martins
ACL/IJCNLP (1)4
2021 Do Context-Aware Translation Models Pay the Right Attention?
abstract
Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, Graham Neubig
ACL/IJCNLP (1)5
2021 SPECTRA: Sparse Structured Text Rationalization
abstract
Selective rationalization aims to produce decisions along with rationales (e.g., text highlights or word alignments between two sentences).Commonly, rationales are modeled as stochastic binary masks, requiring samplingbased gradient estimators, which complicates training and requires careful hyperparameter tuning.Sparse attention mechanisms are a deterministic alternative, but they lack a way to regularize the rationale extraction (e.g., to control the sparsity of a text highlight or the number of alignments).In this paper, we present a unified framework for deterministic extraction of structured explanations via constrained inference on a factor graph, forming a differentiable layer.Our approach greatly eases training and rationale regularization, generally outperforming previous work on what comes to performance and plausibility of the extracted rationales.We further provide a comparative study of stochastic and deterministic methods for rationale extraction for classification and natural language inference tasks, jointly assessing their predictive power, quality of the explanations, and model variability.
Nuno Miguel Guerreiro, André F. T. Martins
EMNLP (1)2
2021 Sparse And Structured Visual Attention
abstract
Visual attention mechanisms are widely used in multimodal tasks, as visual question answering (VQA). One drawback of softmax-based attention mechanisms is that they assign some probability mass to all image regions, regardless of their adjacency structure and of their relevance to the text. In this paper, to better link the image structure with the text, we replace the traditional softmax attention mechanism with two alternative sparsity-promoting transformations: sparsemax, which is able to select only the relevant regions (assigning zero weight to the rest), and a newly proposed Total-Variation Sparse Attention (TVMAX), which further encourages the joint selection of adjacent spatial locations. Experiments in VQA show gains in accuracy as well as higher similarity to human attention, which suggests better interpretability.
Pedro Henrique Martins, Vlad Niculae, Zita Marinho, André F. T. Martins
ICIP4
2021 Smoothing and Shrinking the Sparse Seq2Seq Search Space
abstract
Current sequence-to-sequence models are trained to minimize cross-entropy and use softmax to compute the locally normalized probabilities over target sequences.While this setup has led to strong results in a variety of tasks, one unsatisfying aspect is its length bias: models give high scores to short, inadequate hypotheses and often make the empty string the argmax-the so-called cat got your tongue problem.Recently proposed entmax-based sparse sequence-to-sequence models present a possible solution, since they can shrink the search space by assigning zero probability to bad hypotheses, but their ability to handle word-level tasks with transformers has never been tested.In this work, we show that entmax-based models effectively solve the cat got your tongue problem, removing a major source of model error for neural machine translation.In addition, we generalize label smoothing, a critical regularization technique, to the broader family of Fenchel-Young losses, which includes both cross-entropy and the entmax losses.Our resulting label-smoothed entmax loss models set a new state of the art on multilingual grapheme-to-phoneme conversion and deliver improvements and better calibration properties on cross-lingual morphological inflection and machine translation for 7 language pairs.
Ben Peters, André F. T. Martins
NAACL-HLT2
2020 Revisiting Higher-Order Dependency Parsers
abstract
Neural encoders have allowed dependency parsers to shift from higher-order structured models to simpler first-order ones, making decoding faster and still achieving better accuracy than non-neural parsers.This has led to a belief that neural encoders can implicitly encode structural constraints, such as siblings and grandparents in a tree.We tested this hypothesis and found that neural parsers may benefit from higher-order features, even when employing a powerful pre-trained encoder, such as BERT.While the gains of higher-order features are small in the presence of a powerful encoder, they are consistent for long-range dependencies and long sentences.In particular, higher-order models are more accurate on full sentence parses and on the exact match of modifier lists, indicating that they deal better with larger, more complex structures.
Erick Rocha Fonseca, André F. T. Martins
ACL2
2020 Learning Non-Monotonic Automatic Post-Editing of Translations from Human Orderings
abstract
Recent research in neural machine translation has explored flexible generation orders, as an alternative to left-to-right generation. However, training non-monotonic models brings a new complication: how to search for a good ordering when there is a combinatorial explosion of orderings arriving at the same final result? Also, how do these automatic orderings compare with the actual behaviour of human translators? Current models rely on manually built biases or are left to explore all possibilities on their own. In this paper, we analyze the orderings produced by human post-editors and use them to train an automatic post-editing system. We compare the resulting system with those trained with left-to-right and random post-editing orderings. We observe that humans tend to follow a nearly left-to-right order, but with interesting deviations, such as preferring to start by correcting punctuation or verbs.
António Góis, Kyunghyun Cho, André F. T. Martins
EAMT3
2020 Document-level Neural MT: A Systematic Comparison
abstract
In this paper we provide a systematic comparison of existing and new document-level neural machine translation solutions. As part of this comparison, we introduce and evaluate a document-level variant of the recently proposed Star Transformer architecture. In addition to using the traditional metric BLEU, we report the accuracy of the models in handling anaphoric pronoun translation as well as coherence and cohesion using contrastive test sets. Finally, we report the results of human evaluation in terms of Multidimensional Quality Metrics (MQM) and analyse the correlation of the results obtained by the automatic metrics with human judgments.
António V. Lopes, M. Amin Farajian, Rachel Bawden, André F. T. Martins
EAMT5
2020 DeepSPIN: Deep Structured Prediction for Natural Language Processing
abstract
DeepSPIN is a research project funded by the European Research Council (ERC) whose goal is to develop new neural structured prediction methods, models, and algorithms for improving the quality, interpretability, and data-efficiency of natural language processing (NLP) systems, with special emphasis on machine translation and quality estimation applications.
André F. T. Martins
EAMT1
2020 Project MAIA: Multilingual AI Agent Assistant
abstract
This paper presents the Multilingual Artificial Intelligence Agent Assistant (MAIA), a project led by Unbabel with the collaboration of CMU, INESC-ID and IT Lisbon. MAIA will employ cutting-edge machine learning and natural language processing technologies to build multilingual AI agent assistants, eliminating language barriers. MAIA’s translation layer will empower human agents to provide customer support in real-time, in any language, with human quality.
André F. T. Martins, João Graça, Paulo Dimas, Helena Moniz, Graham Neubig
EAMT1
2020 Sparse Text Generation
abstract
Current state-of-the-art text generators build on powerful language models such as GPT-2, achieving impressive performance.However, to avoid degenerate text, they require sampling from a modified softmax, via temperature parameters or ad-hoc truncation techniques, as in top-k or nucleus sampling.This creates a mismatch between training and testing conditions.In this paper, we use the recently introduced entmax transformation to train and sample from a natively sparse language model, avoiding this mismatch.The result is a text generator with favorable performance in terms of fluency and consistency, fewer repetitions, and n-gram diversity closer to human text.In order to evaluate our model, we propose three new metrics for comparing sparse or truncated distributions: -perplexity, sparsemax score, and Jensen-Shannon divergence.Human-evaluated experiments in story completion and dialogue generation show that entmax sampling leads to more engaging and coherent stories and conversations.
Pedro Henrique Martins, Zita Marinho, André F. T. Martins
EMNLP (1)3
2020 Understanding the Mechanics of SPIGOT: Surrogate Gradients for Latent Structure Learning
abstract
Latent structure models are a powerful tool for modeling language data: they can mitigate the error propagation and annotation bottleneck in pipeline systems, while simultaneously uncovering linguistic insights about the data.One challenge with end-to-end training of these models is the argmax operation, which has null gradient.In this paper, we focus on surrogate gradients, a popular strategy to deal with this problem.We explore latent structure learning through the angle of pulling back the downstream learning objective.In this paradigm, we discover a principled motivation for both the straight-through estimator (STE) as well as the recently-proposed SPIGOT-a variant of STE for structured models.Our perspective leads to new algorithms in the same family.We empirically compare the known and the novel pulled-back estimators against the popular alternatives, yielding new insight for practitioners and revealing intriguing failure cases.
Tsvetomila Mihaylova, Vlad Niculae, André F. T. Martins
EMNLP (1)3
2020 LP-SparseMAP: Differentiable Relaxed Optimization for Sparse Structured Prediction
abstract
Structured predictors require solving a combinatorial optimization problem over a large number of structures, such as dependency trees or alignments. When embedded as structured hidden layers in a neural net, argmin differentiation and efficient gradient computation are further required. Recently, SparseMAP has been proposed as a differentiable, sparse alternative to maximum a posteriori (MAP) and marginal inference. SparseMAP returns an interpretable combination of a small number of structures; its sparsity being the key to efficient optimization. However, SparseMAP requires access to an exact MAP oracle in the structured model, excluding, e.g., loopy graphical models or logic constraints, which generally require approximate inference. In this paper, we introduce LP-SparseMAP, an extension of SparseMAP addressing this limitation via a local polytope relaxation. LP-SparseMAP uses the flexible and powerful language of factor graphs to define expressive hidden structures, supporting coarse decompositions, hard logic constraints, and higher-order correlations. We derive the forward and backward algorithms needed for using LP-SparseMAP as a structured hidden or output layer. Experiments in three structured tasks show benefits versus SparseMAP and Structured SVM.
Vlad Niculae, André F. T. Martins
ICML2
2020 Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
abstract
Training neural network models with discrete (categorical or structured) latent variables can be computationally challenging, due to the need for marginalization over large or combinatorial sets. To circumvent this issue, one typically resorts to sampling-based approximations of the true marginal, requiring noisy gradient estimators (e.g., score function estimator) or continuous relaxations with lower-variance reparameterized gradients (e.g., Gumbel-Softmax). In this paper, we propose a new training strategy which replaces these estimators by an exact yet efficient marginalization. To achieve this, we parameterize discrete distributions over latent assignments using differentiable sparse mappings: sparsemax and its structured counterparts. In effect, the support of these distributions is greatly reduced, which enables efficient marginalization. We report successful results in three tasks covering a range of latent variable modeling applications: a semisupervised deep generative model, a latent communication game, and a generative model with a bit-vector latent representation. In all cases, we obtain good performance while still achieving the practicality of sampling-based approximations.
Gonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. Martins
NeurIPS4
2020 Sparse and Continuous Attention Mechanisms
abstract
Exponential families are widely used in machine learning; they include many distributions in continuous and discrete domains (e.g., Gaussian, Dirichlet, Poisson, and categorical distributions via the softmax transformation). Distributions in each of these families have fixed support. In contrast, for finite domains, there has been recent work on sparse alternatives to softmax (e.g., sparsemax and alpha-entmax), which have varying support, being able to assign zero probability to irrelevant categories. These discrete sparse mappings have been used for improving interpretability of neural attention mechanisms. This paper expands that work in two directions: first, we extend alpha-entmax to continuous domains, revealing a link with Tsallis statistics and deformed exponential families. Second, we introduce continuous-domain attention mechanisms, deriving efficient gradient backpropagation algorithms for alpha in {1,2}. Experiments on attention-based text classification, machine translation, and visual question answering illustrate the use of continuous attention in 1D and 2D, showing that it allows attending to time intervals and compact regions.
André F. T. Martins, António Farinhas, Marcos V. Treviso, Vlad Niculae, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
NeurIPS1
2020 Learning with Fenchel-Young losses
abstract
Over the past decades, numerous loss functions have been been proposed for a variety of supervised learning tasks, including regression, classification, ranking, and more generally structured prediction. Understanding the core principles and theoretical properties underpinning these losses is key to choose the right loss for the right problem, as well as to create new losses which combine their strengths. In this paper, we introduce Fenchel-Young losses, a generic way to construct a convex loss function for a regularized prediction function. We provide an in-depth study of their properties in a very broad setting, covering all the aforementioned supervised learning tasks, and revealing new connections between sparsity, generalized entropies, and separation margins. We show that Fenchel-Young losses unify many well-known loss functions and allow to create useful new ones easily. Finally, we derive efficient predictive and training algorithms, making Fenchel-Young losses appealing both in theory and practice.
Mathieu Blondel, André F. T. Martins, Vlad Niculae
J. Mach. Learn. Res.2
2019 A Simple and Effective Approach to Automatic Post-Editing with Transfer Learning
abstract
Automatic post-editing (APE) seeks to automatically refine the output of a black-box machine translation (MT) system through human post-edits.APE systems are usually trained by complementing human post-edited data with large, artificial data generated through backtranslations, a time-consuming process often no easier than training a MT system from scratch.In this paper, we propose an alternative where we fine-tune pre-trained BERT models on both the encoder and decoder of an APE system, exploring several parameter sharing strategies.By only training on a dataset of 23K sentences for 3 hours on a single GPU we obtain results that are competitive with systems that were trained on 5M artificial sentences.When we add this artificial data, our method obtains state-of-the-art results.
Gonçalo M. Correia, André F. T. Martins
ACL (1)2
2019 Sparse Sequence-to-Sequence Models
abstract
Sequence-to-sequence models are a powerful workhorse of NLP. Most variants employ a softmax transformation in both their attention mechanism and output layer, leading to dense alignments and strictly positive output probabilities. This density is wasteful, making models less interpretable and assigning probability mass to many implausible outputs. In this paper, we propose sparse sequence-to-sequence models, rooted in a new family of 𝛼-entmax transformations, which includes softmax and sparsemax as particular cases, and is sparse for any 𝛼 > 1. We provide fast algorithms to evaluate these transformations and their gradients, which scale well for large vocabulary sizes. Our models are able to produce sparse alignments and to assign nonzero probability to a short list of plausible outputs, sometimes rendering beam search exact. Experiments on morphological inflection and machine translation reveal consistent gains over dense models.
Ben Peters, Vlad Niculae, André F. T. Martins
ACL (1)3
2019 Learning Classifiers with Fenchel-Young Losses: Generalized Entropies, Margins, and Algorithms
abstract
This paper studies Fenchel-Young losses, a generic way to construct convex loss functions from a regularization function. We analyze their properties in depth, showing that they unify many well-known loss functions and allow to create useful new ones easily. Fenchel-Young losses constructed from a generalized entropy, including the Shannon and Tsallis entropies, induce predictive probability distributions. We formulate conditions for a generalized entropy to yield losses with a separation margin, and probability distributions with sparse support. Finally, we derive efficient algorithms, making Fenchel-Young losses appealing both in theory and practice.
Mathieu Blondel, André F. T. Martins, Vlad Niculae
AISTATS2
2019 Adaptively Sparse Transformers
abstract
Gonçalo M. Correia, Vlad Niculae, André F. T. Martins. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Gonçalo M. Correia, Vlad Niculae, André F. T. Martins
EMNLP/IJCNLP (1)3
2019 Translator2Vec: Understanding and Representing Human Post-Editors
António Góis, André F. T. Martins
MTSummit (1)2
2018 Towards Dynamic Computation Graphs via Sparse Latent Structure
abstract
Deep NLP models benefit from underlying structures in the data-e.g., parse treestypically extracted using off-the-shelf parsers.Recent attempts to jointly learn the latent structure encounter a tradeoff: either make factorization assumptions that limit expressiveness, or sacrifice end-to-end differentiability.Using the recently proposed SparseMAP inference, which retrieves a sparse distribution over latent structures, we propose a novel approach for end-to-end learning of latent structure predictors jointly with a downstream predictor.To the best of our knowledge, our method is the first to enable unrestricted dynamic computation graph construction from the global latent structure, while maintaining differentiability.2016).Assigning the conjunction as head instead seems preferable in a Child-Sum TreeLSTM.
Vlad Niculae, André F. T. Martins, Claire Cardie
EMNLP2
2018 SparseMAP: Differentiable Sparse Structured Inference
abstract
Structured prediction requires searching over a combinatorial number of structures. To tackle it, we introduce SparseMAP, a new method for sparse structured inference, together with corresponding loss functions. SparseMAP inference is able to automatically select only a few global structures: it is situated between MAP inference, which picks a single structure, and marginal inference, which assigns probability mass to all structures, including implausible ones. Importantly, SparseMAP can be computed using only calls to a MAP oracle, hence it is applicable even to problems where marginal inference is intractable, such as linear assignment. Moreover, thanks to the solution sparsity, gradient backpropagation is efficient regardless of the structure. SparseMAP thus enables us to augment deep neural networks with generic and sparse structured hidden layers. Experiments in dependency parsing and natural language inference reveal competitive accuracy, improved interpretability, and the ability to capture natural language ambiguities, which is attractive for pipeline systems.
Vlad Niculae, André F. T. Martins, Mathieu Blondel, Claire Cardie
ICML2
2017 Learning What's Easy: Fully Differentiable Neural Easy-First Taggers
abstract
We introduce a novel neural easy-first decoder that learns to solve sequence tagging tasks in a flexible order.In contrast to previous easy-first decoders, our models are end-to-end differentiable.The decoder iteratively updates a "sketch" of the predictions over the sequence.At its core is an attention mechanism that controls which parts of the input are strategically the best to process next.We present a new constrained softmax transformation that ensures the same cumulative attention to every word, and show how to efficiently evaluate and backpropagate over it.Our models compare favourably to BILSTM taggers on three sequence tagging tasks.* This research was partially carried out during an internship at Unbabel.
André F. T. Martins, Julia Kreutzer
EMNLP1
2017 Pushing the Limits of Translation Quality Estimation
abstract
Translation quality estimation is a task of growing importance in NLP, due to its potential to reduce post-editing human effort in disruptive ways. However, this potential is currently limited by the relatively low accuracy of existing systems. In this paper, we achieve remarkable improvements by exploiting synergies between the related tasks of word-level quality estimation and automatic post-editing. First, we stack a new, carefully engineered, neural model into a rich feature-based word-level quality estimation system. Then, we use the output of an automatic post-editing system as an extra feature, obtaining striking results on WMT16: a word-level FMULT1 score of 57.47% (an absolute gain of +7.95% over the current state of the art), and a Pearson correlation score of 65.56% for sentence-level HTER prediction (an absolute gain of +13.36%).
André F. T. Martins, Marcin Junczys-Dowmunt, Fábio N. Kepler, Ramón Fernandez Astudillo, Chris Hokamp, Roman Grundkiewicz
Trans. Assoc. Comput. Linguistics1
2016 Jointly Learning to Embed and Predict with Multiple Languages
abstract
We propose a joint formulation for learning task-specific cross-lingual word embeddings, along with classifiers for that task.Unlike prior work, which first learns the embeddings from parallel data and then plugs them in a supervised learning problem, our approach is oneshot: a single optimization problem combines a co-regularizer for the multilingual embeddings with a task-specific loss.We present theoretical results showing the limitation of Euclidean co-regularizers to increase the embedding dimension, a limitation which does not exist for other co-regularizers (such as the ℓ 1distance).Despite its simplicity, our method achieves state-of-the-art accuracies on the RCV1/RCV2 dataset when transferring from English to German, with training times below 1 minute.On the TED Corpus, we obtain the highest reported scores on 10 out of 11 languages.
Daniel C. Ferreira, André F. T. Martins, Mariana S. C. Almeida
ACL (1)2
2016 Semi-Supervised Learning of Sequence Models with Method of Moments
abstract
We propose a fast and scalable method for semi-supervised learning of sequence models, based on anchor words and moment matching. Our method can handle hidden Markov models with feature-based log-linear emissions. Unlike other semi-supervised methods, no decoding passes are necessary on the unlabeled data and no graph needs to be constructed— only one pass is necessary to collect moment statistics. The model parameters are estimated by solving a small quadratic program for each feature. Experiments on part-of-speech (POS) tagging for Twitter and for a low-resource language (Malagasy) show that our method can learn from very few annotated sentences.
Zita Marinho, André F. T. Martins, Shay B. Cohen, Noah A. Smith
EMNLP2
2016 From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
abstract
We propose sparsemax, a new activation function similar to the traditional softmax, but able to output sparse probabilities. After deriving its properties, we show how its Jacobian can be efficiently computed, enabling its use in a network trained with backpropagation. Then, we propose a new smooth and convex loss function which is the sparsemax analogue of the logistic loss. We reveal an unexpected connection between this new loss and the Huber classification loss. We obtain promising empirical results in multi-label classification problems and in attention-based neural networks for natural language inference. For the latter, we achieve a similar performance as the traditional softmax, but with a selective, more compact, attention focus.
André F. T. Martins, Ramón Fernandez Astudillo
ICML1
2015 Aligning Opinions: Cross-Lingual Opinion Mining with Dependencies
abstract
Mariana S. C. Almeida, Cláudia Pinto, Helena Figueira, Pedro Mendes, André F. T. Martins. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Mariana S. C. Almeida, Cláudia Pinto, Helena Figueira, Pedro Mendes 0004, André F. T. Martins
ACL (1)5
2015 Parsing as Reduction
abstract
Daniel Fernández-González, André F. T. Martins. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Daniel Fernández-González, André F. T. Martins
ACL (1)2
2015 Transferring Coreference Resolvers with Posterior Regularization
abstract
André F. T. Martins. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
André F. T. Martins
ACL (1)1
2015 AD3: alternating directions dual decomposition for MAP inference in graphical models
André F. T. Martins, Mário A. T. Figueiredo, Pedro M. Q. Aguiar, Noah A. Smith, Eric P. Xing
J. Mach. Learn. Res.1
2014 A Joint Model for Quotation Attribution and Coreference Resolution
abstract
We address the problem of automatically attributing quotations to speakers, which has great relevance in text mining and media monitoring applications.While current systems report high accuracies for this task, they either work at mentionlevel (getting credit for detecting uninformative mentions such as pronouns), or assume the coreferent mentions have been detected beforehand; the inaccuracies in this preprocessing step may lead to error propagation.In this paper, we introduce a joint model for entity-level quotation attribution and coreference resolution, exploiting correlations between the two tasks.We design an evaluation metric for attribution that captures all speakers' mentions.We present results showing that both tasks benefit from being treated jointly.
Mariana S. C. Almeida, Miguel B. Almeida, André F. T. Martins
EACL3
2014 Priberam Compressive Summarization Corpus: A New Multi-Document Summarization Corpus for European Portuguese
Miguel B. Almeida, Mariana S. C. Almeida, André F. T. Martins, Helena Figueira, Pedro Mendes 0004, Cláudia Pinto
LREC3
2014 Frame-Semantic Parsing
abstract
Frame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised learning to improve frame disambiguation for targets unseen at training time. The second stage finds the target's locally expressed semantic arguments. At inference time, a fast exact dual decomposition algorithm collectively predicts all the arguments of a frame at once in order to respect declaratively stated linguistic constraints, resulting in qualitatively better structures than naïve local predictors. Both components are feature-based and discriminatively trained on a small set of annotated frame-semantic parses. On the SemEval 2007 benchmark data set, the approach, along with a heuristic identifier of frame-evoking targets, outperforms the prior state of the art by significant margins. Additionally, we present experiments on the much larger FrameNet 1.5 data set. We have released our frame-semantic parser as open-source software.
Dipanjan Das 0001, Desai Chen, André F. T. Martins, Nathan Schneider 0001, Noah A. Smith
Comput. Linguistics3
2013 Fast and Robust Compressive Summarization with Dual Decomposition and Multi-Task Learning
Miguel B. Almeida, André F. T. Martins
ACL (1)2
2013 Combining information theoretic kernels with generative embeddings for classification
Manuele Bicego, Aydin Ulas, Umberto Castellani, Alessandro Perina, Vittorio Murino, André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
Neurocomputing6
2012 Structured Sparsity in Natural Language Processing: Models, Algorithms and Applications
André F. T. Martins, Mário A. T. Figueiredo, Noah A. Smith
HLT-NAACL1
2011 Dual Decomposition with Many Overlapping Components
André F. T. Martins, Noah A. Smith, Mário A. T. Figueiredo, Pedro M. Q. Aguiar
EMNLP1
2011 Structured Sparsity in Structured Prediction
André F. T. Martins, Noah A. Smith, Mário A. T. Figueiredo, Pedro M. Q. Aguiar
EMNLP1
2011 An Augmented Lagrangian Approach to Constrained MAP Inference
André F. T. Martins, Mário A. T. Figueiredo, Pedro M. Q. Aguiar, Noah A. Smith, Eric P. Xing
ICML1
2010 Turbo Parsers: Dependency Parsing by Approximate Variational Inference
André F. T. Martins, Noah A. Smith, Eric P. Xing, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
EMNLP1
2010 Combining free energy score spaces with information theoretic kernels: Application to scene classification
abstract
Most approaches to learn classifiers for structured objects (e.g., images) use generative models in a classical Bayesian framework. However, state-of-the-art classifiers for vectorial data (e.g., support vector machines) are learned discriminatively. A generative embedding is a mapping from the object space into a fixed dimensional score space, induced by a generative model, usually learned from data. The fixed dimensionality of these generative score spaces makes them adequate for discriminative learning of classifiers, thus bringing together the best of the discriminative and generative paradigms. In particular, it was recently shown that this hybrid approach outperforms a classifier obtained directly for the generative model upon which the score space was built. Using a generative embedding involves two steps: (i) defining and learning the generative model and using it to build the embedding; (ii) discriminatively learning a (maybe kernel) classifier on the adopted score space. The literature on generative embeddings is essentially focused on step (i), usually using some standard off-the-shelf tool for step (ii). In this paper, we adopt a different approach, by focusing also on the discriminative learning step. In particular, we combine two very recent and top performing tools in each of the steps: (i) the free energy score space; (ii) non-extensive information theoretic kernels. In this paper, we apply this methodology in scene recognition. Experimental results on two benchmark datasets shows that our approach yields state-of-the-art performance.
Manuele Bicego, Alessandro Perina, Vittorio Murino, André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
ICIP4
2010 2D Shape Recognition Using Information Theoretic Kernels
abstract
In this paper, a novel approach for contour based 2D shape recognition is proposed, using a class of information theoretic kernels recently introduced. This kind of kernels, based on a non-extensive generalization of the classical Shannon information theory, are defined on probability measures. In the proposed approach, chain code representations are first extracted from the contours; then n-gram statistics are computed and used as input to the information theoretic kernels. We tested different versions of such kernels, using support vector machine and nearest neighbor classifiers. An experimental evaluation on the Chicken pieces dataset shows that the proposed approach significantly outperforms the current state-of-the-art methods.
Manuele Bicego, André F. T. Martins, Vittorio Murino, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
ICPR2
2009 Concise Integer Linear Programming Formulations for Dependency Parsing
André F. T. Martins, Noah A. Smith, Eric P. Xing
ACL/IJCNLP1
2009 Polyhedral outer approximations with application to natural language parsing
abstract
Recent approaches to learning structured predictors often require approximate inference for tractability; yet its effects on the learned model are unclear. Meanwhile, most learning algorithms act as if computational cost was constant within the model class. This paper sheds some light on the first issue by establishing risk bounds for max-margin learning with LP relaxed inference and addresses the second issue by proposing a new paradigm that attempts to penalize “timeconsuming” hypotheses. Our analysis relies on a geometric characterization of the outer polyhedra associated with the LP relaxation. We then apply these techniques to the problem of dependency parsing, for which a concise LP formulation is provided that handles non-local output features. A significant improvement is shown over arc-factored models. 1.
André F. T. Martins, Noah A. Smith, Eric P. Xing
ICML1
2009 Nonextensive Information Theoretic Kernels on Measures
André F. T. Martins, Noah A. Smith, Eric P. Xing, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
J. Mach. Learn. Res.1
2008 Stacking Dependency Parsers
André F. T. Martins, Dipanjan Das 0001, Noah A. Smith, Eric P. Xing
EMNLP1
2008 Nonextensive entropic kernels
abstract
Positive definite kernels on probability measures have been recently applied in structured data classification problems. Some of these kernels are related to classic information theoretic quantities, such as mutual information and the Jensen-Shannon divergence. Meanwhile, driven by recent advances in Tsallis statistics, nonextensive generalizations of Shannon's information theory have been proposed. This paper bridges these two trends. We introduce the Jensen-Tsallis q-difference, a generalization of the Jensen-Shannon divergence. We then define a new family of nonextensive mutual information kernels, which allow weights to be assigned to their arguments, and which includes the Boolean, Jensen-Shannon, and linear kernels as particular cases. We illustrate the performance of these kernels on text categorization tasks.
André F. T. Martins, Mário A. T. Figueiredo, Pedro M. Q. Aguiar, Noah A. Smith, Eric P. Xing
ICML1
2008 Tsallis kernels on measures
abstract
Recent approaches to classification of text, images, and other types of structured data, launched the quest for positive definite (p.d.) kernels on probability measures. In particular, kernels based on the Jensen-Shannon (JS) divergence and other information-theoretic quantities have been proposed. We introduce new JS-type divergences, by extending its two building blocks: convexity and Shannon’s entropy. These divergences are then used to define new information-theoretic kernels on measures. In particular, we introduce a new concept of q-convexity, for which a Jensen q-inequality is proved. Based on this inequality, we introduce the Jensen-Tsallis q-difference, a nonextensive generalization of the Jensen-Shannon divergence. Furthermore, we provide denormalization formulae for entropies and divergences, which we use to define a family of nonextensive information-theoretic kernels on measures. This family, grounded in nonextensive entropies, extends Jensen-Shannon divergence kernels, and allows assigning weights to its arguments.
André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
ITW1
2006 Web-Based Support for Resource-Effective E-Learning
Maria Alexandra Rentroia-Bonito, Frederico C. Figueiredo, André F. T. Martins, Vitor Fernandes, Joaquim Jorge 0001
WEBIST (2)3
2005 Orientation in Manhattan: Equiprojective Classes and Sequential Estimation
abstract
The problem of inferring 3D orientation of a camera from video sequences has been mostly addressed by first computing correspondences of image features. This intermediate step is now seen as the main bottleneck of those approaches. In this paper, we propose a new 3D orientation estimation method for urban (indoor and outdoor) environments, which avoids correspondences between frames. The scene property exploited by our method is that many edges are oriented along three orthogonal directions; this is the recently introduced Manhattan world (MW) assumption. The main contributions of this paper are: the definition of equivalence classes of equiprojective orientations, the introduction of a new small rotation model, formalizing the fact that the camera moves smoothly, and the decoupling of elevation and twist angle estimation from that of the compass angle. We build a probabilistic sequential orientation estimation method, based on an MW likelihood model, with the above-listed contributions allowing a drastic reduction of the search space for each orientation estimate. We demonstrate the performance of our method using real video sequences.
André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Design and Implementation of a Semantic Search Engine for Portuguese
Carlos Amaral, Dominique Laurent 0003, André F. T. Martins, Afonso Mendes, Cláudia Pinto
LREC3
2003 Navigating in Manhattan: 3D orientation from video without correspondences
abstract
The problem of inferring 3D orientation of a camera from video sequences has been mostly addressed by first computing correspondences of image features. This intermediate step is now seen as the main bottleneck of those approaches. In this paper, we propose a new 3D orientation estimation method for urban (indoor and outdoor) environments, which avoids correspondences between frames. The basic scene property exploited by our method is that many edges are oriented along three orthogonal directions; this is the recently introduced Manhattan world (MW) assumption. In addition to the novel adoption of the MW assumption for video analysis, we introduce the small rotation (SR) assumption, that expresses the fact that the video camera undergoes a smooth 3D motion. Using these two assumptions, we build a probabilistic estimation approach. We demonstrate the performance of our method using real video sequences.
André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo
ICIP (1)1