VLDB 2026 Research / reviewers in the wild / expert
José Guilherme Camargo de Souza
dblp:66/1087 · also José G. C. de Souza
· DBLP profile ↗
15ranked-venue papers
5as first author
9since 2021 · last 2024
0000-0001-6344-7633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine TranslationabstractSweta Agrawal, José G. C. De Souza, Ricardo Rei, António Farinhas, Gonçalo Faria, Patrick Fernandes, Nuno M Guerreiro, Andre Martins. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sweta Agrawal, José Guilherme Camargo de Souza, Ricardo Rei, António Farinhas, Gonçalo Rui Alves Faria, Patrick Fernandes, Nuno Miguel Guerreiro, André F. T. Martins |
EMNLP | 2 |
| 2024 | QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine TranslationabstractAn important challenge in machine translation (MT) is to generate high-quality and diverse translations.
Prior work has shown that the estimated likelihood from the MT model correlates poorly with translation quality.
In contrast, quality evaluation metrics (such as COMET or BLEURT) exhibit high correlations with human judgments, which has motivated their use as rerankers (such as quality-aware and minimum Bayes risk decoding). However, relying on a single translation with high estimated quality increases the chances of "gaming the metric''.
In this paper, we address the problem of sampling a set of high-quality and diverse translations.
We provide a simple and effective way to avoid over-reliance on noisy quality estimates by using them as the energy function of a Gibbs distribution. Instead of looking for a mode in the distribution, we generate multiple samples from high-density areas through the Metropolis-Hastings algorithm, a simple Markov chain Monte Carlo approach.
The results show that our proposed method leads to high-quality and diverse outputs across multiple language pairs (English$\leftrightarrow$\{German, Russian\}) with two strong decoder-only LLMs (Alma-7b, Tower-7b). Gonçalo Rui Alves Faria, Sweta Agrawal, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, André F. T. Martins |
NeurIPS | 5 |
| 2023 | Empirical Assessment of kNN-MT for Real-World Translation ScenariosabstractThis paper aims to investigate the effectiveness of the k-Nearest Neighbor Machine Translation model (kNN-MT) in real-world scenarios. kNN-MT is a retrieval-augmented framework that combines the advantages of parametric models with non-parametric datastores built using a set of parallel sentences. Previous studies have primarily focused on evaluating the model using only the BLEU metric and have not tested kNN-MT in real world scenarios. Our study aims to fill this gap by conducting a comprehensive analysis on various datasets comprising different language pairs and different domains, using multiple automatic metrics and expert evaluated Multidimensional Quality Metrics (MQM). We compare kNN-MT with two alternate strategies: fine-tuning all the model parameters and adapter-based finetuning. Finally, we analyze the effect of the datastore size on translation quality, and we examine the number of entries necessary to bootstrap and configure the index. Pedro Henrique Martins, João Alves 0003, Tânia Vaz, Madalena Gonçalves, Beatriz Silva, Marianna Buchicchio, José Guilherme Camargo de Souza, André F. T. Martins |
EAMT | 7 |
| 2023 | An Empirical Study of Translation Hypothesis Ensembling with Large Language ModelsabstractLarge language models (LLMs) are becoming a one-fits-many solution, but they sometimes hallucinate or produce unreliable output.In this paper, we investigate how hypothesis ensembling can improve the quality of the generated text for the specific problem of LLM-based machine translation.We experiment with several techniques for ensembling hypotheses produced by LLMs such as ChatGPT, LLaMA, and Alpaca.We provide a comprehensive study along multiple dimensions, including the method to generate hypotheses (multiple prompts, temperaturebased sampling, and beam search) and the strategy to produce the final translation (instructionbased, quality-based reranking, and minimum Bayes risk (MBR) decoding).Our results show that MBR decoding is a very effective method, that translation quality can be improved using a small number of samples, and that instruction tuning has a strong impact on the relation between the diversity of the hypotheses and the sampling temperature.Our code is available at https://github.com/deep-spin/ translation-hypothesis-ensembling. António Farinhas, José Guilherme Camargo de Souza, André F. T. Martins |
EMNLP | 2 |
| 2023 | Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language GenerationabstractAbstract Natural language generation has witnessed significant advancements due to the training of large language models on vast internet-scale datasets. Despite these advancements, there exists a critical challenge: These models can inadvertently generate content that is toxic, inaccurate, and unhelpful, and existing automatic evaluation metrics often fall short of identifying these shortcomings. As models become more capable, human feedback is an invaluable signal for evaluating and improving models. This survey aims to provide an overview of recent research that has leveraged human feedback to improve natural language generation. First, we introduce a taxonomy distilled from existing research to categorize and organize the varied forms of feedback. Next, we discuss how feedback can be described by its format and objective, and cover the two approaches proposed to use feedback (either for training or decoding): directly using feedback or training feedback models. We also discuss existing datasets for human-feedback data collection, and concerns surrounding feedback collection. Finally, we provide an overview of the nascent field of AI feedback, which uses large language models to make judgments based on a set of principles and minimize the need for human intervention. We also release a website of this survey at feedback-gap-survey.info. Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José Guilherme Camargo de Souza, Shuyan Zhou, Sherry Tongshuang Wu, Graham Neubig, André F. T. Martins |
Trans. Assoc. Comput. Linguistics | 7 |
| 2022 | Multi3Generation: Multitask, Multilingual, Multimodal Language GenerationabstractThis paper presents the Multitask, Multilingual, Multimodal Language Generation COST Action – Multi3Generation (CA18231), an interdisciplinary network of research groups working on different aspects of language generation. This “meta-paper” will serve as reference for citations of the Action in future publications. It presents the objectives, challenges and a the links for the achieved outcomes. Anabela Barreiro, José Guilherme Camargo de Souza, Albert Gatt, Mehul Bhatt, Elena Lloret, Aykut Erdem, Dimitra Gkatzia, Helena Moniz, Irene Russo, Fábio N. Kepler, Iacer Calixto, Marcin Paprzycki, François Portet, Isabelle Augenstein, Mirela Alhasani |
EAMT | 2 |
| 2022 | Searching for COMETINHO: The Little Metric That CouldabstractIn recent years, several neural fine-tuned machine translation evaluation metrics such as COMET and BLEURT have been proposed. These metrics achieve much higher correlations with human judgments than lexical overlap metrics at the cost of computational efficiency and simplicity, limiting their applications to scenarios in which one has to score thousands of translation hypothesis (e.g. scoring multiple systems or Minimum Bayes Risk decoding). In this paper, we explore optimization techniques, pruning, and knowledge distillation to create more compact and faster COMET versions. Our results show that just by optimizing the code through the use of caching and length batching we can reduce inference time between 39% and 65% when scoring multiple systems. Also, we show that pruning COMET can lead to a 21% model reduction without affecting the model’s accuracy beyond 0.01 Kendall tau correlation. Furthermore, we present DISTIL-COMET a lightweight distilled version that is 80% smaller and 2.128x faster while attaining a performance close to the original model and above strong baselines such as BERTSCORE and PRISM. Ricardo Rei, Ana C. Farinha, José Guilherme Camargo de Souza, Pedro G. Ramos, André F. T. Martins, Luísa Coheur, Alon Lavie |
EAMT | 3 |
| 2022 | QUARTZ: Quality-Aware Machine TranslationabstractThis paper presents QUARTZ, QUality-AwaRe machine Translation, a project led by Unbabel which aims at developing machine translation systems that are more robust and produce fewer critical errors. With QUARTZ we want to enable machine translation for user-generated conversational content types that do not tolerate critical errors in automatic translations. José Guilherme Camargo de Souza, Ricardo Rei, Ana C. Farinha, Helena Moniz, André F. T. Martins |
EAMT | 1 |
| 2022 | Quality-Aware Decoding for Neural Machine TranslationabstractPatrick Fernandes, António Farinhas, Ricardo Rei, José De Souza, Perez Ogayo, Graham Neubig, Andre Martins. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Patrick Fernandes, António Farinhas, Ricardo Rei, José Guilherme Camargo de Souza, Perez Ogayo, Graham Neubig, André F. T. Martins |
NAACL-HLT | 4 |
| 2018 | Generating E-Commerce Product Titles and Predicting their QualityabstractJosé G. Camargo de Souza, Michael Kozielski, Prashant Mathur, Ernie Chang, Marco Guerini, Matteo Negri, Marco Turchi, Evgeny Matusov. Proceedings of the 11th International Conference on Natural Language Generation. 2018. José Guilherme Camargo de Souza, Michael Kozielski, Prashant Mathur, Ernie Chang, Marco Guerini, Matteo Negri, Marco Turchi, Evgeny Matusov |
INLG | 1 |
| 2015 | Online Multitask Learning for Machine Translation Quality EstimationabstractJosé G. C. de Souza, Matteo Negri, Elisa Ricci, Marco Turchi. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. José Guilherme Camargo de Souza, Matteo Negri, Elisa Ricci 0001, Marco Turchi |
ACL (1) | 1 |
| 2015 | Multitask Learning for Adaptive Quality Estimation of Automatically Transcribed UtterancesabstractJosé G. C. de Souza, Hamed Zamani, Matteo Negri, Marco Turchi, Daniele Falavigna. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. José Guilherme Camargo de Souza, Hamed Zamani, Matteo Negri, Marco Turchi, Daniele Falavigna |
HLT-NAACL | 1 |
| 2014 | Adaptive Quality Estimation for Machine TranslationabstractThe automatic estimation of machine translation (MT) output quality is a hard task in which the selection of the appropriate algorithm and the most predictive features over reasonably sized training sets plays a crucial role.When moving from controlled lab evaluations to real-life scenarios the task becomes even harder.For current MT quality estimation (QE) systems, additional complexity comes from the difficulty to model user and domain changes.Indeed, the instability of the systems with respect to data coming from different distributions calls for adaptive solutions that react to new operating conditions.To tackle this issue we propose an online framework for adaptive QE that targets reactivity and robustness to user and domain changes.Contrastive experiments in different testing conditions involving user and domain changes demonstrate the effectiveness of our approach. Marco Turchi, Antonios Anastasopoulos, José Guilherme Camargo de Souza, Matteo Negri |
ACL (1) | 3 |
| 2014 | Quality Estimation for Automatic Speech Recognition
Matteo Negri, Marco Turchi, José Guilherme Camargo de Souza, Daniele Falavigna |
COLING | 3 |
| 2014 | Machine Translation Quality Estimation Across Domains
José Guilherme Camargo de Souza, Marco Turchi, Matteo Negri |
COLING | 1 |