Francisco Guzmán

dblp:92/1459 · DBLP profile ↗
← Back
37ranked-venue papers
10as first author
12since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 10 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation
abstract
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzmán
ACL (1)8
2022 Alternative Input Signals Ease Transfer in Multilingual Machine Translation
abstract
Simeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán
ACL (1)7
2022 MLQE-PE: A Multilingual Quality Estimation and Post-Editing Dataset
abstract
We present MLQE-PE, a new dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE). The dataset contains annotations for eleven language pairs, including both high- and low-resource languages. Specifically, it is annotated for translation quality with human labels for up to 10,000 translations per language pair in the following formats: sentence-level direct assessments and post-editing effort, and word-level binary good/bad labels. Apart from the quality-related scores, each source-translation sentence pair is accompanied by the corresponding post-edited sentence, as well as titles of the articles where the sentences were extracted from, and information on the neural MT models used to translate the text. We provide a thorough description of the data collection and annotation process as well as an analysis of the annotation distribution for each language pair. We also report the performance of baseline systems trained on the MLQE-PE dataset. The dataset is freely available and has already been used for several WMT shared tasks.
Marina Fomicheva, Erick Rocha Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, André F. T. Martins
LREC7
2022 The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
abstract
Abstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the Flores-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are fully aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.
Naman Goyal 0001, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzmán, Angela Fan
Trans. Assoc. Comput. Linguistics9
2021 Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data
abstract
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab
ACL/IJCNLP (1)6
2021 Improving Zero-Shot Translation by Disentangling Positional Information
abstract
Danni Liu, Jan Niehues, James Cross, Francisco Guzmán, Xian Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jan Niehues, James Cross 0003, Francisco Guzmán, Xian Li 0003
ACL/IJCNLP (1)4
2021 WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia
abstract
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, Francisco Guzmán. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Holger Schwenk, Vishrav Chaudhary, Hongyu Gong, Francisco Guzmán
EACL5
2021 Quality Estimation without Human-labeled Data
abstract
Yi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Francisco Guzmán, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Yi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Francisco Guzmán, Lucia Specia
EACL5
2021 XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment
abstract
Cross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification.While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing.To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data.We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English.We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community.
Ahmed El-Kishky, Adithya Renduchintala, James Cross 0003, Francisco Guzmán, Philipp Koehn
EMNLP (1)4
2021 Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications
abstract
Sentence-level Quality Estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels.Recent QE models have achieved previously-unseen levels of correlation with human judgments, but they rely on large multilingual contextualized language models that are computationally expensive and thus infeasible for many real-world applications.In this work, we evaluate several model compression techniques for QE and find that, despite their popularity in other NLP tasks, they lead to poor performance in this regression setting.We observe that a full model parameterization is required to achieve SoTA results in a regression task.However, we argue that the level of expressiveness of a model in a continuous range is unnecessary given the downstream applications of QE, and show that reframing QE as a classification problem and evaluating QE models using classification metrics would better reflect their actual performance in real-world applications.
Ahmed El-Kishky, Vishrav Chaudhary, James Cross 0003, Lucia Specia, Francisco Guzmán
EMNLP (1)6
2021 Proceedings of the 18th Biennial Machine Translation Summit (Volume 1: Research Track)
Kevin Duh, Francisco Guzmán
MTSummit (1)2
2021 A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
abstract
As modern neural machine translation (NMT) systems have been widely deployed, their security vulnerabilities require close scrutiny. Most recently, NMT systems have been found vulnerable to targeted attacks which cause them to produce specific, unsolicited, and even harmful translations. These attacks are usually exploited in a white-box setting, where adversarial inputs causing targeted translations are discovered for a known target system. However, this approach is less viable when the target system is black-box and unknown to the adversary (e.g., secured commercial systems). In this paper, we show that targeted attacks on black-box NMT systems are feasible, based on poisoning a small fraction of their parallel training data. We show that this attack can be realised practically via targeted corruption of web documents crawled to form the system’s training data. We then analyse the effectiveness of the targeted poisoning in two common NMT training scenarios: the from-scratch training and the pre-train & fine-tune paradigm. Our results are alarming: even on the state-of-the-art systems trained with massive parallel data (tens of millions), the attacks are still successful (over 50% success rate) under surprisingly low poisoning budgets (e.g., 0.006%). Lastly, we discuss potential defences to counter such attacks.
Jun Wang 0126, Francisco Guzmán, Benjamin I. P. Rubinstein, Trevor Cohn
WWW4
2020 Unsupervised Cross-lingual Representation Learning at Scale
abstract
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Alexis Conneau, Kartikay Khandelwal, Naman Goyal 0001, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov
ACL6
2020 Multi-Hypothesis Machine Translation Evaluation
abstract
Reliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem.One of the main challenges is the fact that multiple outputs can be equally valid.Attempts to minimise this issue include metrics that relax the matching of MT output and reference strings, and the use of multiple references.The latter has been shown to significantly improve the performance of evaluation metrics.However, collecting multiple references is expensive and in practice a single reference is generally used.In this paper, we propose an alternative approach: instead of modelling linguistic variation in human reference we exploit the MT model uncertainty to generate multiple diverse translations and use these: (i) as surrogates to reference translations; (ii) to obtain a quantification of translation variability to either complement existing metric scores or (iii) replace references altogether.We show that for a number of popular evaluation metrics our variability estimates lead to substantial improvements in correlation with human judgements of quality by up 15%.
Marina Fomicheva, Lucia Specia, Francisco Guzmán
ACL3
2020 Are we Estimating or Guesstimating Translation Quality?
abstract
Recent advances in pre-trained multilingual language models lead to state-of-the-art results on the task of quality estimation (QE) for machine translation.A carefully engineered ensemble of such models won the QE shared task at WMT19.Our in-depth analysis, however, shows that the success of using pre-trained language models for QE is overestimated due to three issues we observed in current QE datasets: (i) The distributions of quality scores are imbalanced and skewed towards good quality scores; (ii) QE models can perform well on these datasets while looking at only source or translated sentences; (iii) They contain statistical artifacts that correlate well with human-annotated QE labels.Our findings suggest that although QE models might capture fluency of translated sentences and complexity of source sentences, they cannot model adequacy of translations effectively.
Francisco Guzmán, Lucia Specia
ACL2
2020 CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs
abstract
Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other.We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English.In addition to curating this massive dataset, we introduce baseline methods that leverage crosslingual representations to identify aligned documents based on their textual content.Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data.Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages.
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn
EMNLP (1)3
2020 CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
abstract
Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved. In this paper, we describe an automatic pipeline to extract massive high-quality monolingual datasets from Common Crawl for a variety of languages. Our pipeline follows the data processing introduced in fastText (Mikolov et al., 2017; Grave et al., 2018), that deduplicates documents and identifies their language. We augment this pipeline with a filtering step to select documents that are close to high quality corpora like Wikipedia.
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, Edouard Grave
LREC5
2020 Unsupervised Quality Estimation for Neural Machine Translation
abstract
Quality Estimation (QE) is an important component in making Machine Translation (MT) useful in real-world applications, as it is aimed to inform the user on the quality of the MT output at test time. Existing approaches require large amounts of expert annotated data, computation, and time for training. As an alternative, we devise an unsupervised approach to QE where no training or access to additional resources besides the MT system itself is required. Different from most of the current work that treats the MT system as a black box, we explore useful information that can be extracted from the MT system as a by-product of translation. By utilizing methods for uncertainty quantification, we achieve very good correlation with human judgments of quality, rivaling state-of-the-art supervised QE models. To evaluate our approach we collect the first dataset that enables work on both black-box and glass-box approaches to QE.
Marina Fomicheva, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, Lucia Specia
Trans. Assoc. Comput. Linguistics5
2019 Design and Evaluation of a Social Media Writing Support Tool for People with Dyslexia
abstract
People with dyslexia face challenges expressing themselves in writing on social networking sites (SNSs). Such challenges come from not only the technicality of writing, but also the self-representation aspect of sharing and communicating publicly on social networking sites such as Facebook. To empower people with dyslexia-style writing to express them-selves more confidently on SNSs, we designed and implemented Additional Writing Help(AWH) - a writing assistance tool to proofread text produced by users with dyslexia before they post on Facebook. AWH was powered by a neural machine translation (NMT) model that translates dyslexia style to non-dyslexia style writing. We evaluated the performance and the design of AWH through a week-long field study with 19 people with dyslexia and received highly positive feedback. Our field study demonstrated the value of providing better and more extensive writing support on SNSs, and the potential of AI for building a more inclusive Internet.
Shaomei Wu, Lindsay Reynolds, Xian Li 0003, Francisco Guzmán
CHI4
2019 The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English
abstract
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino 0001, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc'Aurelio Ranzato
EMNLP/IJCNLP (1)1
2017 QT2S: A System for Monitoring Road Traffic Via Fine Grounding of Tweets
Noora Al Emadi, Sofiane Abbar, Javier Borge-Holthoefer, Francisco Guzmán, Fabrizio Sebastiani 0001
ICWSM4
2017 Discourse Structure in Machine Translation Evaluation
abstract
In this article, we explore the potential of using sentence-level discourse structure for machine translation evaluation. We first design discourse-aware similarity measures, which use all-subtree kernels to compare discourse parse trees in accordance with the Rhetorical Structure Theory (RST). Then, we show that a simple linear combination with these measures can help improve various existing machine translation evaluation metrics regarding correlation with human judgments both at the segment level and at the system level. This suggests that discourse information is complementary to the information used by many of the existing evaluation metrics, and thus it could be taken into account when developing richer evaluation metrics, such as the WMT-14 winning combined metric DiscoTKparty. We also provide a detailed analysis of the relevance of various discourse elements and relations from the RST parse trees for machine translation evaluation. In particular, we show that (i) all aspects of the RST tree are relevant, (ii) nuclearity is more useful than relation type, and (iii) the similarity of the translation RST tree to the reference RST tree is positively correlated with translation quality.
Shafiq R. Joty, Francisco Guzmán, Lluís Màrquez, Preslav Nakov
Comput. Linguistics2
2017 Machine translation evaluation with neural networks
Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov
Comput. Speech Lang.1
2016 Machine Translation Evaluation for Arabic using Morphologically-enriched Embeddings
abstract
Evaluation of machine translation (MT) into morphologically rich languages (MRL) has not been well studied despite posing many challenges. In this paper, we explore the use of embeddings obtained from different levels of lexical and morpho-syntactic linguistic analysis and show that they improve MT evaluation into an MRL. Specifically we report on Arabic, a language with complex and rich morphology. Our results show that using a neural-network model with different input representations produces results that clearly outperform the state-of-the-art for MT evaluation into Arabic, by almost over 75% increase in correlation with human judgments on pairwise MT evaluation quality task. More importantly, we demonstrate the usefulness of morpho-syntactic representations to model sentence similarity for MT evaluation and address complex linguistic phenomena of Arabic.
Francisco Guzmán, Houda Bouamor, Ramy Baly, Nizar Habash
COLING1
2016 It Takes Three to Tango: Triangulation Approach to Answer Ranking in Community Question Answering
abstract
We address the problem of answering new questions in community forums, by selecting suitable answers to already asked questions.We approach the task as an answer ranking problem, adopting a pairwise neural network architecture that selects which of two competing answers is better.We focus on the utility of the three types of similarities occurring in the triangle formed by the original question, the related question, and an answer to the related comment, which we call relevance, relatedness, and appropriateness.Our proposed neural network models the interactions among all input components using syntactic and semantic embeddings, lexical matching, and domain-specific features.It achieves state-of-the-art results, showing that the three similarities are important and need to be modeled together.Our experiments demonstrate that all feature types are relevant, but the most important ones are the lexical similarity features, the domain-specific features, and the syntactic and semantic embeddings.
Preslav Nakov, Lluís Màrquez, Francisco Guzmán
EMNLP3
2016 Eyes Don't Lie: Predicting Machine Translation Quality Using Eye Movement
abstract
Hassan Sajjad, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina Temnikova, Stephan Vogel. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Hassan Sajjad 0001, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina P. Temnikova, Stephan Vogel
HLT-NAACL2
2015 Pairwise Neural Machine Translation Evaluation
abstract
Francisco Guzmán, Shafiq Joty, Lluís Màrquez, Preslav Nakov. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov
ACL (1)1
2015 Analyzing Optimization for Statistical Machine Translation: MERT Learns Verbosity, PRO Learns Length
abstract
We study the impact of source length and verbosity of the tuning dataset on the performance of parameter optimizers such as MERT and PRO for statistical machine translation.In particular, we test whether the verbosity of the resulting translations can be modified by varying the length or the verbosity of the tuning sentences.We find that MERT learns the tuning set verbosity very well, while PRO is sensitive to both the verbosity and the length of the source sentences in the tuning set; yet, overall PRO learns best from highverbosity tuning datasets.
Francisco Guzmán, Preslav Nakov, Stephan Vogel
CoNLL1
2015 QAT2 - the QCRI advanced transcription and translation system
Ahmed Abdelali, Ahmed Ali 0002, Francisco Guzmán, Felix Stahlberg, Stephan Vogel
INTERSPEECH3
2014 Using Discourse Structure Improves Machine Translation Evaluation
abstract
We present experiments in using discourse structure for improving machine translation evaluation.We first design two discourse-aware similarity measures, which use all-subtree kernels to compare discourse parse trees in accordance with the Rhetorical Structure Theory.Then, we show that these measures can help improve a number of existing machine translation evaluation metrics both at the segment-and at the system-level.Rather than proposing a single new metric, we show that discourse information is complementary to the state-of-the-art evaluation metrics, and thus should be taken into account in the development of future richer evaluation metrics.
Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov
ACL (1)1
2014 Learning to Differentiate Better from Worse Translations
abstract
Francisco Guzmán, Shafiq Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, Massimo Nicosia. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, Massimo Nicosia
EMNLP1
2014 The AMARA Corpus: Building Parallel Language Resources for the Educational Domain
Ahmed Abdelali, Francisco Guzmán, Hassan Sajjad 0001, Stephan Vogel
LREC2
2012 Understanding the Performance of Statistical MT Systems: A Linear Regression Framework
Francisco Guzmán, Stephan Vogel
COLING1
2012 Optimizing for Sentence-Level BLEU+1 Yields Short Translations
Preslav Nakov, Francisco Guzmán, Stephan Vogel
COLING2
2010 EMDC: A Semi-supervised Approach for Word Alignment
Qin Gao, Francisco Guzmán, Stephan Vogel
COLING2
2009 Reassessment of the Role of Phrase Extraction in PBSMT
Francisco Guzmán, Qin Gao, Stephan Vogel
MTSummit1
2008 Translation Paraphrases in Phrase-Based Machine Translation
Francisco Guzmán, Leonardo Garrido
CICLing1