VLDB 2026 Research / reviewers in the wild / expert
Francisco Guzmán
dblp:92/1459
· DBLP profile ↗
37ranked-venue papers
10as first author
12since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 10 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationabstractJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzmán |
ACL (1) | 8 |
| 2022 | Alternative Input Signals Ease Transfer in Multilingual Machine TranslationabstractSimeng Sun, Angela Fan, James Cross, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Simeng Sun, Angela Fan, James Cross 0003, Vishrav Chaudhary, Chau Tran, Philipp Koehn, Francisco Guzmán |
ACL (1) | 7 |
| 2022 | MLQE-PE: A Multilingual Quality Estimation and Post-Editing DatasetabstractWe present MLQE-PE, a new dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE). The dataset contains annotations for eleven language pairs, including both high- and low-resource languages. Specifically, it is annotated for translation quality with human labels for up to 10,000 translations per language pair in the following formats: sentence-level direct assessments and post-editing effort, and word-level binary good/bad labels. Apart from the quality-related scores, each source-translation sentence pair is accompanied by the corresponding post-edited sentence, as well as titles of the articles where the sentences were extracted from, and information on the neural MT models used to translate the text. We provide a thorough description of the data collection and annotation process as well as an analysis of the annotation distribution for each language pair. We also report the performance of baseline systems trained on the MLQE-PE dataset. The dataset is freely available and has already been used for several WMT shared tasks. Marina Fomicheva, Erick Rocha Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, André F. T. Martins |
LREC | 7 |
| 2022 | The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine TranslationabstractAbstract One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the Flores-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are fully aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond. Naman Goyal 0001, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzmán, Angela Fan |
Trans. Assoc. Comput. Linguistics | 9 |
| 2021 | Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataabstractWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab |
ACL/IJCNLP (1) | 6 |
| 2021 | Improving Zero-Shot Translation by Disentangling Positional InformationabstractDanni Liu, Jan Niehues, James Cross, Francisco Guzmán, Xian Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jan Niehues, James Cross 0003, Francisco Guzmán, Xian Li 0003 |
ACL/IJCNLP (1) | 4 |
| 2021 | WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from WikipediaabstractHolger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, Francisco Guzmán. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Holger Schwenk, Vishrav Chaudhary, Hongyu Gong, Francisco Guzmán |
EACL | 5 |
| 2021 | Quality Estimation without Human-labeled DataabstractYi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Francisco Guzmán, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Yi-Lin Tuan, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Francisco Guzmán, Lucia Specia |
EACL | 5 |
| 2021 | XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word AlignmentabstractCross-lingual named-entity lexica are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification.While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing.To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexica from mined web data.We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English.We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community. Ahmed El-Kishky, Adithya Renduchintala, James Cross 0003, Francisco Guzmán, Philipp Koehn |
EMNLP (1) | 4 |
| 2021 | Classification-based Quality Estimation: Small and Efficient Models for Real-world ApplicationsabstractSentence-level Quality Estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels.Recent QE models have achieved previously-unseen levels of correlation with human judgments, but they rely on large multilingual contextualized language models that are computationally expensive and thus infeasible for many real-world applications.In this work, we evaluate several model compression techniques for QE and find that, despite their popularity in other NLP tasks, they lead to poor performance in this regression setting.We observe that a full model parameterization is required to achieve SoTA results in a regression task.However, we argue that the level of expressiveness of a model in a continuous range is unnecessary given the downstream applications of QE, and show that reframing QE as a classification problem and evaluating QE models using classification metrics would better reflect their actual performance in real-world applications. Ahmed El-Kishky, Vishrav Chaudhary, James Cross 0003, Lucia Specia, Francisco Guzmán |
EMNLP (1) | 6 |
| 2021 | Proceedings of the 18th Biennial Machine Translation Summit (Volume 1: Research Track)
Kevin Duh, Francisco Guzmán |
MTSummit (1) | 2 |
| 2021 | A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data PoisoningabstractAs modern neural machine translation (NMT) systems have been widely deployed, their security vulnerabilities require close scrutiny. Most recently, NMT systems have been found vulnerable to targeted attacks which cause them to produce specific, unsolicited, and even harmful translations. These attacks are usually exploited in a white-box setting, where adversarial inputs causing targeted translations are discovered for a known target system. However, this approach is less viable when the target system is black-box and unknown to the adversary (e.g., secured commercial systems). In this paper, we show that targeted attacks on black-box NMT systems are feasible, based on poisoning a small fraction of their parallel training data. We show that this attack can be realised practically via targeted corruption of web documents crawled to form the system’s training data. We then analyse the effectiveness of the targeted poisoning in two common NMT training scenarios: the from-scratch training and the pre-train & fine-tune paradigm. Our results are alarming: even on the state-of-the-art systems trained with massive parallel data (tens of millions), the attacks are still successful (over 50% success rate) under surprisingly low poisoning budgets (e.g., 0.006%). Lastly, we discuss potential defences to counter such attacks. Jun Wang 0126, Francisco Guzmán, Benjamin I. P. Rubinstein, Trevor Cohn |
WWW | 4 |
| 2020 | Unsupervised Cross-lingual Representation Learning at ScaleabstractAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Alexis Conneau, Kartikay Khandelwal, Naman Goyal 0001, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov |
ACL | 6 |
| 2020 | Multi-Hypothesis Machine Translation EvaluationabstractReliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem.One of the main challenges is the fact that multiple outputs can be equally valid.Attempts to minimise this issue include metrics that relax the matching of MT output and reference strings, and the use of multiple references.The latter has been shown to significantly improve the performance of evaluation metrics.However, collecting multiple references is expensive and in practice a single reference is generally used.In this paper, we propose an alternative approach: instead of modelling linguistic variation in human reference we exploit the MT model uncertainty to generate multiple diverse translations and use these: (i) as surrogates to reference translations; (ii) to obtain a quantification of translation variability to either complement existing metric scores or (iii) replace references altogether.We show that for a number of popular evaluation metrics our variability estimates lead to substantial improvements in correlation with human judgements of quality by up 15%. Marina Fomicheva, Lucia Specia, Francisco Guzmán |
ACL | 3 |
| 2020 | Are we Estimating or Guesstimating Translation Quality?abstractRecent advances in pre-trained multilingual language models lead to state-of-the-art results on the task of quality estimation (QE) for machine translation.A carefully engineered ensemble of such models won the QE shared task at WMT19.Our in-depth analysis, however, shows that the success of using pre-trained language models for QE is overestimated due to three issues we observed in current QE datasets: (i) The distributions of quality scores are imbalanced and skewed towards good quality scores; (ii) QE models can perform well on these datasets while looking at only source or translated sentences; (iii) They contain statistical artifacts that correlate well with human-annotated QE labels.Our findings suggest that although QE models might capture fluency of translated sentences and complexity of source sentences, they cannot model adequacy of translations effectively. Francisco Guzmán, Lucia Specia |
ACL | 2 |
| 2020 | CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsabstractCross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other.We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English.In addition to curating this massive dataset, we introduce baseline methods that leverage crosslingual representations to identify aligned documents based on their textual content.Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data.Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages. Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn |
EMNLP (1) | 3 |
| 2020 | CCNet: Extracting High Quality Monolingual Datasets from Web Crawl DataabstractPre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved. In this paper, we describe an automatic pipeline to extract massive high-quality monolingual datasets from Common Crawl for a variety of languages. Our pipeline follows the data processing introduced in fastText (Mikolov et al., 2017; Grave et al., 2018), that deduplicates documents and identifies their language. We augment this pipeline with a filtering step to select documents that are close to high quality corpora like Wikipedia. Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, Edouard Grave |
LREC | 5 |
| 2020 | Unsupervised Quality Estimation for Neural Machine TranslationabstractQuality Estimation (QE) is an important component in making Machine Translation (MT) useful in real-world applications, as it is aimed to inform the user on the quality of the MT output at test time. Existing approaches require large amounts of expert annotated data, computation, and time for training. As an alternative, we devise an unsupervised approach to QE where no training or access to additional resources besides the MT system itself is required. Different from most of the current work that treats the MT system as a black box, we explore useful information that can be extracted from the MT system as a by-product of translation. By utilizing methods for uncertainty quantification, we achieve very good correlation with human judgments of quality, rivaling state-of-the-art supervised QE models. To evaluate our approach we collect the first dataset that enables work on both black-box and glass-box approaches to QE. Marina Fomicheva, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, Lucia Specia |
Trans. Assoc. Comput. Linguistics | 5 |
| 2019 | Design and Evaluation of a Social Media Writing Support Tool for People with DyslexiaabstractPeople with dyslexia face challenges expressing themselves in writing on social networking sites (SNSs). Such challenges come from not only the technicality of writing, but also the self-representation aspect of sharing and communicating publicly on social networking sites such as Facebook. To empower people with dyslexia-style writing to express them-selves more confidently on SNSs, we designed and implemented Additional Writing Help(AWH) - a writing assistance tool to proofread text produced by users with dyslexia before they post on Facebook. AWH was powered by a neural machine translation (NMT) model that translates dyslexia style to non-dyslexia style writing. We evaluated the performance and the design of AWH through a week-long field study with 19 people with dyslexia and received highly positive feedback. Our field study demonstrated the value of providing better and more extensive writing support on SNSs, and the potential of AI for building a more inclusive Internet. Shaomei Wu, Lindsay Reynolds, Xian Li 0003, Francisco Guzmán |
CHI | 4 |
| 2019 | The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-EnglishabstractFrancisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino 0001, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc'Aurelio Ranzato |
EMNLP/IJCNLP (1) | 1 |
| 2017 | QT2S: A System for Monitoring Road Traffic Via Fine Grounding of Tweets
Noora Al Emadi, Sofiane Abbar, Javier Borge-Holthoefer, Francisco Guzmán, Fabrizio Sebastiani 0001 |
ICWSM | 4 |
| 2017 | Discourse Structure in Machine Translation EvaluationabstractIn this article, we explore the potential of using sentence-level discourse structure for machine translation evaluation. We first design discourse-aware similarity measures, which use all-subtree kernels to compare discourse parse trees in accordance with the Rhetorical Structure Theory (RST). Then, we show that a simple linear combination with these measures can help improve various existing machine translation evaluation metrics regarding correlation with human judgments both at the segment level and at the system level. This suggests that discourse information is complementary to the information used by many of the existing evaluation metrics, and thus it could be taken into account when developing richer evaluation metrics, such as the WMT-14 winning combined metric DiscoTKparty. We also provide a detailed analysis of the relevance of various discourse elements and relations from the RST parse trees for machine translation evaluation. In particular, we show that (i) all aspects of the RST tree are relevant, (ii) nuclearity is more useful than relation type, and (iii) the similarity of the translation RST tree to the reference RST tree is positively correlated with translation quality. Shafiq R. Joty, Francisco Guzmán, Lluís Màrquez, Preslav Nakov |
Comput. Linguistics | 2 |
| 2017 | Machine translation evaluation with neural networks
Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
Comput. Speech Lang. | 1 |
| 2016 | Machine Translation Evaluation for Arabic using Morphologically-enriched EmbeddingsabstractEvaluation of machine translation (MT) into morphologically rich languages (MRL) has not been well studied despite posing many challenges. In this paper, we explore the use of embeddings obtained from different levels of lexical and morpho-syntactic linguistic analysis and show that they improve MT evaluation into an MRL. Specifically we report on Arabic, a language with complex and rich morphology. Our results show that using a neural-network model with different input representations produces results that clearly outperform the state-of-the-art for MT evaluation into Arabic, by almost over 75% increase in correlation with human judgments on pairwise MT evaluation quality task. More importantly, we demonstrate the usefulness of morpho-syntactic representations to model sentence similarity for MT evaluation and address complex linguistic phenomena of Arabic. Francisco Guzmán, Houda Bouamor, Ramy Baly, Nizar Habash |
COLING | 1 |
| 2016 | It Takes Three to Tango: Triangulation Approach to Answer Ranking in Community Question AnsweringabstractWe address the problem of answering new questions in community forums, by selecting suitable answers to already asked questions.We approach the task as an answer ranking problem, adopting a pairwise neural network architecture that selects which of two competing answers is better.We focus on the utility of the three types of similarities occurring in the triangle formed by the original question, the related question, and an answer to the related comment, which we call relevance, relatedness, and appropriateness.Our proposed neural network models the interactions among all input components using syntactic and semantic embeddings, lexical matching, and domain-specific features.It achieves state-of-the-art results, showing that the three similarities are important and need to be modeled together.Our experiments demonstrate that all feature types are relevant, but the most important ones are the lexical similarity features, the domain-specific features, and the syntactic and semantic embeddings. Preslav Nakov, Lluís Màrquez, Francisco Guzmán |
EMNLP | 3 |
| 2016 | Eyes Don't Lie: Predicting Machine Translation Quality Using Eye MovementabstractHassan Sajjad, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina Temnikova, Stephan Vogel. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Hassan Sajjad 0001, Francisco Guzmán, Nadir Durrani, Ahmed Abdelali, Houda Bouamor, Irina P. Temnikova, Stephan Vogel |
HLT-NAACL | 2 |
| 2015 | Pairwise Neural Machine Translation EvaluationabstractFrancisco Guzmán, Shafiq Joty, Lluís Màrquez, Preslav Nakov. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
ACL (1) | 1 |
| 2015 | Analyzing Optimization for Statistical Machine Translation: MERT Learns Verbosity, PRO Learns LengthabstractWe study the impact of source length and verbosity of the tuning dataset on the performance of parameter optimizers such as MERT and PRO for statistical machine translation.In particular, we test whether the verbosity of the resulting translations can be modified by varying the length or the verbosity of the tuning sentences.We find that MERT learns the tuning set verbosity very well, while PRO is sensitive to both the verbosity and the length of the source sentences in the tuning set; yet, overall PRO learns best from highverbosity tuning datasets. Francisco Guzmán, Preslav Nakov, Stephan Vogel |
CoNLL | 1 |
| 2015 | QAT2 - the QCRI advanced transcription and translation system
Ahmed Abdelali, Ahmed Ali 0002, Francisco Guzmán, Felix Stahlberg, Stephan Vogel |
INTERSPEECH | 3 |
| 2014 | Using Discourse Structure Improves Machine Translation EvaluationabstractWe present experiments in using discourse structure for improving machine translation evaluation.We first design two discourse-aware similarity measures, which use all-subtree kernels to compare discourse parse trees in accordance with the Rhetorical Structure Theory.Then, we show that these measures can help improve a number of existing machine translation evaluation metrics both at the segment-and at the system-level.Rather than proposing a single new metric, we show that discourse information is complementary to the state-of-the-art evaluation metrics, and thus should be taken into account in the development of future richer evaluation metrics. Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Preslav Nakov |
ACL (1) | 1 |
| 2014 | Learning to Differentiate Better from Worse TranslationsabstractFrancisco Guzmán, Shafiq Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, Massimo Nicosia. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014. Francisco Guzmán, Shafiq R. Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, Massimo Nicosia |
EMNLP | 1 |
| 2014 | The AMARA Corpus: Building Parallel Language Resources for the Educational Domain
Ahmed Abdelali, Francisco Guzmán, Hassan Sajjad 0001, Stephan Vogel |
LREC | 2 |
| 2012 | Understanding the Performance of Statistical MT Systems: A Linear Regression Framework
Francisco Guzmán, Stephan Vogel |
COLING | 1 |
| 2012 | Optimizing for Sentence-Level BLEU+1 Yields Short Translations
Preslav Nakov, Francisco Guzmán, Stephan Vogel |
COLING | 2 |
| 2010 | EMDC: A Semi-supervised Approach for Word Alignment
Qin Gao, Francisco Guzmán, Stephan Vogel |
COLING | 2 |
| 2009 | Reassessment of the Role of Phrase Extraction in PBSMT
Francisco Guzmán, Qin Gao, Stephan Vogel |
MTSummit | 1 |
| 2008 | Translation Paraphrases in Phrase-Based Machine Translation
Francisco Guzmán, Leonardo Garrido |
CICLing | 1 |