Gonzalo Iglesias

dblp:67/2550 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 48% Trustworthy machine learning · 24% Question answering and dialogue systems · 16%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.012026
Benchmarking Deflection and Hallucination in Large Vision-Language Models · ACL (1) 2026
Computer vision › Vision and language › vision-language model
vision-language model evaluation
1.012026
Benchmarking Deflection and Hallucination in Large Vision-Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination
1.012026
Benchmarking Deflection and Hallucination in Large Vision-Language Models · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
table question answering
0.712023
An Inner Table Retriever for Robust Table Question Answering · ACL (1) 2023
Information retrieval › search engines › structured data search
table retrieval
0.712023
An Inner Table Retriever for Robust Table Question Answering · ACL (1) 2023
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › language modeling for speech recognition
lattice rescoring
0.212015
Transducer Disambiguation with Sparse Topological Features · EMNLP 2015
Automata and formal languages › transducers
finite-state transducers
0.212015
Transducer Disambiguation with Sparse Topological Features · EMNLP 2015
Natural language and speech › Machine translation › statistical machine translation
hierarchical phrase-based translation
0.112011
Hierarchical Phrase-based Translation Representations · EMNLP 2011
Natural language and speech › Machine translation
statistical machine translation
0.112011
Hierarchical Phrase-based Translation Representations · EMNLP 2011
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.112015
Transducer Disambiguation with Sparse Topological Features · EMNLP 2015
Programming languages and type systems
grammar formalisms
0.012011
Hierarchical Phrase-based Translation Representations · EMNLP 2011
Programming languages and type systems › grammar formalisms
synchronous grammars
0.012011
Hierarchical Phrase-based Translation Representations · EMNLP 2011

Methods — techniques the papers use, named apart from their topics

transformer language model · 1.3tropical sparse tuple vector semiring · 0.4bilingual neural network language model · 0.4hierarchical phrase-based translation · 0.2
YearPublicationVenuePosition
2026 Benchmarking Deflection and Hallucination in Large Vision-Language Models
abstract
Nicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne, Gonzalo Iglesias. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nicholas Moratelli, Leonardo F. R. Ribeiro, William J. Byrne, Gonzalo Iglesias
ACL (1)5
2023 An Inner Table Retriever for Robust Table Question Answering
abstract
Recent years have witnessed the thriving of pretrained Transformer-based language models for understanding semi-structured tables, with several applications, such as Table Question Answering (TableQA).These models are typically trained on joint tables and surrounding natural language text, by linearizing table content into sequences comprising special tokens and cell information.This yields very long sequences which increase system inefficiency, and moreover, simply truncating long sequences results in information loss for downstream tasks.We propose Inner Table Retriever (ITR), 1 a generalpurpose approach for handling long tables in TableQA that extracts sub-tables to preserve the most relevant information for a question.We show that ITR can be easily integrated into existing systems to improve their accuracy with up to 1.3-4.8%and achieve state-of-the-art results in two benchmarks, i.e., 63.4% in Wik-iTableQuestions and 92.1% in WikiSQL.Additionally, we show that ITR makes TableQA systems more robust to reduced model capacity and to different ordering of columns and rows. * Work done as an intern at Amazon Alexa AI.
Weizhe Lin, Rexhina Blloshmi, William J. Byrne, Adrià de Gispert, Gonzalo Iglesias
ACL (1)5
2016 Speed-Constrained Tuning for Statistical Machine Translation Using Bayesian Optimization
abstract
Daniel Beck, Adrià de Gispert, Gonzalo Iglesias, Aurelien Waite, Bill Byrne. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Daniel Beck, Adrià de Gispert, Gonzalo Iglesias, Aurelien Waite, William J. Byrne
HLT-NAACL3
2015 Transducer Disambiguation with Sparse Topological Features
abstract
We describe a simple and efficient algorithm to disambiguate non-functional weighted finite state transducers (WFSTs), i.e. to generate a new WFST that contains a unique, best-scoring path for each hypothesis in the input labels along with the best output labels.The algorithm uses topological features combined with a tropical sparse tuple vector semiring.We empirically show that our algorithm is more efficient than previous work in a PoStagging disambiguation task.We use our method to rescore very large translation lattices with a bilingual neural network language model, obtaining gains in line with the literature.
Gonzalo Iglesias, Adrià de Gispert, William J. Byrne
EMNLP1
2015 Fast and Accurate Preordering for SMT using Neural Networks
abstract
We propose the use of neural networks to model source-side preordering for faster and better statistical machine translation.The neural network trains a logistic regression model to predict whether two sibling nodes of the source-side parse tree should be swapped in order to obtain a more monotonic parallel corpus, based on samples extracted from the word-aligned parallel corpus.For multiple language pairs and domains, we show that this yields the best reordering performance against other state-of-the-art techniques, resulting in improved translation quality and very fast decoding.
Adrià de Gispert, Gonzalo Iglesias, William J. Byrne
HLT-NAACL2
2014 Pushdown Automata in Statistical Machine Translation
abstract
This article describes the use of pushdown automata (PDA) in the context of statistical machine translation and alignment under a synchronous context-free grammar. We use PDAs to compactly represent the space of candidate translations generated by the grammar when applied to an input sentence. General-purpose PDA algorithms for replacement, composition, shortest path, and expansion are presented. We describe HiPDT, a hierarchical phrase-based decoder using the PDA representation and these algorithms. We contrast the complexity of this decoder with a decoder based on a finite state automata representation, showing that PDAs provide a more suitable framework to achieve exact decoding for larger synchronous context-free grammars and smaller language models. We assess this experimentally on a large-scale Chinese-to-English alignment and translation task. In translation, we propose a two-pass decoding strategy involving a weaker language model in the first-pass to address the results of PDA complexity analysis. We study in depth the experimental conditions and tradeoffs in which HiPDT can achieve state-of-the-art performance for large-scale SMT.
Cyril Allauzen, William J. Byrne, Adrià de Gispert, Gonzalo Iglesias, Michael Riley 0001
Comput. Linguistics4
2013 N-gram posterior probability confidence measures for statistical machine translation: an empirical study
abstract
We report an empirical study of n -gram posterior probability confidence measures for statistical machine translation (SMT). We first describe an efficient and practical algorithm for rapidly computing n -gram posterior probabilities from large translation word lattices. These probabilities are shown to be a good predictor of whether or not the n -gram is found in human reference translations, motivating their use as a confidence measure for SMT. Comprehensive n -gram precision and word coverage measurements are presented for a variety of different language pairs, domains and conditions. We analyze the effect on reference precision of using single or multiple references, and compare the precision of posteriors computed from k -best lists to those computed over the full evidence space of the lattice. We also demonstrate improved confidence by combining multiple lattices in a multi-source translation framework.
Adrià de Gispert, Graeme W. Blackwood, Gonzalo Iglesias, William J. Byrne
Mach. Transl.3
2012 Can Automatic Post-Editing Make MT More Meaningful
Kristen Parton, Nizar Habash, Kathy McKeown, Gonzalo Iglesias, Adrià de Gispert
EAMT4
2011 Hierarchical Phrase-based Translation Representations
Gonzalo Iglesias, Cyril Allauzen, William J. Byrne, Adrià de Gispert, Michael Riley 0001
EMNLP1
2010 Hierarchical Phrase-Based Translation with Weighted Finite-State Transducers and Shallow-n Grammars
abstract
In this article we describe HiFST, a lattice-based decoder for hierarchical phrase-based translation and alignment. The decoder is implemented with standard Weighted Finite-State Transducer (WFST) operations as an alternative to the well-known cube pruning procedure. We find that the use of WFSTs rather than k-best lists requires less pruning in translation search, resulting in fewer search errors, better parameter optimization, and improved translation performance. The direct generation of translation lattices in the target language can improve subsequent rescoring procedures, yielding further gains when applying long-span language models and Minimum Bayes Risk decoding. We also provide insights as to how to control the size of the search space defined by hierarchical rules. We show that shallow-n grammars, low-level rule catenation, and other search constraints can help to match the power of the translation system to specific language pairs.
Adrià de Gispert, Gonzalo Iglesias, Graeme W. Blackwood, Eduardo Rodríguez Banga, William J. Byrne
Comput. Linguistics2
2009 Rule Filtering by Pattern for Efficient Hierarchical Translation
Gonzalo Iglesias, Adrià de Gispert, Eduardo Rodríguez Banga, William J. Byrne
EACL1
2009 Hierarchical Phrase-Based Translation with Weighted Finite State Transducers
Gonzalo Iglesias, Adrià de Gispert, Eduardo Rodríguez Banga, William J. Byrne
HLT-NAACL1
2008 Specific features of the Galician language and implications for speech technology development
Manuel González González, Eduardo Rodríguez Banga, Francisco Campillo Díaz, Francisco Méndez Pazó, Leandro Rodríguez-Liñares, Gonzalo Iglesias
Speech Commun.6